Semantic Conflict Disambiguation Processing Method and Electronic Device Based on Entity Text

By detecting and judging the semantic conflict disambiguation type of entity text, querying the attribute database, avoiding semantic conflicts, the search accuracy problem caused by the ambiguous nature of entity words is solved, and the reliability of semantic search is ensured.

CN113935325BActive Publication Date: 2025-08-05HISENSE VISUAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111074721.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-14
Publication Date
2025-08-05
Estimated Expiration
2041-09-14

AI Technical Summary

Technical Problem

The ambiguous nature of the word entity word leads to a decrease in the accuracy of semantic recognition, affecting the accuracy of search results.

Method used

Detect the semantic conflict disambiguation type of the target entity text, obtain the target attribute to be added, query the pre-noted entity text attribute database and the first conflict attribute database, judge the conflict probability between the marked attribute and the target attribute, and output the semantic strong conflict identifier to avoid attribute addition.

Benefits of technology

This avoids conflicts in semantic annotations of entity texts and ensures the reliability of semantic search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113935325B_ABST
    Figure CN113935325B_ABST
Patent Text Reader

Abstract

The present disclosure provides a semantic conflict disambiguation processing method and an electronic device based on entity text. The method includes: detecting the semantic conflict disambiguation type of a target entity text to be processed, and obtaining a target attribute to be added to the target entity text when the semantic conflict disambiguation type is a first type; querying a pre-annotated entity text attribute database to obtain an annotated attribute corresponding to the target entity text; querying a preset first conflict attribute database to obtain a first conflict attribute set that matches the target attribute, where the conflict probability of the attribute pair combination in the first conflict attribute database is greater than a preset threshold; if it is detected that the first conflict attribute set contains the annotated attribute, outputting a strong semantic conflict identifier and not performing the operation of adding the target attribute to the target entity text. Thus, in the semantic annotation disambiguation processing of entity text, conflicts between annotations are avoided, ensuring the reliability of semantic-based search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This non - disclosure relates to the technical attributes of data processing. More specifically, it relates to a method for semantic conflict disambiguation processing based on entity text and an electronic device. Background Art

[0002] With the popularization of intelligent search, more and more entity semantics are introduced into the database. It is common to perform semantic recognition of entity words based on the database so as to conduct searches according to the recognition results. For example, search for web pages related to entity words according to the semantic recognition results.

[0003] However, the polysemy of entity words poses a greater challenge to the accuracy of semantic recognition. Usually, entity words with different semantics will lead to different search results. Therefore, in order to obtain semantic recognition results that meet the search requirements, it is crucial to perform disambiguation processing on the semantics annotated for entity words. Summary of the Invention

[0004] To solve the above - mentioned technical problems or at least partially solve the above - mentioned technical problems, the present disclosure provides a method, apparatus, device and medium for semantic conflict disambiguation processing based on entity text.

[0005] In a first aspect, the present disclosure provides a method for semantic conflict disambiguation processing based on entity text, including: detecting the semantic conflict disambiguation type of a target entity text to be processed; in the case where the semantic conflict disambiguation type is the first type, obtaining a target attribute to be added to the target entity text; querying a pre - annotated entity text attribute database to obtain an annotated attribute corresponding to the target entity text; querying a preset first conflict attribute database to obtain a first set of conflict attributes matching the target attribute, where the conflict probability of the attribute pairs in the first conflict attribute database is greater than a preset threshold; if it is detected that the first set of conflict attributes contains the annotated attribute, outputting a strong semantic conflict identifier and not performing the operation of adding the target attribute to the target entity text.

[0006] Second aspect, the present disclosure provides an electronic device, the electronic device includes: a processor; a memory configured to store executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement detecting a semantic conflict disambiguation type of a target entity text to be processed, and in a case where the semantic conflict disambiguation type is a first type, obtaining a target attribute to be added to the target entity text; querying a pre-annotated entity text attribute database to obtain an annotated attribute corresponding to the target entity text; querying a preset first conflict attribute database to obtain a first conflict attribute set matching the target attribute, wherein a conflict probability of an attribute pair combination in the first conflict attribute database is greater than a preset threshold; if it is detected that the first conflict attribute set contains the annotated attribute, outputting a semantic strong conflict identifier and not performing an adding operation of the target attribute on the target entity text.

[0007] Third aspect, the present disclosure provides a computer-readable storage medium, the storage medium stores a computer program, and the computer program is used to execute the above-mentioned semantic conflict disambiguation processing method based on entity text.

[0008] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:

[0009] Detecting the semantic conflict disambiguation type of the target entity text to be processed, and in a case where the semantic conflict disambiguation type is a first type, obtaining the target attribute to be added to the target entity text. Further, querying the pre-annotated entity text attribute database to obtain the annotated attribute corresponding to the target entity text, querying the preset first conflict attribute database to obtain the first conflict attribute set matching the target attribute, wherein the conflict probability of the attribute pair combination in the first conflict attribute database is greater than the preset threshold. If it is detected that the first conflict attribute set contains the annotated attribute, outputting a semantic strong conflict identifier and not performing an adding operation of the target attribute on the target entity text. Thus, in the semantic annotation disambiguation processing of entity text, conflicts between annotations are avoided, ensuring the reliability of semantic search. Description of the Drawings

[0010] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 A flowchart of a semantic conflict disambiguation processing method based on entity text proposed by the present disclosure;

[0013] Figure 2 A schematic diagram of a selection interface corresponding to a semantic conflict disambiguation type provided by an embodiment of the present disclosure;

[0014] Figure 3 A schematic diagram of the structure of an entity text attribute database provided by an embodiment of the present disclosure;

[0015] Figure 4 A schematic diagram of the structure of another entity text attribute database provided by an embodiment of the present disclosure;

[0016] Figure 5 A flowchart of another semantic conflict disambiguation processing method based on entity text proposed by the present disclosure;

[0017] Figure 6 A schematic diagram of the structure of another entity text attribute database provided by an embodiment of the present disclosure;

[0018] Figure 7 A flowchart of another semantic conflict disambiguation processing method based on entity text proposed by the present disclosure;

[0019] Figure 8 A flowchart of another semantic conflict disambiguation processing method based on entity text proposed by the present disclosure;

[0020] Figure 9 A schematic diagram of the structure of a knowledge graph proposed by the present disclosure;

[0021] Figure 10 A schematic diagram of the structure of another knowledge graph proposed by the present disclosure;

[0022] Figure 11 A schematic diagram of the structure of another knowledge graph proposed by the present disclosure;

[0023] Figure 12 A flowchart of another semantic conflict disambiguation processing method based on entity text proposed by the present disclosure;

[0024] Figure 13 A schematic diagram of a semantic conflict disambiguation processing scenario based on entity text proposed by the present disclosure;

[0025] Figure 14 A schematic diagram of the structure of a semantic conflict disambiguation processing device based on entity text according to an embodiment of the present disclosure. Detailed implementation manners

[0026] To make the objectives and implementation manners of this application clearer, the following will clearly and completely describe the exemplary implementation manners of this application in combination with the accompanying drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only a part rather than all of the embodiments of this application.

[0027] It should be noted that the brief description of terms in this application is only for facilitating the understanding of the subsequent described implementation manners, rather than intending to limit the implementation manners of this application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.

[0028] In this application, terms such as "first", "second", "third", etc. in the description, claims and the above accompanying drawings are used to distinguish similar or homogeneous objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.

[0029] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclusively include. For example, a product or device comprising a series of components does not necessarily have to be limited to all the clearly listed components, but may include other components not clearly listed or inherent to these products or devices.

[0030] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to this element.

[0031] Figure 1 As shown in the flowchart of a semantic conflict disambiguation processing method based on entity text proposed for this disclosure, Figure 1 as shown, this method includes:

[0032] Step 101, detecting the semantic conflict disambiguation type of the target entity text to be processed. In the case where the semantic conflict disambiguation type is the first type, obtaining the target attribute to be added to the target entity text.

[0033] The target entity text in this embodiment is the entity text to be attribute-labeled, and this entity text can correspond to the text of objectively existing and distinguishable things, including specific people, events, objects, institutions, abstract concept texts, etc.

[0034] Since in the actual execution process, when performing semantic annotation on the target entity text, there may be cases of polysemy, and there are ambiguities between different semantics, resulting in recognition errors in subsequent search scenario applications such as text recognition based on the annotated semantics, thus affecting the accuracy of search results.

[0035] Therefore, in this embodiment, in order to ensure the reliability of the semantics of the target entity text annotation, disambiguation processing is performed on the semantic annotation of the target text. Among them, the disambiguation of the semantic annotation can be examined based on different dimensions, and this dimension is related to the type of semantic conflict disambiguation of the target entity text.

[0036] In different application scenarios, different methods can be used to detect the type of semantic conflict disambiguation of the target entity text to be processed:

[0037] In some possible embodiments, as Figure 2 shown, trigger controls corresponding to different types of semantic conflict disambiguation are provided in advance on the relevant display interface, and the type of semantic conflict disambiguation of the target entity text to be processed is determined based on the user's trigger operation.

[0038] In some other possible embodiments, determine the input text type of the entity word corresponding to each type of semantic conflict disambiguation, and query the preset corresponding relationship based on the input text type to determine the type of semantic conflict disambiguation of the target entity text to be processed.

[0039] For example, when the type of semantic conflict disambiguation is the semantic conflict disambiguation of the attribute type, the input text type of the corresponding entity word is the attribute type, that is, the input is not the target entity text, but the attribute of the target entity text.

[0040] When the type of semantic conflict disambiguation is the syntactic disambiguation of the attribute type, the input text type of the corresponding entity word is the regular expression corresponding to the sentence pattern, etc.

[0041] In this embodiment, in the case where the type of semantic conflict disambiguation is the first type, it can be considered that the target entity text is semantically annotated in the attribute dimension. Thus, the target attribute to be added to the target entity text is obtained. Among them, the attribute can be understood as the characteristic of the target entity word. For example, for the target entity word "Xiaoming", its corresponding characteristics may include "name", "food name", etc.

[0042] Step 102, query the pre-annotated entity text attribute database to obtain the annotated attributes corresponding to the target entity text.

[0043] In this embodiment, the entity text attribute database is pre-annotated, as Figure 3 shown, there are multiple entity texts in this entity text attribute database, and one or more annotated attributes corresponding to each entity text. Among them, the attributes in the entity text attribute database have been subjected to semantic annotation disambiguation processing.

[0044] Thus, in this embodiment, the pre-annotated entity text attribute database can be queried to obtain the annotated attributes corresponding to the target entity text.

[0045] Step 103: Query the preset first conflict attribute database to obtain a set of first conflict attributes that match the target attribute. Among them, the conflict probability of the attribute pair combinations in the first conflict attribute database is greater than a preset threshold.

[0046] It is easy to understand that if there is ambiguity between the target attribute to be labeled of the target entity text and the already labeled attributes, then when performing semantic recognition based on the already labeled attributes and the target attribute, two mutually different ambiguous semantic recognition results may be obtained, resulting in two completely different search results, affecting the accuracy of the search results.

[0047] In this embodiment, whether such semantics is ambiguous is achieved by determining whether the already labeled attributes match a preset set of attributes with a large conflict with the target attribute.

[0048] In this embodiment, when constructing the entity text attribute database in advance, a first conflict attribute database is also constructed. As Figure 4 shown, the first conflict attribute database contains attribute pair combinations composed of the target attribute and the corresponding strongly conflicting attributes, that is, the conflict probability of the attribute pair combinations in the first conflict attribute database is greater than a preset threshold. Among them, the preset threshold can be calibrated according to experimental data.

[0049] Step 104: If it is detected that the set of first conflict attributes contains the already labeled attributes, output a strong semantic conflict identifier and do not perform the operation of adding the target attribute to the target entity text.

[0050] In this embodiment, if it is detected that the set of first conflict attributes contains the already labeled attributes, it indicates that among the already labeled attributes, there are attributes that have a large conflict with the target attribute. If the target attribute is labeled to the target entity text, it may cause semantic ambiguity of the target entity text. For example, if the target attribute of the target entity text A is "novel" and the set of first conflict attributes contains the already labeled attribute "movie", then if "novel" is also labeled as an attribute of the target entity text A, it is obvious that in subsequent semantic recognition, it is impossible to recognize whether the attribute of the target entity text A is "novel" or "movie", which may cause users who want to search for movies of the target entity text A to obtain corresponding novel search results.

[0051] Therefore, in this embodiment, in order to meet the semantic search service of relevant scenarios, if it is detected that the set of first conflict attributes contains the already labeled attributes, output a strong semantic conflict identifier and do not perform the operation of adding the target attribute to the target entity text, thereby avoiding attributes with strong mutual ambiguity in the semantics of the target entity word labeling.

[0052] Among them, the above semantic strong conflict identifier can be an identifier information that pre-agrees on the strong conflict between the target identifier and the annotated attribute. The strong conflict identifier information includes, but is not limited to, any form such as text, picture, code, etc. In this embodiment, after outputting the semantic strong conflict identifier for the target attribute, it is considered that if the target attribute is marked in the attributes of the target entity word text, it will cause a strong conflict with other annotated attributes, affecting the semantic understanding of the target entity word text. Therefore, the addition operation of the target attribute is not performed on the target entity text.

[0053] In some embodiments, referring to Figure 5 , after the above method step 103, the method further includes:

[0054] Step 501, if it is detected and known that the first conflict attribute set does not include the annotated attribute, query the preset second conflict attribute database to obtain the second conflict attribute set matching the target attribute.

[0055] Among them, the conflict probability of the attribute pair combination in the second conflict attribute database is less than or equal to the preset threshold.

[0056] In this embodiment, if it is detected and known that the first conflict attribute set does not include the annotated attribute, it indicates that there is no strong semantic conflict between the target attribute and the annotated attribute. At this time, in order to determine whether there is a weak semantic conflict between the target attribute and the annotated attribute, query the preset second conflict attribute database, where, as Figure 6 shown, the second conflict attribute database stores the attribute pair combination composed of the target attribute and the corresponding attribute with a weak conflict, that is, the conflict probability of the attribute pair combination in the first conflict attribute database is less than or equal to the preset threshold.

[0057] Step 502, if it is detected and known that the second conflict attribute set includes the annotated attribute, output the semantic weak conflict identifier, and perform the addition operation of the target attribute on the target entity text, and mark the target attribute as the second-level attribute.

[0058] In this embodiment, if it is detected and known that the second conflict attribute set includes the annotated attribute, it indicates that there is a semantic conflict between the target attribute and the annotated attribute, but the conflict is weak. Therefore, output the semantic weak conflict identifier, where the weak conflict identifier includes, but is not limited to, any form such as text, picture, code, etc.

[0059] In this embodiment, perform the addition operation of the target attribute on the target entity text, and mark the target attribute as the second-level attribute. The second-level attribute is used to indicate that the target attribute occupies a lower weight of the contribution of semantic understanding relative to the annotated attribute with which it conflicts. Therefore, on the search result, it has a lower presentation probability relative to the annotated attribute with which it conflicts.

[0060] For example, it can be pre - specified that, with respect to the conflicting labeled attributes, the proportion value of the search results obtained according to the attributes corresponding to the second - level attributes in the total search results is less than a certain value. Thus, by reducing the number of results of the second - level attributes corresponding to the second attributes, the accuracy of the search results and the relevance to the second - level attributes are balanced.

[0061] Step 503, if it is detected and known that the second conflict attribute set does not contain the labeled attribute, output a semantic non - conflict flag, and perform an addition operation of the target attribute on the target entity text.

[0062] In this embodiment, if it is detected and known that the second conflict attribute set does not contain the labeled attribute, it indicates that the target attribute has neither a strong conflict nor a weak conflict with the labeled attribute. Output a semantic non - conflict flag, which includes but is not limited to any form such as text, picture, code, etc. When the semantic non - conflict flag is output, perform an addition operation of the target attribute on the target entity text, thereby labeling the target attribute on the target entity text.

[0063] In summary, for the semantic conflict disambiguation processing method based on entity text in the embodiments of the present disclosure, detect the semantic conflict disambiguation type of the target entity text to be processed. In the case where the semantic conflict disambiguation type is the first type, obtain the target attribute to be added to the target entity text. Further, query the pre - labeled entity text attribute database to obtain the labeled attributes corresponding to the target entity text, and query the preset first conflict attribute database to obtain the first conflict attribute set matching the target attribute. Among them, the conflict probability of the attribute pair combination in the first conflict attribute database is greater than the preset threshold. If it is detected and known that the first conflict attribute set contains the labeled attribute, output a semantic strong conflict flag and do not perform the addition operation of the target attribute on the target entity text. Thus, in the semantic annotation disambiguation processing of entity text, conflicts between annotations are avoided, ensuring the reliability of semantic - based search.

[0064] Based on the above description, in order to perform disambiguation processing on the semantic annotation of the target attribute, it is necessary to pre - construct the first conflict attribute data and the second conflict attribute database.

[0065] In an embodiment of the present disclosure, as Figure 7 shown, the method further includes:

[0066] Step 701, according to the labeled attribute sets of multiple sample entity texts, determine the word - frequency coefficients of each attribute pair combination in each sample entity text.

[0067] In this embodiment, the set of attributes already labeled for multiple sample entity texts contains each sample entity text and all the attributes corresponding to that sample entity text. Among them, the attributes corresponding to a sample entity text may be one or more. For example, for the sample entity text A, its corresponding attributes in the set of attributes already labeled for sample entity texts are "novel" and "movie", etc.

[0068] To determine the conflict situation between different attributes, the frequency of two attributes being used as the attributes of a sample entity text simultaneously is counted. The greater this frequency, the more it indicates that these two attributes can be used as the attributes of an entity text simultaneously, and thus the probability of conflict is smaller. Based on this logic, in this embodiment, the word frequency coefficient of each attribute pair combination in each sample entity text can be determined.

[0069] This attribute pair combination refers to all the attribute pairs that appear in the same sample entity text. For example, for the sample entity text B, its corresponding attributes in the set of attributes already labeled for sample entity texts are "a, b, c", then the attribute pair combinations corresponding to the sample entity text B are "a, b", "b, c", and "a, c".

[0070] In this embodiment, the word frequency coefficient is related to the number of times each attribute pair combination appears in the set of attributes already labeled for multiple sample entity texts. In one embodiment of the present disclosure, first, the number of times each attribute pair combination appears is counted, and the maximum number of appearances is statistically obtained according to the number of appearances. Secondly, the ratio of each attribute pair combination to the maximum number of appearances is calculated as the corresponding attribute pair combination.

[0071] That is, in this embodiment, traverse the set of attributes already labeled for multiple sample entity texts, obtain the attribute pair combinations of all sample entity texts, determine the maximum frequency of the attribute pair combinations according to the attribute pair combinations of the sample entity texts, and determine the word frequency coefficient of each attribute pair combination in each sample entity text according to the frequency of each attribute pair combination of the sample entity text and the maximum frequency of the attribute pair combination. For example, determine the word frequency coefficient of each attribute pair combination in each sample entity text according to the ratio of the frequency of each attribute pair combination of the sample entity text to the maximum frequency of the attribute pair combination.

[0072] Step 702: Determine the attribute conflict coefficient between each labeled attribute and the ambiguous attribute in each sample entity text according to the case corpus associated with multiple sample entity texts.

[0073] During the actual execution process, when a sample entity text with a labeled attribute is recognized for attributes, it is possible that the recognized attribute is the same as the labeled attribute, or it is possible that the recognized attribute is inconsistent with the labeled attribute. When the attributes are inconsistent, the recognized attribute is considered an ambiguous attribute.

[0074] In this embodiment, a case corpus containing labeled attributes and identified attributes of sample entity texts is obtained, and an attribute conflict coefficient between each labeled attribute and the ambiguous attribute in each sample entity text is determined.

[0075] In some possible embodiments, a use case corpus containing multiple sample entity texts and preset attributes is obtained, wherein the preset attributes are annotated attributes of the sample entity texts. For example, since "A is a novel" contains both the sample entity text "A" and the annotated attribute "novel", it is considered that "A is a novel" corresponds to the use case corpus of the sample entity text, etc.

[0076] Then, semantic parsing is performed on the associated use case corpus according to the annotated attributes of each sample entity text to obtain the positioning attributes of the sample entity text. The positioning attributes here can be understood as the parsed attributes.

[0077] In this embodiment, when the preset attributes and the positioning attributes are consistent, the corresponding use case corpus is considered to be the positive use case corpus, thereby obtaining the number of positive use case corpora associated with the annotated attributes; when the preset attributes and the positioning attributes are inconsistent, the corresponding use case corpus is considered to be the negative use case corpus, thereby obtaining the number of negative use case corpora associated with the annotated attributes and the parsed ambiguous attributes.

[0078] Furthermore, based on the number of positive example corpora and the number of negative example corpora, the attribute conflict coefficient between each labeled attribute and the ambiguous attribute in each sample entity text is determined.

[0079] For example, the ratio of the number of negative example corpora to the number of positive example corpora of each sample entity text is determined as the attribute conflict coefficient between each labeled attribute and the ambiguous attribute in each sample entity text.

[0080] For example, the sum of the number of inverse example corpora and the number of positive example corpora of each sample entity text is determined, and the ratio of the number of inverse example corpora and the sum of the numbers of positive example corpora of each sample entity text is calculated as the attribute conflict coefficient between each labeled attribute and the ambiguous attribute in each sample entity text.

[0081] Step 703 : determining the conflict probability of each attribute pair combination based on the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each marked attribute and the ambiguous attribute.

[0082] It is easy to understand that the word frequency coefficient of each attribute pair combination represents two different attributes, and at the same time, the possibility of being used as the attributes of the same sample entity text. The attribute conflict coefficient between the labeled attribute and the ambiguous attribute represents the attributes that are likely to have parsing conflicts between the two attributes of each attribute pair combination. Based on the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute, the conflict probability of each attribute pair combination is determined, which not only considers the possibility that the two attributes can be used to label the attributes of the same entity sample at the same time, but also considers the possibility of parsing conflicts, ensuring the accuracy of the conflict probability calculation.

[0083] It should be noted that in different application scenarios, the method of determining the conflict probability of each attribute pair combination according to the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute is different:

[0084] In one embodiment of the present disclosure, as Figure 8 shown, determining the conflict probability of each attribute pair combination according to the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute includes:

[0085] Step 801, obtain the conflict probability between directly associated attribute pair combinations according to the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute.

[0086] In this embodiment, an attribute pair can be two directly associated attributes or two indirectly associated attributes. For example, for the same sample entity text, if it appears twice in the attribute pair combinations corresponding to the labeled attribute set, and the two labeled attribute pair combinations corresponding to the two sample entity texts are "ac" and "ab" respectively, then for "ac" and "ab", they are obviously directly associated attribute pair combinations. For "bc", it is actually indirectly connected through the "a" attribute. Therefore, the attribute pair combination "bc" belongs to an indirectly associated attribute pair combination.

[0087] In one embodiment of the present disclosure, the relationship between the attribute pair combinations of the sample entity text can be known to be indirect or direct by constructing a knowledge graph. Among them, the knowledge graph: essentially an attribute network, which can represent the attribute relationship between entities. In the knowledge graph, attributes are used as vertices or nodes, and relationships are used as edges. The knowledge graph can be constructed in various ways. The focus of the embodiments of the present disclosure is not on how to construct the knowledge graph, so it will not be described in detail here. As a possible implementation, as Figure 9As shown, when the sample entity text D is in the set of labeled attributes, and the labeled attributes are "ab", "cb", "cd", "ef", "fd" respectively, based on constructing the knowledge graph between attributes, the association relationships between different attributes can be intuitively obtained. For example, in Figure 9 it can be intuitively known that "ab", "cb", "cd", "ef", "fd" are directly associated attribute pair combinations, and "ac", "ad", "ae", "af", "bd", etc. are indirectly associated attribute pair combinations.

[0088] In this embodiment, the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute can be input into a deep learning model constructed in advance according to experimental data, and the conflict probability between directly associated attribute pair combinations can be obtained based on the output of the deep learning model.

[0089] In another embodiment of the present disclosure, according to a preset algorithm, the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute are calculated to obtain the conflict probability between directly associated attribute pair combinations.

[0090] Among them, the preset algorithm can be as shown in the following formula (1). In formula (1), p1 ij +p2 ij =1, i and j are respectively two attributes in the directly connected attribute pair combination, α ij is the word frequency coefficient of the attribute pair combination, and β ij is the attribute conflict coefficient that attribute i is misparsed as ambiguous attribute j.

[0091] P ij =p1 ij α ij +p2 ij β ij Formula (1)

[0092] Step 802, according to the preset attenuation factor, the word frequency coefficient of each attribute pair combination, and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute, obtain the conflict probability between indirectly associated attribute pair combinations.

[0093] In this embodiment, for indirectly associated attribute pair combinations, according to the preset attenuation factor, the word frequency coefficient of each attribute pair combination, and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute, to obtain the conflict probability between indirectly associated attribute pair combinations.

[0094] In this embodiment, the preset attenuation factor can be calibrated according to experimental data. [[ID=

[0095] In some possible embodiments, the conflict probability between combinations of directly associated attribute pairs related to combinations of indirectly associated attributes can be calculated, and based on the product value of the conflict probabilities between combinations of directly associated attribute pairs concerned, the conflict probability between combinations of indirectly associated attributes can be determined.

[0096] For example, for the combination of attribute pairs ik that are indirectly associated, as Figure 10 shown, there is one path of ik in the knowledge graph. In this path, the combinations of directly associated attribute pairs corresponding to ik are ij and ik. Therefore, the conflict probability P ij and P jk between the combinations of directly associated attribute pairs are obtained, and the corresponding attenuation factor is λ. Then, the calculation method of the conflict probability of the combination of indirectly associated attribute pairs ik is as shown in the following formula (2):

[0097] P ik = P ij * P jk * λ Formula (2)

[0098] In this embodiment, when there are multiple paths indirectly associated in the corresponding indication graph for the combination of indirectly associated attribute pairs, after calculating the conflict probability corresponding to each path, the sum of the conflict probabilities of all paths can be calculated as the conflict probability of the combination of indirectly associated attribute pairs.

[0099] For example, as Figure 11 mentioned, when there are two paths of ik in the knowledge graph, in one of the paths, the combinations of directly associated attribute pairs corresponding to ik are ij and ik. Therefore, the conflict probability P ij and P jk between the combinations of directly associated attribute pairs are obtained, and the corresponding attenuation factor is λ1. Then, according to the above formula (2), the conflict probability of the combination of indirectly associated attribute pairs ik on this path is calculated. In the other path, the combinations of directly associated attribute pairs corresponding to ik are im and mk. Therefore, the conflict probability P im and P mk between the combinations of directly associated attribute pairs are obtained, and the corresponding attenuation factor is λ2. Then, according to the above formula (2), the conflict probability of the combination of indirectly associated attribute pairs ik on this path is calculated, and the sum of the conflict probabilities on the two paths is calculated as the final conflict probability of ik.

[0100] Step 704, compare the conflict probability of each combination of attribute pairs with a preset threshold.

[0101] Among them, the preset threshold is calibrated according to experimental data.

[0102] In this embodiment, the conflict probability of each attribute pair combination is compared with a preset threshold to determine whether the conflict of each attribute pair combination is strong.

[0103] Step 705: Obtain the attribute pair combinations greater than the threshold to establish a first conflict attribute database, and obtain the attribute pair combinations less than or equal to the threshold to establish a second conflict attribute database.

[0104] In this embodiment, if the conflict probability of an attribute pair combination is greater than the preset threshold, it is considered that the two attributes in the corresponding attribute pair combination have a relatively strong conflict. If two attributes with a strong conflict are used to annotate the entity text, it will lead to an error in the semantic recognition of the entity text. Therefore, in this embodiment, the attribute pair combinations greater than the threshold are obtained to establish a first conflict attribute database, so as to avoid the annotation of strong conflict attributes based on the attribute pairs in the first conflict attribute database.

[0105] If the conflict probability of an attribute pair combination is not greater than the preset threshold, it indicates that although there is a conflict between the corresponding attributes, the conflict is not strong. Therefore, if two attributes without a strong conflict are used to annotate the entity text, it will not lead to an error in the semantic recognition of the entity text. Therefore, in this embodiment, the attribute pair combinations less than or equal to the threshold are obtained to establish a second conflict attribute database, so as to implement the annotation of the attributes of the relevant entity text based on the second conflict attribute database.

[0106] In summary, the semantic conflict disambiguation processing method based on entity text in the embodiments of the present disclosure takes the attribute pair combination as the granularity, and jointly establishes a first conflict attribute database and a second conflict attribute database based on the word frequency coefficient and attribute conflict coefficient between attribute pairs, providing technical support for annotating attributes based on the first conflict attribute database and the second conflict attribute database.

[0107] Based on the above embodiments, when the semantic conflict disambiguation type can also be the syntactic disambiguation of the attribute type, the input text type of the corresponding entity word is the regular expression corresponding to the sentence pattern.

[0108] In this embodiment, when the semantic conflict disambiguation type is the second type, it can be considered that the target entity text is semantically annotated in the syntactic dimension, and thus, the disambiguation processing of semantic annotation is performed in the syntactic dimension.

[0109] In this embodiment, as Figure 12 shown, after detecting the semantic conflict disambiguation type of the target entity text to be processed, it further includes:

[0110] Step 1201: When the semantic conflict disambiguation type of the target entity text is the second type, extract the regular expression of the target entity text.

[0111] Among them, regular expressions are used to indicate the syntactic features of the target entity text, etc. The syntactic features include the attributes of the constituent words and the order of the attributes of the constituent words (for example, for the target entity text "What's the weather like today", its corresponding regular expression is: today (date attribute) weather (weather attribute) what's (question attribute)). Or, the syntactic features can include the constituent keyword texts (for example, for the target entity text "I want to call my dad", its corresponding regular expression is: **want** to call) and so on.

[0112] Step 1202: Detect whether the regular expression of the target entity text is regularly matched with the regular expressions in the preset semantic slot template library.

[0113] In this embodiment, the preset semantic slot template library contains the regular expressions of the corresponding semantic templates. Therefore, it is necessary to detect whether the regular expression of the target entity text is regularly matched with the regular expressions in the preset semantic slot template library.

[0114] In some possible embodiments, extract the attribute regular expressions corresponding to multiple sample entity texts, establish a corresponding semantic slot attribute template library, split the target entity text, label the word segmentation attributes according to the splitting results, determine the target attribute regular expression according to the word segmentation attributes, and then detect whether the target attribute regular expression matches each attribute regular expression in the semantic slot attribute template library. That is, in this embodiment, the syntactic feature is the attribute regular expression feature of the entity text, that is, the feature that the target entity text contains the attributes of words mentioned above, etc.

[0115] In some possible embodiments, obtain the text regular expressions corresponding to multiple sample entity texts, establish a corresponding semantic slot text template library, split the target entity text, extract the target text regular expression according to the splitting results, and detect whether the target text regular expression matches each text regular expression in the semantic slot text template library. That is, in this embodiment, the syntactic feature is the text regular feature, that is, the feature of the text keywords contained in the target entity text mentioned above, etc. Step 1203: If it is detected that the regular expression of the target entity text is regularly matched with the regular expressions in the semantic slot template library, output a semantic conflict flag and do not perform an information addition operation on the target entity text.

[0116] In this embodiment, if it is detected that the regular expression of the target entity text is regularly matched with the regular expressions in the semantic slot template library, it means that the regular expression corresponding to this target entity text has been marked in the semantic slot template library. Therefore, output a semantic conflict flag to not perform an information addition operation on the target entity text.

[0117] In some possible embodiments, when the regular expression is related to the intention of the above-mentioned target entity text, if it is detected that the regular expression of the target entity text exactly matches the target template in the semantic slot template library, semantic parsing is performed based on the associated corpus of the target entity text to obtain the target intention.

[0118] Among them, the associated corpus can be the corpus with a higher frequency screened according to the sentence frequency in the corpus corresponding to each vertical domain. Based on the corpus with a higher frequency, templates in the corresponding semantic slot template library are defined. In this embodiment, the template can be the regular expression of the commonly used corpus sentence pattern in the vertical domain corresponding to the target entity text, etc. Each regular expression is pre-parsed and stored with the intention under the vertical domain. For example, in the communication field, there are intentions such as "make a call" and "hang up the phone".

[0119] Furthermore, obtain the preset original intention corresponding to the target template, and compare whether the original intention and the target intention are consistent. Among them, obtaining the preset original intention corresponding to the target template can be achieved based on semantic recognition technology, etc.

[0120] In this embodiment, if it is determined through comparison that the original intention and the target intention are inconsistent, a semantic content conflict flag is output, and no information addition operation is performed on the target entity text. For example, if the target entity text is "I want to call my dad", the original intention corresponding to the target template is "make a call", while the target intention is "play music", that is, play the song "I want to call my dad", then obviously the original intention and the target intention are inconsistent. Therefore, in order to avoid subsequent semantic recognition ambiguity, no information addition operation is performed on the target entity text.

[0121] In this embodiment, if it is determined through comparison that the original intention and the target intention Figure 1 are consistent, a semantic quantity conflict flag is output, that is, there is already a template in the semantic slot template library that is consistent with the intention of the target entity text. Figure 1 For example, for the target entity text "I want to call my dad" with the target intention of "make a call", there is already a template "call my dad" with the original intention of "make a call" in the semantic slot template library. Therefore, here there is no need to add the target entity text "I want to call my dad" to the corresponding semantic slot template library.

[0122] In another embodiment of the present disclosure, after detecting whether the regular expression of the target entity text exactly matches the preset semantic slot template library, if it is detected that the regular expression of the target entity text does not exactly match the semantic slot template library, it is determined that there is no semantic conflict, an information addition operation is performed on the target entity text, and the target intention is marked for the target entity text.

[0123] In practical applications, such as Figure 13As shown, the entity text attribute database and the semantic slot template database after disambiguation annotation can be stored in a server, which can be a local server or a cloud server, etc. The server communicates with the TV set, obtains the control text input by the user for the TV drama, and after sending the control text to the server, performs semantic recognition on the control text based on the entity text attribute database and the semantic slot template database stored in the server, generates a control instruction according to the semantic recognition result, and controls the TV set to perform relevant operations.

[0124] For example, when the control text is "Please play Mom and Dad", after semantic recognition based on the server, it is found that the semantic recognition result corresponding to this control text in the vertical domain of the TV set is "Play the TV drama 'Mom and Dad'", so the TV set is controlled to play the corresponding TV drama.

[0125] In summary, for the semantic conflict disambiguation processing method based on entity text in the embodiments of the present disclosure, when the semantic conflict disambiguation type of the target entity text is the second type, the regular expression of the target entity text is extracted, and it is detected whether the regular expression of the target entity text is regularly matched with the preset semantic slot template database. If it is detected that the regular expression of the target entity text is regularly matched with the semantic slot template database, a semantic conflict identifier is output, and no information addition operation is performed on the target entity text. Thus, the disambiguation processing of semantic annotation based on the sentence pattern dimension is realized, avoiding semantic conflicts between the marked regular expression and the regular expressions already existing in the semantic slot template database, etc., and ensuring the reliability of semantic search.

[0126] To implement the above embodiments, the present disclosure also proposes a semantic conflict disambiguation processing device based on entity text. Figure 14 It is a schematic structural diagram of a semantic conflict disambiguation processing device based on entity text according to an embodiment of the present disclosure, as Figure 14 shown. The semantic conflict disambiguation processing device based on entity text may include a first acquisition module 1410, a second acquisition module 1420, a third acquisition module 1430, and a marking module 1440, where

[0127] The first acquisition module 1410 is configured to detect the semantic conflict disambiguation type of the target entity text to be processed, and when the semantic conflict disambiguation type is the first type, acquire the target attribute to be added to the target entity text;

[0128] The second acquisition module 1420 is used to query the pre-annotated entity text attribute database and acquire the annotated attributes corresponding to the target entity text;

[0129] The third acquisition module 1430 is used to query the preset first conflict attribute database and acquire the first conflict attribute set matching the target attribute, where the conflict probability of the attribute pairs in the first conflict attribute database is greater than the preset threshold;

[0130] A labeling module 1440, configured to output a semantic strong conflict identifier when it is detected that the first conflict attribute set contains labeled attributes, and not perform the operation of adding target attributes to the target entity text.

[0131] It should be noted that the foregoing explanation of the embodiments of the semantic conflict disambiguation processing method based on entity text also applies to the semantic conflict disambiguation processing device based on entity text in the embodiments of the present disclosure. Their implementation principles are similar and will not be elaborated here.

[0132] To implement the above embodiments, the present disclosure also proposes an electronic device, which may be a television, a computer, etc. In this embodiment, the electronic device includes: a processor; a memory for storing processor-executable instructions; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the semantic conflict disambiguation processing method based on entity text described in the above embodiments.

[0133] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium storing a computer program for executing the above semantic conflict disambiguation processing method based on entity text.

[0134] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

[0135] For the sake of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, many modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.

Claims

1. A semantic conflict disambiguation processing method based on entity text, characterized in that: include: Detecting a semantic conflict disambiguation type of a target entity text to be processed, and when the semantic conflict disambiguation type is a first type, obtaining a target attribute to be added to the target entity text; Querying a pre-annotated entity text attribute database to obtain annotated attributes corresponding to the target entity text; Querying a preset first conflict attribute database to obtain a first conflict attribute set that matches the target attribute, wherein each attribute pair combination in the first conflict attribute database corresponds to the same entity text, and the conflict probability of each attribute pair combination is greater than a preset threshold; If it is detected that the first conflict attribute set includes the marked attribute, a semantically strong conflict flag is output, and the operation of adding the target attribute is not performed on the target entity text.

2. The method according to claim 1, characterized in that Also includes: If it is detected that the first conflict attribute set does not include the marked attribute, querying a preset second conflict attribute database to obtain a second conflict attribute set that matches the target attribute, wherein the conflict probability of the attribute pair combination in the second conflict attribute database is less than or equal to a preset threshold; If it is detected that the second conflict attribute set includes the marked attribute, a semantic weak conflict identifier is output, and an adding operation of the target attribute is performed on the target entity text, and the target attribute is marked as a second-level attribute; If it is detected that the second conflict attribute set does not include the marked attribute, a semantic non-conflict flag is output, and an operation of adding the target attribute is performed on the target entity text.

3. The method according to claim 2, characterized in that Also includes: Determining, based on attribute sets annotated for a plurality of sample entity texts, a word frequency coefficient for each attribute pair combination in each of the sample entity texts; Determining, based on the use case corpus associated with the plurality of sample entity texts, an attribute conflict coefficient between each labeled attribute and the ambiguous attribute in each of the sample entity texts; Determining the conflict probability of each attribute pair combination according to the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute; Comparing the conflict probability of each attribute pair combination with a preset threshold; The first conflicting attribute database is established by acquiring attribute pair combinations greater than the preset threshold, and the second conflicting attribute database is established by acquiring attribute pair combinations less than or equal to the preset threshold.

4. The method according to claim 3, characterized in that Determining the attribute conflict coefficient between each labeled attribute and the ambiguous attribute in each of the sample entity texts based on the use case corpus associated with the plurality of sample entity texts includes: Obtain a use case corpus containing multiple sample entity texts and preset attributes; Performing semantic parsing on the associated use case corpus according to the annotated attributes of each sample entity text to obtain the positioning attributes of the sample entity text; When the preset attribute and the positioning attribute are consistent, the number of positive example corpora associated with the annotated attribute is obtained; and when the preset attribute and the positioning attribute are inconsistent, the number of negative example corpora associated with the annotated attribute and the parsed ambiguous attribute is obtained; According to the number of the positive example corpus and the number of the negative example corpus, an attribute conflict coefficient between each labeled attribute and the ambiguous attribute in each of the sample entity texts is determined.

5. The method according to claim 3, characterized in that Determining the conflict probability of each attribute pair combination according to the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute includes: Obtaining the conflict probability between directly associated attribute pair combinations according to the word frequency coefficient of each attribute pair combination and the attribute conflict coefficient between each labeled attribute and the ambiguous attribute; The conflict probability between indirectly associated attribute pair combinations is obtained according to a preset attenuation factor, a word frequency coefficient of each attribute pair combination, and an attribute conflict coefficient between each labeled attribute and the ambiguous attribute.

6. The method according to any one of claims 1 to 5, characterized in that: After detecting the semantic conflict disambiguation type of the target entity text to be processed, the method further includes: When the semantic conflict disambiguation type of the target entity text is the second type, extracting a regular expression of the target entity text; Detect whether the regular expression of the target entity text matches the regular expression of the preset semantic slot template library; If it is detected that the regular expression of the target entity text matches the regular expression of the semantic slot template library, a semantic conflict identifier is output and no information adding operation is performed on the target entity text.

7. The method according to claim 6, characterized in that If it is detected that the regular expression of the target entity text matches the regular expression of the semantic slot template library, a semantic conflict flag is output, and no information adding operation is performed on the target entity text, including: If it is detected that the regular expression of the target entity text matches the target template in the semantic slot template library, semantic parsing is performed based on the associated corpus with the target entity text to obtain the target intent; Obtaining a preset original intent corresponding to the target template, and comparing whether the original intent is consistent with the target intent; If the comparison shows that the original intent and the target intent are inconsistent, a semantic content conflict flag is output, and no information addition operation is performed on the target entity text; If the comparison shows that the original intent is consistent with the target intent, a semantic quantity conflict flag is output, and no information adding operation is performed on the target entity text.

8. The method according to claim 6, characterized in that After detecting whether the regular expression of the target entity text matches the regular expression of the preset semantic slot template library, the method further includes: If it is detected that the regular expression of the target entity text does not match the regular expression of the semantic slot template library, it is determined that there is no semantic conflict, an information addition operation is performed on the target entity text, and the target intent is marked for the target entity text.

9. The method according to claim 6, characterized in that The detecting whether the regular expression of the target entity text matches the regular expression of the preset semantic slot template library includes: Obtain the text regular expressions corresponding to multiple sample entity texts and establish the corresponding semantic slot text template library; Splitting the target entity text, and extracting the target text regular expression according to the splitting results; Detect whether the target text regular expression matches each text regular expression in the semantic slot text template library.

10. An electronic device, characterized in that: The electronic device includes: a processor; a memory configured to store instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement: Detecting a semantic conflict disambiguation type of a target entity text to be processed, and when the semantic conflict disambiguation type is a first type, obtaining a target attribute to be added to the target entity text; Querying a pre-annotated entity text attribute database to obtain annotated attributes corresponding to the target entity text; Querying a preset first conflict attribute database to obtain a first conflict attribute set that matches the target attribute, wherein each attribute pair combination in the first conflict attribute database corresponds to the same entity text, and the conflict probability of each attribute pair combination is greater than a preset threshold; If it is detected that the first conflict attribute set includes the marked attribute, a semantically strong conflict flag is output, and the operation of adding the target attribute is not performed on the target entity text.

Citation Information

Patent Citations

  • System and method for associating documents with contextual advertisements

    CN1871597A