Knowledge Graph-Based Entity Description Extraction Method, Apparatus, and Computing Device
Through the entity description extraction method based on the knowledge graph, the confidence of the entity description is calculated and the backup entity description is screened out, and the problems of noise and low-quality content in the recommended reason extraction in the prior art are solved, achieving more efficient and reliable knowledge acquisition.
Patent Information
- Application Number
- CN201910435222.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2039-05-23
AI Technical Summary
The prior art has noise problems and low-quality content in mining recommendation reasons, and it is difficult to effectively automatically screen and compare the extracted content, resulting in low-quality knowledge acquisition efficiency and high cost.
Through a knowledge graph-based entity description extraction method, the entity description set is extracted from the knowledge graph database for a given entity, the confidence of each entity description is calculated, and the alternate entity description is filtered out for display.
This method does not need to rely on large-scale annotation data, and is suitable for extraction of unsupervised entity descriptions of large-scale open domains. By considering the content information of entity descriptions, the calculation results are more reliable, and better alternative entity descriptions are selected, which improves the trust of the user experience and search system.
Smart Images

Figure CN111984794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and particularly to a method, an apparatus, and a computing device for extracting entity descriptions based on a knowledge graph. Background Art
[0002] In the process of development, modern search engines continuously improve their product functions to provide users with more comprehensive and convenient information services. Among them, the right-side recommendation system displays entities related to a user query in the form of pictures, corresponding names, and category labels to the user, enabling the user to more conveniently understand and query other relevant knowledge information, thereby obtaining a good user experience. Figure 1 The schematic diagram of the right-side recommendation when there is no recommendation reason for searching for Pomeranians is shown. As Figure 1 shown, the user can obtain other knowledge information related to Pomeranians according to the recommended content in the four sections of "Guessed You Like", "Related Organisms", "Guessed You Follow", and "What Others Also Searched" recommended on the right side. However, most users do not understand the working principle behind the recommendation system and are not particularly familiar with the content recommended by the recommendation system. In this case, the search experience of the user and the trust in the search system can be greatly improved by presenting the recommendation reason in an intuitive form. A relatively typical method is to display the introduction of the recommended entity itself to the user.
[0003] In the prior art, the mining of recommendation reasons is usually implemented based on a template-based approach. Among them, the sources of the templates are mainly the following two types: First, a semi-automatic acquisition method based on high-quality knowledge triple seeds for BootsTrapping. This method assumes that a sentence in which two entity words appear simultaneously describes the relationship between the entity pair. Although it is applicable to large-scale knowledge extraction scenarios, as the number of iterations increases, it is easy to introduce noise and cause a sharp decline in system performance. Second, a method defined by human experts. The results extracted by the templates of this type of method generally have high quality. However, due to the flexibility and diversity of natural language expressions, it is still inevitable that there will be cases of incorrect or low-quality extraction content. Therefore, how to enable the machine to automatically compare and screen the extracted content is a challenge in the field of knowledge extraction for reducing labor costs and improving the efficiency of knowledge acquisition.
[0004] However, in terms of the calculation of knowledge quality, one of the main existing methods is to calculate the support degree of the template and the confidence degree of the knowledge through frequency statistics, and then quantify the quality of the extraction results. However, frequency is only an external factor for measuring knowledge quality, and it cannot comprehensively measure knowledge quality. For example, some knowledge may have a high quality although its frequency of occurrence is low; moreover, this quantification method cannot further distinguish knowledge with the same frequency statistics results. In addition, another main method is to manually label a batch of high-quality data, and then sort the sentences through ranking learning. This method is time-consuming and laborious, with too high a cost, and cannot be well extended in the open domain. Summary of the Invention
[0005] In view of the above problems, the present invention is proposed to provide a method, apparatus and computing device for entity description extraction based on a knowledge graph that overcomes the above problems or at least partially solves the above problems.
[0006] According to one aspect of the present invention, there is provided a method for entity description extraction based on a knowledge graph, including:
[0007] Step S1, for a given entity, extract the entity description set of this entity from the knowledge graph database;
[0008] Step S2, for each non-repeated entity description in the entity description set, calculate the confidence degree of this entity description according to the similarity between each non-repeated entity description and this entity description, and the frequency of each non-repeated entity description appearing in the entity description set;
[0009] Step S3, according to the confidence degrees of each entity description, screen out at least one entity description from the entity description set as the alternative entity description of this entity.
[0010] According to another aspect of the present invention, there is provided a method for pushing search engine recommendation reasons, including:
[0011] Obtain the recommendation results of the search engine;
[0012] Use the recommendation results as the given entity, and use the above-mentioned method for entity description extraction based on a knowledge graph to obtain the alternative entity description corresponding to the recommendation results;
[0013] Select an entity description from the alternative entity descriptions as a recommendation reason and present it on the search result display page.
[0014] According to yet another aspect of the present invention, there is provided an apparatus for entity description extraction based on a knowledge graph, including:
[0015] An extraction module, adapted to extract a set of entity descriptions of a given entity from a knowledge graph database for the given entity;
[0016] A confidence calculation module, adapted to calculate the confidence of each non-repeated entity description in the set of entity descriptions according to the similarity between each non-repeated entity description and the entity description, and the frequency of occurrence of each non-repeated entity description in the set of entity descriptions;
[0017] A screening module, adapted to screen out at least one entity description from the set of entity descriptions as an alternative entity description of the entity according to the confidence of each entity description.
[0018] According to another aspect of the present invention, there is provided a device for pushing search engine recommendation reasons, including:
[0019] An acquisition module, adapted to acquire the recommendation results of a search engine;
[0020] An entity description extraction module, adapted to use the above-mentioned entity description extraction device based on a knowledge graph with the recommendation results as a given entity to obtain an alternative entity description corresponding to the recommendation results;
[0021] A selection module, adapted to select an entity description from the alternative entity descriptions as a recommendation reason for presentation on a search result display page.
[0022] According to one aspect of the present invention, there is provided a computing device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0023] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above-mentioned entity description extraction method based on a knowledge graph.
[0024] According to another aspect of the present invention, there is provided a computing device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0025] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above-mentioned method for pushing search engine recommendation reasons.
[0026] According to still another aspect of the present invention, there is provided a computer storage medium, and at least one executable instruction is stored in the storage medium, and the executable instruction causes a processor to perform operations corresponding to the above-mentioned entity description extraction method based on a knowledge graph.
[0027] According to another aspect of the present invention, there is provided a computer storage medium storing at least one executable instruction, and the executable instruction causes a processor to perform operations corresponding to the above-mentioned push method for search engine recommendation reasons.
[0028] For the entity description extraction method, device and computing device based on a knowledge graph according to the present invention, for a given entity, an entity description set is extracted; for each non-repeated entity description in the entity description set, the confidence of the entity description is calculated according to the similarity between each non-repeated entity description and the entity description, and the frequency of each non-repeated entity description in the entity description set; and an alternative entity description of the given entity is selected according to the calculation result of the confidence. It can be seen that the solution of this embodiment does not need to rely on a large amount of labeled data and can be widely applied to the extraction of large-scale open-domain unsupervised entity descriptions; and, compared with the method of only measuring the quality of entity descriptions based on frequency, the content information of the entity descriptions themselves is further considered, making the calculation result of the confidence more reliable. Correspondingly, higher-quality alternative entity descriptions can be selected for display to users.
[0029] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0031] Figure 1 A schematic diagram showing the right-side recommendation when there is no recommendation reason for searching for Pomeranians is shown;
[0032] Figure 2 A flowchart showing a method for extracting entity descriptions based on a knowledge graph according to an embodiment of the present invention is shown;
[0033] Figure 3 A flowchart showing a method for extracting entity descriptions based on a knowledge graph according to another embodiment of the present invention is shown;
[0034] Figure 4 A flowchart showing a method for pushing search engine recommendation reasons according to an embodiment of the present invention is shown;
[0035] Figure 5Shows a schematic diagram of the right-side recommendations when searching for Pomeranians with recommended reasons;
[0036] Figure 6 Shows a functional block diagram of an entity description extraction device based on a knowledge graph according to an embodiment of the present invention;
[0037] Figure 7 Shows a functional block diagram of a push device for search engine recommended reasons according to another embodiment of the present invention;
[0038] Figure 8 Shows a schematic structural diagram of a computing device according to an embodiment of the present invention;
[0039] Figure 9 Shows a schematic structural diagram of a computing device according to an embodiment of the present invention. Detailed implementation manners
[0040] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0041] Figure 2 Shows a flowchart of a method for extracting an entity description based on a knowledge graph according to an embodiment of the present invention. As Figure 2 shown, the method includes:
[0042] Step S201, for a given entity, extract a set of entity descriptions of the entity from a knowledge graph database.
[0043] The solution of the present invention is used to extract high-quality entity descriptions introducing the given entity, that is, alternative entity descriptions.
[0044] Among them, the given entity refers to any entity that can be displayed as a recommendation on a recommendation page. Taking the example of searching for Pomeranians in a search engine (see Figure 1 ), correspondingly, the entities for which recommendations are displayed include related organisms such as Bichon Frises, Teacup Dogs, and Silver Fox Dogs, and these related organisms can be used as the given entities of the present invention respectively.
[0045] Specifically, an entity description set is extracted from the knowledge graph database, and the entity description set includes multiple entity descriptions. In the present invention, the source of the knowledge graph data is not limited. Optionally, the underlying data of the knowledge graph of a specific search engine can be directly used as the knowledge graph data in the present invention, or multiple knowledge graph data can be integrated to obtain the knowledge graph data in the present invention. Moreover, in the present invention, the specific method for extracting the entity description set is not limited, and any method that can achieve knowledge extraction is included in the scope of the present invention.
[0046] Step S202: For each non-repeated entity description in the entity description set, calculate the confidence of the entity description according to the similarity between each non-repeated entity description and the entity description, and the frequency of each non-repeated entity description appearing in the entity description set.
[0047] In the present invention, each constituent element in the entity description set is denoted as an entity description. Correspondingly, the entity description set is composed of multiple elements, so the entity description set contains multiple entity descriptions. There are duplicates among the multiple entity descriptions, and the non-repeated entity description refers to the remaining entity descriptions after removing the repeatedly appearing entity descriptions. For example, the entity descriptions remaining after removing the same entity descriptions that appear for the second time and later in the entity description set.
[0048] Among them, the frequency of an entity description appearing in the entity description set refers to the number of times it appears in the entity description set.
[0049] In this step, the confidence of each non-repeated entity description is calculated to quantify the quality of the entity description. Specifically, through big data analysis, it is found that the quality of an entity description is not only related to the frequency of the entity description appearing in the entity description set, but also related to the content information of the entity description itself. In the present invention, the similarity between each non-repeated entity description and the entity description is used to represent the content information of the entity description itself, and based on the similarity data representing the content information of the entity description itself and the frequency data representing the external usage times information of the entity description, the confidence of the entity description is comprehensively calculated. Among them, according to the results of big data analysis, it can be determined that the higher the overall similarity between the current entity description and the remaining entity descriptions, the higher the quality of the current entity description, and the higher the frequency of the current entity description appearing, the higher the quality of the current entity description. The present invention calculates the confidence based on this rule, but does not limit the specific calculation method.
[0050] Step S203: According to the confidence of each entity description, at least one entity description is selected from the entity description set as the alternative entity description of the entity.
[0051] Specifically, the confidence of an entity description is a quantification of the quality of the entity description. Based on the confidence level, alternative entity descriptions of the entity are screened from the set of entity descriptions to ensure the quality of the recommended reasons that can be selected and presented on the recommendation page. In the present invention, the specific screening method is not limited. Optionally, a preset number of entity descriptions can be screened in descending order of confidence, or entity descriptions with a confidence higher than a preset confidence threshold can be screened out.
[0052] According to the entity description extraction method based on a knowledge graph provided in this embodiment, for a given entity, a set of entity descriptions is extracted; for each non-duplicate entity description in the set of entity descriptions, the confidence of the entity description is calculated according to the similarity between each non-duplicate entity description and the entity description, and the frequency of each non-duplicate entity description appearing in the set of entity descriptions; and alternative entity descriptions of the given entity are selected according to the calculation result of the confidence. It can be seen that the solution of this embodiment does not rely on a large amount of labeled data and can be widely applied to the extraction of large-scale open-domain unsupervised entity descriptions; moreover, compared with the method of only measuring the quality of entity descriptions based on frequency, the content information of the entity descriptions themselves is further considered, making the calculation result of the confidence more reliable. Correspondingly, higher-quality alternative entity descriptions can be screened out for display to users.
[0053] Figure 3 FIG. shows a flowchart of an entity description extraction method based on a knowledge graph according to another embodiment of the present invention. As Figure 3 shown, the method includes:
[0054] Step S301, for a given entity, extract a set of entity descriptions of the entity from the knowledge graph database.
[0055] Specifically, the knowledge graph data is used as the original data for extracting entity descriptions. Among them, the knowledge graph data includes, but is not limited to, entity profiles, text descriptions, semantic tags, and / or existing descriptions. The existing descriptions usually come from the entry meanings of various encyclopedias and are entity descriptions of the given entity recognized by the industry. According to the general linguistic features of entity descriptions, one or more extraction models are constructed, and one or more extraction models are used to extract a set of entity descriptions of the entity from the knowledge graph database. Among them, each extraction model is used to extract entity descriptions that meet the model features of the extraction model from the entity profile, text description, and / or semantic tags respectively.
[0056] For example, words such as "is a kind of", "is known as", "as" are all general linguistic features of entity descriptions. Regular expressions matching these general linguistic features are written as extraction templates, and then a set of entity descriptions can be extracted.
[0057] Further, in the process of extracting the entity description set using the extraction model, the same or different entity descriptions can be extracted from the knowledge graph data of the entity using different templates. Among them, for the case of extracting the same entity description, for example, the same entity description can be extracted from the entity introduction and the text description using different extraction templates, which makes the entity descriptions included in the entity description set duplicate.
[0058] Step S302, filter out the entity descriptions with extraction errors in the entity description set.
[0059] After the entity description set is extracted, filter out the entity descriptions in the entity description set that obviously do not conform to the characteristics of the entity description, so as to avoid introducing too much noise in the subsequent confidence calculation process. In practice, due to the diversity of natural language, the extraction results in large-scale scenarios are prone to extraction errors, so the extraction results need to be preliminarily filtered.
[0060] Specifically, the entity descriptions with extraction errors in the entity description set can be filtered according to the compositional characteristics of the entity description and / or the linguistic characteristics of non-entity descriptions, where the compositional characteristics include length characteristics and / or punctuation characteristics.
[0061] In some alternative embodiments, for each non-duplicate entity description in the entity description set, determine whether the length of the entity description is within a preset length interval. If not, filter out the entity description from the entity description set. Among them, according to the conventional length of the sentence, a specific range between the conventional length intervals of the sentence is set as the preset length interval. Assuming that the conventional length of the sentence is between 2 characters and 20 characters, the preset length interval is set to 1-19 characters (in this case, the length of the entity is default to be at least 1 character). In this way, the entity descriptions that do not conform to the length characteristics can be filtered out.
[0062] And / or, in some other alternative embodiments, for each non-duplicate entity description in the entity description set, determine whether the entity description contains a preset symbol. If so, filter out the entity description from the entity description set. Among them, the preset symbol can refer to any punctuation mark, including semicolons, commas, periods, question marks, etc. In practice, the entity description is a phrase or a short sentence and usually does not contain punctuation marks. In this way, the entity descriptions that do not match the punctuation characteristics can be filtered out.
[0063] And / or, in some other alternative embodiments, a conflict model conflicting with the extraction model can be constructed according to the linguistic features described by non-entities, and conflict entity descriptions satisfying the model features of the conflict model are extracted from the knowledge graph database by using the conflict model corresponding to one or more extraction models. For example, statements in the sentence patterns of "that is..." and "because..." are not suitable as entity descriptions, so statements that are not suitable as entity descriptions, that is, conflict entity descriptions, can be extracted by using the conflict model. For each non-repeated entity description in the entity description set, it is determined whether the entity description is the same as the conflict entity description. If so, the entity description is filtered out from the entity description set. In this way, statements that are not suitable as entity descriptions can be filtered out.
[0064] It should be noted that the present invention is not limited to the three filtering methods listed above. In specific implementation, those skilled in the art can also use other methods for filtering. Optionally, filtering is performed according to the character content of the entity description. If the entity description further includes a given entity, filtering is performed.
[0065] Furthermore, for each non-repeated entity description in the entity description set, filtering out the entity description from the entity description set means filtering out all entity descriptions in the entity description set that have the same text as the entity description. For example, if the entity description a appears 3 times in the entity description set, then the entity description a at the corresponding 3 element positions is filtered out.
[0066] Step S303, for each non-repeated entity description in the entity description set, calculate the confidence of the entity description according to the similarity between each non-repeated entity description and the entity description, and the frequency of each non-repeated entity description appearing in the entity description set.
[0067] In this embodiment, the process of calculating the confidence of the entity description includes the following steps:
[0068] Step one, construct a probability transition matrix.
[0069] According to the frequency of each non-repeated entity description appearing in the entity description set and the similarity between each non-repeated entity description, construct an N*N probability transition matrix; where N is the number of non-repeated entity descriptions in the entity description set.
[0070] Specifically, for each non-repeated entity description, the frequency of the entity description appearing in the entity description set is statistically obtained; and the similarity between each non-repeated entity description can be obtained by using an unsupervised sentence similarity calculation method, such as jaccard similarity, edit distance, vector space model, average or weighted summation of word vectors / character vectors, and similarity calculation combining an external language knowledge base (such as HowNet).
[0071] For example, the set of entity descriptions is {a, b, c, b, c, c, d, a, b, c}. Among them, the non-repeated entity descriptions include a, b, c, d. The frequencies of these 4 entity descriptions appearing in the set of entity descriptions are counted as 2, 3, 4, 1 respectively, and the following 10 similarity values S are calculated aa ,S ab ,S ac ,S ad ,S bb ,S bc ,S bd ,S cc ,S cd ,S dd 。Correspondingly, a 4×4 probability transition matrix can be constructed.
[0072] Furthermore, in the process of constructing the N×N probability transition matrix, the number of rows corresponds to the non-repeated entity descriptions. For example, the first row corresponds to the non-repeated entity description a, the second row corresponds to the non-repeated entity description b... Among them, the probability values of the N elements in each row are related to the following three values: the similarity between each non-repeated entity description and the non-repeated entity description corresponding to this row, the frequency of the entity description corresponding to this row, and the maximum frequency of any entity description appearing in the set of entity descriptions. For example, the probability value of the element in the first row and the second column is related to S ab , the frequency 2 of the entity description a appearing, and the maximum frequency 4 of the entity description c appearing. The probability transition matrix constructed in this way makes the N probability values corresponding to each non-repeated entity description related not only to the frequency but also to the similarity between each non-repeated entity description and this entity description, and thus in subsequent calculations, for each non-repeated entity description, two pieces of information, namely frequency and content, can be comprehensively referred to.
[0073] In a specific embodiment of the present invention, the probability transition matrix is set as M, where M [i][j] is obtained through the following steps: calculate the sum of the similarity between the i-th non-repeated entity description and the j-th non-repeated entity description and the frequency of the i-th non-repeated entity description appearing in the set of entity descriptions to obtain the summation result; calculate the ratio of the summation result to the highest frequency of any non-repeated entity description appearing in the set of entity descriptions to obtain the value of M [i][j] . Still taking the set of entity descriptions {a, b, c, b, c, c, d, a, b, c} as an example for illustration, the order of the non-repeated entity descriptions is a, b, c, d in sequence, then a 4×4 probability transition matrix can be constructed, and M [2][4] is: the sum of the similarity S bd between the entity description b and the entity description d and the frequency 3 of the entity description b appearing in the set of entity descriptions (Sbd +3), the ratio to the highest frequency 4 corresponding to the entity description c, i.e., (S bd +3) / 4.
[0074] In addition, in some alternative embodiments, if the knowledge graph database includes existing descriptions of a given entity, the existing descriptions can be added to the entity description set so that the entity description set contains the existing descriptions of the entity. In practice, the existing descriptions are usually descriptions that have been generally accepted or recognized. Adding the existing descriptions to the entity description combination can provide a relatively adaptive comparison benchmark for the reliability of the extraction results. For example, if the confidence level is higher than that of the existing description, it is determined as an alternative entity description. Specifically, check whether there is a specific field in the data corresponding to the given entity. For example, the specific field is the entry sense field. If the specific field is found, extract the field content of the specific field as the existing description of the given entity. For this existing description, when calculating the statistical frequency, calculate the average frequency value of the frequencies of multiple non-repeating entity descriptions appearing in the entity description set, and determine the frequency of the existing description according to the average frequency value. If the frequency of the existing description appearing in the entity description set is low, it is obviously unreasonable to use this low frequency as the frequency of the existing description. In these alternative embodiments, the frequency of the existing description is determined by calculating the average frequency value of the frequencies of multiple non-repeating entity descriptions other than the existing description in the entity description set. For example, use the average frequency value as the frequency of the existing description. In this way, the frequency of the determined existing description can be made more reasonable.
[0075] Step two, determine the initial state vector S0.
[0076] In this embodiment, to calculate the confidence levels of each non-repeating entity description by recursion, it is necessary to determine the initial state vector S0 of the recursion.
[0077] Specifically, the initial parameter is specifically the initial state vector S0. Set the state vector S0 as the column vector [1 / N, 1 / N,..., 1 / N], where the number of rows is N. That is, initially, the same initial confidence value 1 / N is assigned to each non-repeating entity description.
[0078] After determining the initial state vector S0, before performing the recursion in the following step three, normalize the probability transition matrix M row by row.
[0079] Step three, according to the transpose matrix of the probability transition matrix and the state vector S t-1 Recursively obtain the state vector S t ; calculate the differences between the elements in the state vector S t and the corresponding elements in the state vector S t-1 to obtain multiple difference results.
[0080] Among them, at the first recurrence, the initial state vector S0 is used as the state vector S t-1 , and in subsequent recurrences, the state vector S t obtained by recurrence is used as the state vector S t-1 .
[0081] Specifically, the state vector S t is obtained according to the state vector S t-1 and the transposed matrix M T of the probability transition matrix M. It can be obtained by setting a recurrence expression and substituting the state vector S t-1 and the transposed matrix M T into the expression, and recursively obtaining the state vector S t . Among them, the recurrence expression can be set as follows:
[0082] S t = β·M T S t -1+(1-β)·S t-1 ; where β is a constant set fixedly.
[0083] Furthermore, the state vector S t obtained by recursively calculating through the above expression is a column vector with N rows. Among them, each element corresponds to the confidence of an entity description. After obtaining the state vector S t , subtract the elements in the same row of the state vector S t and the state vector S t-1 to obtain the difference of the corresponding elements, so as to judge whether the recurrence end condition is satisfied.
[0084] Step 4, judge whether the sum of the absolute values of multiple difference results is less than a preset target value. If so, determine each element in the state vector S t as the confidence of each non-repeated entity description; if not, assign the state vector S t to the state vector S t-1 , and repeat Step 3 and Step 4.
[0085] Specifically, sum the absolute values of the calculated multiple difference results to obtain the sum of the absolute values, and judge whether the sum of the absolute values is less than the preset target value. If it is less than the preset target value, the recurrence end condition is satisfied, and the confidence of each non-repeated entity description can be determined. Among them, according to the order of the non-repeated entity descriptions corresponding to each row when constructing the probability transition matrix, determine each element in the state vector St as the confidence of each non-repeated entity description.
[0086] Still taking the example where the non-repeated entity descriptions above are in the order of a, b, c, d in sequence, correspondingly, the first row, the second row, the third row, and the fourth row in the probability transition matrix respectively represent the probabilities corresponding to a, b, c, d, then the 4 elements in the state vector S t are the confidence levels of a, b, c, d in sequence.
[0087] In addition, if the sum of absolute values is greater than or equal to the preset target value, then assign the state vector S t to the state vector S t-1 and repeat steps three and four until the sum of absolute values is less than the preset target value, then the confidence levels of each non-repeated entity description are obtained.
[0088] It should be noted here that in steps one to four above, the method of calculating the confidence level of entity descriptions based on frequency and similarity is only a feasible implementation method, but the present invention is not limited thereto. In some other alternative embodiments of the present invention, the probability-biased random walk algorithm integrating external features can be directly used to implement the processes of steps one to four above, that is, on the basis of the existing random walk algorithm, the probability transition matrix is constructed by combining the frequencies of each non-repeated entity description. Among them, the input of the algorithm is N non-repeated entity descriptions and their corresponding frequencies, and the output of the algorithm is N entity descriptions and their corresponding confidence levels. Since the probability transition matrix M satisfies the following three conditions of the Markov process, M is a stochastic matrix (all elements of M are greater than or equal to 0, and the sum of elements in each column is 1), M is irreducible, and M is aperiodic, the convergence of the algorithm modeled into the Markov process is also theoretically guaranteed. By calculating the confidence level in this way, reliability modeling can be performed on entity description extraction, including combining the similarity between entity descriptions, statistical frequencies, existing sense descriptions as a benchmark, and combining the external information of the extraction results to adaptively measure the reliability of the extraction results. Or, in some other embodiments of the present invention, the confidence level value can also be calculated by setting weights for frequency and similarity and performing weighted summation.
[0089] Step S304, according to the confidence levels of each entity description, screen out at least one entity description from the entity description set as the alternative entity description of the entity.
[0090] Among them, the confidence level of an entity description is a quantification of the quality of the entity description.
[0091] Specifically, according to the high and low confidence level values, screen out the alternative entity description from the entity description combination to obtain the alternative recommendation reasons for the given entity, so as to give high-quality recommendation reasons on the recommendation page to guide the user's understanding of the given entity.
[0092] In some alternative embodiments of the present invention, spare entity descriptions are screened based on the average confidence value. Calculate the average confidence value of the confidence levels of multiple non-repeating entity descriptions in the entity description set, and screen out at least one entity description with a confidence level higher than the average confidence value from the entity description set as the spare entity description of the entity. For example, the confidence levels of entity descriptions a, b, c, and d are 0.2, 0.1, 0.25, and 0.12 respectively. The calculated average confidence value is (0.2 + 0.1 + 0.25 + 0.12) / 4 = 0.1675. Then, a and c are screened out as spare entity descriptions.
[0093] In some other alternative embodiments of the present invention, if the entity description set contains an existing description of the entity, screen out at least one entity description with a confidence level higher than the confidence level of the existing description from the entity description set as the spare entity description of the entity. Or, on this basis, the existing description can be added as a spare entity description. In the previous example, if a is the existing description, then c is screened out as the spare entity description, or a and c can also be used as spare entity descriptions.
[0094] According to the entity description extraction method based on the knowledge graph provided in this embodiment, filter out the entity descriptions with extraction errors in the extracted entity description set to avoid introducing excessive noise into the confidence level calculation process due to extraction errors; and, construct a probability transition matrix based on the frequency of each non-repeating entity description in the entity description set and the similarity between each non-repeating entity description, and calculate the confidence level of each non-repeating entity description according to the probability transition matrix, so that the calculation of the confidence level no longer depends only on the frequency parameter, but also on the similarity between entity descriptions, making the calculation result of the confidence level more reliable; and, add the existing description of the entity to the entity description set for confidence level calculation, which can make the reliability of the extraction result have a relatively adaptive benchmark. It can be seen that the solution of this embodiment does not depend on a large amount of labeled data and is applicable to large-scale open-domain unsupervised entity description extraction; and, compared with the method of only measuring the quality of entity descriptions based on frequency, it further considers the content information of the entity descriptions themselves, making the calculation result of the confidence level more reliable. Correspondingly, higher-quality spare entity descriptions can be screened out for display to users.
[0095] Figure 4 The flowchart of the method for pushing search engine recommendation reasons according to an embodiment of the present invention is shown. As Figure 4 shown, the method includes:
[0096] Step S401, obtain the recommendation result of the search engine.
[0097] Among them, the recommendation result refers to the recommended content matched by the search engine according to the search conditions. Refer to theFigure 1 When a user searches for a Pomeranian, the Bichon Frise, Teacup Dog, etc. recommended in the right-side recommendation column are all recommendation results.
[0098] Step S402: Use the above-mentioned entity description extraction method based on the knowledge graph to obtain the alternative entity description corresponding to the recommendation result by taking the recommendation result as the given entity.
[0099] Taking the recommendation result as the given entity, that is, as the object for which the recommendation reason is required, the alternative entity description corresponding to the given entity is extracted through the entity description extraction method based on the knowledge graph in the above method embodiment.
[0100] Step S403: Select an entity description from the alternative entity descriptions as the recommendation reason and present it on the search result display page.
[0101] Specifically, if there are multiple alternative entity descriptions, a preset number of entity descriptions are selected from them as the recommendation reasons for display. Optionally, one entity description is selected as the recommendation reason. For example, the entity description with the highest confidence is selected as the recommendation reason, or when the same entity is recommended to the same user multiple times, different entity descriptions are sequentially selected from the alternative entity descriptions as the recommendation reasons to avoid repeated recommendation reasons for consecutive times.
[0102] Figure 5 Shows a schematic diagram of the right-side recommendation when there is a recommendation reason for searching for a Pomeranian. As Figure 5 shown, when making a right-side recommendation, the recommendation reason is displayed under each recommendation result. For example, the recommendation reason for the Bichon Frise is "adorable, naughty and cute".
[0103] According to the search engine recommendation reason pushing method provided in this embodiment, taking the search result of the search engine as the given entity, extracting the alternative entity description according to the entity description extraction method based on the knowledge graph in the above embodiment, and selecting the recommendation reason from the alternative entity descriptions; presenting the recommendation reason and the recommendation result to the user together to guide the user's understanding of the recommendation result. It can be seen that the solution of the present invention can display high-quality recommendation reasons on the search result display page, thereby improving the user experience.
[0104] Figure 6 Shows a functional block diagram of an entity description extraction device based on a knowledge graph according to an embodiment of the present invention. As Figure 6 shown, the device includes:
[0105] An extraction module 601, adapted to extract the entity description set of a given entity from the knowledge graph database;
[0106] A confidence calculation module 602, which is adapted to calculate the confidence of each non-repeated entity description in the entity description set according to the similarity between each non-repeated entity description and this entity description, and the frequency of each non-repeated entity description appearing in the entity description set;
[0107] A screening module 603, which is adapted to screen out at least one entity description from the entity description set as the alternative entity description of the entity according to the confidence of each entity description.
[0108] In an alternative embodiment, the extraction module is further adapted to: extract the entity description set of the entity from the knowledge graph database by using one or more extraction models.
[0109] In an alternative embodiment, the device further includes: a filtering module, which is adapted to, for each non-repeated entity description in the entity description set, determine whether the length of the entity description is within a preset length range, and if not, filter out the entity description from the entity description set; and / or,
[0110] Determine whether the entity description contains a preset symbol, and if so, filter out the entity description from the entity description set.
[0111] In an alternative embodiment, the device further includes: a filtering module, which is adapted to extract, from the knowledge graph database, conflict entity descriptions that meet the model features of the conflict model by using the conflict model corresponding to the one or more extraction models;
[0112] For each non-repeated entity description in the entity description set, determine whether the entity description is the same as the conflict entity description, and if so, filter out the entity description from the entity description set.
[0113] In an alternative embodiment, the confidence calculation module is further adapted to:
[0114] Construct an N*N probability transition matrix according to the frequency of each non-repeated entity description appearing in the entity description set and the similarity between each non-repeated entity description; where N is the number of non-repeated entity descriptions in the entity description set;
[0115] The probability transition matrix is set as M, where M [i][j] is obtained through the following steps: calculate the sum of the similarity between the i-th non-repeated entity description and the j-th non-repeated entity description and the frequency of the i-th non-repeated entity description appearing in the entity description set to obtain a summation result; calculate the ratio of the summation result to the highest frequency of any non-repeated entity description appearing in the entity description set to obtain the value of M [i][j] ;
[0116] Determine the initial state vector S0;
[0117] According to the transpose matrix of the probability transition matrix and the state vector S t-1 Recursively obtain the state vector S t ; Calculate the state vector S t The difference between each element in and the corresponding element in the state vector S t-1 is obtained to get multiple difference results;
[0118] Determine whether the sum of the absolute values of the multiple difference results is less than a preset target value. If so, determine each element in the state vector S t as the confidence level of each non-repetitive entity description; if not, assign the state vector S t to the state vector S t-1 and repeat steps S33 and S34.
[0119] In an alternative embodiment, the screening module is further adapted to:
[0120] Calculate the average confidence level value of the confidence levels of multiple non-repetitive entity descriptions in the entity description set, and screen out at least one entity description with a confidence level higher than the average confidence level value from the entity description set as the alternative entity description of the entity.
[0121] In an alternative embodiment, the entity description set includes the existing descriptions of the entity;
[0122] The screening module is further adapted to: screen out at least one entity description with a confidence level higher than the confidence level of the existing description from the entity description set as the alternative entity description of the entity.
[0123] In an alternative embodiment, the entity description set includes the existing descriptions of the entity;
[0124] The device further includes: a frequency calculation module, adapted to calculate the average frequency value of the frequencies at which multiple non-repetitive entity descriptions appear in the entity description set, and determine the frequency of the existing description according to the average frequency value.
[0125] Figure 7 Shows a functional block diagram of a push device for search engine recommendation reasons according to another embodiment of the present invention. As Figure 7 shown, the device includes:
[0126] An acquisition module 701, adapted to acquire the recommendation results of a search engine;
[0127] The entity description extraction module 702 is adapted to use the recommendation result as a given entity and utilize the knowledge graph-based entity description extraction device in the above device embodiments to obtain an alternative entity description corresponding to the recommendation result;
[0128] The selection module 703 is adapted to select an entity description from the alternative entity descriptions as a recommendation reason for presentation on the search result display page.
[0129] An embodiment of the present application provides a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction, and the computer executable instruction can execute the knowledge graph-based entity description extraction method in any of the above method embodiments.
[0130] An embodiment of the present application provides a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction, and the computer executable instruction can execute the push method of the search engine recommendation reason in any of the above method embodiments.
[0131] Figure 8 A schematic structural diagram of a computing device according to an embodiment of the present invention is shown, and the specific implementation of the computing device is not limited in the specific embodiments of the present invention.
[0132] As Figure 8 shown, the computing device may include: a processor 802, a communication interface 804, a memory 806, and a communication bus 808.
[0133] Wherein:
[0134] The processor 802, the communication interface 804, and the memory 806 communicate with each other through the communication bus 808.
[0135] The communication interface 804 is used for communicating with network elements of other devices such as clients or other servers.
[0136] The processor 802 is used to execute the program 810, and specifically can execute the relevant steps in the above method embodiments of the knowledge graph-based entity description extraction method.
[0137] Specifically, the program 810 may include program code, and the program code includes computer operation instructions.
[0138] The processor 802 may be a central processing unit (CPU), or a specific application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0139] A memory 806 for storing a program 810. The memory 806 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0140] The program 810 may specifically be used to cause the processor 802 to perform the following operations:
[0141] Step S1: For a given entity, extract a set of entity descriptions of the entity from the knowledge graph database;
[0142] Step S2: For each non-repeated entity description in the set of entity descriptions, calculate the confidence of the entity description according to the similarity between each non-repeated entity description and the entity description, and the frequency of each non-repeated entity description appearing in the set of entity descriptions;
[0143] Step S3: According to the confidence of each entity description, screen out at least one entity description from the set of entity descriptions as the alternative entity description of the entity.
[0144] In an alternative embodiment, the program 810 may specifically be further used to cause the processor 802 to perform the following operations:
[0145] Use one or more extraction models to extract a set of entity descriptions of the entity from the knowledge graph database.
[0146] In an alternative embodiment, the program 810 may specifically be further used to cause the processor 802 to perform the following operations:
[0147] For each non-repeated entity description in the set of entity descriptions, determine whether the length of the entity description is within a preset length range. If not, filter out the entity description from the set of entity descriptions; and / or,
[0148] Determine whether the entity description contains a preset symbol. If so, filter out the entity description from the set of entity descriptions.
[0149] In an alternative embodiment, the program 810 may specifically be further configured to cause the processor 802 to perform the following operations:
[0150] Using a conflict model corresponding to the one or more extraction models, extract conflict entity descriptions that satisfy the model features of the conflict model from the knowledge graph database;
[0151] For each non-repeated entity description in the entity description set, determine whether the entity description is the same as the conflict entity description. If so, filter out the entity description from the entity description set.
[0152] In an alternative embodiment, the program 810 may specifically be further configured to cause the processor 802 to perform the following operations:
[0153] Step S31, construct an N*N probability transition matrix according to the frequency of occurrence of each non-repeated entity description in the entity description set and the similarity between each non-repeated entity description; where N is the number of non-repeated entity descriptions in the entity description set;
[0154] The probability transition matrix is denoted as M, where M [i][j] is obtained through the following steps: calculate the sum of the similarity between the i-th non-repeated entity description and the j-th non-repeated entity description and the frequency of occurrence of the i-th non-repeated entity description in the entity description set to obtain a summation result; calculate the ratio of the summation result to the highest frequency of occurrence of any non-repeated entity description in the entity description set to obtain the value of M [i][j] ;
[0155] Step S32, determine the initial state vector S0;
[0156] Step S33, recursively obtain the state vector S according to the transpose matrix of the probability transition matrix and the state vector S t-1 ; calculate the difference between each element in the state vector S t and the corresponding element in the state vector S t to obtain a plurality of difference results; t-1
[0157] Step S34, determine whether the sum of the absolute values of the plurality of difference results is less than a preset target value. If so, determine each element in the state vector S t as the confidence of each non-repeated entity description; if not, assign the state vector S t to the state vector S t-1 , and repeat steps S33 and S34.
[0158] In an alternative embodiment, the program 810 may specifically be further configured to cause the processor 802 to perform the following operations:
[0159] Calculate the average confidence value of the confidence levels of multiple non-duplicate entity descriptions in the entity description set, and filter out at least one entity description with a confidence level higher than the average confidence value from the entity description set as the alternative entity description of the entity.
[0160] In an alternative embodiment, the entity description set includes the existing description of the entity;
[0161] The program 810 can specifically be further configured to cause the processor 802 to perform the following operations:
[0162] Filter out at least one entity description with a confidence level higher than the confidence level of the existing description from the entity description set as the alternative entity description of the entity.
[0163] In an alternative embodiment, the entity description set includes the existing description of the entity;
[0164] The program 810 can specifically be further configured to cause the processor 802 to perform the following operations:
[0165] Calculate the average frequency value of the frequencies at which multiple non-duplicate entity descriptions appear in the entity description set, and determine the frequency of the existing description according to the average frequency value.
[0166] Figure 9 FIG. shows a schematic structural diagram of a computing device according to an embodiment of the present invention. The specific implementation of the computing device is not limited in the specific embodiments of the present invention.
[0167] As Figure 9 shown, the computing device may include: a processor 902, a communications interface 904, a memory 906, and a communications bus 908.
[0168] Wherein:
[0169] The processor 902, the communications interface 904, and the memory 906 communicate with each other through the communications bus 908.
[0170] The communications interface 904 is used to communicate with network elements of other devices such as clients or other servers.
[0171] The processor 902 is used to execute the program 910, and can specifically execute the relevant steps in the embodiment of the method for pushing search engine recommendation reasons described above.
[0172] Specifically, the program 910 may include program code, and the program code includes computer operation instructions.
[0173] The processor 902 may be a central processing unit (CPU), or a specific application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the computing device may be of the same type, such as one or more CPUs; or may be of different types, such as one or more CPUs and one or more ASICs.
[0174] A memory 906 for storing a program 910. The memory 906 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0175] The program 910 may specifically be configured to cause the processor 902 to perform the following operations:
[0176] Obtain the recommended results of a search engine;
[0177] Use the recommended results as given entities, and utilize the entity description extraction method based on the knowledge graph described in the above method embodiments to obtain the alternative entity descriptions corresponding to the recommended results;
[0178] Select entity descriptions from the alternative entity descriptions as recommended reasons and present them on the search result display page.
[0179] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used in conjunction with the teachings provided herein. The structure required to construct such systems will be apparent from the above description. Additionally, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.
[0180] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0181] Similarly, it should be understood that, for the purpose of streamlining the present disclosure and facilitating the understanding of one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present invention.
[0182] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from those of the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Any combination can be adopted for all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed, except that at least some of such features and / or processes or units are mutually exclusive. Each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose, unless otherwise expressly stated.
[0183] In addition, those skilled in the art will be able to understand that, although some of the embodiments described herein include certain features included in other embodiments but not others, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0184] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the entity description extraction device based on the knowledge graph and the push device for search engine recommendation reasons according to the embodiments of the present invention. The present invention can also be implemented as a device or device program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0185] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0186] The present invention discloses: A1. A method for extracting entity descriptions based on a knowledge graph, comprising:
[0187] Step S1, for a given entity, extracting a set of entity descriptions of the entity from a knowledge graph database;
[0188] Step S2, for each non-repeated entity description in the set of entity descriptions, calculating the confidence of the entity description according to the similarity between each non-repeated entity description and the entity description, and the frequency of each non-repeated entity description appearing in the set of entity descriptions;
[0189] Step S3, screening out at least one entity description from the set of entity descriptions as the alternative entity description of the entity according to the confidence of each entity description.
[0190] A2. The method according to A1, wherein the extracting the set of entity descriptions of the entity from a knowledge graph database further includes:
[0191] Extract the entity description set of the entity from the knowledge graph database using one or more extraction models.
[0192] A3. The method according to A1 or A2, wherein, after extracting the entity description set of the entity from the knowledge graph database, the method further includes:
[0193] For each non-duplicate entity description in the entity description set, determine whether the length of the entity description is within a preset length interval, and if not, filter out the entity description from the entity description set; and / or,
[0194] Determine whether the entity description contains a preset symbol, and if so, filter out the entity description from the entity description set.
[0195] A4. The method according to A2, wherein, after extracting the entity description set of the entity from the knowledge graph database, the method further includes:
[0196] Extract the conflicting entity descriptions that meet the model features of the conflict model from the knowledge graph database using the conflict model corresponding to the one or more extraction models;
[0197] For each non-duplicate entity description in the entity description set, determine whether the entity description is the same as the conflicting entity description, and if so, filter out the entity description from the entity description set.
[0198] A5. The method according to any one of A1-A4, wherein the step S2 further includes:
[0199] Step S31, construct an N*N probability transition matrix according to the frequency of each non-duplicate entity description appearing in the entity description set and the similarity between each non-duplicate entity description; where N is the number of non-duplicate entity descriptions in the entity description set;
[0200] The probability transition matrix is set as M, where M [i][j] is obtained through the following steps: calculate the sum of the similarity between the i-th non-duplicate entity description and the j-th non-duplicate entity description and the frequency of the i-th non-duplicate entity description appearing in the entity description set to obtain a summation result; calculate the ratio of the summation result to the highest frequency of any non-duplicate entity description appearing in the entity description set to obtain the value of M [i][j] value;
[0201] Step S32, determine the initial state vector S0;
[0202] Step S33, according to the transpose matrix of the probability transition matrix and the state vector S t-1Recursively obtain the state vector S t ; Calculate the state vector S t for each element in t-1 the state vector S
[0203] Step S34, determine whether the sum of the absolute values of multiple difference results is less than a preset target value. If so, determine each element in the state vector S t as the confidence of each non-repeated entity description; if not, assign the state vector S t to the state vector S t-1 , and repeat steps S33 and S34.
[0204] A6. According to the method described in A5, wherein, further comprising: screening out at least one entity description from the entity description set as the alternative entity description of the entity according to the confidence of each entity description
[0205] Calculate the average confidence value of the confidences of multiple non-repeated entity descriptions in the entity description set, and screen out at least one entity description with a confidence higher than the average confidence value from the entity description set as the alternative entity description of the entity.
[0206] A7. According to the method described in A5, wherein the entity description set contains the existing description of the entity;
[0207] Further comprising: screening out at least one entity description from the entity description set as the alternative entity description of the entity according to the confidence of each entity description
[0208] Screen out at least one entity description with a confidence higher than the confidence of the existing description from the entity description set as the alternative entity description of the entity.
[0209] A8. According to the method described in any one of A1 - A7, wherein the entity description set contains the existing description of the entity;
[0210] The method further comprises: calculating the average frequency value of the frequencies of multiple non-repeated entity descriptions appearing in the entity description set, and determining the frequency of the existing description according to the average frequency value.
[0211] The present invention also discloses: B9. A method for pushing search engine recommendation reasons, comprising:
[0212] Obtain the recommendation result of the search engine;
[0213] Take the recommendation result as the given entity, and use the method described in any one of A1 - A8 to obtain the alternative entity description corresponding to the recommendation result;
[0214] Select an entity description from the described alternative entity descriptions as a recommended reason and present it on the search result display page.
[0215] The present invention also discloses: C10. An entity description extraction device based on a knowledge graph, including:
[0216] An extraction module, adapted to extract a set of entity descriptions of a given entity from a knowledge graph database;
[0217] A confidence calculation module, adapted to calculate the confidence of each non-repeated entity description in the set of entity descriptions according to the similarity between each non-repeated entity description and this entity description, and the frequency of each non-repeated entity description appearing in the set of entity descriptions;
[0218] A screening module, adapted to screen out at least one entity description from the set of entity descriptions as an alternative entity description of the entity according to the confidence of each entity description.
[0219] C11. The device according to C10, wherein the extraction module is further adapted to: extract a set of entity descriptions of the entity from a knowledge graph database by using one or more extraction models.
[0220] C12. The device according to C10 or C11, wherein the device further includes: a filtering module, adapted to, for each non-repeated entity description in the set of entity descriptions, judge whether the length of this entity description is within a preset length range, and if not, filter out this entity description from the set of entity descriptions; and / or,
[0221] Judge whether the entity description contains a preset symbol, and if so, filter out this entity description from the set of entity descriptions.
[0222] C13. The device according to C11, wherein the device further includes: a filtering module, adapted to extract conflicting entity descriptions that meet the model features of the conflict model from a knowledge graph database by using a conflict model corresponding to the one or more extraction models;
[0223] For each non-repeated entity description in the set of entity descriptions, judge whether this entity description is the same as the conflicting entity description, and if so, filter out this entity description from the set of entity descriptions.
[0224] C14. The device according to any one of C10-C13, wherein the confidence calculation module is further adapted to:
[0225] Construct an N*N probability transition matrix according to the frequencies at which each non-repeating entity description appears in the set of entity descriptions and the similarities between each non-repeating entity description; where N is the number of non-repeating entity descriptions in the set of entity descriptions;
[0226] Let the probability transition matrix be M, where M [i][j] is obtained through the following steps: Calculate the sum of the similarity between the i-th non-repeating entity description and the j-th non-repeating entity description and the frequency at which the i-th non-repeating entity description appears in the set of entity descriptions to obtain a summation result; Calculate the ratio of the summation result to the highest frequency at which any non-repeating entity description appears in the set of entity descriptions to obtain the value of M [i][j] ;
[0227] Determine the initial state vector S0;
[0228] According to the transpose matrix of the probability transition matrix and the state vector S t-1 recursively obtain the state vector S t ; Calculate the differences between the elements in the state vector S t and the corresponding elements in the state vector S t-1 to obtain a plurality of difference results;
[0229] Determine whether the sum of the absolute values of the plurality of difference results is less than a preset target value. If so, determine the elements in the state vector S t as the confidence levels of each non-repeating entity description; if not, assign the state vector S t to the state vector S t-1 , and repeat steps S33 and S34.
[0230] C15. The apparatus according to C14, wherein the screening module is further adapted to:
[0231] Calculate the average confidence level value of the confidence levels of a plurality of non-repeating entity descriptions in the set of entity descriptions, and screen out at least one entity description with a confidence level higher than the average confidence level value from the set of entity descriptions as the alternative entity description of the entity.
[0232] C16. The apparatus according to C14, wherein the set of entity descriptions includes the existing descriptions of the entity;
[0233] The screening module is further adapted to: Screen out at least one entity description with a confidence level higher than the confidence level of the existing description from the set of entity descriptions as the alternative entity description of the entity.
[0234] C17. The apparatus according to any one of C10-C16, wherein the set of entity descriptions includes the existing descriptions of the entity;
[0235] The device further includes: a frequency calculation module, adapted to calculate an average frequency value of the frequencies at which a plurality of non-repeating entity descriptions appear in the entity description set, and determine the frequency of the existing descriptions according to the average frequency value.
[0236] The present invention also discloses: D18. A push device for search engine recommendation reasons, including:
[0237] An acquisition module, adapted to acquire the recommendation results of a search engine;
[0238] An entity description extraction module, adapted to use the recommendation results as given entities, and obtain the corresponding alternative entity descriptions of the recommendation results by using the device according to any one of B10 - B17;
[0239] A selection module, adapted to select entity descriptions from the alternative entity descriptions as recommendation reasons for presentation on the search result display page.
[0240] The present invention also discloses: E19. A computing device, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0241] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the knowledge graph-based entity description extraction method according to any one of A1 - A8.
[0242] The present invention also discloses: F20. A computing device, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0243] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the push method for search engine recommendation reasons according to B9.
[0244] The present invention also discloses: G21. A computer storage medium, wherein at least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to perform the operations corresponding to the knowledge graph-based entity description extraction method according to any one of A1 - A8.
[0245] The present invention also discloses: H22. A computer storage medium, wherein at least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to perform the operations corresponding to the push method for search engine recommendation reasons according to B9.
Claims
1. A method for extracting entity descriptions based on a knowledge graph, comprising: Step S1, for a given entity, extracting a set of entity descriptions of the entity from a knowledge graph database; Step S2, for each non-repeated entity description in the set of entity descriptions, calculating the confidence of the entity description according to the similarity between each non-repeated entity description and the entity description, and the frequency of each non-repeated entity description appearing in the set of entity descriptions, where the entity description is a phrase or a short sentence; Step S3, screening out at least one entity description from the set of entity descriptions as the alternative entity description of the entity according to the confidence of each entity description; wherein, the set of entity descriptions contains the existing descriptions of the entity; The method further includes: calculating the average frequency value of the frequencies of multiple non-repeated entity descriptions appearing in the set of entity descriptions, and determining the frequency of the existing description according to the average frequency value.
2. The method according to claim 1, wherein The further step of extracting the set of entity descriptions of the entity from the knowledge graph database includes: Using one or more extraction models to extract the set of entity descriptions of the entity from the knowledge graph database.
3. The method according to claim 1, wherein, After extracting the set of entity descriptions of the entity from the knowledge graph database, the method further includes: For each non-repeated entity description in the set of entity descriptions, determining whether the length of the entity description is within a preset length interval, and if not, filtering out the entity description from the set of entity descriptions; and / or, Determining whether the entity description contains a preset symbol, and if so, filtering out the entity description from the set of entity descriptions.
4. The method according to claim 2, wherein, After extracting the set of entity descriptions of the entity from the knowledge graph database, the method further includes: Using a conflict model corresponding to the one or more extraction models to extract conflict entity descriptions that meet the model features of the conflict model from the knowledge graph database; For each non-repeated entity description in the set of entity descriptions, determining whether the entity description is the same as the conflict entity description, and if so, filtering out the entity description from the set of entity descriptions.
5. The method according to claim 1, wherein, The step S2 further includes: Step S31, constructing an N*N probability transition matrix according to the frequency of each non-repeated entity description appearing in the set of entity descriptions and the similarity between each non-repeated entity description; where N is the number of non-repeated entity descriptions in the set of entity descriptions; The probability transition matrix is set as M, where M[i][j] is obtained through the following steps: calculating the sum of the similarity between the i-th non-repeated entity description and the j-th non-repeated entity description and the frequency of the i-th non-repeated entity description appearing in the set of entity descriptions to obtain a summation result; calculating the ratio of the summation result to the highest frequency of any non-repeated entity description appearing in the set of entity descriptions to obtain the value of M[i][j]; Step S32, determining the initial state vector S0; Step S33: Recursively obtain the state vector St based on the transpose matrix of the probability transition matrix and the state vector St-1; calculate the differences between the elements in the state vector St and the corresponding elements in the state vector St-1 to obtain multiple difference results. Step S34: Determine whether the sum of the absolute values of the multiple difference results is less than a preset target value. If so, determine the elements in the state vector St as the confidence levels of the respective non-duplicate entity descriptions; if not, assign the state vector St to the state vector St-1, and repeat Step S33 and Step S34.
6. The method according to claim 5, wherein, The further step of screening out at least one entity description from the entity description set as the alternate entity description of the entity according to the confidence levels of the respective entity descriptions includes: Calculating the average confidence level value of the confidence levels of the multiple non-duplicate entity descriptions in the entity description set, and screening out at least one entity description with a confidence level higher than the average confidence level value from the entity description set as the alternate entity description of the entity.
7. The method according to claim 5, wherein, The entity description set contains the existing descriptions of the entity. The further step of screening out at least one entity description from the entity description set as the alternate entity description of the entity according to the confidence levels of the respective entity descriptions includes: Screening out at least one entity description with a confidence level higher than the confidence level of the existing description from the entity description set as the alternate entity description of the entity.
8. A method for pushing search engine recommendation reasons, including: Obtaining the recommendation results of the search engine. Using the recommendation results as the given entity, and obtaining the alternate entity descriptions corresponding to the recommendation results by using the method according to any one of claims 1-7. Selecting an entity description from the alternate entity descriptions as a recommendation reason and presenting it on the search result display page.
9. An apparatus for extracting entity descriptions based on a knowledge graph, including: An extraction module, adapted to extract the entity description set of a given entity from a knowledge graph database. A confidence level calculation module, adapted to calculate the confidence level of each non-duplicate entity description in the entity description set according to the similarity between each non-duplicate entity description and the entity description, and the frequency of each non-duplicate entity description appearing in the entity description set, where the entity description is a phrase or a short sentence. A screening module, adapted to screen out at least one entity description from the entity description set as the alternate entity description of the entity according to the confidence levels of the respective entity descriptions. Wherein, the entity description set contains the existing descriptions of the entity. The apparatus further includes: a frequency calculation module, adapted to calculate the average frequency value of the frequencies of multiple non-duplicate entity descriptions appearing in the entity description set, and determine the frequency of the existing description according to the average frequency value.
10. The apparatus according to claim 9, wherein, The extraction module is further adapted to: extract the entity description set of the entity from the knowledge graph database by using one or more extraction models.
11. The device according to claim 9, wherein, The device further includes: a filtering module, adapted to, for each non-duplicate entity description in the entity description set, determine whether the length of the entity description is within a preset length range, and if not, filter the entity description from the entity description set; and / or, determine whether the entity description contains a preset symbol, and if so, filter the entity description from the entity description set.
12. The apparatus according to claim 10, wherein, The device further includes: a filtering module, adapted to extract, from the knowledge graph database, conflict entity descriptions that meet the model features of the conflict model by using the conflict model corresponding to the one or more extraction models; for each non-duplicate entity description in the entity description set, determine whether the entity description is the same as the conflict entity description, and if so, filter the entity description from the entity description set.
13. The apparatus according to claim 9, wherein, The confidence calculation module is further adapted to: construct an N*N probability transition matrix according to the frequency of occurrence of each non-duplicate entity description in the entity description set and the similarity between each non-duplicate entity description; where N is the number of non-duplicate entity descriptions in the entity description set; The probability transition matrix is set as M, where M[i][j] is obtained through the following steps: calculate the sum of the similarity between the i-th non-duplicate entity description and the j-th non-duplicate entity description and the frequency of occurrence of the i-th non-duplicate entity description in the entity description set to obtain a sum result; calculate the ratio of the sum result to the highest frequency of occurrence of any non-duplicate entity description in the entity description set to obtain the value of M[i][j]; determine the initial state vector S0; recursively obtain the state vector St according to the transposed matrix of the probability transition matrix and the state vector St-1; calculate the difference between each element in the state vector St and the corresponding element in the state vector St-1 to obtain a plurality of difference results; determine whether the sum of the absolute values of the plurality of difference results is less than a preset target value, and if so, determine each element in the state vector St as the confidence of each non-duplicate entity description; if not, assign the state vector St to the state vector St-1, and repeat steps S33 and S34.
14. The apparatus according to claim 13, wherein, The screening module is further adapted to: calculate the average confidence value of the confidences of a plurality of non-duplicate entity descriptions in the entity description set, and screen out at least one entity description with a confidence higher than the average confidence value from the entity description set as the alternative entity description of the entity.
15. The apparatus according to claim 13, wherein, The entity description set contains the existing descriptions of the entity; The screening module is further adapted to: screen out at least one entity description with a confidence higher than the confidence of the existing description from the entity description set as the alternative entity description of the entity.
16. A device for pushing search engine recommendation reasons, including: an acquisition module, adapted to acquire the recommendation results of the search engine; an entity description extraction module, adapted to use the device according to any one of claims 9-15 to obtain the alternative entity description corresponding to the recommendation result by using the recommendation result as a given entity. A selection module, adapted to select an entity description from the backup entity descriptions as a recommended reason for presentation on a search result display page.
17. A computing device, comprising: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method for extracting entity descriptions based on a knowledge graph according to any one of claims 1-7.
18. A computing device, comprising: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method for pushing recommended reasons of a search engine according to claim 8.
19. A computer storage medium, in which at least one executable instruction is stored, and the executable instruction causes a processor to perform operations corresponding to the method for extracting entity descriptions based on a knowledge graph according to any one of claims 1-7.
20. A computer storage medium, in which at least one executable instruction is stored, and the executable instruction causes a processor to perform operations corresponding to the method for pushing recommended reasons of a search engine according to claim 8.
Citation Information
Patent Citations
Search result display method and device based on profound questioning and answering
CN106649761A