Method and apparatus for determining text relevance, storage medium, and electronic device

HK40029143BActive Publication Date: 2026-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Patents
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2020-10-14
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in calculating the relevance of long texts, especially in terms of effectively ranking relevant documents in term search scenarios.

Method used

By introducing a pre-defined knowledge base and a self-attention mechanism, the entities associated with the text and their relevance are determined, and the attention values ​​of words in the text are calculated, thereby improving the accuracy of text relevance calculation.

Benefits of technology

It improves the accuracy of text relevance calculation, especially in long text scenarios, effectively retaining useful information and eliminating useless information, thereby improving the ranking accuracy of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The application discloses a text relevance determination method and device, a storage medium and an electronic device. The method comprises the following steps: determining a first set of entities associated with a first text and a second set of entities associated with a second text based on a knowledge base, wherein the knowledge base comprises knowledge representation constituted by entities, relationships between the entities and entity attributes; determining entity relevance between the first set of entities and the second set of entities according to the knowledge representation; determining attention values of each word in the first text, each word in the second text and each word in the first text and each word in the second text according to the association relationship between the words; and determining text relevance of the first text and the second text according to at least the attention values and the entity relevance. In the scheme, the relationship between the words in the text and between the texts is focused on during the text relevance calculation, and then useful information is focused on and useless information is ignored, so that the accuracy of the text relevance calculation result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, specifically to a method, apparatus, storage medium, and electronic device for determining text relevance. Background Technology

[0002] Text relevance, also known as the degree of text matching, is necessary to determine the relevance between different texts in many scenarios. For example, in term search scenarios, it's typically necessary to determine the relevance of the text in each document to the terms in the search query, and then rank the relevant documents on the search results page based on the degree of relevance. Determining text relevance is based on understanding the text; it depends not only on the semantic similarity between two texts but also on the degree of matching between them. Especially for long texts, the problem of information diffusion can easily lead to lower accuracy in calculating text relevance. Summary of the Invention

[0003] This application provides a method, apparatus, storage medium, and electronic device for determining text relevance, which can improve the accuracy of text relevance calculation results.

[0004] This application provides a method for determining text relevance, including:

[0005] Based on a preset knowledge base, a first group of entities associated with the first text and a second group of entities associated with the second text are determined. The preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes.

[0006] Determine the entity relevance between the first group of entities and the second group of entities based on the knowledge representation;

[0007] Based on the relationships between each word in the first text, the relationships between each word in the second text, and the relationships between words in the first text and words in the second text, an attention value is determined for each word in the first text and the second text with respect to other words, wherein the attention value is used to reflect the degree of attention each word in the first text and the second text pays to other words;

[0008] The text relevance between the first text and the second text is determined based at least on the attention value and the entity relevance.

[0009] Accordingly, embodiments of this application also provide a device for determining text relevance, including:

[0010] An entity determination unit is used to determine, based on a preset knowledge base, a first set of entities associated with a first text and a second set of entities associated with a second text, wherein the preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes.

[0011] The first relevance determination unit is used to determine the entity relevance between the first group of entities and the second group of entities based on the knowledge representation;

[0012] An attention determination unit is used to determine the attention value of each word in the first text and the second text about other words based on the association between each word in the first text, the association between each word in the second text, and the association between words in the first text and words in the second text, wherein the attention value is used to reflect the degree of attention each word in the first text and the second text pays to other words;

[0013] The second relevance determination unit is used to determine the text relevance between the first text and the second text based at least on the attention value and the entity relevance.

[0014] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute the text relevance determination method as described above.

[0015] Accordingly, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the text relevance determination method as described above.

[0016] In this embodiment, a first group of entities associated with the first text and a second group of entities associated with the second text are determined based on a preset knowledge base. The preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes. The entity relevance between the first group of entities and the second group of entities is determined based on the knowledge representation. Based on the relationships between each word in the first text, between each word in the second text, and between words in the first text and words in the second text, the attention value of each word in the first and second texts with respect to other words is determined. At least based on the attention value and the entity relevance, the text relevance between the first text and the second text is determined. In this scheme, the relationships between words within and between texts are considered when calculating text relevance, thus focusing on useful information while ignoring useless information, thereby improving the accuracy of the text relevance calculation results. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the method for determining text relevance provided in the embodiments of this application.

[0019] Figure 2 This is a schematic diagram of the model architecture provided in the embodiments of this application.

[0020] Figure 3 This is a structural diagram of the application scenario provided in the embodiments of this application.

[0021] Figure 4 This is a schematic diagram of the structure of the text relevance determination device provided in the embodiments of this application.

[0022] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.

[0023] Figure 6 This is a schematic diagram of the server structure provided in an embodiment of this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] This application provides a method, apparatus, storage medium, and electronic device for determining text relevance. Specifically, the apparatus for determining text relevance can be integrated into electronic devices such as tablet PCs (Personal Computers) and mobile phones that have storage units and are equipped with microprocessors and thus have computing capabilities.

[0026] Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, acquire knowledge, and use that knowledge to obtain optimal results, enabling machines to have the functions of perception, reasoning, and decision-making.

[0027] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing and semantic understanding.

[0028] This solution introduces a self-attention mechanism, which mimics the internal processes of biological observation—aligning internal experience with external sensation to increase the precision of observation in certain areas. Through this mechanism, important features of sparse data in text can be quickly extracted, and the internal correlations of data or features can be captured for analysis and processing, thereby achieving intelligent text processing.

[0029] The following sections provide detailed descriptions. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments. Please refer to... Figure 1 , Figure 1 This is a flowchart illustrating the method for determining text relevance provided in this application embodiment. The specific flow of this method for determining text relevance can be as follows:

[0030] 101. Based on a preset knowledge base, determine the first group of entities associated with the first text and the second group of entities associated with the second text. The preset knowledge base includes knowledge representations consisting of entities, relationships between entities, and entity attributes.

[0031] In this embodiment, the first text and the second text can both be relatively long texts. The first text can be the text to be retrieved, and the second text can be candidate texts that need to be matched with the text to be retrieved. Specifically, taking a question-and-answer scenario as an example, the first text can be the question text input by the user through an electronic device, and the second text can be the answer text pre-set in the answer database for that question text.

[0032] For example, the question text might be "Who is Zhang San's wife?", while the answer text could be "Zhang San's wife is Li Si," "Li Si's husband is Zhang San," or "She was born in February 1990, a teacher, and formerly..." or "Zhang San's wife is Wang Wu, a singer, born in May 1880 in City S..." and so on. It's clear that there are many choices of answer text for a single question. To save users time in accurately finding the most relevant answer, sorting the answer texts based on their relevance to the question becomes particularly important.

[0033] A pre-defined knowledge base, also known as a knowledge graph, is defined as follows: Entities refer to various objects and concepts existing in the real world, such as people, geographical locations, organizations, brands, professions, dates, etc. Relationships between entities refer to the associations between two entities; entity attributes refer to the properties of the entity itself. For example, a person's attributes could include profession, birthday, representative works, age, height, weight, gender, etc. Entity attributes can sometimes be considered a noun-like relationship between entities; therefore, the knowledge base describes one or more relationships between various entities. Using the question and answer text examples above, entities could include the people "Zhang San" and "Li Si," entity attributes could include the profession "teacher," the date "February 1990," and the relationship between entities could include the marital relationship between Zhang San and Li Si.

[0034] To facilitate computer processing and understanding, knowledge in a predefined knowledge base can be represented in the form of a triple, "Subject-Predication-Object (SPO)," such as: (First Entity, Relation / Attribute, Second Entity). For example, the knowledge "Zhang San's wife is Li Si" can be represented as the triple (Zhang San, Wife, Li Si). In this paper, relations or attributes (such as wife) are also referred to as "predicates," and the two entities with corresponding relations or attributes can serve as "subjects" or "objects." If we consider an entity as a node and the relations and attributes between entities as edges, then a knowledge base containing a large number of triples forms a vast knowledge graph. By associating knowledge elements such as entities, relations / attributes, etc., the corresponding knowledge can be easily retrieved from the knowledge base.

[0035] 102. Determine the entity relevance between the first group of entities and the second group of entities based on the knowledge representation.

[0036] Entity relevance is a quantitative representation of the degree of entity matching between the first group of entities and the second group of entities. It can be expressed as the similarity between each entity in the first group and each entity in the second group and each entity in the other group. In this embodiment, this similarity can be specifically determined by the word co-occurrence and essential co-occurrence of entities. The word co-occurrence represents the word overlap rate between the first and second groups of entities, while the essential co-occurrence represents the overlap rate of entity identifiers corresponding to the same entity words in the first and second groups. In practical applications, the word co-occurrence and essential co-occurrence of entities can be calculated manually and used as shallow features for calculating the relevance between the first and second texts.

[0037] In this embodiment, text entity association technology can be used to identify entities in the first and second texts and connect them to corresponding nodes in the knowledge graph. Considering that the entities included in the knowledge graph cannot guarantee complete coverage, the relevance between the question and the answer can be described simultaneously using entity words (i.e., entity mentions) identified from the text and entity identifiers (i.e., entity IDs) in a preset knowledge base. That is, when determining the entity relevance between the first group of entities and the second group of entities based on the knowledge representation, the following process can be included:

[0038] (11) Determine the first number of entities with the same name in the first group of entities and the second group of entities;

[0039] (12) Determine the second number of entities in the first group of entities and the second group of entities that have the same identifier in the knowledge base, wherein the identifier of an entity uniquely identifies the entity in the preset knowledge base;

[0040] (13) Determine the entity relevance based on the first number and the second number.

[0041] Specifically, the second number of entities with the same identifier in the knowledge base can be determined based on the knowledge representation in the pre-defined knowledge base. That is, entities in the first and second groups of entities that share the same entity relationships and attributes are identified through knowledge representation.

[0042] Taking a question-and-answer scenario as an example, for the entity mention, pre-trained word vector mention embeddings can be used to calculate the entity mention similarity between the text to be retrieved and the candidate texts respectively; for the entity ID, the entity words in the text to be retrieved and the candidate texts can be directly matched, and the matching result is used as the entity-level similarity. That is, in some embodiments, the first text is the text to be retrieved, and the second text is the candidate text. Then the step "determine the entity relevance based on the first number and the second number" can include the following process:

[0043] Based on the first number and the number of entities in the first group, the entity word similarity between the text to be retrieved and the candidate text is determined;

[0044] Based on the number of entities in the second group and the number of entities in the first group, the entity identifier similarity between the text to be retrieved and the candidate text is determined.

[0045] Entity relevance is determined based on the similarity of entity words and entity identifiers.

[0046] When determining entity relevance based on entity word similarity and entity identifier similarity, entity word similarity and entity identifier similarity can be used together as entity relevance; alternatively, entity word similarity and entity identifier similarity can be weighted based on corresponding weight information to obtain entity relevance.

[0047] The formula for calculating entity mention similarity is as follows:

[0048]

[0049] in:

[0050] Where `mention_qi` is the vector representation of the i-th entity word in the first text, `mention_dj` is the vector representation of the j-th entity word in the second text, `n` represents the number of entity words in the first text, and `m` represents the number of entity words in the second text, where `n` and `m` are both integers greater than or equal to 1. The above formula indicates that, for any one of the vector representations of entity words in the first text, the difference between it and the vector representations of entity words in the second text is determined, and then the maximum difference value is selected. For the vector representations of all entity words in the first text, the sum of the selected maximum difference values ​​is calculated, and the average is taken over the number of entity words in the first text. The average value is used as the entity `mention` similarity between the first and second texts.

[0051] The formula for calculating entity ID similarity is as follows:

[0052]

[0053] in:

[0054] Where id_qi is the i-th entity in the first group, id_dj is the j-th entity in the second group, n represents the number of entities in the first group, and m represents the number of entities in the second group, where n and m are both integers greater than or equal to 1. The above formula indicates that, for any entity in the first group, it determines whether there exists an entity with the same identifier in the second group. Then, the ratio of the number of entities with the same identifier in the first group to the total number of entities n in the first group is used to indicate the similarity of entity IDs. This can be understood as determining the similarity between the two groups of entities at the identifier level.

[0055] For example, the text to be retrieved includes: Xiao A (ID1), University E (ID2); Candidate text 1 includes: Xiao A (ID3), Xiao B (ID4), Xiao C (ID5), Xiao D (ID6); Candidate text 2 includes: Xiao A (ID1), Professor (ID7), University E (ID8).

[0056] Therefore, regarding the mention similarity between the text to be retrieved and candidate text 1: with the number of mentions in the text to be retrieved being 2 as the denominator, and the intersection of the mentions in the text to be retrieved and the mentions in candidate text 1 having 1 (i.e., small A) as the numerator, the entity mention similarity is: 1 / 2;

[0057] Regarding the ID similarity between the text to be retrieved and candidate text 1: since the number of IDs in the text to be retrieved is 2, which is the denominator, and there is no intersection between the IDs of the text to be retrieved and the IDs of candidate text 1, the numerator is 0, resulting in an entity ID similarity of 0 / 2.

[0058] Regarding the mention similarity between the text to be retrieved and candidate text 2: the number of mentions in the text to be retrieved is 2, which is used as the denominator; the mentions in the text to be retrieved and the mentions in candidate text 1 have an intersection of 1 (i.e., small A), which is used as the numerator, resulting in an entity mention similarity of 1 / 2;

[0059] Regarding the ID similarity between the text to be retrieved and candidate text 2: Since the number of IDs in the text to be retrieved is 2, which is the denominator, and there is one intersection between the IDs in the text to be retrieved and the IDs in candidate text 1 (i.e., ID1), which is the numerator, the entity ID similarity is 1 / 2.

[0060] 103. Based on the relationships between each word in the first text, the relationships between each word in the second text, and the relationships between words in the first text and words in the second text, determine the attention value of each word in the first text and the second text with respect to other words. The attention value is used to reflect the degree of attention each word in the first text and the second text pays to other words.

[0061] In some embodiments, when performing deep representation of text, an RNN is introduced into the representation layer to process the text. However, since the RNN only focuses on the part of speech of the words themselves, there will be a problem of information diffusion for longer sentences.

[0062] For example, consider a question-and-answer scenario where a person's wife is a person named Qian. The question is, "Who is a person's wife?" The answer is a longer text: "In 1984, she and her sisters won third place in a competition in a certain country, and then went to that country to study cosmetology; from 1985 to 1987, she worked as a print model; on June 23, 2008, she registered her marriage with a person named Hua in a certain place." It is clear that "marriage" and the person named Qian (who is also a person's wife) are quite distant, and RNNs cannot model long-distance relationships between entities. Therefore, this solution introduces a self-attention mechanism into this scenario. Since the self-attention mechanism can focus on the dependencies between words within the text and learn the characteristics of sentence structure, it can effectively solve the problem of long-distance word dependencies. In other words, this scheme can determine the dependencies between words by learning the relationships between words within the text and between words in different texts, and determine the importance (i.e. attention) of each word to the text. When outputting the deep representation of the text, the scheme increases the attention to words with high attention values ​​and reduces the attention to words with low attention values, thereby retaining the "useful information" that is highly relevant to the text itself and removing the "useless information" that is less relevant.

[0063] In some embodiments, when processing the concatenated matrix based on the relationships between words in the first text, the relationships between words in the second text, and the relationships between words in the first and second texts to obtain a processed matrix, the specific steps may be as follows: Calculate the relevance between each word in the first text and other words based on the relationships between words in the first text and the relationships between words in the first and second texts; calculate the relevance between each word in the second text and other words based on the relationships between words in the second text and the relationships between words in the first and second texts. Finally, determine the attention value of each word in both the first and second texts regarding other words based on the relevance values ​​of each word in the first and second texts with respect to other words.

[0064] In this embodiment, a self-attention mechanism is introduced to find the most relevant context for different words or entities in the text, and then weighted to obtain the final hidden layer representation. In the self-attention mechanism, each word has three different vectors: a query vector (Q), a key vector (K), and a value vector (V), each with a length of 64. These vectors are obtained by multiplying the word's embedding vector X by three different weight matrices W. Q W K WV We obtain the three weight matrices W. Q W K W V All dimensions are the same; for example, the dimensions can be 512 x 64.

[0065] In practice, the input words or entities can be converted into embedding vectors, and then three vectors, Q, K, and V, can be obtained from the embedding vectors. A relevance score (representing relevance) is calculated for each word or entity to other words or entities, i.e., socre = Q * K. To stabilize the gradient, the softmax activation function can be used to normalize the value of each score. The normalized value is then multiplied by the value vector V of each word or entity to obtain a weighted score V for each input vector. These scores are summed to obtain the final output: Z = sum(V), which serves as the attention vector for the input words or entities. Processing this attention vector yields the attention value for each word or entity.

[0066] It can be seen that by calculating the correlation between different words, the attention between words can be represented, thus retaining "useful information" and removing "useless information". The introduction of this self-attention mechanism can effectively represent long texts.

[0067] refer to Figure 2 , Figure 2 This is a schematic diagram of the model architecture provided in an embodiment of this application. In some embodiments, it is necessary to pre-construct feature matrices corresponding to the first text and the second text, respectively, to obtain a first feature matrix and a second feature matrix; the first feature matrix and the second feature matrix are then concatenated to obtain a concatenated matrix. Then, when determining the attention value of each word in the first and second texts with respect to other words based on the relevance between each word in the first text and other words, and the relevance between each word in the second text and other words, specifically, the relevance between each word in the first text and other words, and the relevance between each word in the second text and other words can be normalized. The concatenated matrix is ​​then weighted based on the normalized relevance to obtain a weighted matrix. The attention value of each word in the first and second texts with respect to other words is then determined based on the weighted matrix.

[0068] In some embodiments, the process of constructing the feature matrix corresponding to the first text and the feature matrix corresponding to the second text to obtain the first feature matrix and the second feature matrix may specifically include the following steps:

[0069] (21) Perform word segmentation on the first text and the second text to obtain the first group of words associated with the first text and the second group of words associated with the second text;

[0070] (22) Based on each word in the first group of words and the position of each word in the first text, construct the first vector representation of each word in the first group of words;

[0071] (23) Based on each word in the second group of words and the position of each word in the second text, construct the second vector representation of each word in the second group of words;

[0072] (24) Determine the first feature matrix based at least on the constructed first vector representation, and determine the second feature matrix based at least on the constructed second vector representation.

[0073] Specifically, for the first and second texts, word segmentation technology (such as Yaha segmentation, Jieba segmentation, etc.) can be used to segment the first and second texts, and operations such as removing stop words, removing punctuation marks, and converting emoticons can be performed to divide the first and second texts into individual words, resulting in the first group of words associated with the first text and the second group of words associated with the second text.

[0074] In this embodiment, a vector representation can be constructed by combining the word's part of speech and its position in the text to better express the word's actual semantics in the text. That is, both the first and second vector representations of the above-mentioned components include: word embeddings and position embeddings. Specifically, the word embeddings are vector representations of each word in the text mapped onto the real number field; the position embeddings are vector representations of the position of each word in the text mapped onto the real number field.

[0075] Specifically, for word embeddings, we first utilize the entity SPO triple information stored in the knowledge graph to train word-level embeddings using a CBOW approach, aiming to ensure that entities with a relationship will have similar embeddings. This scheme, by utilizing entity SPO triple information, better characterizes the similarity between entities with similar relationships. Since the input SPO information is relatively short, the word window can be set to 1 in practical implementation.

[0076] Position embedding is introduced to model word order. For example, to map a word's position p to a dpos-dimensional vector PE, the calculation method is as follows:

[0077]

[0078] In this context, the value of the i-th element of the identifier vector PE in the above formula is PE.i (p), where PE 2i For an even number of digits, PE 2i+1 (For odd-numbered positions). In practical applications, since sin(α+β)=sinαcosβ+cosαsinβ and cos(α+β)=cosαcosβ-sinαsinβ, in this embodiment, the vector at position p+k can be represented as a linear transformation of the vector at position p through this expression method, providing the possibility of expressing relative position information.

[0079] Subsequently, a preset model (reference) can be used. Figure 2 The deep representation layer in the text performs operations such as constructing vector representations of the segmented words, constructing feature matrices, matrix transformations, and extracting feature vectors from the matrices. For different words or entities in the text, it finds the most relevant context and outputs the final hidden layer representation with weights.

[0080] refer to Figure 2 and Figure 3 In some embodiments, determining the first feature matrix based at least on the constructed first vector representation may specifically include the following process:

[0081] The constructed first vector representations are concatenated to obtain the first submatrix;

[0082] Based on a pre-defined knowledge base, the first entity word is identified from the first group of words. The first knowledge element related to the first entity word is determined from the pre-defined knowledge base. According to the position of the first entity word in the first text and the knowledge representation formed by the first knowledge element, the vector representation of the first knowledge element is concatenated to obtain the second sub-matrix. The first knowledge element includes: the first related entity that has a relationship with the first target entity in the pre-defined knowledge base corresponding to the first entity word, the relationship between the first target entity and the first related entity, and / or the entity attribute of the first target entity.

[0083] The first feature matrix is ​​determined based on the first and second submatrices.

[0084] Specifically, when performing entity word recognition, entity linking technology can be used to identify entity words from the first group of words. When determining the first feature matrix based on the first and second sub-matrices, the first feature matrix can be obtained by directly concatenating the first and second sub-matrices. This first feature matrix integrates the word embeddings of each word in the first text, the positions of each word in the first text, and the entity embeddings of each entity word in the first text. It should be noted that the entity embedding vector is a vector representation of the relevant knowledge elements of each entity word in the first text mapped onto the real number field.

[0085] In this scheme, the SPO of an entity is modeled as an additive relationship (i.e., S + P = O) for model training. The goal is to make the sum of the embeddings of S and P as equal as possible to the embedding of O. After training in this way, the embeddings of S, P, and O describe the holding relationship, that is, the entity is characterized by the vector representations of its related knowledge elements, and the vector representation resulting from the combination of the vector representations of the related knowledge elements is as close as possible to the vector representation of the entity itself.

[0086] For example, taking the SPO triple knowledge representation as (Zhang San, couple, Li Si), if the entity word in the text is "Zhang San", then the vector representation of the relevant knowledge elements of the entity word "Zhang San" (i.e., the relevant entity "Li Si" and the relationship between entities "couple") can be obtained as the entity embedding of the entity word "Zhang San".

[0087] Continue to refer to Figure 2 and Figure 3 When determining the second feature matrix based at least on the constructed second vector representation, the specific process may include the following:

[0088] The constructed second vector representation is concatenated to obtain the third submatrix;

[0089] Based on the preset knowledge base, the second entity word is identified from the second group of words, the second knowledge element related to the second entity word is determined from the preset knowledge base, and the vector representation of the second knowledge element is concatenated according to the position of the second entity word in the second text and the knowledge representation constituted by the second knowledge element to obtain the fourth sub-matrix. The second knowledge element includes: the second related entity that has a relationship with the second target entity in the preset knowledge base corresponding to the second entity word, the relationship between the second target entity and the second related entity, and / or the entity attribute of the second target entity.

[0090] The second characteristic matrix is ​​determined based on the third and fourth submatrices.

[0091] Similarly, when performing entity word recognition, entity linking technology can be used to identify entity words from the second group of words. When determining the second feature matrix based on the third and fourth sub-matrices, it can be directly obtained by concatenating the third and fourth sub-matrices. This second feature matrix integrates the word embeddings of each word in the second text, the positions of each word in the second text, and the entity embeddings of each entity word in the second text. It should be noted that the entity embedding vector is a vector representation of the relevant knowledge elements of each entity word in the second text mapped onto the real number field.

[0092] As can be seen, the self-attention value in this scheme integrates the vector representations of deep features such as word embeddings, positional embeddings, and entity embeddings of each word in the first and second texts. The self-attention mechanism introduced in the deep representation layer transforms and calculates the feature matrix composed of these deep feature vector representations, outputting the final hidden layer representation, which is the attention vector. This attention vector can be a one-dimensional vector of size 1xn. This one-dimensional vector contains the attention value of each word in the first and second texts with respect to the other words.

[0093] 104. Determine the text relevance between the first text and the second text based at least on the attention value and entity relevance.

[0094] In practical applications, the more words from the question text appear in the answer text, the more relevant the answer and question are to each other. Since questions are relatively shorter than answers, the proportion of question words covered in the answer text can be used to construct a word co-occurrence feature between questions and answers. This feature can then be combined with the word co-occurrence between the first and second texts to calculate text relevance.

[0095] In other words, before determining the relevance between the first and second texts, a third number of identical words in both texts can be determined. Based on this third number and the number of words in the first text, the word relevance between the first and second texts can be determined. Therefore, when determining the text relevance between the first and second texts based on attention values ​​and entity relevance, the text relevance can specifically be determined using attention values, entity relevance, and word relevance.

[0096] In other embodiments, features related to the statistical information of the first and second texts themselves can be introduced when calculating text relevance. For example, relevance features may include the character length and word length of the first text, the character length and word length of the second text, the source confidence of the answer text (i.e., the second text), and the similarity between the classification of the question text (i.e., the first text) and the classification of the answer text (i.e., the second text), which can be defined according to actual needs.

[0097] For details, please refer to [link / reference]. Figure 3When calculating text relevance, shallow features of the first and second texts can be extracted (i.e., word co-occurrence, entity co-occurrence, word length, character length, confidence score, classification similarity, etc.). A feature vector (usually a one-dimensional feature vector) is then constructed from these shallow features. Subsequently, the constructed feature vector is concatenated with the feature vector output from the model's deep representation layer (which represents the attention value of each word in the first and second texts with respect to other words), resulting in a concatenated vector. Finally, the concatenated vector is normalized using the softmax activation function to obtain the text relevance between the question and answer texts.

[0098] In practical applications, taking the scenario of searching for relevant answers to a question in a certain search database as an example, after calculating the text relevance between the question text and each answer text, the answer texts can be sorted and displayed based on the magnitude of the text relevance. The answer texts with higher text relevance are displayed first, and the answer texts with lower relevance are displayed later, thereby increasing the exposure of accurate answers.

[0099] The text relevance determination method provided in this application's embodiments determines a first group of entities associated with a first text and a second group of entities associated with a second text based on a preset knowledge base. The preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes. The method determines the entity relevance between the first group of entities and the second group of entities based on the knowledge representation. It also determines the attention value of each word in the first and second texts with respect to other words based on the relationships between words in the first and second texts, between words in the second text, and between words in the first and second texts. Finally, it determines the text relevance between the first and second texts based at least on the attention value and the entity relevance. In this solution, the text relevance calculation focuses on the relationships between words within a text and between words in other texts, and based on these relationships, it increases the weight of useful information and decreases the weight of useless information, thereby improving the accuracy of the text relevance calculation results.

[0100] To facilitate better implementation of the text relevance determination method provided in the embodiments of this application, the embodiments of this application also provide an apparatus based on the above-described text relevance determination method. The meanings of the terms used are the same as in the above-described text relevance determination method, and specific implementation details can be found in the descriptions in the method embodiments.

[0101] Please see Figure 4 , Figure 4This is a schematic diagram of the structure of a text relevance determination device provided in an embodiment of this application. The processing device may include: an entity determination unit 301, a first relevance determination unit 302, an attention determination unit 303, and a second relevance determination unit 304. Specifically, it may be as follows:

[0102] The entity determination unit 301 is used to determine a first group of entities associated with the first text and a second group of entities associated with the second text based on a preset knowledge base. The preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes.

[0103] The first relevance determination unit 302 is used to determine the entity relevance between the first group of entities and the second group of entities based on the knowledge representation;

[0104] Attention determination unit 303 is used to determine the attention value of each word in the first text and the second text about other words based on the association between each word in the first text, the association between each word in the second text, and the association between words in the first text and words in the second text, wherein the attention value is used to reflect the degree of attention each word in the first text and the second text pays to other words;

[0105] The second relevance determination unit 304 is used to determine the text relevance between the first text and the second text based at least on the attention value and the entity relevance.

[0106] In some embodiments, the attention determination unit 303 may be used for:

[0107] Based on the relationships between each word in the first text and the relationships between words in the first text and words in the second text, calculate the relevance of each word in the first text to other words;

[0108] Based on the relationships between each word in the second text and the relationships between words in the first text and words in the second text, the relevance of each word in the second text to other words is calculated.

[0109] Based on the relevance of each word in the first text to other words, and the relevance of each word in the second text to other words, determine the attention value of each word in the first text and the second text with respect to other words.

[0110] In some embodiments, the device may further include:

[0111] The construction unit is used to construct the feature matrix corresponding to the first text and the feature matrix corresponding to the second text respectively, to obtain the first feature matrix and the second feature matrix;

[0112] The splicing unit is used to splice the first feature matrix and the second feature matrix to obtain a spliced ​​matrix;

[0113] The attention unit 303 can also be used for:

[0114] The relevance between each word in the first text and other words, and the relevance between each word in the second text and other words are normalized.

[0115] The spliced ​​matrix is ​​weighted according to the relevance after normalization to obtain a weighted matrix;

[0116] The attention value of each word in the first and second texts with respect to other words is determined based on the weighted matrix.

[0117] In some embodiments, the building unit can be specifically used for:

[0118] The first text and the second text are segmented to obtain a first group of words associated with the first text and a second group of words associated with the second text.

[0119] Based on each word in the first group of words and the position of each word in the first text, construct the first vector representation of each word in the first group of words;

[0120] Based on each word in the second group of words and the position of each word in the second text, construct a second vector representation of each word in the second group of words;

[0121] The first feature matrix is ​​determined at least based on the constructed first vector representation, and the second feature matrix is ​​determined at least based on the constructed second vector representation.

[0122] In some embodiments, the building unit can further be used for:

[0123] The constructed first vector representations are concatenated to obtain the first submatrix;

[0124] Based on the preset knowledge base, a first entity word is identified from the first group of words. First knowledge elements related to the first entity word are determined from the preset knowledge base. Then, according to the position of the first entity word in the first text and the knowledge representation constituted by the first knowledge elements, the vector representations of the first knowledge elements are concatenated to obtain a second sub-matrix. The first knowledge elements include: a first related entity that has a relationship with the first target entity corresponding to the first entity word in the preset knowledge base; the relationship between the first target entity and the first related entities; and / or the entity attributes of the first target entity.

[0125] The constructed second vector representation is concatenated to obtain the third submatrix;

[0126] Based on the preset knowledge base, the second entity word is identified from the second group of words, the second knowledge element related to the second entity word is determined from the preset knowledge base, and the vector representation of the second knowledge element is concatenated according to the position of the second entity word in the second text and the knowledge representation constituted by the second knowledge element to obtain the fourth sub-matrix. The second knowledge element includes: the second related entity that has a relationship with the second target entity in the preset knowledge base corresponding to the second entity word, the relationship between the second target entity and the second related entity, and / or the entity attribute of the second target entity.

[0127] The second feature matrix is ​​determined based on the third and fourth submatrices.

[0128] In some embodiments, the first relevance determination unit 302 may be used to:

[0129] Determine the first number of entities with the same name in the first group of entities and the second group of entities;

[0130] Determine a second number of entities in the first group of entities and the second group of entities that have the same identifier in the knowledge base, wherein the identifier of an entity uniquely identifies the entity in the preset knowledge base;

[0131] The entity relevance is determined based on the first number and the second number.

[0132] In some embodiments, the first text is the text to be retrieved, and the second text is the candidate text; the first relevance determination unit 302 can further be used to:

[0133] Based on the first number and the number of entities in the first group of entities, the entity word similarity between the text to be retrieved and the candidate text is determined;

[0134] Based on the second number and the number of entities in the first group, the entity identifier similarity between the text to be retrieved and the candidate text is determined;

[0135] The entity relevance is determined based on the entity word similarity and the entity identifier similarity.

[0136] In some embodiments, the apparatus may further include a third correlation determination unit, configured to:

[0137] Determine a third number of identical words in the first text and the second text;

[0138] Based on the third number and the number of words in the first text, the word relevance between the first text and the second text is determined;

[0139] The second relevance unit can specifically be used for:

[0140] The text relevance between the first text and the second text is determined based on the attention value, the entity relevance, and the word relevance.

[0141] The text relevance determination method provided in this solution determines a first group of entities associated with a first text and a second group of entities associated with a second text based on a pre-set knowledge base. The pre-set knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes. The method then determines the entity relevance between the first and second groups of entities based on the knowledge representation. Based on the relationships between words in the first and second texts, between words in the second text, and between words in the first and second texts, the method determines the attention value of each word in both texts with respect to other words. Finally, based at least on the attention value and entity relevance, the method determines the text relevance between the first and second texts. This solution focuses on the relationships between words within a text and between words in other texts during text relevance calculation, and based on these relationships, increases the weight of useful information and decreases the weight of useless information, thereby improving the accuracy of the text relevance calculation results.

[0142] This application also provides an electronic device, which may specifically be a smartphone, tablet computer, or other terminal device. Figure 5 As shown, the electronic device may include a radio frequency (RF) circuit 601, a memory 602 including one or more computer-readable storage media, an input unit 603, a display unit 604, a sensor 605, an audio circuit 606, a wireless Fidelity (WiFi) module 607, a processor 608 including one or more processing cores, and a power supply 609, among other components. Those skilled in the art will understand that... Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0143] RF circuit 601 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 608 for processing; additionally, it transmits uplink data to the base station. Typically, RF circuit 601 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 601 can also communicate wirelessly with networks and other devices. The wireless communication can use any communication standard or protocol, including but not limited to GSM, GPRS, CDMA, WCDMA, LTE, email, and SMS.

[0144] The memory 602 can be used to store software programs and modules. The processor 608 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, telephone directory, etc.). In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide access to the memory 602 for the processor 608 and the input unit 603.

[0145] The input unit 603 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, in one embodiment, the input unit 603 may include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display or touchpad, can collect user touch operations on or near it (e.g., user operations using fingers, styluses, or any suitable object or accessory on or near the touch-sensitive surface), and drive corresponding connection devices according to a pre-set program. Optionally, the touch-sensitive surface may include a touch detection device and a touch controller. The touch detection device detects the user's touch orientation and the signal generated by the touch operation, transmitting the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 608, and can receive and execute commands from the processor 608. Furthermore, various types of touch-sensitive surfaces, such as resistive, capacitive, infrared, and surface acoustic wave, can be used. In addition to the touch-sensitive surface, the input unit 603 may also include other input devices. Specifically, other input devices may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0146] Display unit 604 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic devices. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 604 may include a display panel, optionally configured as a liquid crystal display (LCD), organic light-emitting diode (OLED), or similar form. Furthermore, a touch-sensitive surface may cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it transmits the information to processor 608 to determine the type of touch event. Subsequently, processor 608 provides corresponding visual output on the display panel according to the type of touch event. Although in Figure 5 In this context, the touch-sensitive surface and the display panel are two separate components for implementing input and output functions. However, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve both input and output functions.

[0147] The electronic device may also include at least one sensor 605, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel according to the ambient light level, and the proximity sensor can turn off the display panel and / or backlight when the electronic device is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the electronic device, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0148] Audio circuitry 606, a speaker, and a microphone provide an audio interface between the user and the electronic device. Audio circuitry 606 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 606, converted back into audio data, and processed by processor 608. The processed data is then transmitted via RF circuitry 601 to, for example, another electronic device, or output to memory 602 for further processing. Audio circuitry 606 may also include an earphone jack to facilitate communication between external headphones and the electronic device.

[0149] WiFi is a short-range wireless transmission technology. Electronic devices using the WiFi module 607 can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 5 WiFi module 607 is shown, but it is understood that it is not a necessary component of an electronic device and can be omitted as needed without changing the nature of the invention.

[0150] The processor 608 is the control center of the electronic device. It connects various parts of the phone via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 602, and by calling data stored in the memory 602, thereby providing overall monitoring of the phone. Optionally, the processor 608 may include one or more processing cores; preferably, the processor 608 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 608.

[0151] The electronic device also includes a power supply 609 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 608 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 609 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0152] Although not shown, the electronic device may also include a camera, Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 608 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 602 according to the following instructions, and the processor 608 runs the applications stored in the memory 602 to realize various functions:

[0153] Based on a preset knowledge base, a first group of entities associated with a first text and a second group of entities associated with a second text are determined. The preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes. The entity relevance between the first group of entities and the second group of entities is determined based on the knowledge representation. Based on the relationships between words in the first text, the relationships between words in the second text, and the relationships between words in the first text and the second text, an attention value is determined for each word in the first and second texts regarding other words. The attention value reflects the degree of attention each word in the first and second texts pays to other words. The text relevance between the first text and the second text is determined based at least on the attention value and the entity relevance.

[0154] The electronic device provided in this application embodiment can focus on the relationship between each word in the text and between words in the text when calculating text relevance, and based on this relationship, increase the weight of useful information and decrease the weight of useless information, thereby improving the accuracy of text relevance calculation results.

[0155] This application also provides a server, which may specifically be an application server. For example... Figure 6 As shown, the server may include radio frequency (RF) circuitry 701, a memory 702 including one or more computer-readable storage media, a processor 704 including one or more processing cores, and a power supply 703, among other components. Those skilled in the art will understand that... Figure 6 The server architecture shown does not constitute a limitation on the server and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. Wherein:

[0156] RF circuit 701 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 704 for processing; additionally, it transmits uplink data to the base station. Typically, RF circuit 701 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 701 can also communicate wirelessly with networks and other devices. The wireless communication can use any communication standard or protocol, including but not limited to GSM, GPRS, CDMA, WCDMA, LTE, email, and SMS.

[0157] The memory 702 can be used to store software programs and modules. The processor 704 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the server (such as audio data, telephone directory, etc.). In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide access to the memory 702 for the processor 704 and the input unit 703.

[0158] The processor 704 is the control center of the server, connecting various parts of the mobile phone through various interfaces and lines. It performs various server functions and processes data by running or executing software programs and / or modules stored in the memory 702, and by calling data stored in the memory 702, thereby providing overall monitoring of the mobile phone. Optionally, the processor 704 may include one or more processing cores; preferably, the processor 704 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 704.

[0159] The server also includes a power supply 703 (such as a battery) to power various components. Preferably, the power supply can be logically connected to the processor 704 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0160] Specifically, in this embodiment, the processor 704 in the server loads the executable files corresponding to one or more application processes into the memory 702 according to the following instructions, and the processor 704 runs the applications stored in the memory 702 to achieve various functions:

[0161] Based on a preset knowledge base, a first group of entities associated with a first text and a second group of entities associated with a second text are determined. The preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes. The entity relevance between the first group of entities and the second group of entities is determined based on the knowledge representation. Based on the relationships between words in the first text, the relationships between words in the second text, and the relationships between words in the first text and the second text, an attention value is determined for each word in the first and second texts regarding other words. The attention value reflects the degree of attention each word in the first and second texts pays to other words. The text relevance between the first text and the second text is determined based at least on the attention value and the entity relevance.

[0162] The electronic device provided in this application embodiment can focus on the relationship between each word in the text and between words in the text when calculating text relevance, and based on this relationship, increase the weight of useful information and decrease the weight of useless information, thereby improving the accuracy of text relevance calculation results.

[0163] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0164] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the text relevance determination methods provided in embodiments of this application. For example, the instructions can execute the following steps:

[0165] Based on a preset knowledge base, a first group of entities associated with a first text and a second group of entities associated with a second text are determined. The preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes. The entity relevance between the first group of entities and the second group of entities is determined based on the knowledge representation. Based on the relationships between words in the first text, the relationships between words in the second text, and the relationships between words in the first text and the second text, an attention value is determined for each word in the first and second texts regarding other words, where the attention value reflects the degree of attention each word in the first and second texts pays to other words. At least based on the attention value and the entity relevance, the text relevance between the first text and the second text is determined.

[0166] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0167] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0168] Since the instructions stored in the storage medium can execute the steps in any of the text relevance determination methods provided in the embodiments of this application, the beneficial effects that any of the text relevance determination methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0169] The foregoing has provided a detailed description of a method, apparatus, storage medium, and electronic device for determining text relevance according to embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for determining text relevance, characterized in that, include: Based on a preset knowledge base, a first group of entities associated with the first text and a second group of entities associated with the second text are determined. The preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes. Determine the entity relevance between the first group of entities and the second group of entities based on the knowledge representation; Based on the relationships between each word in the first text and the relationships between words in the first text and words in the second text, calculate the relevance of each word in the first text to other words; Based on the relationships between each word in the second text and the relationships between words in the first text and words in the second text, the relevance of each word in the second text to other words is calculated. Construct the feature matrix corresponding to the first text and the feature matrix corresponding to the second text respectively to obtain the first feature matrix and the second feature matrix; The first feature matrix and the second feature matrix are concatenated to obtain a concatenated matrix; The relevance between each word in the first text and other words, and the relevance between each word in the second text and other words are normalized. The concatenated matrix is ​​weighted according to the relevance after normalization to obtain a weighted matrix; based on the weighted matrix, the attention value of each word in the first text and the second text with respect to other words is determined, wherein the attention value is used to reflect the attention of each word in the first text and the second text to other words; The text relevance between the first text and the second text is determined based at least on the attention value and the entity relevance.

2. The method for determining text relevance according to claim 1, characterized in that, The step of constructing the feature matrix corresponding to the first text and the feature matrix corresponding to the second text respectively, to obtain the first feature matrix and the second feature matrix, includes: The first text and the second text are segmented to obtain a first group of words associated with the first text and a second group of words associated with the second text. Based on each word in the first group of words and the position of each word in the first text, construct the first vector representation of each word in the first group of words; Based on each word in the second group of words and the position of each word in the second text, construct a second vector representation of each word in the second group of words; The first feature matrix is ​​determined at least based on the constructed first vector representation, and the second feature matrix is ​​determined at least based on the constructed second vector representation.

3. The method for determining text relevance according to claim 2, characterized in that, The step of concatenating at least the constructed first vector representation to obtain the first feature matrix includes: The constructed first vector representations are concatenated to obtain the first submatrix; Based on the preset knowledge base, a first entity word is identified from the first group of words, a first knowledge element related to the first entity word is determined from the preset knowledge base, and the vector representations of the first knowledge elements are concatenated according to the position of the first entity word in the first text and the knowledge representation constituted by the first knowledge elements to obtain a second sub-matrix. The first knowledge element includes: a first related entity that has a relationship with the first target entity in the preset knowledge base corresponding to the first entity word, the relationship between the first target entity and the first related entity, and / or the entity attributes of the first target entity. The first feature matrix is ​​determined based on the first sub-matrix and the second sub-matrix; Determining the second feature matrix based at least on the constructed second vector representation includes: The constructed second vector representation is concatenated to obtain the third submatrix; Based on the preset knowledge base, the second entity word is identified from the second group of words, the second knowledge element related to the second entity word is determined from the preset knowledge base, and the vector representation of the second knowledge element is concatenated according to the position of the second entity word in the second text and the knowledge representation constituted by the second knowledge element to obtain the fourth sub-matrix. The second knowledge element includes: the second related entity that has a relationship with the second target entity in the preset knowledge base corresponding to the second entity word, the relationship between the second target entity and the second related entity, and / or the entity attribute of the second target entity. The second feature matrix is ​​determined based on the third and fourth submatrices.

4. The method for determining text relevance according to claim 1, characterized in that, Determining the entity relevance between the first group of entities and the second group of entities based on the knowledge representation includes: Determine the first number of entities with the same name in the first group of entities and the second group of entities; Determine a second number of entities in the first group of entities and the second group of entities that have the same identifier in the knowledge base, wherein the identifier of an entity uniquely identifies the entity in the preset knowledge base; The entity relevance is determined based on the first number and the second number.

5. The method for determining text relevance according to claim 4, characterized in that, The first text is the text to be searched, and the second text is the candidate text; Determining the entity relevance based on the first number and the second number includes: Based on the first number and the number of entities in the first group of entities, the entity word similarity between the text to be retrieved and the candidate text is determined; Based on the second number and the number of entities in the first group, the entity identifier similarity between the text to be retrieved and the candidate text is determined; The entity relevance is determined based on the entity word similarity and the entity identifier similarity.

6. The method for determining text relevance according to claim 5, characterized in that, Also includes: Determine a third number of identical words in the first text and the second text; Based on the third number and the number of words in the first text, the word relevance between the first text and the second text is determined; Determining the text relevance between the first text and the second text based at least on the attention value and the entity relevance includes: The text relevance between the first text and the second text is determined based on the attention value, the entity relevance, and the word relevance.

7. A device for determining text relevance, characterized in that, include: An entity determination unit is used to determine, based on a preset knowledge base, a first set of entities associated with a first text and a second set of entities associated with a second text, wherein the preset knowledge base includes a knowledge representation consisting of entities, relationships between entities, and entity attributes. The first relevance determination unit is used to determine the entity relevance between the first group of entities and the second group of entities based on the knowledge representation; An attention determination unit is used to calculate the relevance between each word in the first text and other words based on the correlation between each word in the first text and the correlation between words in the first text and words in the second text. Based on the relationships between each word in the second text and the relationships between words in the first text and words in the second text, the relevance of each word in the second text to other words is calculated. The construction unit is used to construct the feature matrix corresponding to the first text and the feature matrix corresponding to the second text respectively, to obtain the first feature matrix and the second feature matrix; The splicing unit is used to splice the first feature matrix and the second feature matrix to obtain a spliced ​​matrix; The attention determination unit is further configured to: normalize the relevance between each word in the first text and other words, and the relevance between each word in the second text and other words; weight the concatenation matrix according to the normalized relevance to obtain a weighted matrix; and determine the attention value of each word in the first text and the second text with respect to other words based on the weighted matrix, wherein the attention value is used to reflect the attention of each word in the first text and the second text to other words; The second relevance determination unit is used to determine the text relevance between the first text and the second text based at least on the attention value and the entity relevance.

8. The device for determining text relevance according to claim 7, characterized in that, The first relevance determination unit is used for: Determine the first number of entities with the same name in the first group of entities and the second group of entities; Determine a second number of entities in the first group of entities and the second group of entities that have the same identifier in the knowledge base, wherein the identifier of an entity uniquely identifies the entity in the preset knowledge base; The entity relevance is determined based on the first number and the second number.

9. The device for determining text relevance according to claim 8, characterized in that, The first text is the text to be retrieved, and the second text is the candidate text; the first relevance determination unit is further used for: Based on the first number and the number of entities in the first group of entities, the entity word similarity between the text to be retrieved and the candidate text is determined; Based on the second number and the number of entities in the first group, the entity identifier similarity between the text to be retrieved and the candidate text is determined; The entity relevance is determined based on the entity word similarity and the entity identifier similarity.

10. A computer-readable storage medium, characterized in that, The storage medium stores multiple instructions adapted for loading by a processor to execute a method for determining text relevance as described in any one of claims 1-6.

11. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a method for determining text relevance as described in any one of claims 1-6.