Training and Representation Methods of Entity Representation Model, Electronic Device, and Storage Medium

By acquiring multiple entities in an entity set and their association relationships, and optimizing entity representation using the entity representation model, the problem of low availability of entity representation in the prior art is solved, and efficient optimization and widely applicable entity representation model are achieved in unsupervised situations.

CN112632986BActive Publication Date: 2025-06-20HEFEI IFLYTEK TOYCLOUD TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011529521.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-22
Publication Date
2025-06-20
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

The available entity characterization obtained by existing entity characterization methods is low.

Method used

By obtaining multiple entities in the entity set and their association relationships, a training example is formed, and the entity representation model is used to obtain the representation of entities. Based on these representations, the model parameters are adjusted to optimize the representation effect.

Benefits of technology

Gradually optimize the entity representation model without supervision, improve the availability of entity representation, and is applicable to all types of entities, with high versatility and reducing human resource consumption during training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112632986B_ABST
    Figure CN112632986B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for an entity representation model, an entity representation method, an electronic device, and a storage medium. The training method includes: obtaining an entity set, where the entity set includes a plurality of entities and the association relationships between different entities, and the association relationships between different entities are used to indicate whether the corresponding entities are relevant; extracting a first entity and a second entity from the entity set to form a training example; sending the training example into the entity representation model; respectively using the entity representation model to obtain a first representation of the first entity and a first representation of the second entity; based on the first representation of the first entity and the first representation of the second entity, determining whether the first entity and the second entity are relevant; comparing the determination result with the association relationship between the first entity and the second entity, and adjusting the parameters of the entity representation model according to the comparison result. Through the above method, the availability of the representation of the entity obtained by the entity representation model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and particularly to a method for training an entity representation model, an entity representation method, an electronic device, and a storage medium. Background Art

[0002] An entity refers to an object that exists in the real world and can be distinguished from other objects. In the field of natural language processing, anything with an entry link in a search engine can be called an entity. The most typical ones include various personal names, place names, organization names, etc. In addition, words such as the Chinese words "happy" and "embarrassed" can also be regarded as entities.

[0003] The representation of an entity refers to the vector expression of the entity in the feature space. Obtaining the representation of an entity is of great significance for many downstream tasks in natural language processing. For example, the obtained entity representation can be used to construct the triple relationship in the knowledge base; or the obtained entity representation can be used as additional information in named entity recognition to improve the accuracy of slotting; or the obtained entity representation can also be used to assist in improving the effect of document classification tasks, etc.

[0004] However, the availability of the entity representations obtained by existing entity representation methods is relatively low. Summary of the Invention

[0005] This application provides a method for training an entity representation model, an entity representation method, an electronic device, and a storage medium, which can solve the problem that the availability of the entity representations obtained by existing entity representation methods is relatively low.

[0006] To solve the above technical problems, a technical solution adopted in this application is: to provide a method for training an entity representation model. The method includes: obtaining an entity set, where the entity set includes multiple entities and the association relationships between different entities, and the association relationships between different entities are used to indicate whether the corresponding entities are relevant; extracting a first entity and a second entity from the entity set to form a training example; sending the training example into the entity representation model; respectively using the entity representation model to obtain a first representation of the first entity and a first representation of the second entity; based on the first representation of the first entity and the first representation of the second entity, determining whether the first entity and the second entity are relevant; comparing the determination result with the association relationship between the first entity and the second entity, and adjusting the parameters of the entity representation model according to the comparison result.

[0007] To solve the above technical problems, another technical solution adopted in this application is: to provide an entity representation method, which includes: obtaining the label field, attribute field, and summary field of the target entity; using the entity representation model to obtain the representation of the target entity based on the label field, attribute field, and summary field of the target entity, where the entity representation model is trained by the foregoing method.

[0008] To solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, which includes a processor and a memory connected to the processor, where the memory stores program instructions; the processor is configured to execute the program instructions stored in the memory to implement the above method.

[0009] To solve the above technical problems, another technical solution adopted in this application is: to provide a storage medium storing program instructions, which can implement the above method when executed.

[0010] In the above manner, this application obtains the first representation of the first entity and the first representation of the second entity by using the entity representation model; based on the first representation of the first entity and the first representation of the second entity, obtains the association relationship between the first entity and the second entity; measures the representation effect of the entity representation model according to the difference between the obtained association relationship between the first entity and the second entity and the actual association relationship between the first entity and the second entity. Thus, it is possible to gradually optimize the entity representation model without supervision, so as to improve the availability of the representation of the entity obtained by the entity representation model. Secondly, the entity representation model obtained by the training method of this application is applicable to various entities (including sparse entities), and has high generality. In addition, since the training of the entity representation model in this application is carried out without supervision, it can reduce the human resources consumed in the training process. Description of the Drawings

[0011] Figure 1 is a schematic flowchart of the first embodiment of the training method of the entity representation model of this application;

[0012] Figure 2 is a schematic flowchart of the second embodiment of the training method of the entity representation model of this application;

[0013] Figure 3 is a schematic diagram of the page information of the entity "The Forbidden City" in this application;

[0014] Figure 4 is the address information corresponding to the page information of the documentary "The Forbidden City";

[0015] Figure 5 is the address information corresponding to the page information of the location of the Forbidden City in Beijing;

[0016] Figure 6 is Figure 2 The specific process schematic diagram of S113 in

[0017] Figure 7 is the process schematic diagram of the third embodiment of the training method of the entity representation model of the present application;

[0018] Figure 8 is the schematic diagram of page information where both the name fields "Ziwei Palace" and "Forbidden City" exist;

[0019] Figure 9 is the process schematic diagram of the fourth embodiment of the training method of the entity representation model of the present application;

[0020] Figure 10 is Figure 9 The specific process schematic diagram of S141 in

[0021] Figure 11 is Figure 9 The specific process schematic diagram of S142 in

[0022] Figure 12 is the process schematic diagram of the fourth embodiment of the training method of the entity representation model of the present application;

[0023] Figure 13 is Figure 12 The specific process schematic diagram of S153 in

[0024] Figure 14 is the process schematic diagram of the first embodiment of the entity representation method of the present application;

[0025] Figure 15 is Figure 14 The specific process schematic diagram of S21 in

[0026] Figure 16 is Figure 14 The specific process schematic diagram of S22 in

[0027] Figure 17 is Figure 16 The specific process schematic diagram of S221 in

[0028] Figure 18 is Figure 16 The specific process schematic diagram of S222 in

[0029] Figure 19 is the schematic diagram of the model structure relied on in the training stage of the present application;

[0030] Figure 20 is the schematic diagram of the entity representation model structure of the present application.

[0031] Figure 21 is the schematic diagram of the structure of the first embodiment of the electronic device of the present application;

[0032] Figure 22 It is a schematic structural diagram of an embodiment of the storage medium of the present application. Specific implementation manners

[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0034] The terms "first", "second", and "third" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0035] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that, without conflict, the embodiments described herein may be combined with other embodiments.

[0036] Figure 1 It is a schematic flowchart of the first embodiment of the training method of the entity representation model of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. As Figure 1 shown, this embodiment may include:

[0037] S11: Obtain an entity set.

[0038] Among them, the entity set includes a plurality of entities and the association relationships between different entities, and the association relationships between different entities are used to indicate whether the corresponding entities are related.

[0039] S12: Extract a first entity and a second entity from the entity set to form a training example.

[0040] Multiple first entities and second entity pairs can be extracted from the entity set to form multiple training examples. Among them, if the first entity and the second entity are related, the training example formed by the first entity and the second entity can be called a positive example. If the first entity and the second entity are not related, the training example formed by the first entity and the second entity can be called a negative example.

[0041] S13: Feed the training examples into the entity representation model.

[0042] S14: Use the entity representation model to obtain the first representation of the first entity and the first representation of the second entity respectively.

[0043] For the specific method of obtaining the representation, please refer to the following embodiments.

[0044] S15: Use the entity representation model to judge whether the first entity and the second entity are related based on the first representation of the first entity and the first representation of the second entity.

[0045] The output value (judgment result) of the entity representation model can be 0 or 1. 0 represents that the first entity and the second entity are not related, and 1 represents that the first entity and the second entity are related.

[0046] S16: Compare the judgment result with the association relationship between the first entity and the second entity, and adjust the parameters of the entity representation model according to the comparison result.

[0047] If the judgment result is that the first entity and the second entity are not related, while the association relationship between the first entity and the second entity indicates that the first entity and the second entity are related, the comparison result is incorrect, which means that the first representation extracted by the entity representation model cannot be used to correctly measure the relationship between the first entity and the second entity.

[0048] For each training example, its corresponding comparison result can be obtained. The parameters of the entity representation model can be adjusted based on the proportion of incorrect / correct comparison results corresponding to each training example.

[0049] Through the implementation of this embodiment, the present application obtains the first representation of the first entity and the first representation of the second entity using the entity representation model; based on the first representation of the first entity and the first representation of the second entity, the association relationship between the first entity and the second entity is obtained; the representation effect of the entity representation model is measured according to the difference between the obtained association relationship between the first entity and the second entity and the actual association relationship between the first entity and the second entity. Thus, it is possible to gradually optimize the entity representation model without supervision, so as to improve the availability of the representation of the entity obtained by the entity representation model. Secondly, the entity representation model obtained by the training method of the present application is applicable to various entities (including sparse entities), and has high generality. In addition, since the training of the entity representation model in the present application is carried out without supervision, it can reduce the human resources consumed in the training process.

[0050] Figure 2 It is a schematic flowchart of the second embodiment of the training method of the entity representation model of the present application. It should be noted that if there are substantially the same results, this embodiment does not Figure 2 be limited to the shown process sequence. This embodiment is a further expansion of S11. In this embodiment, the entity set may further include multiple explanatory fields of each entity. For example Figure 2 as shown, this embodiment may include:

[0051] S111: Obtain the second data records of each entity in the database.

[0052] The database can be various search engines. In the following, the present application takes Baidu Encyclopedia as an example for illustration. In the case where the database is Baidu Encyclopedia, the form of the second data record can be the page information of Baidu Encyclopedia, that is, the second data record can be the page information of the entity obtained by entering the name of the entity in the search box of Baidu Encyclopedia and performing subsequent searches. Figure 3 The middle is an example of the page information of the entity "The Forbidden City".

[0053] S112: Extract multiple explanatory fields corresponding to the entity from the second data records of each entity.

[0054] The multiple explanatory fields of the entity can be used to explain / describe the entity. Among them, the multiple explanatory fields may include a name field, multiple keyword fields, and an anchor link field. The multiple keyword fields may include a tag field, an attribute field, and a summary field.

[0055] For example, as Figure 3 shown, a contains a name field, b contains a summary field, c contains an attribute field, d contains a tag field, and e contains an anchor link field.

[0056] It should be noted that in combination with Figure 3It can be seen that there may be different senses for the same entity name. For example, "The Palace Museum" can either refer to a location (the Forbidden City in Beijing) or a documentary named "The Palace Museum". Its sense can be determined by the id corresponding to its data record. Figure 4 The data-lemmaId = "9044683" in the box shown is the id of the documentary "The Palace Museum". Figure 5 The data-lemmaId = "345415" in the box shown is the id of the Forbidden City in Beijing.

[0057] S113: Obtain the association relationship between the corresponding entity and other entities based on the anchor link field of each entity.

[0058] The anchor link field of each entity can be the name field of the entity related to the corresponding entity. Therefore, it can be considered that each entity is related to the entity pointed to by its anchor link field. For example, "The Palace Museum" is related to its anchor link field "Ziwei Palace". Therefore, the association relationship between "The Palace Museum" and "Ziwei Palace" can be set as related. Conversely, it is considered that each entity is not related to the entity other than the entity pointed to by its anchor link field.

[0059] It can be understood that since there are many anchor link fields for an entity, if it is considered that each entity pointed to by its anchor link field is related to it, then there may be too many related entity pairs (positive examples) in the entity set, which will affect the subsequent training of the entity model. Therefore, it can be done through sub-steps (S1131 - S1134) as shown Figure 6 to control the number of related entities in the entity set.

[0060] S1131: Obtain the third data record of the entity pointed to by the anchor link field in the database.

[0061] S1132: Obtain the first like count of the third data record of the entity pointed to by the anchor link field, and the second like count of the second data record of the corresponding entity.

[0062] Figure 3 The field (22615) included in f is an example of the second like count of the second data record of "The Palace Museum".

[0063] S1133: Determine whether the difference between the first like count and the second like count is less than a preset threshold.

[0064] If it is less than, then execute S1134.

[0065] S1134: Consider the entity pointed to by the anchor link field as related to the corresponding entity.

[0066] It can be understood that the difference between the first like count and the second like count can be used to measure the degree of correlation between each entity and the entity pointed to by its anchor link field. When the difference between the first like count and the second like count is smaller, it means that the degree of correlation between each entity and the entity pointed to by its anchor link field is higher; on the contrary, it means that the degree of correlation between each entity and the entity pointed to by its anchor link field is lower.

[0067] Figure 7 It is a schematic flowchart of the third embodiment of the training method of the entity representation model of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 7 the shown process sequence. This embodiment is a further expansion of S12. In this embodiment, the training examples can include multiple explanatory fields of the first entity and multiple explanatory fields of the second entity, as well as the co-occurrence text field of the first entity and the second entity. As Figure 7 shown, this embodiment can include the following steps:

[0068] S121: Extract multiple explanatory fields of the first entity and multiple explanatory fields of the second entity from the entity set.

[0069] The multiple explanatory fields include the name field.

[0070] S122: Obtain the first data record that simultaneously has the name field of the first entity and the name field of the second entity in the database.

[0071] The name fields of the first entity and the second entity can be simultaneously input into the search box of Baidu Encyclopedia to obtain the page information that simultaneously has the name field of the first entity and the name field of the second entity. Figure 8 It is an example of the page information that simultaneously has the name fields "Ziwei Palace" and "Forbidden City".

[0072] S123: Extract the co-occurrence text field of the first entity and the second entity from the first data record to form a training example.

[0073] For example, Figure 8 the co-occurrence text field of "Ziwei Palace" and "Forbidden City" in [example] is: Luoyang Ziwei City is the imperial palace of the Sui and Tang dynasties in Luoyang. The Forbidden City in Beijing is the Forbidden City of the Ming and Qing dynasties in Beijing.... Grand, imposing. Obviously, neither the Daming Palace nor the Ziwei Palace can compare with the Forbidden City. Simply put, it is small and exquisite. The Daming Palace and...

[0074] An example of the training example can be as shown in the following table. This training example is formed by the first entity "Forbidden City" and the second entity "Ziwei Palace". Since the first entity "Forbidden City" and the second entity "Ziwei Palace" are related, this training example is a positive example.

[0075]

[0076]

[0077] Figure 9 This is a schematic flowchart of the fourth embodiment of the training method for the entity representation model of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 9 the shown process sequence. This embodiment is a further expansion of S14. As Figure 9 shown, this embodiment may include:

[0078] S141: Use the entity representation model to obtain the vectors corresponding to each key field of the first entity respectively, and use the entity representation model to obtain the vectors corresponding to each key field of the second entity respectively.

[0079] With reference to Figure 10 , S141 may include the following sub-steps:

[0080] S1411: Use the entity representation model to obtain multiple sub-vectors corresponding to each key field of the first entity respectively, and use the entity representation model to obtain multiple sub-vectors corresponding to each key field of the second entity respectively.

[0081] The multiple key fields may include an attribute field, a label field, and an abstract field.

[0082] For example, use the entity representation model to obtain multiple 200-dimensional label sub-vectors corresponding to the label field and multiple 200-dimensional attribute sub-vectors corresponding to the attribute field.

[0083] S1412: Use the first attention mechanism of the entity representation model to process the multiple sub-vectors corresponding to the attribute field and the multiple sub-vectors corresponding to the label field respectively to obtain the vector corresponding to the attribute field and the vector corresponding to the label field, and use the second attention mechanism of the entity representation model to process the multiple sub-vectors corresponding to the abstract field to obtain the vector corresponding to the abstract field.

[0084] The first attention mechanism may be self-attention, and the second attention mechanism may be TransformerXL. In other embodiments, the second attention mechanism may also be bilstm, bert, etc.

[0085] It can be understood that since the label field and the attribute field are finite in number and can be enumerated, while the abstract field cannot be enumerated and is relatively long. Therefore, different attention mechanisms are used to obtain the vectors corresponding to the label field and the attribute field and the vector corresponding to the abstract field.

[0086] Among them, the first attention mechanism (self-attention) of the entity representation model can be used to process multiple sub-vectors corresponding to the label field to obtain the vector corresponding to the label field. The (self-attention) of the entity representation model can be used to process multiple sub-vectors corresponding to the attribute field to obtain the vector corresponding to the attribute field.

[0087] Taking the acquisition of the vector corresponding to the attribute field as an example for illustration. The way to process the n 200-dimensional label sub-vectors corresponding to the attribute field can be as follows:

[0088]

[0089]

[0090] Among them, p represents the label vector corresponding to the attribute field, p i represents the i-th attribute sub-vector, and r represents the weight distribution corresponding to the i-th attribute sub-vector.

[0091] S142: Use the entity representation model to obtain the first representation of the first entity based on the vectors corresponding to each key field of the first entity, and use the entity representation model to obtain the first representation of the second entity based on the vectors corresponding to each key field of the second entity.

[0092] The vectors corresponding to each key field of the first entity can be concatenated and then sent to a multi-layer perceptron (MLP) layer for processing to obtain the first representation E1 of the first entity. Similarly, the first representation E2 of the second entity is obtained.

[0093] With reference to Figure 11 , S142 may include the following sub-steps:

[0094] S1421: Use the entity representation model to concatenate the vectors corresponding to each key field of the first entity to obtain a first concatenation result, and use the entity representation model to concatenate the vectors corresponding to each key field of the second entity to obtain a second concatenation result.

[0095] S1422: Use the multi-layer perceptron of the entity representation model to process the first concatenation result to obtain the first representation of the first entity, and use the multi-layer perceptron of the entity representation model to process the second concatenation result to obtain the first representation of the second entity.

[0096] Figure 12 is a schematic flowchart of the fourth embodiment of the training method of the entity representation model of the present application. It should be noted that if there are substantially the same results, this embodiment does not take Figure 12 the shown process sequence as a limit. This embodiment is a further expansion of S15. In this embodiment, asFigure 12 As shown in the figure, this embodiment may include:

[0097] S151: Obtain the co-occurrence text vector corresponding to the co-occurrence text field by using the entity representation model.

[0098] There may be one or more co-occurrence text vectors corresponding to the co-occurrence text field.

[0099] S152: Process the first representation of the first entity and the co-occurrence text vector by using the first attention mechanism of the entity representation model to obtain the second representation of the first entity, and process the first representation of the second entity and the co-occurrence text vector by using the first attention mechanism of the entity representation model to obtain the second representation of the second entity.

[0100] The first representation E1 of the first entity and the co-occurrence text vector M can be processed by using the first attention mechanism of the entity representation model to obtain the second representation E1' of the first entity. The first representation E2 of the first entity and the co-occurrence text vector M can be processed by using the first attention mechanism of the entity representation model to obtain the second representation E2' of the first entity.

[0101] Taking the second representation E1' of the first entity as an example for illustration. The formula for obtaining the second representation E1' of the first entity can be as follows:

[0102]

[0103]

[0104] where M j is the j-th co-occurrence text vector, and α(M j , E1) is the relevance / contribution degree of M j and E1.

[0105] S153: Based on the first representation and the second representation of the first entity, and the first representation and the second representation of the second entity, determine whether the first entity and the second entity are relevant.

[0106] With reference to Figure 13 , S153 may include the following sub-steps:

[0107] S1531: Perform a first process on the first representation of the first entity and the first representation of the second entity to obtain a first processing result, and perform a second process on the second representation of the first entity and the second representation of the second entity to obtain a second processing result.

[0108] In a specific embodiment, one of the first process and the second process can be an addition process, and the other can be a multiplication process. In another specific embodiment, both the first process and the second process can be addition processes or both can be multiplication processes.

[0109] Taking the first process as a multiplication process and the second process as an addition process as an example for illustration. The first representation E1 of the first entity and the first representation E2 of the second entity can be multiplied correspondingly to obtain a first processing result; the second representation E1' of the first entity and the second representation E2' of the second entity can be added correspondingly to obtain a second processing result.

[0110] S1532: Based on the first processing result and the second processing result, determine whether the first entity and the second entity are relevant.

[0111] The first processing result and the second processing result can be connected to a fully connected layer, and then softmax classification is performed to obtain a judgment result 0 / 1.

[0112] Figure 14 It is a schematic flowchart of an embodiment of the entity representation method of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 14 the shown process sequence. In this embodiment, as Figure 14 shown, this embodiment may include:

[0113] S21: Obtain the label field, attribute field, and summary field of the target entity.

[0114] With reference to Figure 15 , S21 may include the following sub-steps:

[0115] S211: Obtain the data record of the target entity in the database.

[0116] S212: Extract the label field, attribute field, and summary field of the target entity from the data record of the target entity.

[0117] S22: Use the entity representation model to obtain the representation of the target entity based on the label field, attribute field, and summary field of the target entity.

[0118] With reference to Figure 16 , S22 may include the following sub-steps:

[0119] S221: Respectively use the entity representation model to obtain the vector corresponding to the label field, the vector corresponding to the attribute field, and the vector corresponding to the summary field.

[0120] Among them, the entity representation model can but is not limited to be obtained by the foregoing training method.

[0121] With reference toFigure 17 , S221 may include the following sub-steps:

[0122] S2211: respectively use the entity representation model to obtain multiple sub-vectors corresponding to the label field, multiple sub-vectors corresponding to the attribute field, and multiple sub-vectors corresponding to the summary field.

[0123] S2212: respectively use the first attention mechanism of the entity representation model to process the multiple sub-vectors corresponding to the label field and the multiple sub-vectors corresponding to the attribute field to obtain the vector corresponding to the label field and the vector corresponding to the attribute field, and use the second attention mechanism of the entity representation model to process the multiple sub-vectors corresponding to the summary field to obtain the vector corresponding to the summary field.

[0124] The first attention mechanism may be self-attention, and the second attention mechanism may be transformerXL. In other embodiments, the second attention mechanism may also be bilstm, bert, etc.

[0125] The first attention mechanism of the entity representation model can be used to process the multiple sub-vectors corresponding to the label field to obtain the vector corresponding to the label field; the first attention mechanism of the entity representation model can be used to process the multiple sub-vectors corresponding to the attribute field to obtain the vector corresponding to the attribute field. Transformer XL, bilstm, bert, etc. can be used to process the multiple sub-vectors corresponding to the summary field to obtain the vector corresponding to the summary field.

[0126] S222: use the entity representation model to obtain the representation of the target entity based on the vector corresponding to the label field, the vector corresponding to the attribute field, and the vector corresponding to the summary field.

[0127] Refer to Figure 18 , S222 may include the following sub-steps:

[0128] S2221: use the entity representation model to splice the vector corresponding to the label field, the vector corresponding to the attribute field, and the vector corresponding to the summary field to obtain the third splicing result.

[0129] S2222: use the multi-layer perceptron of the entity representation model to process the third splicing result to obtain the representation of the target entity.

[0130] For the detailed description of the steps in this embodiment, please refer to the previous embodiments and will not be repeated here.

[0131] Next, in combination with Figure 19 , the training method of the entity representation model of the present application will be described in the form of Example 1.

[0132] AsFigure 19 As shown, there is currently a training example formed by entity 1 and entity 2. The entity representation model is used to obtain the co-occurrence text vectors M corresponding to the entry tags (tag fields), attribute names (attribute fields), descriptions (abstract fields), and positive example co-occurrence texts (co-occurrence text fields) of entity 1 and entity 2 respectively.

[0133] The self-att (first attention mechanism) of the entity representation model is used to process multiple sub-vectors corresponding to the entry tags and attribute names respectively, to obtain vectors corresponding to the entry tags and attribute names; the transformer XL (second attention mechanism) of the entity representation model is used to process multiple sub-vectors corresponding to the description, to obtain a vector corresponding to the description.

[0134] The vectors corresponding to the entry tags, attribute names, and descriptions of entity 1 are obtained as E1 after concat and MLP of the entity representation model. The vectors corresponding to the entry tags, attribute names, and descriptions of entity 2 are obtained as E2 after concat and MLP.

[0135] The entity representation model is used to perform attention (weighted processing) on E1 and M to obtain E1'. The entity representation model is used to perform attention (weighted processing) on E2 and M to obtain E2'.

[0136] Perform '*' processing on E1 and E2 to obtain a first processing result, perform '+' processing on E1' and E2' to obtain a second processing result; based on the first processing result and the second processing result, determine whether entity 1 and entity 2 are relevant to obtain a judgment result; based on the judgment result and the actual association relationship between entity 1 and entity 2, adjust the parameters of the entity correlation model.

[0137] The following combines Figure 20 , and takes the form of Example 2 to illustrate the entity representation method of the entity representation model obtained by the above training method.

[0138] Input entity 1 into the entity representation model, and respectively use the entity representation model to obtain multiple sub-vectors corresponding to the entry tags, attribute names, and descriptions of entity 1; respectively use the self-att of the entity representation model to process multiple sub-vectors corresponding to the entry tags and attribute names to obtain vectors corresponding to the entry tags and attribute names, and use the transformXL of the entity representation model to process multiple sub-vectors corresponding to the description to obtain a vector corresponding to the description; after using the entity representation model to perform concat and MLP on multiple sub-vectors corresponding to the entry tags, attribute names, and descriptions, the representation of entity 1 can be obtained.

[0139] Figure 21 It is a schematic structural diagram of an embodiment of an electronic device in the present application. As Figure 21As shown, the electronic device includes a processor 31 and a memory 32 coupled to the processor 31.

[0140] Among them, the memory 32 stores program instructions for implementing the method of any of the above embodiments; the processor 31 is configured to execute the program instructions stored in the memory 32 to implement the steps of the above method embodiments. Among them, the processor 31 may also be referred to as a CPU (Central Processing Unit). The processor 31 may be an integrated circuit chip with signal processing capabilities. The processor 31 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0141] Figure 22 It is a schematic structural diagram of an embodiment of the storage medium of the present application. As Figure 22 shown, the computer-readable storage medium 40 of the embodiment of the present application stores program instructions 41, and when the program instructions 41 are executed, the method provided in the above embodiments of the present application is implemented. Among them, the program instructions 41 may form a program file and be stored in the above computer-readable storage medium 40 in the form of a software product, so that a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods of various embodiments of the present application. And the aforementioned computer-readable storage medium 40 includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes, or a computer, a server, a mobile phone, a tablet, or other terminal devices.

[0142] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces, and the indirect coupling or communication connection of devices or units may be in an electrical, mechanical, or other form.

[0143] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. The above is only the implementation manner of the present application, and does not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A natural language processing method, characterized in that, Including: Extracting the label field, attribute field, and summary field of the target text entity from the data record of the target text entity; Using the entity representation model to obtain the representation of the target text entity based on the label field, attribute field, and summary field of the target text entity; Among them, the training steps of the entity representation model include: Obtaining a text entity set, where the text entity set includes multiple text entities and the association relationships between different text entities, and the association relationships between different text entities are used to indicate whether the corresponding text entities are relevant; Extracting a first text entity and a second text entity from the text entity set to form a training example; where the training example includes multiple key fields of the first text entity and multiple key fields of the second text entity, and the multiple key fields include an attribute field, a label field, and a summary field; Feeding the training example into the entity representation model; Respectively using the entity representation model to obtain multiple sub-vectors corresponding to each key field of the first text entity, and respectively using the entity representation model to obtain multiple sub-vectors corresponding to each key field of the second text entity; Respectively using the first attention mechanism of the entity representation model to process the multiple sub-vectors corresponding to the attribute field and the multiple sub-vectors corresponding to the label field to obtain the vector corresponding to the attribute field and the vector corresponding to the label field, and using the second attention mechanism of the entity representation model to process the multiple sub-vectors corresponding to the summary field to obtain the vector corresponding to the summary field; Using the entity representation model to obtain the first representation of the first text entity based on the vector corresponding to each key field of the first text entity, and using the entity representation model to obtain the first representation of the second text entity based on the vector corresponding to each key field of the second text entity; Based on the first representation of the first text entity and the first representation of the second text entity, determining whether the first text entity and the second text entity are relevant; Comparing the judgment result with the association relationship between the first text entity and the second text entity, and adjusting the parameters of the entity representation model according to the comparison result.

2. The method according to claim 1, characterized in that, The training example includes multiple interpretation fields of the first text entity and multiple interpretation fields of the second text entity, and the multiple interpretation fields include the multiple key fields.

3. The method according to claim 2, characterized in that, The using the entity representation model to obtain the first representation of the first text entity based on the vector corresponding to each key field of the first text entity, and using the entity representation model to obtain the first representation of the second text entity based on the vector corresponding to each key field of the second text entity includes: Using the entity representation model to splice the vectors corresponding to each key field of the first text entity to obtain a first splicing result, and using the entity representation model to splice the vectors corresponding to each key field of the second text entity to obtain a second splicing result; Process the first splicing result using the multi-layer perceptron of the entity representation model to obtain the first representation of the first text entity, and process the second splicing result using the multi-layer perceptron of the entity representation model to obtain the first representation of the second text entity.

4. The method according to claim 2, characterized in that, The training example further includes the co-occurrence text fields of the first text entity and the second text entity. The extracting the first text entity and the second text entity from the text entity set to form a training example includes: Extract multiple explanation fields of the first text entity and multiple explanation fields of the second text entity from the text entity set, where the multiple explanation fields include a name field; Obtain a first data record in the database that simultaneously has the name field of the first text entity and the name field of the second text entity; Extract the co-occurrence text fields of the first text entity and the second text entity from the first data record to form the training example.

5. The method according to claim 4, characterized in that, The determining whether the first text entity and the second text entity are relevant based on the first representation of the first text entity and the first representation of the second text entity includes: Use the entity representation model to obtain the co-occurrence text vector corresponding to the co-occurrence text field; Process the first representation of the first text entity and the co-occurrence text vector using the first attention mechanism of the entity representation model to obtain the second representation of the first text entity, and process the first representation of the second text entity and the co-occurrence text vector using the first attention mechanism of the entity representation model to obtain the second representation of the second text entity; Determine whether the first text entity and the second text entity are relevant based on the first representation and the second representation of the first text entity, and the first representation and the second representation of the second text entity.

6. The method according to claim 5, characterized in that, The determining whether the first text entity and the second text entity are relevant based on the first representation and the second representation of the first text entity, and the first representation and the second representation of the second text entity includes: Perform a first process on the first representation of the first text entity and the first representation of the second text entity to obtain a first processing result, and perform a second process on the second representation of the first text entity and the second representation of the second text entity to obtain a second processing result; Determine whether the first text entity and the second text entity are relevant based on the first processing result and the second processing result; Wherein, one of the first process and the second process is an addition process, and the other is a multiplication process.

7. The method according to claim 1 or 5, characterized in that, The first attention mechanism is self-attention, and the second attention mechanism is transformer XL.

8. The method according to claim 1, characterized in that, The text entity set further includes multiple explanation fields of each text entity. The obtaining the text entity set includes: Obtain second data records of each text entity in the database; Extract multiple explanation fields corresponding to each text entity from the second data records of each text entity, where the multiple explanation fields include an anchor link field; Obtain the association relationship between the text entity and other text entities corresponding to the anchor link field of each of the text entities.

9. The method according to claim 8, wherein, The obtaining of the association relationship between the text entity and other text entities corresponding to the anchor link field of each of the text entities includes: Each of the text entities is related to the text entity pointed to by the corresponding anchor link field; Or, the obtaining of the association relationship between the text entity and other text entities corresponding to the anchor link field of each of the text entities includes: Obtain the third data record of the text entity pointed to by the anchor link field in the database; Obtain the first like count of the third data record of the text entity pointed to by the anchor link field, and the second like count of the second data record of the corresponding text entity; If the difference between the first like count and the second like count is less than a preset threshold, it is considered that the text entity pointed to by the anchor link field is related to the corresponding text entity.

10. The method according to claim 1, wherein, Using the entity representation model to obtain the representation of the target text entity based on the tag field, attribute field, and summary field of the target text entity, includes: Respectively use the entity representation model to obtain the vector corresponding to the tag field, the vector corresponding to the attribute field, and the vector corresponding to the summary field; Using the entity representation model to obtain the representation of the target text entity based on the vector corresponding to the tag field, the vector corresponding to the attribute field, and the vector corresponding to the summary field.

11. The method according to claim 10, wherein, The using the entity representation model to obtain the representation of the target text entity based on the vector corresponding to the tag field, the vector corresponding to the attribute field, and the vector corresponding to the summary field includes: Use the entity representation model to splice the vector corresponding to the tag field, the vector corresponding to the attribute field, and the vector corresponding to the summary field to obtain a third splicing result; Use the multi-layer perceptron of the entity representation model to process the third splicing result to obtain the representation of the target text entity.

12. The method according to claim 10, wherein, The using the entity representation model to obtain the vector corresponding to the tag field, the vector corresponding to the attribute field, and the vector corresponding to the summary field includes: Respectively use the entity representation model to obtain multiple sub-vectors corresponding to the tag field, multiple sub-vectors corresponding to the attribute field, and multiple sub-vectors corresponding to the summary field; Respectively use the first attention mechanism of the entity representation model to process the multiple sub-vectors corresponding to the tag field and the multiple sub-vectors corresponding to the attribute field to obtain the vector corresponding to the tag field and the vector corresponding to the attribute field, and use the second attention mechanism of the entity representation model to process the multiple sub-vectors corresponding to the summary field to obtain the vector corresponding to the summary field.

13. The method according to claim 12, wherein, The first attention mechanism is self-attention, and the second attention mechanism is transformer XL.

14. The method according to claim 1, wherein, The extracting the tag field, attribute field, and summary field of the target text entity from the data record of the target text entity includes: Obtain the data record of the target text entity in the database; Extract the label field, attribute field, and summary field of the target text entity from the data record of the target text entity.

15. An electronic device, wherein, Comprising a processor and a memory connected to the processor, wherein, The memory stores program instructions; The processor is configured to execute the program instructions stored in the memory to implement the method according to any one of claims 1-14.

16. A storage medium, wherein, The storage medium stores program instructions, which when executed implement the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • A heterogeneous social network user entity anchor link identification method

    CN109949174A

  • Cross-language knowledge graph entity alignment method based on GCN twinning network

    CN110472065A