Training Method, Device, Electronic Device, and Storage Medium for Entity Representation Model
By splitting the sample statements into word sequences according to character granularity, and matching them with reference entities in the knowledge graph, we determine the sample entities for training entity representation, and the problem of low training accuracy and completeness of the entity representation model caused by incomplete knowledge graph is solved, and higher training accuracy and semantic expression adequacy are achieved.
Patent Information
- Application Number
- CN202210161016.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-02-22
AI Technical Summary
In the case of incomplete knowledge graphs in the prior art, it is difficult to effectively characterize independent entities outside the field, resulting in low training accuracy and completeness of entity representation models.
The sample entity is determined by splitting the sample statement into word sequences according to the character granularity and matching it with reference entities in the preset knowledge graph, so as to train entity representation.
The training accuracy of the entity representation model is improved, the semantic expression sufficiency of entity representation is ensured, and errors caused by training of independent entities outside the domain are avoided.
Smart Images

Figure CN114519396B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method, device, electronic device, and storage medium for training an entity representation model. Background Art
[0002] With the development of artificial intelligence technology, the application of natural language processing (NLP) is becoming more and more extensive, involving many entity-related tasks, such as named entity recognition, relation classification, question answering systems, etc. The key to solving these problems lies in effectively representing the input statements. The common practice in the industry is to represent an entity and its own semantic information with a vector of a fixed dimension. The richer the information covered by the vector about the entity, the more conducive it is to the development of subsequent tasks. In the prior art, an entity representation model usually combines a domain knowledge graph and a graph neural network to generate entity representations. Before prediction, the sample statement needs to be split into multiple independent entities according to entity granularity, and multiple adjacent independent entities are used as training data for model training. However, in the case where the domain entities in the knowledge graph are incomplete, it is very likely that an independent entity is not within the domain of the knowledge graph. During the training process, it is difficult to effectively represent an independent entity outside the domain, and the training of the entity representation model is likely to fail, resulting in low accuracy and integrity of the entity representation model. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0004] Embodiments of the present invention provide a method, device, electronic device, and storage medium for training an entity representation model, which can combine character-level word sequences and sample entities in a knowledge graph to train entity representations and improve the training accuracy of the entity representation model.
[0005] In a first aspect, an embodiment of the present invention provides a method for training an entity representation model, including:
[0006] Obtain a sample statement, and split the sample statement into multiple sample words according to character granularity to obtain a word sequence composed of multiple sample words;
[0007] Obtain a preset knowledge graph, in which multiple reference entities are preset, and each reference entity is labeled with reference information;
[0008] Match the word sequence with the reference information, and determine at least one sample entity from the reference entities;
[0009] A sample sequence is obtained by splicing the sample entity and the word sequence, and the sample sequence is input into an entity representation model for entity representation training.
[0010] In some embodiments, before obtaining the preset knowledge graph, the method further includes:
[0011] Configure a plurality of the reference entities in the knowledge graph;
[0012] Annotate the reference information for the reference entities according to a preset data set.
[0013] In some embodiments, the determining at least one sample entity from the reference entities according to the word sequence and the reference information includes:
[0014] Continuously select at least two of the sample words from the word sequence to obtain a sample phrase;
[0015] Match at least one of the sample entities from the reference entities according to the sample phrase and the reference information.
[0016] In some embodiments, the entity representation model includes a RoBERTa model, and the inputting the sample sequence into the entity representation model for entity representation training includes:
[0017] Perform semantic encoding on the sample sequence through the RoBERTa model to obtain a first token corresponding to the sample word and a second token corresponding to the sample entity;
[0018] Perform entity representation training on the sample sequence according to the first token, the second token, to obtain a semantic representation vector of the sample entity.
[0019] In some embodiments, the entity representation model further includes a Transformer model, and the performing entity representation training on the sample sequence according to the first token, the second token, to obtain a semantic representation vector of the sample entity includes:
[0020] Input the first token, the second token and the sample sequence into the Transformer model;
[0021] Determine a first attention matrix through the Transformer model, and the first attention matrix represents the attention relationship between a plurality of the first tokens;
[0022] Determine a second attention matrix through the Transformer model, where the second attention matrix represents the attention relationship between the second token and the first token;
[0023] Obtain a first feature vector according to the word sequence and the first attention matrix, and obtain a second feature vector according to the sample entity and the second attention matrix;
[0024] Obtain the semantic representation vector according to the first feature vector and the second feature vector.
[0025] In some embodiments, the determining the second attention matrix through the Transformer model includes:
[0026] Obtain the start position embedding information and the end position embedding information corresponding to the sample entity, where the start position embedding information is the position embedding information in the sample word with the earliest order corresponding to the sample entity, and the end position embedding information is the position embedding information in the sample word with the latest order corresponding to the sample entity;
[0027] Determine the target position embedding information of the sample entity according to the start position embedding information and the end position embedding information;
[0028] Determine the second attention matrix according to the target position embedding information, the first token, and the second token.
[0029] In some embodiments, the obtaining the semantic representation vector according to the first feature vector and the second feature vector includes:
[0030] Obtain a preset loss weight;
[0031] Perform loss calculations on the first feature vector and the second feature vector respectively according to the loss weight;
[0032] Merge the feature vectors obtained by the loss calculation to obtain the semantic representation vector.
[0033] In a second aspect, an embodiment of the present invention provides a training device for an entity representation model, including:
[0034] A word sequence acquisition unit, configured to acquire a sample statement, split the sample statement into a plurality of sample words according to the character granularity, and obtain a word sequence composed of the plurality of sample words;
[0035] A knowledge graph acquisition unit, configured to acquire a preset knowledge graph, where a plurality of reference entities are preset in the knowledge graph, and each reference entity is labeled with reference information;
[0036] An entity acquisition unit, configured to match the word sequence with the reference information, and determine at least one sample entity from the reference entities;
[0037] A training unit, configured to obtain a sample sequence by splicing the sample entity and the word sequence, and input the sample sequence into an entity representation model for training entity representation.
[0038] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the training method of the entity representation model as described in the first aspect is implemented.
[0039] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program for executing the training method of the entity representation model as described in the first aspect.
[0040] Embodiments of the present invention include: obtaining a sample statement, splitting the sample statement into a plurality of sample words according to character granularity to obtain a word sequence composed of a plurality of the sample words; obtaining a preset knowledge graph, where a plurality of reference entities are preset in the knowledge graph, and each reference entity is labeled with reference information; matching the word sequence with the reference information to determine at least one sample entity from the reference entities; obtaining a sample sequence by splicing the sample entity and the word sequence, and inputting the sample sequence into an entity representation model for training entity representation. According to the technical solution of this embodiment, the sample statement is split into a word sequence composed of sample words according to character granularity, and sample entities in the field are obtained from the knowledge graph according to the word sequence, ensuring the sufficiency of semantic expression in entity representation training. For independent entities outside the knowledge graph field, indirect representation is performed through sample words, effectively avoiding errors caused by training with independent entities outside the field and improving the accuracy of entity representation model training.
[0041] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings are used to provide a further understanding of the technical solutions of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation to the technical solutions of the present invention.
[0043] Figure 1It is a flowchart of a method for training an entity representation model provided by an embodiment of the present invention;
[0044] Figure 2 It is a flowchart of annotating a knowledge graph provided by another embodiment of the present invention;
[0045] Figure 3 It is a flowchart of selecting sample entities provided by another embodiment of the present invention;
[0046] Figure 4 It is a flowchart of calculating tokens provided by another embodiment of the present invention;
[0047] Figure 5 It is a flowchart of entity representation training provided by another embodiment of the present invention;
[0048] Figure 6 It is a flowchart of obtaining target position embedding information provided by another embodiment of the present invention;
[0049] Figure 7 It is a flowchart of loss calculation provided by another embodiment of the present invention;
[0050] Figure 8 It is a structural diagram of a training device for an entity representation model provided by another embodiment of the present invention;
[0051] Figure 9 It is a device diagram of an electronic device provided by another embodiment of the present invention. Detailed implementation manners
[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description, claims or the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0054] The present invention provides a method, an apparatus, an electronic device, and a storage medium for training an entity representation model. The method includes: obtaining a sample statement, splitting the sample statement into a plurality of sample words according to the character granularity to obtain a word sequence composed of the plurality of sample words; obtaining a preset knowledge graph, in which a plurality of reference entities are preset, and each reference entity is labeled with reference information; matching the word sequence with the reference information to determine at least one sample entity from the reference entities; obtaining a sample sequence by splicing the sample entity and the word sequence, and inputting the sample sequence into the entity representation model for training of entity representation. According to the technical solution of this embodiment, the sample statement is split into a word sequence composed of sample words according to the character granularity, and the sample entity in the domain is obtained from the knowledge graph according to the word sequence, ensuring the sufficiency of semantic expression in entity representation training. For independent entities outside the knowledge graph domain, indirect representation is performed through sample words, effectively avoiding errors caused by training with independent entities outside the domain and improving the accuracy of entity representation model training.
[0055] Embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application devices.
[0056] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction devices, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0057] The terminal mentioned in the embodiments of the present invention may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a vehicle-mounted computer, a smart home, a wearable electronic device, a VR (Virtual Reality) / AR (Augmented Reality) device, etc.; the server may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms, etc.
[0058] It should be noted that the data of the embodiments of the present invention can be stored in a server. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
[0059] Natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language commonly used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0060] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specializes in studying how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0061] As Figure 1 shown, Figure 1 is a flowchart of a method for training an entity representation model provided by an embodiment of the present invention. The method for training the entity representation model includes but is not limited to the following steps:
[0062] Step S110, obtain a sample sentence, and split the sample sentence into multiple sample words according to the character granularity to obtain a word sequence composed of multiple sample words;
[0063] Step S120, obtain a preset knowledge graph. There are multiple reference entities preset in the knowledge graph, and each reference entity is labeled with reference information;
[0064] Step S130, match the word sequence with the reference information to determine at least one sample entity from the reference entities;
[0065] Step S140, obtain a sample sequence by splicing the sample entity and the word sequence, and input the sample sequence into the entity representation model for entity representation training.
[0066] It should be noted that after obtaining the sample sentence, multiple sample words are obtained by splitting at the character granularity. Using the character-level words as training data, since the splitting granularity is smaller than the independent entity granularity, there is no need to associate independent entities in the knowledge graph during the training process. Instead, the indirect representation of entities is achieved through character-level words, which can effectively avoid training errors caused by entities not being within the scope of the knowledge graph. For example, if the input sample sentence is "Metformin is used to treat diabetes", then at the character granularity, each character is used as a sample word.
[0067] It should be noted that the reference entities of the knowledge graph can be set according to different domains. For example, for the knowledge graph in the medical field, the reference entities can be disease names or drug names, and the reference information can be the abbreviations, scientific names, or definitions of diseases or drugs, as long as the annotation of the reference entities can be achieved. This embodiment does not make further limitations on this.
[0068] It should be noted that since the word sequence is composed of multiple sample words, it is possible to match each sample word with the reference information one by one, or to continuously select multiple sample words to form a phrase for query. For example, in the word sequence "Er, Jia, Shuang, Gua, Yong, Yu, Zhi, Liao, Tang, Niao, Bing" (Two, A, Double, Gua, Used, For, Treat, Sugar, Urine, Disease), it is possible to query the sample entity through "Gua" (Gua), or to query through "Erjia Shuanggua" (Metformin). Since "Erjia Shuanggua" (Metformin) is a drug, the corresponding sample entity can be found in the knowledge graph of the medical field.
[0069] It should be noted that the sample entity can be directly concatenated after the word sequence. For example, for the word sequence "Er, Jia, Shuang, Gua, Yong, Yu, Zhi, Liao, Tang, Niao, Bing" (Two, A, Double, Gua, Used, For, Treat, Sugar, Urine, Disease), the found sample entity "Erjia Shuanggua" (Metformin), then the resulting sample sequence after concatenation is "Er, Jia, Shuang, Gua, Yong, Yu, Zhi, Liao, Tang, Niao, Bing, Erjia Shuanggua" (Two, A, Double, Gua, Used, For, Treat, Sugar, Urine, Disease, Metformin). Since the input sentence is split into character-level sample words, in this embodiment, based on the word sequence, the sample entity is obtained from the knowledge graph, which can ensure that the sample entity in the sample sequence belongs to the scope of the knowledge graph. And, the entities in the sample sentence can be directly subjected to semantic modeling, realizing the sufficiency of entity semantic expression. In addition, compared with indirectly representing entities by performing character-level fine-grained splitting through independent entities, it avoids the transformation of the semantic space, reduces the accumulation of errors, and maintains the semantic integrity of the entity.
[0070] In addition, with reference to Figure 2 , in one embodiment, before performing Figure 1 the step S120 of the embodiment shown, it further includes but is not limited to the following steps:
[0071] Step S210, configuring multiple reference entities in the knowledge graph;
[0072] Step S220, annotate reference information for the reference entity according to a preset dataset.
[0073] It should be noted that the reference entities of the knowledge graph can be set according to specific fields and actual needs before training. For example, in the medical field, reference entities can be established based on all drug names or disease names, as long as the comprehensiveness of the knowledge graph can be ensured.
[0074] It should be noted that to ensure the comprehensiveness of annotation, the reference information of the reference entity can be annotated through a preset dataset, such as wikipedia. Of course, other datasets can also be used. By selecting a large number of basic datasets for model pre-training, the model can learn more rich domain knowledge and has better generalization ability. No more limitations are made here.
[0075] In addition, referring to Figure 3 , in one embodiment, Figure 1 Step S130 of the illustrated embodiment further includes but is not limited to the following steps:
[0076] Step S310, continuously select at least two sample words from the word sequence to obtain a sample phrase;
[0077] Step S320, according to the sample phrase and the reference information, match at least one sample entity from the reference entities.
[0078] It should be noted that since the sample words in the word sequence are split at the character level, for Chinese usage scenarios, the meaning represented by only one character is relatively rich, and it is difficult to accurately match the corresponding sample entity. For example, in the above word sequence, if only "sugar" is used to search for sample entities, there are very many concepts related to the word "sugar", such as "glucose" in medicine and "diabetes" in disease, which are completely different concepts. If both are used, it will have a greater impact on the accuracy of entity representation. To improve the accuracy and efficiency of sample entity query, at least two sample words can be selected to form a sample phrase for query.
[0079] It should be noted that since sample sentences are usually coherent sentences, for example, the four sample words "two", "methyl", "double", and "guanidine" are all split from the phrase "metformin", so when selecting sample words to form a sample phrase, multiple sample words can be continuously selected to form a phrase with a specific meaning for query, which can effectively improve the accuracy of query. The specific number of continuously selected sample words can be adjusted according to actual needs, and in the case of matching failure, the number of sample words can be further increased or decreased until successful matching.
[0080] In addition, in one embodiment, the entity representation model includes the RoBERTa model, referring to Figure 4 , Figure 1Step S140 of the illustrated embodiment further includes, but is not limited to, the following steps:
[0081] Step S410, semantically encoding the sample sequence through the RoBERTa model to obtain a first token corresponding to the sample word and a second token corresponding to the sample entity;
[0082] Step S420, training the sample sequence for entity representation according to the first token and the second token to obtain a semantic representation vector of the sample entity.
[0083] It should be noted that the RoBERTa model has the characteristic of generality. After pre-training, it has good performance in natural language processing, especially good semantic representation ability. Therefore, in this embodiment, the RoBERTa model is used as the pre-training model of the entity representation model to semantically encode the input sample sequence, and the accuracy of the trained semantic representation vector can be higher. The specific configuration and training process of the RoBERTa model are well-known technologies to those skilled in the art, and will not be elaborated here for the sake of simplicity.
[0084] It should be noted that after the sample sequence is input into the RoBERTa model, the sample word and the sample entity can be semantically encoded to obtain their respective tokens. Since the sample entity is expanded on the original vocabulary basis, when pre-training through the RoBERTa model, the first token and the second token can be treated as a single token, and for data diversity, random token replacement can be performed. The sample entity can be replaced with the sample word, or the sample word can be replaced with the sample entity. Those skilled in the art know how to perform token replacement in the RoBERTa model, and the specific operations will not be elaborated here.
[0085] In addition, in one embodiment, the entity representation model further includes a Transformer model. Refer to Figure 5 , Figure 4 Step S420 of the illustrated embodiment further includes, but is not limited to, the following steps:
[0086] Step S510, inputting the first token, the second token, and the sample sequence into the Transformer model;
[0087] Step S520, determining a first attention matrix through the Transformer model, where the first attention matrix represents the attention relationship between multiple first tokens;
[0088] Step S530: Determine the second attention matrix through the Transformer model. The second attention matrix represents the attention relationship between the second token and the first token.
[0089] Step S540: Obtain the first feature vector based on the word sequence and the first attention matrix, and obtain the second feature vector based on the sample entity and the second attention matrix.
[0090] Step S550: Obtain the semantic representation vector based on the first feature vector and the second feature vector.
[0091] It should be noted that the Transformer model is widely used in the field of natural language processing, such as machine translation, question answering systems, text summarization, and speech recognition, etc. The Transformer model based on the self-attention mechanism is currently the most advanced neural network architecture, including an encoder and a decoder. The encoder is responsible for extracting the feature information of the text, extracting a feature vector for each word in the text, so as to obtain the feature vector of the entire text. The decoder is responsible for using the feature vector extracted by the encoder to generate keywords that conform to the feature information as the output. On this basis, after the first token, the second token, and the sample sequence are input into the Transformer model, calculate the first attention matrix between the first tokens, and the second attention matrix between the first token and the second token. The first attention matrix represents the attention relationship between the sample words, and the second attention matrix represents the attention relationship between the sample words and the sample entity, which can effectively improve the accuracy of entity representation in the subsequent training process.
[0092] It should be noted that for the Transformer model based on the self-attention mechanism, after obtaining the attention matrix, the feature vector can be obtained through the attention matrix and the input information. In this embodiment, through the first attention matrix and the sample words, the first feature vector of the word sequence can be extracted. Similarly, through the second attention matrix and the sample entity, the second feature vector of the sample entity can be extracted. In the case of having the attention matrix, those skilled in the art are familiar with how to perform feature extraction, and will not elaborate here.
[0093] It should be noted that after the first feature vector and the second feature vector, simple feature fusion can be performed to obtain the semantic representation vector, or it can be obtained through inference by the inference layer. This embodiment does not make specific limitations on this.
[0094] In addition, referring to Figure 6 , in an embodiment, Figure 5 Step S530 of the embodiment shown further includes, but is not limited to, the following steps:
[0095] Step S610: Obtain the start position embedding information and end position embedding information corresponding to the sample entity. The start position embedding information is the position embedding information in the sample word with the earliest order corresponding to the sample entity, and the end position embedding information is the position embedding information in the sample word with the latest order corresponding to the sample entity.
[0096] Step S620: Determine the target position embedding information of the sample entity according to the start position embedding information and the end position embedding information.
[0097] Step S630: Determine the second attention matrix according to the target position embedding information, the first token, and the second token.
[0098] It should be noted that the embedding position information is important information for each sample data in the semantic recognition process. As the position identifier of each sample data in the sentence, it is usually associated with the sample data in the form of a flag vector. The sample words are obtained by splitting the sample sentence. Therefore, the embedding position information of each sample word is known. For example, in the word sequence "two, metformin, used, for, treating, diabetes", the embedding position information of each sample word can be represented according to its order in the word sequence, such as C2 to C12 in sequence, where C1 is the start identifier of the word sequence and C13 is the end identifier of the sequence. The sample entity is not the content in the word sequence but is obtained by matching the sample phrase. Therefore, in order to establish the association relationship between the sample entity and the sample phrase, so that the attention between the sample entity and the sample phrase can be enhanced in the subsequent training process, thereby improving the accuracy of entity representation, the target position embedding information can be obtained by calculating the mean sum according to the boundary characters of the sample entity. For example, in the above word sequence, the matched sample entity is "metformin", and its corresponding boundary characters are "two" and "guanidine", the start position embedding information is C2, and the end position embedding information is C6. Then the target position embedding information of the sample entity "metformin" is (C2 + C6) / 2.
[0099] It should be noted that the target position embedding information can reflect the attention relationship between the first token and the second token. For example, if the target position embedding information of the sample entity "metformin" is (C2 + C6) / 2, then the embedding position information of its associated sample words can be determined as C2 to C6, and the second attention matrix is calculated according to the first token corresponding to the sample words from C2 to C6 and the second token corresponding to the sample entity.
[0100] It should be noted that after determining the information embedded in the target position, it can be associated with the sample entity. When the sample entity is input into the entity representation model to calculate the attention, the model can determine the associated sample words according to the information embedded in the target position, so that the training of the entity representation not only focuses on the semantic expression of a certain part, but can more comprehensively express the semantic relationship between entities from the semantic level in combination with the sample entity.
[0101] In addition, in one embodiment, with reference to Figure 7 , Figure 5 the step S550 of the illustrated embodiment further includes but is not limited to the following steps:
[0102] Step S710, obtaining a preset loss weight;
[0103] Step S720, calculating the losses of the first feature vector and the second feature vector respectively according to the loss weight;
[0104] Step S730, merging the feature vectors obtained by the loss calculation to obtain a semantic representation vector.
[0105] It should be noted that in order to realize the training of the entity representation, a common inference layer can be set in the entity representation model, and inferences are made based on the first feature vector and the second feature vector as the data basis. The calculation of the loss function is a common inference method. In this embodiment, the losses of the first feature vector and the second feature vector can be calculated with the same loss weight, and the specific loss function can be selected according to actual needs.
[0106] It should be noted that through the loss calculation, the final prediction result can be obtained according to the first feature vector and the second feature vector. For example, the feature words obtained after calculating the loss of the first feature vector are "guanidine" and "sugar", and the feature words obtained after calculating the loss of the second feature vector are "diabetes". The semantics corresponding to the obtained semantic representation vector are "metformin" and "diabetes", which can be used for operations such as entity recognition and entity linking in downstream related tasks.
[0107] In addition, with reference to Figure 8 , the embodiment of the present invention provides a training device for an entity representation model. The training device 800 for the entity representation model includes:
[0108] A word sequence acquisition unit 810, configured to acquire a sample statement, split the sample statement into a plurality of sample words according to the character granularity, and obtain a word sequence composed of a plurality of sample words;
[0109] A knowledge graph acquisition unit 820, configured to acquire a preset knowledge graph, where a plurality of reference entities are preset in the knowledge graph, and each reference entity is labeled with reference information;
[0110] An entity acquisition unit 830, configured to match a word sequence with reference information to determine at least one sample entity from reference entities;
[0111] A training unit 840, configured to obtain a sample sequence by concatenating a sample entity and a word sequence, and input the sample sequence into an entity representation model for training of entity representation.
[0112] In addition, referring to Figure 9 , an embodiment of the present invention further provides an electronic device 900, which includes: a memory 910, a processor 920, and a computer program stored on the memory 910 and executable on the processor 920.
[0113] The processor 920 and the memory 910 may be connected through a bus or other means.
[0114] The non-transitory software program and instructions required to implement the training method of the entity representation model in the above embodiment are stored in the memory 910. When executed by the processor 920, the training method of the entity representation model applied to the device in the above embodiment is executed. For example, the method steps S110 to S140 described above are executed, Figure 1 the method steps S210 to S220 in Figure 2 the method steps S310 to S320 in Figure 3 the method steps S410 to S420 in Figure 4 the method steps S510 to S550 in Figure 5 the method steps S610 to S630 in Figure 6 the method steps S710 to S730 in Figure 7 The above-described device embodiments are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0115] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. The computer program is executed by a processor or a controller, for example, executed by a processor in the above electronic device embodiment, so that the processor can execute the training method of the entity representation model in the above embodiment. For example, the method steps S110 to S140 described above are executed,
[0116] the method steps S210 to S220 in Figure 1 the method steps S310 to S320 in Figure 2 the method steps S210 to S220 inFigure 3 Steps S310 to S320 in Figure 4 Steps S410 to S420 in Figure 5 Steps S510 to S550 in Figure 6 Steps S610 to S630 in Figure 7 Steps S710 to S730 in. Those of ordinary skill in the art can understand that all or some of the steps and devices in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, which can include a computer storage medium (or a non-transitory storage medium) and a communication storage medium (or a transitory storage medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable storage media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other storage medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication storage medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery storage medium.
[0117] This application can be used in many general-purpose or special-purpose computer device environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multiprocessor devices, microprocessor-based devices, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above devices or equipment, and so on. This application can be described in the general context of a computer program executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more programs for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based device that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0119] The units described in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0120] It should be noted that although several modules or units of devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0121] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0122] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
[0123] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0124] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above-mentioned embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A training method for an entity representation model, characterized in that, it includes: Obtain a sample sentence, split the sample sentence into multiple sample words according to the character granularity, and obtain a word sequence composed of multiple sample words; Configure multiple reference entities in a preset knowledge graph; Annotate reference information for the reference entities according to a preset data set; Continuously select at least two of the sample words from the word sequence to obtain a sample phrase; Match at least one sample entity from the reference entities according to the sample phrase and the reference information; Obtain a sample sequence by splicing the sample entity and the word sequence, and input the sample sequence into the entity representation model for entity representation training.
2. The training method for the entity representation model according to claim 1, characterized in that, the entity representation model includes a RoBERTa model, and the inputting the sample sequence into the entity representation model for entity representation training includes: Performing semantic encoding on the sample sequence through the RoBERTa model to obtain a first token corresponding to the sample word and a second token corresponding to the sample entity; Performing entity representation training on the sample sequence according to the first token, the second token, and obtaining a semantic representation vector of the sample entity.
3. The training method for the entity representation model according to claim 2, characterized in that, the entity representation model further includes a Transformer model, and the performing entity representation training on the sample sequence according to the first token, the second token, and obtaining a semantic representation vector of the sample entity includes: Inputting the first token, the second token, and the sample sequence into the Transformer model; Determining a first attention matrix through the Transformer model, where the first attention matrix represents the attention relationship between multiple first tokens; Determining a second attention matrix through the Transformer model, where the second attention matrix represents the attention relationship between the second token and the first tokens; Obtaining a first feature vector according to the word sequence and the first attention matrix, and obtaining a second feature vector according to the sample entity and the second attention matrix; Obtaining the semantic representation vector according to the first feature vector and the second feature vector.
4. The training method for the entity representation model according to claim 3, characterized in that, the determining the second attention matrix through the Transformer model includes: Obtaining start position embedding information and end position embedding information corresponding to the sample entity, where the start position embedding information is the position embedding information in the sample word with the earliest order corresponding to the sample entity, and the end position embedding information is the position embedding information in the sample word with the latest order corresponding to the sample entity; Determine the target position embedding information of the sample entity according to the start position embedding information and the end position embedding information; Determine the second attention matrix according to the target position embedding information, the first token, and the second token.
5. The training method of the entity representation model according to claim 4, wherein, The obtaining the semantic representation vector according to the first feature vector and the second feature vector includes: Obtain a preset loss weight; Calculate losses for the first feature vector and the second feature vector respectively according to the loss weight; Merge the feature vectors obtained by loss calculation to obtain the semantic representation vector.
6. An entity representation model training device, wherein, Comprising: A word sequence acquisition unit, configured to acquire a sample sentence, split the sample sentence into a plurality of sample words according to character granularity, and obtain a word sequence composed of the plurality of sample words; A knowledge graph acquisition unit, configured to configure a plurality of reference entities in a preset knowledge graph; Annotate reference information for the reference entities according to a preset data set; An entity acquisition unit, configured to continuously select at least two of the sample words from the word sequence to obtain a sample phrase; match at least one sample entity from the reference entities according to the sample phrase and the reference information; A training unit, configured to obtain a sample sequence by splicing the sample entity and the word sequence, and input the sample sequence into an entity representation model for entity representation training.
7. An electronic device, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the training method of the entity representation model according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium storing a computer program, wherein, The computer program is used to execute the training method of the entity representation model according to any one of claims 1 to 5.
Citation Information
Patent Citations
Named entity recognition model training method and named entity recognition method and device
CN110705294A
Entity linking method based on RoBERTa and heuristic algorithm
CN111125380A
Intention recognition method and device, equipment and medium
CN113360751A