Method for training information extraction model, information extraction method and device
By extracting entity and relation type fields from the original text to generate prompt data and negative sample data, and training the information extraction model, the problems of time-consuming, labor-intensive, and false recall in existing technologies are solved, and efficient and accurate information extraction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2023-02-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies are time-consuming and labor-intensive in information extraction, requiring a significant amount of time to refine the rule base, and the phenomenon of false recall occurs frequently during model training.
By extracting entity, entity type, and relation type fields from the original text data as labeled data, prompt data is generated and negative sample data is constructed to train the information extraction model and improve model performance.
This improved the training performance of the information extraction model, reduced false recalls, and enhanced the accuracy and efficiency of information extraction.
Smart Images

Figure CN116069785B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the field of natural language processing technology. Background Technology
[0002] The rapid development of the internet has driven the advancement of information processing technology, enabling information dissemination and sharing. The internet provides an inexhaustible information carrier, and the core value of information extraction lies in solving the problems of information being widely dispersed, difficult to share, and hard to extract desired content from vast information resources. Information extraction refers to the automatic extraction of specified types of entity, relationship, and event factual information from natural language text and the resulting structured data output. Information extraction technology is widely used in numerous industries such as finance, government affairs, law, and healthcare, processing large amounts of document information into digital, structured data.
[0003] Knowledge engineering methods can be used for information extraction. Knowledge engineering primarily relies on manually writing extraction rules to enable the system to handle information extraction problems within a specific knowledge domain. This method requires the knowledge engineer who develops the rules to have a deep understanding of that knowledge domain. Therefore, the development process can be very time-consuming and labor-intensive, requiring a significant amount of time to refine the rule base. Summary of the Invention
[0004] This disclosure provides a method for training an information extraction model, an information extraction method, an apparatus, a device, a storage medium, and a program product.
[0005] According to one aspect of this disclosure, a method for training an information extraction model is provided, comprising: extracting multiple entity fields, multiple entity type fields, and at least one relation type field from original text data as labeled data, wherein the multiple entity fields represent multiple entities in the original text data, the multiple entity type fields represent the types of the multiple entities, and the at least one relation type field represents the relationship between the multiple entities; determining multiple prompt data based on the labeled data; determining multiple negative sample data based on the multiple prompt data; and training the information extraction model based on the multiple negative sample data to obtain a target information extraction model.
[0006] According to another aspect of this disclosure, an information extraction method is provided, comprising: acquiring target text data; and inputting the target text data into an information extraction model to obtain a target information extraction result, wherein the information extraction model is trained according to the method described in the embodiments of this disclosure.
[0007] According to another aspect of this disclosure, an apparatus for training an information extraction model is provided, comprising: an annotation module for extracting multiple entity fields, multiple entity type fields, and at least one relation type field from original text data as annotation data, wherein the multiple entity fields represent multiple entities in the original text data, the multiple entity type fields represent the types of the multiple entities, and the at least one relation type field represents the relationship between the multiple entities; a prompt data determination module for determining multiple prompt data based on the annotation data; a negative sample data determination module for determining multiple negative sample data based on the multiple prompt data; and a training module for training the information extraction model based on the multiple negative sample data to obtain a target information extraction model.
[0008] According to another aspect of this disclosure, an information extraction apparatus is provided, comprising: an acquisition module for acquiring target text data; and an input module for inputting the target text data into an information extraction model to obtain a target information extraction result, wherein the information extraction model is trained according to the method described in the embodiments of this disclosure.
[0009] Another aspect of this disclosure provides an electronic device including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods shown in embodiments of this disclosure.
[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the methods shown in the embodiments of the present disclosure.
[0011] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instructions, characterized in that, when the computer program / instructions are executed by a processor, they implement the steps of the method shown in the embodiments of the present disclosure.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0014] Figure 1 This is an exemplary system architecture of the method, information extraction method, and apparatus for training information extraction models according to embodiments of this disclosure;
[0015] Figure 2 A flowchart illustrating a method for extracting training information from a model according to an embodiment of the present disclosure is shown schematically.
[0016] Figure 3 A flowchart illustrating a method for determining multiple prompt data based on annotation data according to an embodiment of the present disclosure is shown schematically.
[0017] Figure 4 The illustration shows a schematic diagram of a method for determining a plurality of prompt data based on annotation data according to another embodiment of the present disclosure;
[0018] Figure 5 A flowchart illustrating a method for determining multiple negative sample data based on multiple cue data according to an embodiment of the present disclosure is shown schematically.
[0019] Figure 6 The flowchart illustrating a method for training an information extraction model based on a plurality of negative sample data according to an embodiment of the present disclosure is shown in the illustration.
[0020] Figure 7 A flowchart illustrating an information extraction method according to an embodiment of the present disclosure is shown schematically.
[0021] Figure 8 A block diagram of an apparatus for extracting training information from a model according to an embodiment of the present disclosure is shown schematically.
[0022] Figure 9 A block diagram of an information extraction apparatus according to an embodiment of the present disclosure is shown schematically;
[0023] Figure 10 A block diagram of an example electronic device that can be used to implement embodiments of this disclosure is shown schematically. Detailed Implementation
[0024] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0025] The following will combine Figure 1 The application scenarios of the training information extraction model, information extraction method and device provided in this disclosure are described.
[0026] Figure 1This is an exemplary system architecture 100 of a method for training an information extraction model, an information extraction method, and an apparatus according to embodiments of this disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0028] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0029] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0030] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0031] It should be noted that the training information extraction model method and information extraction method provided in this disclosure embodiment can generally be executed by server 105. Correspondingly, the training information extraction model apparatus and information extraction apparatus provided in this disclosure embodiment can generally be located in server 105. The training information extraction model method and information extraction method provided in this disclosure embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the training information extraction model apparatus and information extraction apparatus provided in this disclosure embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0032] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0033] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.
[0034] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0035] The following will combine Figure 2 The method for extracting training information from the model provided in this disclosure is described.
[0036] Figure 2 A flowchart illustrating a method for extracting training information from a model according to an embodiment of the present disclosure is shown schematically.
[0037] like Figure 2 As shown, the method 200 includes operation S210, extracting multiple entity fields, multiple entity type fields and at least one relation type field from the original text data as annotation data.
[0038] According to embodiments of this disclosure, the original text data may include multiple entities. Multiple entity fields may represent multiple entities in the original text data, with each entity field corresponding to one of the multiple entities. Multiple entity type fields may represent the types of multiple entities, with each entity type field corresponding to one of the multiple entities. At least one relationship type field may represent the relationship between the multiple entities, and the multiple entities may have at least one relationship, with each relationship type field corresponding to one of the at least one relationship.
[0039] For example, entities may include specific people, objects, events, concepts, etc. Entity types may include, for example, person name types, job name types, legal types, article number types, clause number types, etc. Relationship type fields may include, for example, "established at," "name is," "occupation is," "legal type is," "article number is," "clause number is," etc.
[0040] It should be noted that, in addition to entity fields, entity type fields, and relation type fields, annotation data may also include other fields, which are not specifically limited in this disclosure.
[0041] Then, in operation S220, multiple prompt data are determined based on the labeled data.
[0042] According to embodiments of this disclosure, the prompt data may include, for example, a Prompt. The Prompt is additional prompt information appended to text data, and it can be used to train an information extraction model.
[0043] In operation S230, multiple negative sample data are determined based on multiple prompt data.
[0044] According to embodiments of this disclosure, negative sample data is text data that does not contain target information, wherein the target information is the object to be extracted by the information extraction model.
[0045] In operation S240, the information extraction model is trained based on multiple negative sample data to obtain the target information extraction model.
[0046] According to embodiments of this disclosure, an information extraction model can be used to extract target information from text data. The target information may include, for example, entities, entity types, entity relationships, etc.
[0047] According to embodiments of this disclosure, if labeled data is directly used as training data without considering the inclusion of negative samples in the training data, false recalls may occur during prediction. Therefore, in this embodiment, by converting labeled data into cue data, constructing negative sample data based on the cue data, and using training data with added negative sample data to train the model, the training effect of the model can be improved, and false recalls can be avoided.
[0048] For example, the original text data could be "The chairman of Bank A is Zhang San". Therefore, we can determine that the entities in the original text data include "Bank A", "Chairman", and "Zhang San". "Bank A", "Chairman", and "Zhang San" can be labeled as entity fields. The entity type field corresponding to "Bank A" can be labeled as "Company Name". The entity type field corresponding to "Chairman" can be labeled as "Position". The entity type field corresponding to "Zhang San" can be labeled as "Person Name". The relationship type field between "Bank A" and "Chairman" can be labeled as "Chairman is". The relationship type field between "Chairman" and "Zhang San" can be labeled as "Person Name is".
[0049] The following will combine Figure 3 The method for determining multiple prompt data based on labeled data, as provided in this disclosure, is described.
[0050] Figure 3 A flowchart illustrating a method for determining multiple prompt data based on labeled data according to an embodiment of the present disclosure is shown schematically.
[0051] like Figure 3 As shown, method 320 includes operation S321, generating multiple nodes based on multiple entity fields and multiple entity type fields.
[0052] According to embodiments of this disclosure, each node may correspond to an entity, and the node may include the entity's entity fields and entity type.
[0053] In operation S322, edges between multiple nodes are generated based on at least one relation type field to obtain a directed graph.
[0054] According to embodiments of this disclosure, the directed graph includes multiple nodes and corresponding edges. For example, nodes with relationships can be determined based on at least one relationship type field, and these nodes with relationships can be connected by edges. The relationship indicated by the relationship type field is directional, i.e., from the relationship subject to the relationship object. For example, in "The Chairman is Zhang San", "Chairman" is the relationship subject, and "Zhang San" is the relationship object. Based on this, the direction of the edges can be set according to the direction of the relationship, i.e., from the relationship subject to the relationship object.
[0055] In operation S323, the path information between the root node and each other node in the directed graph is determined, resulting in multiple path information.
[0056] According to embodiments of this disclosure, path information may include, for example, the nodes and edges between the root node and the target node. The target node can be any node in the directed graph other than the root node.
[0057] In operation S324, for each path information among multiple path information, the prompt data is determined based on the entity field, entity type field, and relation type field corresponding to the path information.
[0058] According to embodiments of this disclosure, for example, multiple levels of prompt fields can be determined based on entity fields and relationship type fields corresponding to path information. Multiple levels of type fields are then determined based on the entity type fields and relationship type fields corresponding to path information. Finally, the multiple levels of prompt fields and the multiple levels of type fields are determined as prompt data.
[0059] According to embodiments of this disclosure, the prompt data can be in the form of key-value pairs. For example, the key in the prompt data may include a type field, and the value may include a prompt field.
[0060] For example, in this embodiment, prompt fields can be determined based on the entity fields corresponding to each node in the path information, resulting in multiple prompt fields. The hierarchy of these multiple prompt fields is set based on the relationship type fields corresponding to the edges in the path information. Alternatively, type fields can be determined based on the entity type fields corresponding to each node in the path information, resulting in multiple type fields. The hierarchy of these multiple type fields is set based on the relationship type fields corresponding to the edges in the path information.
[0061] The following is for reference. Figure 4 The method for determining multiple prompt data based on labeled data, as described above, will be further explained in conjunction with specific embodiments. Those skilled in the art will understand that the following example embodiments are only for understanding this disclosure, and this disclosure is not limited thereto.
[0062] Figure 4 The illustration shows a schematic diagram of a method for determining a plurality of prompt data based on labeled data according to another embodiment of the present disclosure.
[0063] like Figure 4 As shown, exemplarily, in this embodiment, the entity fields include "Bank A", "Chairman", and "Zhang San". The entity type field corresponding to "Bank A" is "Company Name". The entity type field corresponding to "Chairman" is "Position". The entity type field corresponding to "Zhang San" is "Person Name". The relationship type field between "Bank A" and "Chairman" is "Chairman is". The relationship type field between "Chairman" and "Zhang San" is "Person Name is".
[0064] Based on this, node 411 can be determined according to "Bank A" and "Company Name". Node 412 can be determined according to "Chairman" and "Position". Node 413 can be determined according to "Zhang San" and "Person's Name". Then, according to the relation type field "Chairman is", edge 421 between nodes 411 and 412 can be determined, and according to the relation type field "Person's Name is", edge 422 between nodes 412 and 413 can be determined.
[0065] Then, the first path information from root node 411 to node 412 can be determined, including node 411, edge 421, and node 412. The second path information from root node 411 to node 413 includes node 411, edge 421, node 412, edge 422, and node 413. Therefore, the prompt data can be determined based on the first path information. For example, in the first path information, edge 421 indicates that node 411 points to node 412. Therefore, the entity field corresponding to node 411 can be used as the first-level prompt field, and the entity type field corresponding to node 411 can be used as the first-level type field. The entity field corresponding to node 412 can be used as the second-level prompt field, and the entity type field corresponding to node 412 can be used as the second-level type field. Thus, the prompt data shown in Table 1 can be obtained.
[0066] hierarchy Type field Prompt field 1 Company Name Bank A 2 Position Chairman
[0067] Table 1
[0068] Alternatively, the prompt data can be determined based on the second path information. For example, in the second path information, edge 421 indicates that node 411 points to node 412, and edge 422 indicates that node 413 points to node 414. Therefore, the entity field corresponding to node 411 can be used as the first-level prompt field, and the entity type field corresponding to node 411 can be used as the first-level type field. The entity field corresponding to node 412 can be used as the second-level prompt field, and the entity type field corresponding to node 412 can be used as the second-level type field. The entity field corresponding to node 413 can be used as the third-level prompt field, and the entity type field corresponding to node 413 can be used as the third-level type field. Thus, the prompt data shown in Table 2 can be obtained.
[0069]
[0070]
[0071] Table 2
[0072] The following will combine Figure 5 The method for determining multiple negative sample data based on multiple prompt data provided in this disclosure is described.
[0073] Figure 5 A flowchart illustrating a method for determining multiple negative sample data based on multiple cue data according to an embodiment of the present disclosure is shown.
[0074] like Figure 5 As shown, the method 530 includes performing operations S531 to S533 for each of the multiple prompt data.
[0075] In operation S531, determine the intermediate fields based on the prompt fields of the prompt data for all levels except the last level.
[0076] In operation S532, determine the relation field based on the type field of the last level in the prompt data.
[0077] In operation S533, negative sample data is determined based on intermediate fields and relational fields.
[0078] For example, taking the prompt data shown in Table 2 above as an example, the last level in this prompt data is level 3. Based on this, the prompt field "Bank A" from level 1 and the prompt field "Chairman" from level 2 can be concatenated to obtain "Chairman of Bank A", which serves as the intermediate field. Then, based on the type field "Name" from the last level, the relation field "Name" can be determined. Next, the intermediate field and the relation field can be concatenated to obtain "Name of the Chairman of Bank A". This "Name of the Chairman of Bank A" can then be used as a negative sample for model training.
[0079] According to embodiments of this disclosure, positive sample data can also be obtained. For example, labeled data can be used as positive sample data. Intermediate information and relational fields are concatenated to obtain a concatenated result. Then, it is determined whether the concatenated result conflicts with the positive sample data. If the concatenated result does not conflict with the positive sample data, the concatenated result is determined as negative sample data.
[0080] The following will combine Figure 6 This disclosure describes the method provided for training an information extraction model based on multiple negative sample data to obtain a target information extraction model.
[0081] Figure 6 The flowchart illustrating a method for training an information extraction model based on a plurality of negative sample data according to an embodiment of the present disclosure is shown.
[0082] like Figure 6 As shown, the method 640 includes performing operations S641 to S643 for each of the multiple negative sample data.
[0083] In operation S641, negative sample data is input into the information extraction model to obtain information extraction results.
[0084] In operation S642, the loss value is determined based on the information extraction results.
[0085] According to embodiments of this disclosure, the loss value can be used to represent the difference between the information extraction result and the correct extraction result. For example, the loss value can be calculated based on a loss function, which may include, for example, a cross-entropy loss function, a hinge loss function, and an exponential loss function.
[0086] When operating S643, adjust the parameters of the information extraction model based on the loss value.
[0087] According to embodiments of this disclosure, the above training operations can be repeated until the information extraction result of the information extraction model meets predetermined requirements. These predetermined requirements can be set according to user needs; for example, they can be set to the convergence of the information extraction result.
[0088] The following will combine Figure 7 The information extraction method provided in this disclosure is described.
[0089] Figure 7 A flowchart illustrating an information extraction method according to an embodiment of the present disclosure is shown.
[0090] like Figure 7 As shown, the information extraction method 700 obtains target text data in operation S710.
[0091] According to embodiments of this disclosure, the target text data may be text data intended for information extraction.
[0092] When operating the S720, the target text data is input into the information extraction model to obtain the target information extraction result.
[0093] According to embodiments of this disclosure, the target information extraction result may include, for example, entity fields, entity type fields, and relationship type fields. Entity fields may represent entities in the target text data, entity type fields may represent the type of the entities, and relationship type fields may represent the relationships between the entities.
[0094] The information extraction model was trained using the method described above.
[0095] The method, information extraction method, and apparatus for training information extraction models according to embodiments of this disclosure can be applied to the business scenarios of financial institutions and financial regulatory agencies.
[0096] The business scenarios for financial institutions may include, for example, inquiries into internal and external regulations, penalty notices, internal compliance management, and public opinion monitoring. The business scenarios for financial regulatory agencies may include, for example, regulation management, penalty notice management, and public opinion monitoring of financial institutions.
[0097] The method for extracting training information from the model described above will be further explained below with reference to specific embodiments. Those skilled in the art will understand that the following examples are for illustrative purposes only and are not limited thereto.
[0098] Fines are widely used in regulatory scenarios. To effectively manage fines, it is necessary to structure and digitize the data, construct relationship graphs, and enable functions such as public opinion monitoring and compliance management. For example, for fines in web-based format, information extraction models can be used to extract target information, which can then be structured for easier subsequent data processing.
[0099] For example, the structured information required in a penalty notice includes: information about the penalized party, information about the penalty result, and information about the basis for the penalty. For the penalized party information, if the penalized party is an individual, the penalized party's name, their employer and position, and the corresponding penalty result can be extracted. If the penalty involves an amount, the specific amount can be extracted. For the penalty basis information, the name of the law on which the penalty is based, the date of its promulgation, and the specific clauses on which it is based can be extracted.
[0100] For example, in this embodiment, the original text data corresponding to the basis for the penalty can be: "Article 7 and Article 42, Paragraph (3) of the Interim Measures for the Administration of Personal Loans". Based on this, the labeled data can be generated as shown in Table 3.
[0101]
[0102]
[0103] Table 3
[0104] Here, `key` is the entity field, `key_str` is the entity type field, `value_positions` is the entity's position in the original text data, `global_offset` represents the entity's offset in the original text data, `value_strs` is the entity's content, and `tag_id` is the entity's identifier.
[0105] Then, based on the above labeled data, multiple prompt data can be determined, and based on the multiple prompt data, multiple negative sample data can be determined as shown in Table 4.
[0106]
[0107] Table 4
[0108] Next, the information extraction model can be trained using the data shown in Table 3 as positive samples and the data shown in Table 4 as negative samples.
[0109] According to embodiments of this disclosure, adding negative sample data can significantly improve the extraction performance of the information extraction model. Since Article 7 of the "Interim Measures for the Administration of Personal Loans" does not include any specific loan amount, without training with negative sample data, the information extraction model could easily mistake the item number in Article 42 for the item number in Article 7, leading to false positives. Furthermore, this paragraph does not mention loan numbers, while both loan amounts and items appear after the article; therefore, without training with negative sample data, the information extraction model can easily confuse loan amounts and items, resulting in false positives. Adding negative sample data improves the recall accuracy of the information extraction model.
[0110] The following will combine Figure 8 The apparatus for extracting training information provided in this disclosure is described.
[0111] Figure 8 A block diagram of an apparatus for extracting training information from a model according to an embodiment of the present disclosure is shown schematically.
[0112] like Figure 8 As shown, the apparatus 800 for training information extraction models includes a labeling module 810, a prompt data determination module 820, a negative sample data determination module 830, and a training module 840.
[0113] The annotation module 810 is used to extract multiple entity fields, multiple entity type fields and at least one relationship type field from the original text data as annotation data. The multiple entity fields represent multiple entities in the original text data, the multiple entity type fields represent the types of multiple entities, and the at least one relationship type field represents the relationship between multiple entities.
[0114] The prompt data determination module 820 is used to determine multiple prompt data based on the labeled data.
[0115] The negative sample data determination module 830 is used to determine multiple negative sample data based on multiple prompt data.
[0116] Training module 840 is used to train the information extraction model based on multiple negative sample data to obtain the target information extraction model.
[0117] According to embodiments of this disclosure, the prompt data determination module may include: a node generation submodule, used to generate multiple nodes based on multiple entity fields and multiple entity type fields; an edge generation submodule, used to generate edges between multiple nodes based on at least one relation type field, to obtain a directed graph; a path information determination module, used to determine path information from the root node to each other node in the directed graph, to obtain multiple path information; and a prompt data determination module, used to determine prompt data for each path information based on the entity field, entity type field, and relation type field corresponding to the path information.
[0118] According to embodiments of this disclosure, the prompt data determination module may include: a prompt field determination submodule, used to determine multiple levels of prompt fields based on entity fields and relationship type fields corresponding to path information; a type field determination submodule, used to determine multiple levels of type fields based on entity type fields and relationship type fields corresponding to path information; and a first determination submodule, used to determine multiple levels of prompt fields and multiple levels of type fields as prompt data.
[0119] According to embodiments of this disclosure, the negative sample data determination module may include: an intermediate field determination submodule, configured to determine an intermediate field for each of the multiple prompt data based on prompt fields of the prompt data at levels other than the last level; a relation field determination submodule, configured to determine a relation field based on the type field of the last level in the prompt data; and a second determination submodule, configured to determine negative sample data based on the intermediate field and the relation field.
[0120] According to embodiments of this disclosure, the second determining submodule may include: a splicing unit for splicing intermediate information and relational fields to obtain a splicing result; a matching unit for determining whether the splicing result conflicts with positive sample data; and a third determining unit for determining the splicing result as negative sample data if the splicing result does not conflict with the positive sample data.
[0121] According to embodiments of this disclosure, the training module may include: an input submodule, used to input negative sample data into an information extraction model for each negative sample data in a plurality of negative sample data to obtain an information extraction result; a loss determination submodule, used to determine a loss value based on the information extraction result; and an adjustment submodule, used to adjust the parameters of the information extraction model based on the loss value.
[0122] The following will combine Figure 9 The information extraction apparatus provided in this disclosure is described.
[0123] Figure 9 A block diagram of an information extraction apparatus according to an embodiment of the present disclosure is shown schematically.
[0124] like Figure 9 As shown, the information extraction device 900 includes an acquisition module 910 and an input module 920.
[0125] The acquisition module 910 is used to acquire target text data.
[0126] The input module 920 is used to input target text data into the information extraction model to obtain target information extraction results, wherein the information extraction model is trained according to the method shown in the embodiments of this disclosure.
[0127] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0128] Figure 10 The illustration shows an example electronic device 1000 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0129] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0130] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 10010, such as network card, modem, wireless transceiver, etc. Communication unit 10010 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0131] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as methods for training information extraction models and information extraction methods. For example, in some embodiments, the methods for training information extraction models and information extraction methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 10010. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the methods for training information extraction models and information extraction methods described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured in any other suitable manner (e.g., by means of firmware) to perform a method for training an information extraction model or an information extraction method.
[0132] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0133] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0134] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0135] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0136] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0137] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0138] A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system. It solves the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"), such as high management difficulty and weak business scalability. A server can also be a server for a distributed system, or a server that incorporates blockchain technology.
[0139] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0140] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for training an information extraction model, comprising: Multiple entity fields, multiple entity type fields, and at least one relationship type field are extracted from the original text data as annotation data. The multiple entity fields represent multiple entities in the original text data, the multiple entity type fields represent the types of the multiple entities, and the at least one relationship type field represents the relationship between the multiple entities. Based on the labeled data, multiple prompt data are determined; Based on the aforementioned multiple prompts, multiple negative sample data are determined; and Based on the multiple negative sample data, the information extraction model is trained to obtain the target information extraction model. The step of determining multiple prompt data based on the labeled data includes: Multiple nodes are generated based on the multiple entity fields and the multiple entity type fields; Based on the at least one relationship type field, edges are generated between the plurality of nodes to obtain a directed graph; The path information from the root node to other nodes in the directed graph is determined to obtain multiple path information. For the path information among the multiple path information, based on the entity field and relation type field corresponding to the path information, multiple levels of prompt fields are determined; Based on the entity type field and relationship type field corresponding to the path information, determine the type fields at multiple levels; and The prompt fields and type fields of the multiple levels are determined as the prompt data, and The step of determining multiple negative sample data based on the multiple prompt data includes: For the prompt data in the plurality of prompt data, the intermediate field is determined based on the prompt fields of the other levels in the prompt data except for the last level; Determine the relationship field based on the type field of the last level in the provided data; and The negative sample data is determined based on the intermediate field and the relationship field.
2. The method according to claim 1, wherein, The step of determining the negative sample data based on the intermediate field and the relation field includes: The intermediate field and the relation field are concatenated to obtain the concatenated result; Determine whether the splicing result conflicts with the positive sample data; and If the splicing result does not conflict with the positive sample data, the splicing result is determined as the negative sample data.
3. The method according to claim 1, wherein, The step of training the information extraction model based on the multiple negative sample data to obtain the target information extraction model includes: For the negative sample data among the multiple negative sample data, The negative sample data is input into the information extraction model to obtain the information extraction result; Based on the information extraction results, the loss value is determined; and The parameters of the information extraction model are adjusted based on the loss value.
4. An information extraction method, comprising: Obtain the target text data; as well as The target text data is input into the information extraction model to obtain the target information extraction result, wherein the information extraction model is trained by the method according to any one of claims 1-3.
5. An apparatus for training an information extraction model, comprising: The annotation module is used to extract multiple entity fields, multiple entity type fields, and at least one relationship type field from the original text data as annotation data. The multiple entity fields represent multiple entities in the original text data, the multiple entity type fields represent the types of the multiple entities, and the at least one relationship type field represents the relationship between the multiple entities. The prompt data determination module is used to determine multiple prompt data based on the labeled data; A negative sample data determination module is used to determine multiple negative sample data based on the multiple prompt data; and The training module is used to train the information extraction model based on the multiple negative sample data to obtain the target information extraction model. The prompt data determination module includes: The node generation submodule is used to generate multiple nodes based on the multiple entity fields and the multiple entity type fields; An edge generation submodule is used to generate edges between the plurality of nodes based on the at least one relation type field, thereby obtaining a directed graph; The path information determination module is used to determine the path information from the root node to all other nodes in the directed graph, thereby obtaining multiple path information items; and The prompt data determination module is used to determine the prompt data based on the path information among the multiple path information, according to the entity field, entity type field, and relation type field corresponding to the path information. The prompt data determination module includes: The prompt field determination submodule is used to determine prompt fields at multiple levels based on the entity fields and relationship type fields corresponding to the path information; The type field determination submodule is used to determine multiple levels of type fields based on the entity type field and relation type field corresponding to the path information; and The first determining submodule is used to determine the prompt fields and type fields of the multiple levels as the prompt data. The negative sample data determination module includes: The intermediate field determination submodule is used to determine the intermediate field for the prompt data in the multiple prompt data based on the prompt fields of the other levels in the prompt data except for the last level; The relation field determination submodule is used to determine the relation field based on the type field of the last level in the prompt data; and The second determining submodule is used to determine the negative sample data based on the intermediate field and the relationship field.
6. The apparatus according to claim 5, wherein, The second determining submodule includes: A splicing unit is used to splice the intermediate field and the relation field to obtain a splicing result; A matching unit is used to determine whether the splicing result conflicts with positive sample data; and The third determining unit is used to determine the splicing result as the negative sample data if the splicing result does not conflict with the positive sample data.
7. The apparatus according to claim 5, wherein, The training module includes: The input submodule is used to input the negative sample data from the plurality of negative sample data into the information extraction model to obtain the information extraction result; The loss determination submodule is used to determine the loss value based on the information extraction results; and The adjustment submodule is used to adjust the parameters of the information extraction model based on the loss value.
8. An information extraction device, comprising: The acquisition module is used to acquire target text data; as well as An input module is used to input the target text data into an information extraction model to obtain the target information extraction result, wherein the information extraction model is trained by the method according to any one of claims 1-3.
9. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-4.
11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1-4.