Methods for determining entities and entity relationships and training methods for information extraction models

By adopting the determination method of entity and entity relationship in legal consultation scenarios, and using vectorization processing, feature extraction and graph neural network technology, the problem of difficulty in accurately extracting legal entity information and relationships in the existing technology is solved, and the accuracy of extraction and the effect of legal event detection are improved.

CN119398160BActive Publication Date: 2025-05-23NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510005472.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-23
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately extract legal entity information and entity relationships in legal consultation scenarios, especially when the legal information content is long and the form of discontinuous characters is present in the consultation information input by users. The traditional naming entity recognition and relationship extraction model cannot be effectively implemented.

Method used

A method of determining entities and entity relationships is adopted to obtain the user-input question information and corresponding event information, vectorized processing and feature extraction are performed, and the entity and entity relationship in the question information is determined by combining the attention mechanism and graph neural network.

Benefits of technology

It improves the accuracy of entity and entity relationship extraction in legal consultation scenarios, can more accurately represent the entity information and relationships in the questioning information, and enhances the accuracy of legal event detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119398160B_ABST
    Figure CN119398160B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technology, and discloses a method for determining entities and entity relationships and a method for training an information extraction model, the method comprising: vectorizing the first event information corresponding to the question information to obtain a supervision vector corresponding to the first event information; extracting features from the question information to obtain multiple element features; determining multiple first entities based on the supervision vector and multiple element features corresponding to the first event information; determining the entity relationships between the first entities based on the multiple first entities; determining the second event information corresponding to the question information based on the entity relationships between the multiple first entities and the first entities; when the second event information is the same as the first event information, determining the multiple first entities as the multiple target entities corresponding to the question information, and determining the entity relationships between the first entities as the entity relationships between the target entities. The present application can improve the accuracy of extracting entities and the relationships between the entities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for determining entities and entity relationships and a method for training an information extraction model. Background Art

[0002] In the legal consultation scenario, users can consult information related to legal cases through the legal consultation system. When the legal consultation system generates a corresponding response based on the consultation information input by the user, it needs to detect the legal events involved in the consultation information. Among them, legal events refer to the legal relationships that legal subjects may be involved in. Therefore, the legal consultation system can further process the consultation information based on the determined legal events (such as retrieving relevant cases and legal basis, etc.) to provide users with accurate consulting opinions. Summary of the invention

[0003] The embodiment of the present application provides a method for determining entities and entity relationships and a method for training an information extraction model, which can accurately obtain the core information of each entity in the question information input by the user and the core relationship between the entities, thereby improving the accuracy of determining the legal events in the question information. Specifically, the embodiment of the present application discloses the following technical solutions:

[0004] A first aspect of an embodiment of the present application provides a method for determining entities and entity relationships, the method comprising: obtaining question information input by a user and first event information corresponding to the question information; performing vectorization processing on the first event information to obtain a supervision vector corresponding to the first event information; performing feature extraction on legal elements in the question information to obtain multiple element features corresponding to the question information; determining multiple first entities corresponding to the question information based on the supervision vector corresponding to the first event information and the multiple element features; determining the entity relationship between each of the multiple first entities based on the multiple first entities; determining the second event information corresponding to the question information based on the entity relationship between the multiple first entities and each of the first entities; when the second event information is the same as the first event information, determining the multiple first entities as multiple target entities corresponding to the question information, and determining the entity relationship between each of the first entities as the entity relationship between each of the multiple target entities corresponding to the question information.

[0005] In some embodiments, the feature extraction of the legal elements in the question information to obtain multiple element features corresponding to the question information includes: feature extraction of multiple types of legal elements in the question information through multiple element extraction models to obtain the multiple element features; wherein each of the multiple element models is used to extract different types of element features.

[0006] In some embodiments, the above-mentioned determination of the multiple first entities corresponding to the above-mentioned question information based on the supervision vector corresponding to the above-mentioned first event information and the above-mentioned multiple element features includes: concatenating or dot-multiplying the supervision vector corresponding to the above-mentioned first event information with each of the above-mentioned element features to obtain the above-mentioned multiple first element features; using an attention mechanism to adjust the weight of each of the above-mentioned multiple first element features to obtain multiple second element features after weight adjustment; performing feature fusion processing on the above-mentioned multiple second element features to obtain content features, and determining the above-mentioned multiple first entities based on the above-mentioned content features.

[0007] In some embodiments, the above-mentioned determination of the above-mentioned multiple first entities based on the above-mentioned content features includes: performing sequence annotation on the above-mentioned content features to obtain multiple label sequences corresponding to the above-mentioned content features; wherein each of the above-mentioned label sequences includes multiple label information; based on the multiple label information corresponding to each of the above-mentioned label sequences, determining the transition probability corresponding to each of the above-mentioned label sequences; wherein the above-mentioned transition probability is used to characterize the dependency relationship between the multiple label information corresponding to each of the above-mentioned label sequences; based on the transition probability corresponding to each of the above-mentioned label sequences, determining the conditional probability corresponding to each of the above-mentioned label sequences; based on the conditional probability corresponding to each of the above-mentioned label sequences, determining the target label sequence in the above-mentioned multiple label sequences, and determining the above-mentioned multiple first entities according to the above-mentioned target label sequence.

[0008] In some embodiments, the above-mentioned determining the second event information corresponding to the above-mentioned question information based on the above-mentioned multiple first entities and the entity relationships between each of the above-mentioned first entities includes: determining question information features corresponding to the above-mentioned question information and initial features corresponding to each of the above-mentioned first entities based on the above-mentioned question information and the above-mentioned multiple first entities; wherein the initial features corresponding to each of the above-mentioned first entities are used to characterize the context information corresponding to each of the above-mentioned first entities; determining core information corresponding to the above-mentioned question information based on the above-mentioned multiple first entities, the entity relationships between each of the above-mentioned first entities and the initial features corresponding to each of the above-mentioned first entities; determining the above-mentioned second event information based on the above-mentioned question information features and the core information corresponding to the above-mentioned question information.

[0009] In some embodiments, the core information corresponding to the question information includes the core information corresponding to each of the first entities in the question information and the core relationship between each of the first entities; the above-mentioned determining the core information corresponding to the question information based on the multiple first entities, the entity relationships between the first entities and the initial features corresponding to the first entities includes: determining a syntactic graph with each of the first entities as a node and the entity relationships between the first entities as an edge; updating the initial features corresponding to each of the first entities based on the syntactic graph and the initial features corresponding to the first entities through multiple neural network layers to obtain updated features of each of the nodes determined by each of the multiple neural network layers; determining the updated features of each of the nodes determined by the last neural network layer in the multiple neural network layers as the core information corresponding to the first entities and the core relationship between the first entities.

[0010] In some embodiments, the initial features corresponding to each of the above-mentioned first entities are updated through the multiple neural network layers based on the above-mentioned syntax graph and the initial features corresponding to each of the above-mentioned first entities, so as to obtain updated features of each of the above-mentioned nodes determined by each of the above-mentioned neural network layers in the above-mentioned multiple neural network layers, including: obtaining the first features of each of the above-mentioned nodes determined by the above-mentioned first neural network layer through the first neural network layer in the above-mentioned multiple neural network layers based on the initial features corresponding to each of the above-mentioned nodes, the adjacency matrix corresponding to the above-mentioned syntax graph, the self-connection matrix and the weight matrix corresponding to the above-mentioned first neural network layer; wherein the above-mentioned first neural network layer is the first neural network layer in the above-mentioned multiple neural network layers; obtaining the second features of each of the above-mentioned nodes determined by the above-mentioned second neural network layer through the second neural network layer in the above-mentioned multiple neural network layers based on the first features of each of the above-mentioned nodes, the adjacency matrix corresponding to the above-mentioned syntax graph, the self-connection matrix and the weight matrix corresponding to the above-mentioned second neural network layer; wherein the updated features of each of the above-mentioned nodes include the first features of each of the above-mentioned nodes and the second features of each of the above-mentioned nodes.

[0011] In some embodiments, the second event information is determined based on the question information feature and the core information corresponding to the question information, including: fusing the question information feature and the core information corresponding to the question information to obtain a fused feature; and classifying the legal events corresponding to the fused feature to obtain the second event information.

[0012] In some embodiments, the method further includes: in a case where the second event information is different from the first event information, determining multiple second entities corresponding to the question information and the entity relationships between each of the multiple second entities based on the second event information; determining third event information corresponding to the question information based on the multiple second entities and the entity relationships between each of the second entities; in a case where the third event information is the same as the second event information, determining the multiple second entities as multiple target entities corresponding to the question information, and determining the entity relationships between each of the second entities as the entity relationships between each of the multiple target entities corresponding to the question information.

[0013] A second aspect of an embodiment of the present application provides a training method for an information extraction model, which is applied to an information extraction model to be trained, and the method includes: obtaining a training text, a target entity corresponding to the training text, and legal event information; vectorizing the legal event information to obtain a supervision vector, and extracting features of the legal elements in the training text to obtain multiple element features corresponding to the training text; determining multiple predicted entities in the training text based on the supervision vector and the multiple element features; determining a first loss function corresponding to the information extraction model based on the multiple predicted entities and the multiple target entities; and optimizing the parameters in the first loss function to determine the trained information extraction model.

[0014] In some embodiments, the above method also includes: obtaining the true label sequence corresponding to the above training text, and determining the conditional probability corresponding to the above true label sequence based on the above true label sequence; determining the second loss function corresponding to the above information extraction model based on the above conditional probability; the above optimization of the parameters in the above first loss function to determine the trained information extraction model includes: optimizing the parameters in the above first loss function and the above second loss function to determine the above trained information extraction model.

[0015] According to a third aspect of an embodiment of the present application, there is provided a device for determining entities and entity relationships, the device comprising: a first acquisition module, configured to obtain question information input by a user and first event information corresponding to the question information; a first processing module, configured to perform vectorization processing on the first event information to obtain a supervision vector corresponding to the first event information; a feature extraction module, configured to perform feature extraction on legal elements in the question information to obtain multiple element features corresponding to the question information; a first determination module, configured to determine multiple first entities corresponding to the question information based on the supervision vector corresponding to the first event information and the multiple element features, and determine the entity relationship between each of the first entities in the multiple first entities based on the multiple first entities; a second determination module, configured to determine the second event information corresponding to the question information based on the multiple first entities and the entity relationship between each of the first entities; a third determination module, configured to determine the multiple first entities as multiple target entities corresponding to the question information when the second event information is the same as the first event information, and determine the entity relationship between each of the first entities as the entity relationship between each of the target entities in the multiple target entities corresponding to the question information.

[0016] The fourth aspect of an embodiment of the present application provides a training device for an information extraction model, which is configured on the information extraction model to be trained, and the device includes: a second acquisition module, configured to acquire a training text, a target entity corresponding to the training text, and legal event information; a second processing module, configured to vectorize the legal event information to obtain a supervision vector, and perform feature extraction on the legal elements in the training text to obtain multiple element features corresponding to the training text; a fourth determination module, configured to determine multiple predicted entities in the training text based on the supervision vector and the multiple element features; a fifth determination module, configured to determine a first loss function corresponding to the information extraction model based on the multiple predicted entities and the multiple target entities; an optimization module, configured to optimize the parameters in the first loss function to determine the trained information extraction model.

[0017] A fifth aspect of an embodiment of the present application provides an electronic device, comprising: one or more processors and a memory, the memory being configured to: store one or more programs; wherein, when the one or more programs are executed by one or more processors, the one or more processors implement the method for determining entities and entity relationships described in the first aspect above, or implement the method for training the information extraction model described in the second aspect above.

[0018] The sixth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores computer program instructions. When the computer reads the instructions, it executes the method for determining entities and entity relationships described in the first aspect, or implements the training method for the information extraction model described in the second aspect.

[0019] A seventh aspect of an embodiment of the present application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the method for determining entities and entity relationships described in the first aspect, or implements the training method for the information extraction model described in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 A schematic diagram of an information extraction model provided for some embodiments of the present application;

[0022] Figure 2 A flowchart of a method for determining entities and entity relationships provided in some embodiments of the present application;

[0023] Figure 3 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application;

[0024] Figure 4 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application;

[0025] Figure 5 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application;

[0026] Figure 6 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application;

[0027] Figure 7 A schematic diagram of a syntax diagram provided for some embodiments of the present application;

[0028] Figure 8 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application;

[0029] Fig. 9 A flowchart of a method for training an information extraction model provided in some embodiments of the present application;

[0030] Fig.10 A flowchart of another information extraction model training method provided in some embodiments of the present application;

[0031] Fig.11 A schematic diagram of a device for determining entities and entity relationships provided in some embodiments of the present application;

[0032] Fig.12 A schematic diagram of a training device for an information extraction model provided in some embodiments of the present application;

[0033] Fig.13 A schematic diagram of an electronic device according to some embodiments of the present application. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention and to make the above-mentioned purposes, features and advantages of the embodiments of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention are further described in detail below with reference to the accompanying drawings.

[0035] When extracting entities and entity relationships, existing named entity recognition (NER) models and relation extraction (RE) models can usually only focus on common entity information such as names of people, places, and institutions. However, in the legal consultation scenario, legal entity information needs to be extracted, and the legal entity information in different legal cases is also different. For example, the legal entity information involved in civil tort cases may include tortious acts, damage facts, causal relationships, subjective faults, etc.; the legal entity information involved in criminal cases may include criminal objects, objective aspects of crimes, criminal subjects, subjective aspects of crimes, etc. Therefore, the process of extracting legal entity information using NER models and RE models will be more complicated, and it is impossible to use supervised learning to let the model learn the entities and entity relationships to be extracted; in addition, the legal information content in the consultation information entered by the user is usually long and does not appear in the form of continuous characters, that is, the legal entity information usually has different degrees of overlap, which will further lead to the traditional NER model and RE model being unable to well extract legal entity information and entity relationships.

[0036] Based on the above technical problems, the present application provides a method for determining entities and entity relationships and a method for training information extraction models, which can improve the accuracy of entity and entity relationship extraction in legal consulting scenarios through remote supervision and legal element extraction.

[0037] It should be noted that the training samples used in the training process of the neural network model involved in this application are all from authorized legal documents, judgments, case descriptions, etc., and the conclusions obtained by the method for determining entities and entity relationships and the information extraction model provided in this application are only used to form consulting opinions for users for reference.

[0038] Figure 1 A schematic diagram of an information extraction model provided in some embodiments of the present application. Figure 1 As shown, the information extraction model 100 may include an entity extraction layer 10 and an entity relationship extraction layer 11. The entity extraction layer 10 includes a supervision model 110 and a content model 120; the entity relationship extraction layer 11 includes a graph neural network (GNN) model 130. The entity extraction layer 10 is used to determine the entities in the question information input by the user; the entity relationship extraction layer 11 is used to determine the entity relationship between entities based on the entities in the question information.

[0039] Exemplarily, after a user inputs legal consulting information (hereinafter referred to as question information) in a legal consulting system (hereinafter referred to as a central control system), the central control system may call corresponding legal tools to process the question information to obtain processing results, thereby summarizing the processing results obtained by each legal tool, obtaining the final consulting opinion and outputting it to the user. Among them, the legal tool may include a legal event detection model 140 and an information extraction model 100. The legal event detection model 140 may identify the legal event information involved in the question information input by the user, so that the central control system may search for legal cases or legal provisions based on the identified legal event information to provide legal consulting opinions to the user. The information extraction model 100 may extract entities and entity relationships in the question information when the legal event detection model 140 performs legal event detection on the question information input by the user, so that the legal event detection model 140 performs legal event detection based on the entities and entity relationships in the question information.

[0040] Exemplarily, after the central control system calls the information extraction model 100 and the legal event detection model 140, the information extraction model 100 first obtains the question information input by the user and the legal event detection result (i.e., the first event information) obtained in the previous legal event detection process, and extracts features of the first event information through the supervision model 110 in the entity extraction layer 10 to obtain a supervision vector. Next, the legal elements in the question information are feature extracted through the content model 120 in the entity extraction layer 10 to obtain multiple element features; then, the content features are determined based on the multiple element features and the supervision vector through the content model 120; wherein the content features are the entity information corresponding to the question information. Afterwards, the content features are subjected to entity relationship extraction through the GNN model 130 in the entity relationship extraction layer 11 to obtain the entity relationships between the entities.

[0041] For example, taking the question information input by the user as: "Zhang San borrowed 100,000 yuan from Li Si, and the two parties agreed to repay the loan one year later. One year later, Zhang San failed to repay the loan on time and ignored Li Si's reminders for payment", after the information extraction model 100 performs entity extraction on the question information, the multiple entities obtained may include: Zhang San (person), Li Si (person), 100,000 yuan (loan amount), one year (loan time); after performing entity relationship extraction on the above multiple entities, the entity relationships between the entities obtained may include: Zhang San borrowed money from Li Si (loan relationship), Zhang San failed to repay the loan on time (breach of contract), and Li Si urges for payment (collection behavior).

[0042] Exemplarily, the first entity corresponding to the question information obtained by the information extraction model 100 and the entity relationship between the first entities are further input into the legal event detection model 140. The legal event detection model 140 can determine the legal event information (i.e., the second event information) corresponding to the question information based on the above-mentioned first entity and the entity relationship between the first entities.

[0043] In the case where the second event information is the same as the first event information, the multiple first entities can be determined as multiple real entities (i.e., target entities) corresponding to the question information, and the entity relationship between the first entities can be determined as the real entity relationship corresponding to the question information (i.e., the entity relationship between the target entities in the multiple target entities). In addition, the central control system can further search for legal cases or legal provisions based on the second event information to provide legal advice to users.

[0044] In the case where the second event information is different from the first event information, the information extraction model 100 can extract legal information (including entities and entity relationships) from the question information again based on the second event information, and determine the legal event information corresponding to the question information (i.e., the third legal event information) based on the second entity and the entity relationship between the second entities obtained by the information extraction through the legal event detection model 140. In the case where the third event information is the same as the second event information, multiple second entities are determined as multiple target entities corresponding to the question information, and the entity relationship between each second entity is determined as the entity relationship between each target entity corresponding to the question information; in the case where the third event information is different from the second event information, continue to extract legal information based on the third event information through the information extraction model 100... and so on, until the legal event information determined by the previous legal event detection is the same as the legal event information determined by the next legal event detection.

[0045] Through the above scheme, different types of legal element features in the question information can be extracted through the content model in the information extraction model, and multiple types of element features can be remotely supervised through the supervision model in the information extraction model to obtain content features. The content features can more accurately represent the entity information corresponding to the question information, thereby improving the accuracy of entity and entity relationship extraction in legal consultation scenarios.

[0046] The following is a detailed description of the method for determining entities and entity relationships provided in this application.

[0047] Figure 2 A flowchart of a method for determining entities and entity relationships provided in some embodiments of the present application, such as Figure 2 As shown, the method for determining entities and entity relationships may include steps 210 to 260.

[0048] Step 210: Acquire question information input by the user and first event information corresponding to the question information.

[0049] In some embodiments, the information extraction model can obtain the question information input by the user and the first event information corresponding to the question information determined by the legal event detection model. When the legal event detection model processes the question information for the first time, the first event information can be the legal event detection result obtained by performing legal event detection on other question information. When the legal event detection model processes the question information not for the first time, the first event information can be the legal event detection result obtained by performing the previous round of legal event detection on the question information.

[0050] Step 220: perform vectorization processing on the first event information to obtain a supervision vector corresponding to the first event information.

[0051] In some embodiments, the supervision model in the information extraction model can vectorize the first event information to obtain a supervision vector corresponding to the first event information. The supervision model can be composed of a BERT model or a Transformer model; the supervision vector is used to assist the content model to more accurately determine the entity information in the question information.

[0052] Exemplarily, when the supervision model determines the supervision vector corresponding to the first event information, it can determine the legal information type corresponding to the first event information, and perform feature extraction (ie, vectorization processing) on ​​the legal information type to obtain the supervision vector.

[0053] In some examples, the question information input by the user is as follows: "A chemical plant was newly built upstream of my home this year. The output of my fish pond this year has dropped a lot compared to previous years. Some people in the village said that this is related to the illegal discharge of pollutants by the chemical plant, but we went to the factory to ask, and the person in charge insisted that they were discharging legally, and the change in fish pond output was caused by the climate. We don’t have evidence now, what should we do?" In the field of civil torts, the constituent elements of general tort incidents (i.e., legal information types) include tortious acts, damage facts, causal relationships, subjective faults, etc. When the legal event detection model determines that the first event information corresponding to the above question information is "environmental pollution tort", the supervision model can determine that the legal information types corresponding to "environmental pollution tort" include tortious acts, damage facts, and causal relationships, and determine the supervision vectors corresponding to the above legal information types. In addition, since the burden of proof is reversed in the causal relationship of "environmental pollution infringement" incidents, when a user suffers losses due to environmental pollution and becomes the injured party, there is no need to extract information related to the fault of the infringer during the legal information extraction process, and the attention to the causal relationship should be low, or the causal relationship may not be paid attention to, that is, the supervision model can mark the causal relationship in the legal information type corresponding to the above-mentioned first event information as having low attention.

[0054] Step 230 , extracting features of the legal elements in the question information to obtain multiple element features corresponding to the question information.

[0055] In some embodiments, the content model in the information extraction model can extract different types of legal elements in the question information to obtain multiple element features corresponding to the question information. That is, each element feature in the multiple element features corresponds to a different type. Among them, legal elements refer to common content existing in different legal fields. For example, multiple types of legal elements can include subject, object, behavior, object, subjective intention, etc.

[0056] Continuing with the above example, after feature extraction of the question information input by the user in the content model, multiple different element features can be obtained, such as user-infringed party, chemical plant-infringer, environmental pollution infringement, chemical plant legality, climate-induced changes in production, etc. These element features cover all possible entities in the question information and the relationships between entities.

[0057] Step 240: determine multiple first entities corresponding to the question information based on the supervision vector corresponding to the first event information and multiple element features, and determine entity relationships between the multiple first entities based on the multiple first entities.

[0058] In some embodiments, after obtaining the multiple element features and supervision vectors corresponding to the question information, the content model can further fuse the multiple element features and supervision vectors to obtain content features corresponding to the question information. The content features represent the multiple first entities corresponding to the question information.

[0059] Following the above example, after obtaining multiple element features and supervision vectors corresponding to the above question information, the content model can remotely supervise each element feature through the supervision vector, and assign weights corresponding to each element feature through the feature fusion layer in the content model, such as giving a larger weight to the element features related to the infringement and damage facts, and giving a smaller weight to the element features related to the causal relationship, and the remaining features can be ignored. Finally, the multiple element features after weight adjustment are fused to obtain the content features. The obtained content features focus on the two element features of the chemical plant's pollution discharge and the user's fish pond property loss, while the characterization of the element features of the causal relationship of climate causing product changes is relatively weak, indicating that there is no need to pay too much attention to the causal relationship in the subsequent reasoning process. In the above way, the content features output by the content model represent the key constituent elements in the environmental pollution infringement relationship between the user and the chemical plant, and exclude irrelevant entities and entity relationships such as the chemical plant's self-defense, so as to further improve the accuracy of subsequent legal event detection.

[0060] In some embodiments, after determining multiple first entities corresponding to the question information through the supervision model and the content model in the information extraction model, the entity relationship between the first entities is determined based on the multiple first entities corresponding to the question information through the graph neural network GNN in the information extraction model.

[0061] Exemplarily, the graph neural network GNN can construct a corresponding graph representation based on multiple first entities in the question information and the entity relationships between the first entities. Among them, each node in the graph representation is used to represent each first entity, and each edge in the graph representation is used to represent the entity relationship between the first entities. The graph neural network GNN can use the information transmission mechanism to update the features of each node (entity) in the graph representation one by one, that is, based on the neighbor nodes corresponding to each node (that is, the nodes connected to it by edges) and the relationship between the neighbor nodes, the features of each node are updated. After multiple information transmissions, the global features of each node are finally obtained; the graph neural network GNN can predict the relationship between the first entities based on the global features of the first entities to obtain the entity relationship between the first entities.

[0062] Step 250: Determine second event information corresponding to the question information based on the multiple first entities and the entity relationships between the first entities.

[0063] In some embodiments, after obtaining multiple first entities corresponding to the question information and the entity relationships between the first entities through the information extraction model, the multiple first entities and the entity relationships between the first entities can be further input into the legal event detection model, so that the legal event detection model can be used to perform legal event detection on the question information again based on the multiple first entities and the entity relationships between the first entities to determine the second event information corresponding to the question information.

[0064] Exemplarily, when the legal event detection model performs legal event detection on question information, it can extract the contextual relationship corresponding to each entity (such as the first entity) in the question information (that is, the initial features corresponding to each entity); then, it updates the initial features corresponding to each entity by adopting hierarchical propagation and aggregation of neighboring nodes to obtain the core information corresponding to each entity and the core relationship between each entity; then, based on the core information corresponding to each entity and the core relationship between each entity, it determines the legal event information (such as the second event information) corresponding to the question information.

[0065] Step 260, when the second event information is the same as the first event information, multiple first entities are determined as multiple target entities corresponding to the question information, and the entity relationship between the first entities is determined as the entity relationship between the target entities in the multiple target entities corresponding to the question information.

[0066] In some embodiments, when the second event information is the same as the first event information, the multiple first entities can be determined as multiple target entities corresponding to the question information, and the entity relationship between the first entities can be determined as the entity relationship between the target entities in the multiple target entities corresponding to the question information. In addition, the central control system can further search for legal cases or legal provisions based on the second event information to provide legal advice to the user.

[0067] In the case where the second event information is different from the first event information, the information extraction model can be used to extract legal information (including entities and entity relationships) from the question information again based on the second event information, and the third legal event information corresponding to the question information can be determined based on the second entity and the entity relationship between the second entities obtained by the information extraction through the legal event detection model. In the case where the third event information is the same as the second event information, multiple second entities are determined as multiple target entities corresponding to the question information, and the entity relationship between each second entity is determined as the entity relationship between each target entity in the multiple target entities corresponding to the question information; in the case where the third event information is different from the second event information, the information extraction model is used to continue to extract legal information based on the third event information... and so on, until the legal event information determined by the previous legal event detection is the same as the legal event information determined by the next legal event detection.

[0068] Through the above scheme, different types of legal element features in the question information can be extracted through the content model in the information extraction model, and various types of element features can be remotely supervised through the supervision model in the information extraction model to obtain content features. The content features can more accurately represent the entity information corresponding to the question information, thereby improving the accuracy of entity and entity relationship extraction in the legal consultation scenario. In addition, the GNN model in the information extraction model can realize the reasoning and extraction of complex entity relationships, improving the accuracy and effectiveness of legal information extraction. In addition, by establishing an iterative relationship between the information extraction model and the legal information detection model, the accuracy and effectiveness of legal information extraction can be further ensured by comparing the results of the previous and subsequent legal event detections.

[0069] In some embodiments, the above step 230 includes: extracting features of various types of legal elements in the question information through multiple element extraction models to obtain multiple element features.

[0070] In some embodiments, the content model in the information extraction model may include multiple element extraction models. Each element extraction model in the multiple element extraction models is used to extract different types of legal elements in the question information to obtain multiple different types of element features corresponding to the question information.

[0071] For example, in the process of classifying the legal elements in the question information, each element model can automatically complete the classification based on the training of a large amount of training data, so that each element model focuses on specific element features. On this basis, the identification of the element features that need to be identified can be set for each element model through manual annotation to enhance the classification effect of the model.

[0072] Through the above scheme, by training multiple feature extraction models dedicated to extracting corresponding types of feature features, different types of feature features in question information can be extracted through different feature extraction models, thereby improving the accuracy of feature feature extraction.

[0073] Figure 3 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application, such as Figure 3 As shown, in the above step 240 , “determining multiple first entities corresponding to the question information based on the supervision vector corresponding to the first event information and multiple element features” may include steps 310 to 330 .

[0074] Step 310 , concatenate or dot-multiply the supervision vector corresponding to the first event information with each element feature to obtain a plurality of first element features.

[0075] In some embodiments, after obtaining the above-mentioned supervision vector and multiple element features, the content model can concatenate or dot-multiply each element feature with the supervision vector to obtain multiple first element features, thereby realizing remote supervision of each element feature through the supervision vector. Among them, concatenating each element feature with the supervision vector can retain the original element feature on the basis of adding the supervision vector; dot-multiplying each element feature with the supervision vector can strengthen or weaken each element feature through the supervision vector, that is, increase or decrease the weight of each element feature.

[0076] For example, if the question information entered by the user is "The discharge of pollutants from the upstream chemical plant caused a sharp drop in the output of my fish pond, but the chemical plant claims that the discharge is legal", if the first event information is "environmental pollution infringement", the supervision vector will vectorize the legal event type corresponding to "environmental pollution infringement" (including infringement behavior, damage facts and causal relationship), and the obtained supervision vector S =[ s 1 , s 2 ,…, s n ] If the element features obtained by each element extraction model after processing the question information are: Chemical Plant F 1 =[ f 11 , f 12 ,…, f 1 n ] , fish pond F 2 =[ f 21 , f 22 ,…, f 2 n ] , sewage discharge F 3 =[ f 31 , f 32 ,…, f 3 n ] etc., then for the above-mentioned supervision vectors and each element feature perform splicing processing, and the spliced first element features can be obtained F i ' =[ F i ; S ] ; for the above-mentioned supervision vectors and each element feature perform dot product processing, and the dot producted first element features can be obtained F i ' =[ F i ∙ S ] .

[0077] Step 320, adopt the attention mechanism to adjust the weights of each first element feature among multiple first element features, and obtain multiple second element features with adjusted weights.

[0078] In some embodiments, the content model may further include a feature fusion layer. The feature fusion layer can adopt the attention mechanism to assign weights to each first element feature according to the importance degree of each first element feature, and obtain multiple second element features with adjusted weights.

[0079] Exemplarily, the feature fusion layer can adopt the attention mechanism to assign higher weights to the first element features with higher importance degrees among each first element feature (such as infringement acts and damage facts), and assign lower weights to the first element features with lower importance degrees (such as causal relationships), so as to improve the accuracy of legal information extraction and further improve the accuracy of legal event detection.

[0080] Step 330 , performing feature fusion processing on the multiple second element features to obtain content features, and determining multiple first entities according to the content features.

[0081] In some embodiments, after obtaining the above-mentioned multiple second factor features, the multiple second factor features are subjected to feature fusion processing through a feature fusion layer to obtain content features. The content features can characterize multiple first entities in the question information, such as "upstream chemical plant", "sharp drop in fish pond production", "pollution discharge", "legal discharge", etc.

[0082] Through the above scheme, after obtaining multiple element features corresponding to the question information through multiple element extraction models, the weight of each element feature can be further adjusted through the supervision vector and attention mechanism to strengthen or weaken the representation of each element feature, thereby improving the accuracy of legal information extraction.

[0083] Figure 4 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application, such as Figure 4 As shown, in the above step 330 , “determining multiple first entities according to content characteristics” may include steps 410 to 440 .

[0084] Step 410 , sequence labeling is performed on the content features to obtain a plurality of label sequences corresponding to the content features.

[0085] In some embodiments, based on the above-mentioned feature fusion layer, a conditional random field (CRF) layer can also be set in the content model. After obtaining the content features through the feature fusion layer, the content features can also be sequenced through the CRF layer to obtain multiple different sequence labeling results, that is, multiple different label sequences. Each label sequence includes multiple label information.

[0086] For example, taking the content feature "the discharge of pollutants from the upstream chemical plant caused a sharp drop in the output of the fish pond, but the chemical plant claimed that the discharge was legal" as an example, if the preset label set includes B-subject (the entity of the category subject is marked as B), B-behavior, B-damage fact, I-behavior (internal supplementary information of the behavior) and O-non-entity, then the CRF layer can obtain multiple label sequences by sequentially labeling the above content features according to the above preset label set, including:

[0087] Y 1 ={B-subject (upstream chemical plant), B-behavior (discharge of pollutants), O-non-entity (cause), B-damage fact (sharp drop in fish pond output), B-subject (upstream chemical plant), O-non-entity (claim), I-behavior (legal discharge)};

[0088] Y2 ={B-subject (upstream chemical plant), B-subject (discharge), B-behavior (cause), B-damage fact (sharp drop in fish pond output), B-subject (upstream chemical plant), B-subject (claim), I-behavior (legal discharge)};

[0089] Y 3 ={B-behavior (upstream chemical plant), B-subject (pollution discharge), O-non-entity (cause), I-behavior (sharp drop in fish pond output), B-subject (upstream chemical plant), O-non-entity (claim), B-damage fact (legal discharge)}, etc.

[0090] Step 420: Determine the transition probability corresponding to each label sequence based on the plurality of label information corresponding to each label sequence.

[0091] In some embodiments, after obtaining multiple label sequences, the CRF layer can determine the transition probability (or transition score) between multiple label information corresponding to each label sequence. The transition probability is used to characterize the dependency relationship between multiple label information corresponding to each label sequence. The transition probability reflects the semantic logical dependency between each label information in the sequence labeling task by quantifying this dependency relationship. For example, if the transition probability corresponding to two label information is high, it means that the probability of the two label information appearing continuously is high, that is, it has a strong semantic association; if the transition probability corresponding to the two label information is low, it means that the continuity of the two label information is low, and usually there is no logical consistency.

[0092] For example, the transition probability between multiple label information can be determined based on the trained transition probability matrix, where the transition probability matrix is ​​used to represent the transition rules between the label information. The transition probability matrix A is a K×K matrix, where K is the number of preset label information in the preset label set; the element A[i,j] of the transition probability matrix represents the transition from label information y t-1 =i transfer to label information y t =j’s transfer score. For example, if the label sequence is Y 1 ={B-subject (upstream chemical plant), B-behavior (pollutant discharge), B-damage fact (sharp drop in fish pond output), B-subject (upstream chemical plant), I-behavior (legal discharge)}, then A[B-subject, B-behavior] represents the probability of transfer from B-subject (upstream chemical plant) to B-behavior (pollutant discharge).

[0093] For example, the transition probability matrix is:

[0094] For example, the label sequence Y 1={B-subject (upstream chemical plant), B-behavior (discharge), O-non-entity (cause), B-damage fact (sharp drop in fish pond output), B-subject (upstream chemical plant), O-non-entity (claim), I-behavior (legal discharge)} The corresponding transfer probability is: A[B-subject, B-behavior]+A[B-behavior, O-non-entity]+A[O-non-entity, B-damage fact]+A[B-damage fact, B-subject]+A[B-subject, O-non-entity]+A[O-non-entity, I-behavior]==0.8+0.1+0.2+0.0+0.1+0.0=1.2;

[0095] Tag sequence Y 2 ={B-subject (upstream chemical plant), B-subject (discharge), B-behavior (cause), B-damage fact (sharp drop in fish pond output), B-subject (upstream chemical plant), B-subject (claim), I-behavior (legal discharge)} The corresponding transfer probability is: A[B-subject, B-subject]+A[B-subject, B-behavior]+A[B-behavior, B-damage fact]+A[B-damage fact, B-subject]+[B-subject, B-subject]+A[B-subject, I-behavior]=0.0+0.8+0.0+0.0+0.0+0.0=0.8;

[0096] Tag sequence Y 3 ={B-behavior (upstream chemical plant), B-subject (pollution discharge), O-non-entity (cause), I-behavior (sharp drop in fish pond output), B-subject (upstream chemical plant), O-non-entity (claim), B-damage fact (legal discharge)} The corresponding transfer probability is: A[B-behavior, B-subject]+A[B-subject, O-non-entity]+A[O-non-entity, I-behavior]+A[I-behavior, B-subject]+A[B-subject, O-non-entity]+A[O-non-entity, B-damage fact]=0.5+0.1+0.0+0.0+0.1+0.2=0.9.

[0097] Step 430 : determining the conditional probability corresponding to each label sequence based on the transition probability corresponding to each label sequence.

[0098] In some embodiments, after determining the transition probability corresponding to each label sequence, the CRF layer can determine the conditional probability corresponding to each label sequence based on the transition probability corresponding to each label sequence. Exemplarily, the score function corresponding to each label sequence can be determined based on the transition probability corresponding to each label sequence; then the score function corresponding to each label sequence is log-likelihood estimated to determine the conditional probability corresponding to each label sequence. The score function is used to quantify the rationality of the label sequence relative to the input sequence (i.e., the content feature), that is, the degree of matching between the label sequence and the content feature.

[0099] In some examples, the conditional probability corresponding to each tag sequence can be determined according to formula (1):

[0100] ; (1)

[0101] in, and is the label information in the label sequence, For content characteristics. That is, the label sequence Y={y 1 ,y 2 ,…,y T}The corresponding score function Score(X,Y). is a characteristic function, which can be determined based on the transition probability corresponding to the label sequence Y. Exponential function The log-likelihood used to calculate the score function uses an exponential function to ensure that all probability values ​​are non-negative, and the logarithmic property can be used to simplify the calculation of the CRF layer.

[0102] is the total score of all label sequences corresponding to the content features, It means that the total score is normalized to ensure the legitimacy of the conditional probability, that is, to ensure that the sum of the conditional probabilities of all label sequences is 1.

[0103] Step 440 : determining a target tag sequence from a plurality of tag sequences based on the conditional probabilities corresponding to the tag sequences, and determining a plurality of first entities according to the target tag sequence.

[0104] In some embodiments, after determining the conditional probabilities corresponding to each tag sequence, the CRF layer may use the Viterbi Algorithm to determine the target tag sequence with the largest conditional probability among multiple tag sequences. ∗ In the process, the CRF layer calculates the label sequence with the highest conditional probability based on the transition dependency (i.e., transition probability) between each label information in the label sequence and the global information of the content features using an efficient dynamic programming method (such as the forward-backward algorithm). The target label sequence can represent multiple first entities corresponding to the question information.

[0105] Through the above process, for legal information extraction, the entities in the input consulting text often have complex contextual relationships and dependencies. Since the CRF layer can effectively utilize the dependencies between label information, it can greatly improve the accuracy and consistency of entity information extraction.

[0106] Figure 5A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application, such as Figure 5 As shown, the above step 250 may include steps 510 to 530.

[0107] Step 510: Based on the question information and the multiple first entities, determine the question information features corresponding to the question information and the initial features corresponding to the first entities.

[0108] In some embodiments, the context information corresponding to each first entity in the question information can be extracted by the BERT model in the legal event detection model, and the context information corresponding to each first entity can be determined as the initial features corresponding to each first entity. The initial features corresponding to each first entity determined by the BERT model usually cover a wide range, that is, they usually cover all possible contextual relationships related to each first entity in the question information.

[0109] Using the above example, the initial features corresponding to the first entities in the question information through the BERT model can be: Zhang San: [0.1, 0.2, 0.3, ...], Li Si: [0.2, 0.3, 0.4, ...], 100,000 RMB: [0.3, 0.4, 0.5, ...], one year: [0.4, 0.5, 0.6, ...]. It should be noted that the values ​​in the initial features corresponding to the first entities are only examples.

[0110] In some embodiments, the question information may also be vectorized using a BERT model to obtain question information features corresponding to the question information.

[0111] Step 520 , based on the multiple first entities, the entity relationships between the first entities, and the initial features corresponding to the first entities, determine the core information corresponding to the question information.

[0112] In some embodiments, since the initial features corresponding to each first entity obtained by the BERT model usually cover a wide range, the initial features corresponding to each first entity can be processed by a multi-layer graph convolutional network model (Graph Convolutional Networks, GCN) in the legal event detection model, so that the initial features corresponding to each first entity gradually converge to the core information corresponding to each first entity and the core relationship between each first entity.

[0113] Exemplarily, the initial features corresponding to each first entity can be updated layer by layer through each layer of GCN based on the syntactic graph and the initial features corresponding to each first entity, by means of hierarchical propagation and aggregation of neighboring nodes. The syntactic graph can be determined based on the first entity and the entity relationship between each first entity. Finally, the updated features corresponding to each first entity obtained by the last layer of GCN are determined as the core information corresponding to the question information (i.e., the core information corresponding to each first entity and the core relationship between each first entity).

[0114] Step 530: Determine the second event information based on the question information feature and the core information corresponding to the question information.

[0115] In some embodiments, the second event information corresponding to the question information can be determined through a classification model in a legal event detection model based on the question information features determined by a BERT model and the core information corresponding to the question information determined by GCN.

[0116] Exemplarily, the core information corresponding to the question information and the question information features corresponding to the question information determined by the BERT model can be fused to obtain the fused features and input them into the classification model; then the fused features are classified by the classification model, and the classification result obtained is the legal event information corresponding to the question information (i.e., the second event information).

[0117] Through the above scheme, the contextual relationship corresponding to each first entity in the question information input by the user (that is, the initial features corresponding to each first entity) can be extracted, and the initial features corresponding to each first entity can be updated by means of hierarchical propagation and aggregation of neighboring nodes, so as to finally obtain the core information corresponding to each first entity and the core relationship between each first entity, and then, based on the core information corresponding to each first entity and the core relationship between each first entity, the legal event information corresponding to the question information can be accurately determined.

[0118] Figure 6 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application, such as Figure 6 As shown, the above step 520 may include steps 610 to 630.

[0119] Step 610 , determining a syntactic graph using the first entities as nodes and the entity relationships between the first entities as edges.

[0120] In some embodiments, after obtaining the above-mentioned multiple first entities and the entity relationships between each first entity, a model can be constructed through a syntactic graph in a legal event detection model, and a syntactic graph can be constructed based on multiple first entities and the entity relationships between each first entity.

[0121] Exemplarily, when constructing the syntax graph, each first entity is used as a node of the syntax graph, and the entity relationship between each first entity is used as an edge between each node. Figure 7 A schematic diagram of a syntax diagram is shown in Figure 7 As shown in the figure, taking the question information "Zhang San and Li Si are neighbors. Zhang San borrowed RMB 100,000 from Li Si, and the two parties agreed to repay the loan in one year. One year later, Zhang San did not repay the loan on time and ignored Li Si's reminders", the corresponding multiple nodes in the syntactic graph are: Zhang San, Li Si, RMB 100,000, and one year; the edges in the syntactic graph include: Zhang San---Neighbor---> Li Si, Zhang San---Loan---> RMB 100,000-Li Si, Zhang San---Breach of Contract---> One Year-Li Si, Li Si---Reminders---> Zhang San, Zhang San---Compensation Request---> Li Si.

[0122] Exemplarily, the syntactic graph can represent all the relationships between the first entities (nodes), and there may be some non-core relationships among them, such as "Zhang San and Li Si are neighbors", etc. Therefore, the entity relationships between the first entities obtained can be subsequently processed by the third neural network model, so that the entity relationships between the first entities gradually converge to the core relationships between the first entities, so as to exclude the non-core relationships between the first entities.

[0123] Step 620, updating the initial features corresponding to each first entity based on the syntactic graph and the initial features corresponding to each first entity through multiple neural network layers, and obtaining updated features of each node determined by each neural network layer in the multiple neural network layers.

[0124] In some embodiments, after obtaining the above-mentioned syntactic graph and the initial features corresponding to each first entity determined by the BERT model, the initial features corresponding to each first entity can be updated based on the syntactic graph and the initial features corresponding to each first entity through a multi-layer graph convolutional network (GCN) in the legal event detection model.

[0125] Exemplarily, the features corresponding to each first entity can be updated by using a multi-layer GCN by hierarchical propagation and aggregation of neighbor node features. Taking a target node among multiple nodes as an example, the current features corresponding to the target node can be updated by using the previous layer of GCN in the multi-layer GCN based on the current features corresponding to the target node and the current features corresponding to the neighbor nodes (i.e., adjacent nodes) of the target node, to obtain the updated features corresponding to the target node; and so on, until the updated features corresponding to the target node are obtained by using the last layer of GCN in the multi-layer GCN. Among them, the current features corresponding to the target node in the first layer of GCN are the initial features corresponding to the target node.

[0126] Step 630, determining the updated features of each node determined by the last neural network layer in the multiple neural network layers as the core information corresponding to each first entity and the core relationship between each first entity.

[0127] In some embodiments, the updated features corresponding to each node obtained by the last layer GCN in the above-mentioned multi-layer GCN are the core information corresponding to each node and the core relationship between each first entity, that is, the core information corresponding to each first entity and the core relationship between each first entity (that is, the corresponding core information).

[0128] Through the above scheme, the initial features corresponding to each first entity output by the BERT model are updated based on the syntactic graph through multi-layer GCN, which can exclude non-core first entity information and non-core relationships between the first entities, so that better classification results can be obtained when legal events are subsequently classified based on the first entity information and entity relationships.

[0129] Figure 8 A flowchart of another method for determining entities and entity relationships provided in some embodiments of the present application, such as Figure 8 As shown, the above step 620 may include steps 810 to 820.

[0130] Step 810, obtaining the first feature of each node determined by the first neural network layer based on the initial features corresponding to each node, the adjacency matrix corresponding to the syntactic graph, the self-connection matrix and the weight matrix corresponding to the first neural network layer through the first neural network layer among the multiple neural network layers.

[0131] In some embodiments, taking a target node among multiple nodes and the first two layers of GCN in a multi-layer GCN as an example, the first layer of GCN (i.e., the first neural network layer) first updates the initial features corresponding to the target node based on the initial features corresponding to the target node and the initial features corresponding to the adjacent nodes of the target node, and obtains the updated initial features corresponding to the target node (i.e., the first features).

[0132] Exemplarily, the first layer GCN can update the initial features corresponding to the target node based on the adjacency matrix corresponding to the syntactic graph, the self-loop matrix, the weight matrix corresponding to the first layer GCN, and the nonlinear activation function to obtain the first features corresponding to the target node.

[0133] In some examples, the method of updating the features of each node through multi-layer GCN can refer to formula (1):

[0134] ; (1)

[0135] in, For the The current features corresponding to each node determined by the layer GCN ( That is, the initial features corresponding to each node); For the The updated features corresponding to each node of the layer.

[0136] , is the N×N adjacency matrix corresponding to the syntactic graph, which is used to represent the connection between each node and other nodes; if there is an edge between node i and node j, then [i][j]=1, if there is no edge between node i and node j, then [i][j]=0; is an N×N identity matrix, where N is the number of nodes; It means that each node in the syntactic graph is connected to itself, that is, the characteristics of the node itself are taken into account.

[0137] yes The degree matrix is ​​a diagonal matrix where each diagonal element [i][i] represents the degree of node i, that is, the number of edges connected to node i.

[0138] is the degree matrix The inverse square root of Normalization is performed to avoid the impact of large differences in node degrees on model training.

[0139] For the The weight matrix corresponding to the layer GCN is a trainable parameter matrix used to linearly transform node features.

[0140] A non-linear activation function, such as the Rectified Linear Unit (ReLU), enables the third neural network model to capture complex patterns.

[0141] It means that the first The current feature corresponding to the node determined by the layer GCN To disseminate and aggregate; Indicates that the node features after propagation and aggregation are linearly transformed to obtain the node features after linear transformation; the updated node features are compared with Multiplying means applying a nonlinear activation function to the linearly transformed node features to increase the expressive power of the model, thereby obtaining the updated node features .

[0142] Step 820, obtaining the second feature of each node determined by the second neural network layer based on the first feature of each node, the adjacency matrix corresponding to the syntactic graph, the self-connection matrix and the weight matrix corresponding to the second neural network layer through the second neural network layer among the multiple neural network layers.

[0143] In some embodiments, after the first layer GCN obtains the first feature corresponding to the target node, the second layer GCN (i.e., the second neural network layer) updates the first feature corresponding to the target node based on the first feature corresponding to the target node and the first features corresponding to the adjacent nodes of the target node to obtain the updated first feature corresponding to the target node (i.e., the second feature).

[0144] Through the above scheme, the multi-layer GCN structure can realize the layer-by-layer update of the relationship between entities represented by the syntactic graph and the entity context relationship represented by the initial features of the entity, thereby improving the convergence effect of the core information and core relationship of the entity.

[0145] In some embodiments, the above step 530 includes: fusing the question information feature and the core information corresponding to the question information to obtain a fused feature; and classifying the legal events corresponding to the fused feature to obtain second event information.

[0146] In some embodiments, after obtaining the core information corresponding to each first entity and the core features between each first entity (i.e., the core information corresponding to the question information) through multi-layer GCN, the question information features corresponding to the question information determined based on the BERT model and the core information corresponding to the question information can be fused to obtain a fused feature representation. Exemplarily, the fusion of the question information features and the core information corresponding to the question information can be achieved through concatenation, weighted summation, or attention mechanism. The specific fusion method is not limited in this embodiment.

[0147] In some embodiments, after determining the above fusion features, the legal events corresponding to the fusion features can be classified through a classification model to obtain one or more target event information. The classification model can be implemented based on a fully connected (FC) layer, etc., which is not limited in this embodiment.

[0148] For example, for the question information input by the user, "Zhang San and Li Si are neighbors. Zhang San borrowed 100,000 yuan from Li Si, and the two parties agreed to repay the loan in one year. One year later, Zhang San did not repay the loan on time and ignored Li Si's reminder for payment", after being processed by the legal event detection model, it can be obtained that the target event information corresponding to the question information is "not repaying debts".

[0149] Through the above scheme, the question information features corresponding to the question information and the core information corresponding to the question information are fused, so that the fourth neural network model can use the question information as a reference in the process of classification prediction, avoiding the problem of not being able to obtain accurate legal event detection results when the output of the multi-layer GCN deviates from the question information.

[0150] In some embodiments, the above method also includes: when the second event information is different from the first event information, determining multiple second entities corresponding to the question information and the entity relationship between each second entity among the multiple second entities based on the second event information; determining third event information corresponding to the question information based on the multiple second entities and the entity relationship between each second entity; when the third event information is the same as the second event information, determining the multiple second entities as multiple target entities corresponding to the question information, and determining the entity relationship between each second entity as the entity relationship between each target entity among the multiple target entities corresponding to the question information.

[0151] In some embodiments, when the second event information is different from the first event information, the information extraction model can extract legal information (including entities and entity relationships) from the question information again based on the second event information, so as to determine the legal event information corresponding to the question information (i.e., the third legal event information) based on the second entity and the entity relationship between the second entities obtained by the information extraction through the legal event detection model again; and when the third event information is the same as the second event information, multiple second entities are determined as multiple target entities corresponding to the question information, and the entity relationship between the second entities is determined as the entity relationship between the target entities in the multiple target entities corresponding to the question information; when the third event information is different from the second event information, continue to extract legal information based on the third event information through the information extraction model... and so on, until the event information determined by the previous legal event detection is the same as the event information determined by the next legal event detection.

[0152] Through the above scheme, the information extraction results obtained by the information extraction model can be fed back to the legal event detection model to re-determine the legal event information, so as to compare the legal event information determined in the previous legal event detection process with the legal event information determined in the current legal event detection process, thereby improving the accuracy of legal event detection.

[0153] Fig. 9 A flowchart of a method for training an information extraction model provided in some embodiments of the present application, wherein: Fig. 9 The method shown can be applied to Figure 1 The information extraction model shown in Fig. 9As shown, the training method of the information extraction model may include steps 910 to 950.

[0154] Step 910, obtaining the training text, the target entity corresponding to the training text, and the legal event information.

[0155] In some embodiments, when training the information extraction model, case facts in public documents such as legal consultation, legal documents, and teaching materials can be obtained as training texts, and the real entity information (i.e., target entity) corresponding to the training texts can be determined as training labels for the case facts. In addition, legal event information obtained by performing legal event detection on other question information can also be obtained.

[0156] Step 920, vectorize the legal event information to obtain a supervision vector, and extract features of the legal elements in the training text to obtain multiple element features corresponding to the training text.

[0157] In some embodiments, after the training text, the target entity corresponding to the training text, and the legal event information are input into the information extraction model, the information extraction model vectorizes the legal event information through the supervision model to obtain a supervision vector, and performs feature extraction on the legal elements in the training text through multiple element extraction models in the content model to obtain multiple element features corresponding to the training text.

[0158] It can be understood that the implementation of step 920 can refer to the description of step 230 and will not be repeated here.

[0159] Step 930 , determining multiple predicted entities in the training text based on the supervision vector and multiple element features.

[0160] In some embodiments, the information extraction model then predicts entities in the training text based on the above-mentioned supervision vector and multiple element features through the content model to obtain multiple predicted entities corresponding to the training text.

[0161] It can be understood that the implementation of step 930 can refer to the description of steps 310 to 330, which will not be repeated here.

[0162] Step 940: Determine a first loss function corresponding to the information extraction model based on the multiple prediction entities and the multiple target entities.

[0163] In some embodiments, thereafter, based on the above-mentioned multiple predicted entities and target entities, the binary cross entropy loss between the predicted legal information and the actual legal information can be calculated to determine the loss function (ie, the first loss function) corresponding to the content model.

[0164] Exemplarily, the first loss function corresponding to the content model can be determined according to formula (2):

[0165] L 1 =- N ∑ i =1 N [ y i log y i ̂ +(1- y i ) l og (1- y i ̂ )] ; (2)

[0166] in, is the first loss function; N is the number of sentences in the question information; is the target entity corresponding to the i-th sentence; is the predicted entity corresponding to the i-th sentence.

[0167] Step 950, optimizing the parameters in the first loss function to determine the trained information extraction model.

[0168] In some embodiments, after obtaining the first loss function corresponding to the above-mentioned content model, the relevant parameters in the first loss function can be optimized to obtain the trained content model.

[0169] Fig.10 A flowchart of another information extraction model training method provided in some embodiments of the present application, such as Fig.10 As shown, the above method also includes steps 1010 to 1030.

[0170] Step 1010, obtaining a true label sequence corresponding to the training text, and determining a conditional probability corresponding to the true label sequence based on the true label sequence.

[0171] In some embodiments, in addition to training the content model in the information extraction model, the CRF layer in the information extraction model can also be trained. When training the CRF layer, first obtain the real label sequence corresponding to the training text, and based on the target label sequence corresponding to the training text, determine the conditional probability corresponding to the target label sequence The conditional probability corresponding to the target tag sequence may be determined based on the transition probability corresponding to the target tag sequence. It is understandable that the implementation of step 1010 may refer to the description of steps 410 to 430, which will not be described in detail here.

[0172] Step 1020: Determine a second loss function corresponding to the information extraction model based on the conditional probability.

[0173] In some embodiments, the second loss function corresponding to the CRF layer can be determined according to formula (3):

[0174] (3)

[0175] in, is the second loss function; is a true label sequence, For training text; is the conditional probability corresponding to the true label sequence.

[0176] In some embodiments, the above step 950 includes: optimizing parameters in the first loss function and the second loss function to determine the trained information extraction model.

[0177] Exemplarily, after determining the first loss function corresponding to the content model and the second loss function corresponding to the CRF layer, the overall loss function can be determined based on the first loss function and the second loss function. The overall loss function can be determined by the first loss function, the second loss function and the corresponding weights. The weight corresponding to the first loss function and the weight corresponding to the second loss function can be adjusted according to the performance of the information extraction model on the validation set, thereby optimizing the information extraction model information training to obtain the trained information extraction model.

[0178] By applying the above technical solution, different types of legal element features in the question information can be extracted through the content model in the information extraction model, and multiple types of element features can be remotely supervised through the supervision model in the information extraction model to obtain content features. The content features can more accurately represent the entity information corresponding to the question information, thereby improving the accuracy of entity and entity relationship extraction in legal consultation scenarios.

[0179] Fig.11 A schematic diagram of a device for determining entities and entity relationships provided in some embodiments of the present application. Fig.11 As shown, the apparatus 1100 for determining entities and entity relationships includes a first acquisition module 1110 , a first processing module 1120 , a feature extraction module 1130 , a first determination module 1140 , a second determination module 1150 and a third determination module 1160 .

[0180] The first acquisition module 1110 is configured to acquire question information input by the user and first event information corresponding to the question information.

[0181] The first processing module 1120 is configured to perform vectorization processing on the first event information to obtain a supervision vector corresponding to the first event information.

[0182] The feature extraction module 1130 is configured to extract features of the legal elements in the question information to obtain multiple element features corresponding to the question information.

[0183] The first determination module 1140 is configured to determine multiple first entities corresponding to the question information based on the supervision vector corresponding to the first event information and multiple element features, and determine the entity relationship between each first entity in the multiple first entities based on the multiple first entities.

[0184] The second determination module 1150 is configured to determine the second event information corresponding to the question information based on the multiple first entities and the entity relationships between the first entities.

[0185] The third determination module 1160 is configured to determine the multiple first entities as the multiple target entities corresponding to the question information when the second event information is the same as the first event information, and determine the entity relationship between the first entities as the entity relationship between the target entities in the multiple target entities corresponding to the question information.

[0186] In some embodiments, the feature extraction module 1130 is specifically configured to: extract features of various types of legal elements in the question information through multiple element extraction models in the information extraction model to obtain multiple element features; wherein each element model in the multiple element models is used to extract different types of element features.

[0187] In some embodiments, the first determination module 1140 is specifically configured to: concatenate or dot-multiply the supervision vector and each element feature to obtain multiple first element features; use an attention mechanism to adjust the weight of each first element feature in the multiple first element features to obtain multiple second element features after weight adjustment; perform feature fusion processing on the multiple second element features to obtain content features, and determine multiple first entities based on the content features.

[0188] In some embodiments, the first determination module 1140 is specifically configured to: perform sequence annotation on content features to obtain multiple label sequences corresponding to the content features; wherein each label sequence in the multiple label sequences includes multiple label information; based on the multiple label information corresponding to each label sequence, determine the transition probability corresponding to each label sequence; wherein the transition probability is used to characterize the dependency relationship between the multiple label information corresponding to each label sequence; based on the transition probability corresponding to each label sequence, determine the conditional probability corresponding to each label sequence; based on the conditional probability corresponding to each label sequence, determine the target label sequence in the multiple label sequences, and determine multiple first entities according to the target label sequence.

[0189] In some embodiments, the second determination module 1150 is specifically configured to: determine the question information features corresponding to the question information and the initial features corresponding to each first entity based on the question information and multiple first entities; wherein the initial features corresponding to each first entity are used to characterize the context information corresponding to each first entity; determine the core information corresponding to the question information based on the multiple first entities, the entity relationships between the first entities and the initial features corresponding to the first entities; determine the second event information based on the question information features and the core information corresponding to the question information.

[0190] In some embodiments, the core information corresponding to the question information includes the core information corresponding to each first entity in the question information and the core relationship between each first entity; the second determination module 1150 is specifically configured to: determine a syntactic graph with each first entity as a node and the entity relationship between each first entity as an edge; update the initial features corresponding to each first entity based on the syntactic graph and the initial features corresponding to each first entity through multiple neural network layers, and obtain updated features of each node determined by each neural network layer in the multiple neural network layers; determine the updated features of each node determined by the last neural network layer in the multiple neural network layers as the core information corresponding to each first entity and the core relationship between each first entity.

[0191] In some embodiments, the second determination module 1150 is specifically configured as follows: obtaining the first feature of each node determined by the first neural network layer through the first neural network layer among multiple neural network layers based on the initial features corresponding to each node, the adjacency matrix corresponding to the syntactic graph, the self-connection matrix and the weight matrix corresponding to the first neural network layer; wherein the first neural network layer is the first neural network layer among the multiple neural network layers; obtaining the second feature of each node determined by the second neural network layer through the second neural network layer among the multiple neural network layers based on the first features of each node, the adjacency matrix corresponding to the syntactic graph, the self-connection matrix and the weight matrix corresponding to the second neural network layer; wherein the updated features of each node include the first feature of each node and the second feature of each node.

[0192] In some embodiments, the second determination module 1150 is specifically configured to: fuse the question information feature and the core information corresponding to the question information to obtain a fused feature; and classify the legal events corresponding to the fused feature to obtain second event information.

[0193] In some embodiments, the third determination module 1160 is also configured to: when the second event information is different from the first event information, determine the multiple second entities corresponding to the question information and the entity relationship between each second entity among the multiple second entities based on the second event information; the second determination module 1150 is also configured to: determine the third event information corresponding to the question information based on the multiple second entities and the entity relationship between each second entity; the third determination module 1160 is also configured to: when the third event information is the same as the second event information, determine the multiple second entities as the multiple target entities corresponding to the question information, and determine the entity relationship between the second entities as the entity relationship between each target entity among the multiple target entities corresponding to the question information.

[0194] Fig.12 A schematic diagram of a training device for an information extraction model provided in some embodiments of the present application. Fig.12As shown, the training device 1200 of the information extraction model can be configured in Figure 1 The information extraction model shown in FIG. 1200 includes a second acquisition module 1210 , a second processing module 1220 , a fourth determination module 1230 , a fifth determination module 1240 and an optimization module 1250 .

[0195] The second acquisition module 1210 is configured to acquire the training text, the first entity corresponding to the training text, and the legal event information.

[0196] The second processing module 1220 is configured to perform vectorization processing on the legal event information to obtain a supervision vector, and perform feature extraction on the legal elements in the training text to obtain multiple element features corresponding to the training text.

[0197] The fourth determination module 1230 is configured to determine multiple predicted entities in the training text based on the supervision vector and multiple element features.

[0198] The fifth determination module 1240 is configured to determine a first loss function corresponding to the information extraction model based on the multiple prediction entities and the multiple first entities.

[0199] The optimization module 1250 is configured to optimize the parameters in the first loss function to determine the trained information extraction model.

[0200] like Fig.12 As shown, the information extraction model training device 1200 also includes a sixth determination module 1260.

[0201] In some embodiments, the second acquisition module 1210 is also configured to: obtain the true label sequence corresponding to the training text, and determine the conditional probability corresponding to the true label sequence based on the true label sequence; the sixth determination module 1260 is configured to: determine the second loss function corresponding to the information extraction model based on the conditional probability; the optimization module 1250 is specifically configured to: optimize the parameters in the first loss function and the second loss function to determine the trained information extraction model.

[0202] Fig.13 A schematic diagram of an electronic device provided for some embodiments of the present application. In some embodiments, the electronic device includes one or more processors and a memory. The memory is configured to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining entities and entity relationships or the method for training an information extraction model in the above-mentioned embodiments.

[0203] like Fig.13As shown, the electronic device 1300 includes: a processor (processor) 1301 and a memory 1302. Exemplarily, the electronic device 1300 may also include: a communication interface (Communications Interface) 1303 and a communication bus 1304.

[0204] The processor 1301, the memory 1302 and the communication interface 1303 communicate with each other via the communication bus 1304. The communication interface 1303 is used to communicate with other devices such as client or other server network elements.

[0205] In some embodiments, the processor 1301 is used to execute the program 1305, which can specifically execute the relevant steps in the above-mentioned entity and entity relationship determination method or information extraction model training method embodiment. Specifically, the program 1305 can include program code, which includes computer executable instructions.

[0206] Exemplarily, the processor 1301 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. The electronic device 1300 may include one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0207] In some embodiments, the memory 1302 is used to store the program 1305. The memory 1302 may include a high-speed RAM memory, and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0208] Program 1305 can be specifically called by processor 1301 to enable electronic device 1300 to execute a method for determining entities and entity relationships or a method for training an information extraction model.

[0209] Some embodiments of the present application provide a computer-readable storage medium storing at least one executable instruction. When the executable instruction is executed on the electronic device 1300, the electronic device 1300 executes the method for determining entities and entity relationships or the method for training information extraction models in the above-mentioned embodiments.

[0210] The executable instructions can be specifically used to enable the electronic device 1300 to execute a method for determining entities and entity relationships or a method for training an information extraction model.

[0211] For example, the computer readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0212] The beneficial effects that can be achieved by the computer-readable storage medium provided in some embodiments of the present application can be referred to the beneficial effects of the corresponding entity and entity relationship determination method or information extraction model training method provided above, and will not be repeated here.

[0213] It should be noted that, in the application, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0214] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0215] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with these instruction execution systems, apparatuses or devices.

[0216] For the purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by an instruction execution system, apparatus, or device or in conjunction with such instruction execution system, apparatus, or device.

[0217] More specific examples (a non-exhaustive list) of computer readable media include the following: an electrical connection having one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM).

[0218] In addition, the computer readable medium can even be a paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting or processing in other suitable ways as necessary, and then storing it in the computer memory. It should be understood that various parts of the present application can be implemented in hardware, software, firmware or a combination thereof.

[0219] In the above embodiments, multiple steps or methods may be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it may be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0220] The above-described embodiments of the present application do not constitute a limitation on the protection scope of the present application.

Claims

1. A method for determining entities and entity relationships, characterized in that: include: Acquire question information input by the user and first event information corresponding to the question information; Performing vectorization processing on the first event information to obtain a supervision vector corresponding to the first event information; Extracting features of legal elements in the question information to obtain multiple element features corresponding to the question information; Determine a plurality of first entities corresponding to the question information based on the supervision vector corresponding to the first event information and the plurality of element features, and determine an entity relationship between each of the plurality of first entities based on the plurality of first entities; Determining second event information corresponding to the question information based on the multiple first entities and the entity relationships between the first entities; When the second event information is the same as the first event information, the multiple first entities are determined as multiple target entities corresponding to the question information, and the entity relationship between the first entities is determined as the entity relationship between the target entities in the multiple target entities corresponding to the question information.

2. The method according to claim 1, characterized in that The feature extraction of the legal elements in the question information to obtain multiple element features corresponding to the question information includes: The plurality of element features are obtained by extracting features of the plurality of types of legal elements in the question information through a plurality of element extraction models; wherein each of the plurality of element extraction models is used to extract different types of element features.

3. The method according to claim 1, characterized in that The determining the multiple first entities corresponding to the question information based on the supervision vector corresponding to the first event information and the multiple element features includes: Performing concatenation or dot multiplication processing on the supervision vector corresponding to the first event information and each of the element features to obtain a plurality of first element features; Using an attention mechanism to adjust the weight of each of the plurality of first element features to obtain a plurality of second element features after the weight adjustment; The plurality of second element features are subjected to feature fusion processing to obtain content features, and the plurality of first entities are determined according to the content features.

4. The method according to claim 3, characterized in that The determining the plurality of first entities according to the content feature comprises: Performing sequence labeling on the content features to obtain a plurality of label sequences corresponding to the content features; wherein each of the plurality of label sequences includes a plurality of label information; Based on the multiple tag information corresponding to each of the tag sequences, determining the transition probability corresponding to each of the tag sequences; wherein the transition probability is used to characterize the dependency relationship between the multiple tag information corresponding to each of the tag sequences; Determine the conditional probability corresponding to each of the label sequences based on the transition probability corresponding to each of the label sequences; Based on the conditional probabilities corresponding to the label sequences, a target label sequence is determined from the multiple label sequences, and the multiple first entities are determined according to the target label sequence.

5. The method according to any one of claims 1 to 4, characterized in that The determining, based on the multiple first entities and the entity relationships between the first entities, the second event information corresponding to the question information includes: Based on the question information and the multiple first entities, determine the question information feature corresponding to the question information and the initial feature corresponding to each of the first entities; wherein the initial feature corresponding to each of the first entities is used to characterize the context information corresponding to each of the first entities; Determining core information corresponding to the question information based on the multiple first entities, entity relationships between the first entities, and initial features corresponding to the first entities; The second event information is determined based on the question information feature and core information corresponding to the question information.

6. The method according to claim 5, characterized in that The core information corresponding to the question information includes core information corresponding to each of the first entities in the question information and core relationships between the first entities; determining the core information corresponding to the question information based on the multiple first entities, the entity relationships between the first entities, and the initial features corresponding to the first entities includes: Determine a syntactic graph using the first entities as nodes and entity relationships between the first entities as edges; Based on the syntactic graph and the initial features corresponding to the first entities, the initial features corresponding to the first entities are updated through multiple neural network layers to obtain updated features of the nodes determined by each of the neural network layers in the multiple neural network layers; The updated features of each of the nodes determined by the last neural network layer in the multiple neural network layers are determined as the core information corresponding to each of the first entities and the core relationship between each of the first entities.

7. The method according to claim 6, characterized in that The updating of the initial features corresponding to the first entities by the multiple neural network layers based on the syntax graph and the initial features corresponding to the first entities to obtain the updated features of the nodes determined by the neural network layers in the multiple neural network layers includes: Obtaining, by a first neural network layer among the multiple neural network layers, a first feature of each of the nodes determined by the first neural network layer based on the initial features corresponding to each of the nodes, the adjacency matrix corresponding to the syntactic graph, the self-connection matrix, and the weight matrix corresponding to the first neural network layer; wherein the first neural network layer is the first neural network layer among the multiple neural network layers; The second feature of each of the nodes determined by the second neural network layer is obtained by a second neural network layer among the multiple neural network layers based on the first feature of each of the nodes, the adjacency matrix corresponding to the syntactic graph, the self-connection matrix and the weight matrix corresponding to the second neural network layer; wherein the updated feature of each of the nodes includes the first feature of each of the nodes and the second feature of each of the nodes.

8. The method according to claim 5, characterized in that The determining the second event information based on the question information feature and the core information corresponding to the question information includes: Fusing the question information feature with the core information corresponding to the question information to obtain a fused feature; The legal events corresponding to the fusion features are classified to obtain the second event information.

9. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: In a case where the second event information is different from the first event information, determining a plurality of second entities corresponding to the question information and an entity relationship between the second entities in the plurality of second entities based on the second event information; Determining third event information corresponding to the question information based on the entity relationships between the multiple second entities and each of the second entities; When the third event information is the same as the second event information, the multiple second entities are determined as multiple target entities corresponding to the question information, and the entity relationship between the second entities is determined as the entity relationship between the target entities in the multiple target entities corresponding to the question information.

10. A training method for an information extraction model, characterized in that: Applied to the information extraction model to be trained, the method comprises: Acquire a training text, a target entity corresponding to the training text, and legal event information; Vectorizing the legal event information to obtain a supervision vector, and extracting features of the legal elements in the training text to obtain multiple element features corresponding to the training text; Determine a plurality of predicted entities in the training text based on the supervision vector and the plurality of element features; Determining a first loss function corresponding to the information extraction model based on the multiple predicted entities and the target entity corresponding to the training text; The parameters in the first loss function are optimized to determine the trained information extraction model.

11. The method according to claim 10, characterized in that The method further comprises: Obtaining a true label sequence corresponding to the training text, and determining a conditional probability corresponding to the true label sequence based on the true label sequence; Determine a second loss function corresponding to the information extraction model based on the conditional probability; The step of optimizing the parameters in the first loss function to determine the trained information extraction model includes: Optimize the parameters in the first loss function and the second loss function to determine the trained information extraction model.

12. A device for determining entities and entity relationships, characterized in that: include: A first acquisition module is configured to acquire question information input by a user and first event information corresponding to the question information; A first processing module is configured to perform vectorization processing on the first event information to obtain a supervision vector corresponding to the first event information; A feature extraction module is configured to extract features of legal elements in the question information to obtain multiple element features corresponding to the question information; A first determination module is configured to determine a plurality of first entities corresponding to the question information based on the supervision vector corresponding to the first event information and the plurality of element features, and determine an entity relationship between each of the plurality of first entities based on the plurality of first entities; A second determining module is configured to determine second event information corresponding to the question information based on the multiple first entities and the entity relationship between each of the first entities; The third determination module is configured to, when the second event information is the same as the first event information, determine the multiple first entities as the multiple target entities corresponding to the question information, and determine the entity relationship between the first entities as the entity relationship between the target entities in the multiple target entities corresponding to the question information.

13. A training device for an information extraction model, characterized in that: Configured in an information extraction model to be trained, the device comprises: A second acquisition module is configured to acquire a training text, a target entity corresponding to the training text, and legal event information; A second processing module is configured to perform vectorization processing on the legal event information to obtain a supervision vector, and perform feature extraction on the legal elements in the training text to obtain a plurality of element features corresponding to the training text; a fourth determination module, configured to determine a plurality of predicted entities in the training text based on the supervision vector and the plurality of element features; a fifth determination module, configured to determine a first loss function corresponding to the information extraction model based on the multiple predicted entities and the target entity corresponding to the training text; The optimization module is configured to optimize the parameters in the first loss function to determine the trained information extraction model.

14. An electronic device, characterized in that: include: one or more processors; and A memory configured to: store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining entities and entity relationships according to any one of claims 1-9, or implement the training method for the information extraction model according to claim 10 or 11.

15. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, it implements the method for determining entities and entity relationships according to any one of claims 1-9, or implements the training method for the information extraction model according to claim 10 or 11.

Citation Information

Patent Citations

  • Event extraction model and military event type prediction method

    CN118277574A

  • Legal document information extraction method

    CN118627619A