Entity Relationship Pair Recognition Method, Apparatus, Device and Computer Readable Storage Medium
By classifying business texts and identifying entity relationships, and building a sample set of resampling and processing, the problem of insufficient accuracy of entity relationship recognition is solved, and the in-depth mining and accuracy of implicit information is achieved.
Patent Information
- Application Number
- CN202310276088.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-03-20
AI Technical Summary
In the prior art, entity relationship identification is insufficient inadequate, especially inadequate in the utilization of implicit information.
By text classification of business text sets, use the BERT model to identify entities and relationships, build positive and negative example sample sets, and obtain target entity relationship pairs through resampling processing.
It improves the accuracy of entity relationship recognition, fully explores the entity relationship of triples, and enhances the utilization of invisible information.
Smart Images

Figure CN116340538B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, apparatus, electronic device, and computer-readable storage medium for identifying entity relationship pairs. Background Art
[0002] With the rise of artificial intelligence, various data in the digital age show an exponential growth in scale, and data mining and relationship mining have become increasingly important.
[0003] In the prior art, by combining machine learning methods with rules based on artificial intelligence technology, the method of extracting entity information relationships has gradually matured, which can effectively help people quickly extract useful information required. However, the entity recognition technology based on rules can only explicitly recognize knowledge entities, and a large amount of implicit information cannot be used, resulting in inaccurate entity relationship recognition. Summary of the Invention
[0004] The present invention provides a method, apparatus, electronic device, and readable storage medium for identifying entity relationship pairs, and its main purpose is to improve the accuracy of entity relationship recognition.
[0005] To achieve the above object, a method for identifying entity relationship pairs provided by the present invention includes:
[0006] Obtain a set of business texts, and classify the set of business texts based on the text categories in the set of business texts to obtain a set of classified texts;
[0007] Perform entity recognition and relationship recognition on the texts in the set of classified texts to obtain a set of entities and a set of relationships;
[0008] Construct a set of positive example samples and a set of negative example samples based on the set of entities and the set of relationships;
[0009] Perform resampling processing on the set of negative example samples based on the set of positive example samples to obtain target entity relationship pairs.
[0010] Optionally, the step of classifying the set of business texts based on the text categories in the set of business texts to obtain a set of classified texts includes:
[0011] Traverse the business texts in the set of business texts, and calculate the number of words in the traversed business text;
[0012] When the number of words meets a preset word threshold, determine the traversed business text as a short text, and use a pre-constructed short text classification model to classify the short text;
[0013] When the number of words does not meet the preset word threshold, determine that the traversed business text is a long text, and use a pre-constructed long text classification model to classify the long text;
[0014] Summarize all classified business texts to obtain the classified text set.
[0015] Optionally, performing entity recognition and relationship recognition on the texts in the classified text set to obtain an entity set and a relationship set, including:
[0016] Use a pre-constructed BERT model to vectorize the classified texts in the classified text set to obtain sentence feature vectors;
[0017] Use a pre-constructed entity recognition model to perform entity recognition and extraction from the sentence feature vectors to obtain an entity set;
[0018] Determine keywords from the classified texts based on a preset number of entities in the entity set, and determine the relationships between different numbers of entities according to the preset attributes of the keywords. Summarize all recognized relationships to obtain a relationship set.
[0019] Optionally, constructing a positive example sample set and a negative example sample set based on the entity set and the relationship set, including:
[0020] Select a target relationship from the relationship set, and select a target entity pair from the entity set based on the target relationship;
[0021] Calculate the similarity score of the target entity pair and the target relationship, and use the entity relationship pair composed of the target entity pair and the target relationship with a similarity score greater than the preset similarity threshold as the positive example sample of the target relationship. Summarize all positive example samples to obtain a positive example sample set;
[0022] Use the entity relationship pair composed of the target entity pair and the target relationship with a similarity score less than or equal to the preset similarity threshold as the negative example sample of the target relationship. Summarize all negative example samples to obtain a negative example sample set.
[0023] Optionally, performing resampling processing on the negative example sample set based on the positive example sample set to obtain target entity relationship pairs, including:
[0024] Obtain the relevant relationships of the target relationship in the positive example samples, and use the relevant relationships to replace the target relationship in the corresponding negative example samples of the positive example samples to obtain replaced entity relationship pairs;
[0025] Calculate the similarity score of the replaced entity relationship pairs, and use the replaced entity relationship pairs with a similarity score greater than the preset similarity threshold as the positive example samples of the relevant relationships;
[0026] Use all positive example samples as target entity-relationship pairs.
[0027] Optionally, after using the replacement entity-relationship pairs with similarity scores greater than a preset similarity threshold as positive example samples of the relevant relationship, the method further includes:
[0028] Use the replacement entity-relationship pairs with similarity scores less than or equal to the preset similarity threshold as negative example samples of the relevant relationship, and resample the negative example samples of the relevant relationship using the positive example samples of the relevant relationship.
[0029] To solve the above problems, the present invention also provides an entity-relationship pair recognition device, which includes:
[0030] A text classification module, configured to obtain a business text set, classify the business text set based on the text categories in the business text set, and obtain a classified text set;
[0031] An entity-relationship recognition module, configured to perform entity recognition and relationship recognition on the texts in the classified text set to obtain an entity set and a relationship set;
[0032] A sample construction module, configured to construct a positive example sample set and a negative example sample set based on the entity set and the relationship set;
[0033] An entity-relationship pair extraction module, configured to resample the negative example sample set based on the positive example sample set to obtain target entity-relationship pairs.
[0034] To solve the above problems, the present invention also provides an electronic device, which includes:
[0035] A memory, storing at least one computer program; and
[0036] A processor, executing the computer program stored in the memory to implement the above-mentioned entity-relationship pair recognition method.
[0037] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned entity-relationship pair recognition method.
[0038] In this embodiment, by classifying the service text set, a classified text set is obtained. Entities and relationships are recognized from the texts in the classified text set to obtain an entity set and a relationship set. Based on the entity set and the relationship set, a positive example sample set and a negative example sample set are constructed. Finally, resampling processing is performed on the negative example sample set based on the positive example sample set to obtain target entity relationship pairs, which can continuously and deeply mine the entity relationships of triples and make full use of the implicit information of entities, improving the accuracy of entity relationship recognition. Therefore, the entity relationship pair recognition method, device, electronic device, and computer-readable storage medium proposed by the present invention can improve the accuracy rate of entity relationship recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of the entity relationship pair recognition method provided by an embodiment of the present invention;
[0040] Figure 2 It is a functional module diagram of the entity relationship pair recognition device provided by an embodiment of the present invention;
[0041] Figure 3 It is a structural diagram of an electronic device for implementing the entity relationship pair recognition method provided by an embodiment of the present invention.
[0042] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] An embodiment of the present application provides an entity relationship pair recognition method. The execution subject of the entity relationship pair recognition method includes, but is not limited to, at least one of an electronic device such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the entity relationship pair recognition method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0045] Refer to Figure 1As shown in the figure, it is a schematic flowchart of the entity relationship pair recognition method provided by an embodiment of the present invention. In this embodiment, the entity relationship pair recognition method includes:
[0046] S1. Obtain a business text set, and perform text classification on the business text set based on the text categories in the business text set to obtain a classified text set.
[0047] In an embodiment of the present invention, the business text set may be text documents crawled from different fields, such as enterprise evaluation reports, enterprise equity data, user resumes, etc. in the financial field. Since the lengths of different texts are different and the classification methods are also different, classifying different types of texts through text categories can improve the accuracy and efficiency of text classification.
[0048] Specifically, the performing text classification on the business text set based on the text categories in the business text set to obtain a classified text set includes:
[0049] Traverse the business texts in the business text set, and calculate the number of words in the traversed business text;
[0050] When the number of words meets a preset word threshold, determine the traversed business text as a short text, and use a pre-constructed short text classification model to perform text classification on the short text;
[0051] When the number of words does not meet the preset word threshold, determine the traversed business text as a long text, and use a pre-constructed long text classification model to perform text classification on the long text;
[0052] Summarize all the classified business texts to obtain the classified text set.
[0053] In an optional embodiment of the present invention, the text classification model may be a pre-trained natural language model. The short text classification model may be a TextCNN model, etc., and the long text classification model may be models such as FastText, HAN, XLNet, etc. For example, the word threshold may be 1000. When the number of texts in a text document is less than 1000, it is a short text, and it is classified by the TextCNN model. When the number of texts in a text document is greater than or equal to 1000, it is a long text, and it is classified by the XLNet model.
[0054] S2. Perform entity recognition and relationship recognition on the texts in the classified text set to obtain an entity set and a relationship set.
[0055] Specifically, the performing entity recognition and relationship recognition on the texts in the classified text set to obtain an entity set and a relationship set includes:
[0056] Use a pre - built BERT model to vectorize the classified texts in the classified text set to obtain sentence feature vectors;
[0057] Use a pre - built entity recognition model to perform entity recognition and extraction from the sentence feature vectors to obtain an entity set;
[0058] Based on a preset number of entities in the entity set, determine keywords from the classified texts, and determine the relationships between different numbers of entities according to the preset attributes of the keywords, and summarize all the recognized relationships to obtain a relationship set.
[0059] In an alternative embodiment of the present invention, after the BERT model performs embedding, sentence feature vectors are obtained, and the part - of - speech features of Jieba word segmentation are added and concatenated as the input of the model. The entity extraction task is realized through the Bi - LSTM + CRF model structure (entity recognition model). At the same time, the relationship between two entities that appear in the same sentence can be used for relationship recognition. For example, if A works in B company, the entities are "A" and "B company" respectively, the keyword is determined as "works in", and the corresponding preset attribute is "employee", and the corresponding relationship is determined as an "employee" relationship.
[0060] S3. Construct a positive example sample set and a negative example sample set based on the entity set and the relationship set.
[0061] In the embodiment of the present invention, a positive example sample represents a triple with a high matching degree between an entity and a relationship, and a negative example sample represents a triple with a low matching degree between an entity and a relationship.
[0062] Specifically, constructing the positive example sample set and the negative example sample set based on the entity set and the relationship set includes:
[0063] Select a target relationship from the relationship set, and select a target entity pair from the entity set based on the target relationship;
[0064] Calculate the similarity score of the target entity pair and the target relationship, and use the entity - relationship pair composed of the target entity pair and the target relationship with a similarity score greater than a preset similarity threshold as the positive example sample of the target relationship, and summarize all the positive example samples to obtain a positive example sample set;
[0065] Use the entity - relationship pair composed of the target entity pair and the target relationship with a similarity score less than or equal to the preset similarity threshold as the negative example sample of the target relationship, and summarize all the negative example samples to obtain a negative example sample set.
[0066] S4. Resample the negative example sample set based on the positive example sample set to obtain target entity - relationship pairs.
[0067] In the embodiments of the present invention, for a certain target relationship (such as "colleague relationship"), although the positive example samples are a kind of high-quality knowledge triples (entity relationship pairs), a large amount of information in the negative example samples has not been mined (for example, it may contain "superior-subordinate relationship", etc.). By resampling the negative example samples with the positive example samples, the relationships between entities in the triples can be further mined, and the recognition depth and accuracy of the entity relationship pairs can be improved.
[0068] Specifically, the resampling process of the negative example sample set based on the positive example sample set to obtain the target entity relationship pair includes:
[0069] Obtain the relevant relationships of the target relationship in the positive example samples, and use the relevant relationships to replace the target relationship in the negative example samples corresponding to the positive example samples to obtain replacement entity relationship pairs;
[0070] Calculate the similarity scores of the replacement entity relationship pairs, and use the replacement entity relationship pairs with similarity scores greater than the preset similarity threshold as the positive example samples of the relevant relationships;
[0071] Use all the positive example samples as the target entity relationship pairs.
[0072] In another alternative embodiment of the present invention, after using the replacement entity relationship pairs with similarity scores greater than the preset similarity threshold as the positive example samples of the relevant relationships, the method further includes:
[0073] Use the replacement entity relationship pairs with similarity scores less than or equal to the preset similarity threshold as the negative example samples of the relevant relationships, and resample the negative example samples of the relevant relationships with the positive example samples of the relevant relationships.
[0074] In the embodiments of the present invention, for example, for the entity relationship recognition of all employees in a company. The target relationship is "finance". After finding all relevant employee entities and department entities, corresponding entities can be continuously searched from the negative example samples according to the relevant relationship "sales", so as to further mine the knowledge triples.
[0075] In the embodiments of the present invention, the process of resampling the negative example samples of the relevant relationships with the positive example samples of the relevant relationships is similar to the steps of constructing the positive example sample set and the negative example sample set based on the target relationship in S3-S4 above, and will not be elaborated here.
[0076] In this embodiment, by classifying the service text set, a classified text set is obtained. Entities and relationships are recognized from the texts in the classified text set to obtain an entity set and a relationship set. Based on the entity set and the relationship set, a positive example sample set and a negative example sample set are constructed. Finally, resampling processing is performed on the negative example sample set based on the positive example sample set to obtain target entity-relationship pairs, which can continuously and deeply mine the entity relationships of triples and fully utilize the implicit information of entities, improving the accuracy of entity relationship recognition. Therefore, the entity-relationship pair recognition method proposed by the present invention can improve the accuracy of entity relationship recognition.
[0077] As Figure 2 shown, it is a functional module diagram of an entity-relationship pair recognition device provided by an embodiment of the present invention.
[0078] The entity-relationship pair recognition device 100 of the present invention can be installed in an electronic device. According to the functions achieved, the entity-relationship pair recognition device 100 may include a text classification module 101, an entity relationship recognition module 102, a sample construction module 103, and an entity-relationship pair extraction module 104. The modules of the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0079] In this embodiment, the functions of each module / unit are as follows:
[0080] The text classification module 101 is configured to obtain a service text set and classify the service text set based on the text categories in the service text set to obtain a classified text set;
[0081] The entity relationship recognition module 102 is configured to perform entity recognition and relationship recognition on the texts in the classified text set to obtain an entity set and a relationship set;
[0082] The sample construction module 103 is configured to construct a positive example sample set and a negative example sample set based on the entity set and the relationship set;
[0083] The entity-relationship pair extraction module 104 is configured to perform resampling processing on the negative example sample set based on the positive example sample set to obtain target entity-relationship pairs.
[0084] Specifically, the specific implementation manners of each module of the entity-relationship pair recognition device 100 are as follows:
[0085] Step 1: The text classification module 101 obtains a service text set and classifies the service text set based on the text categories in the service text set to obtain a classified text set.
[0086] In the embodiments of the present invention, the business text set may be text documents crawled from different fields, such as enterprise evaluation reports, enterprise equity data, user resumes, etc. in the financial field. Since the lengths of different texts are different and the classification methods are also different, different types of texts are classified by text categories, improving the accuracy and efficiency of text classification.
[0087] Specifically, performing text classification on the business text set based on the text categories in the business text set to obtain a classified text set includes:
[0088] Traverse the business texts in the business text set and calculate the number of words in the traversed business text;
[0089] When the number of words meets a preset word threshold, determine that the traversed business text is a short text, and use a pre-constructed short text classification model to perform text classification on the short text;
[0090] When the number of words does not meet the preset word threshold, determine that the traversed business text is a long text, and use a pre-constructed long text classification model to perform text classification on the long text;
[0091] Summarize all the classified business texts to obtain the classified text set.
[0092] In an alternative embodiment of the present invention, the text classification model may be a pre-trained natural language model. The short text classification model may be a TextCNN model, etc., and the long text classification model may be models such as FastText, HAN, XLNet, etc. For example, the word threshold may be 1000. When the number of texts in a text document is less than 1000, it is a short text and is classified by the TextCNN model. When the number of texts in a text document is greater than or equal to 1000, it is a long text and is classified by the XLNet model.
[0093] Step 2: The entity relationship recognition module 102 performs entity recognition and relationship recognition on the texts in the classified text set to obtain an entity set and a relationship set.
[0094] Specifically, performing entity recognition and relationship recognition on the texts in the classified text set to obtain an entity set and a relationship set includes:
[0095] Use a pre-constructed BERT model to perform vectorization processing on the classified texts in the classified text set to obtain sentence feature vectors;
[0096] Use a pre-constructed entity recognition model to perform entity recognition and extraction from the sentence feature vectors to obtain an entity set;
[0097] Determine keywords from the classification text based on a preset number of entities in the entity set, and determine the relationships between different numbers of entities according to the preset attributes of the keywords, and summarize all the recognized relationships to obtain a relationship set.
[0098] In an alternative embodiment of the present invention, after the bert model performs embedding, obtain the sentence feature vector, add the part-of-speech features of jieba segmentation, splice them as the input of the model, and implement the entity extraction task through the Bi-LSTM+CRF model structure (entity recognition model). At the same time, the relationship between two entities that appear in the same sentence can be used for relationship recognition. For example, A works for B company, the entities are "A" and "B company" respectively, the determined keyword is "works for", and the corresponding preset attribute is "employee", and the determined corresponding relationship is the "employee" relationship.
[0099] Step 3: The sample construction module 103 constructs a positive example sample set and a negative example sample set based on the entity set and the relationship set.
[0100] In the embodiment of the present invention, a positive example sample represents a triple with a high degree of matching between an entity and a relationship, and a negative example sample represents a triple with a low degree of matching between an entity and a relationship.
[0101] Specifically, the constructing a positive example sample set and a negative example sample set based on the entity set and the relationship set includes:
[0102] Select a target relationship from the relationship set, and select a target entity pair from the entity set based on the target relationship;
[0103] Calculate the similarity score of the target entity pair and the target relationship, and use the entity relationship pair formed by combining the target entity pair with a similarity score greater than the preset similarity threshold and the target relationship as the positive example sample of the target relationship, and summarize all the positive example samples to obtain a positive example sample set;
[0104] Use the entity relationship pair formed by combining the target entity pair with a similarity score less than or equal to the preset similarity threshold and the target relationship as the negative example sample of the target relationship, and summarize all the negative example samples to obtain a negative example sample set.
[0105] Step 4: The entity relationship pair extraction module 104 performs resampling processing on the negative example sample set based on the positive example sample set to obtain a target entity relationship pair.
[0106] In the embodiments of the present invention, for a certain target relationship (such as "colleague relationship"), although positive example samples are high-quality knowledge triples (entity relationship pairs), a large amount of information in the negative example samples has not been mined (for example, it may contain "superior-subordinate relationship", etc.). By resampling the negative example samples with the positive example samples, the relationships between entities in the triples can be further mined, and the recognition depth and accuracy of the entity relationship pairs can be improved.
[0107] Specifically, the resampling process of the negative example sample set based on the positive example sample set to obtain the target entity relationship pair includes:
[0108] Obtain the relevant relationships of the target relationship in the positive example samples, and use the relevant relationships to replace the target relationship in the negative example samples corresponding to the positive example samples to obtain replaced entity relationship pairs;
[0109] Calculate the similarity scores of the replaced entity relationship pairs, and use the replaced entity relationship pairs with similarity scores greater than a preset similarity threshold as the positive example samples of the relevant relationships;
[0110] Use all the positive example samples as the target entity relationship pairs.
[0111] In another optional embodiment of the present invention, after using the replaced entity relationship pairs with similarity scores greater than a preset similarity threshold as the positive example samples of the relevant relationships, the entity relationship pair extraction module 104 further executes:
[0112] Use the replaced entity relationship pairs with similarity scores less than or equal to the preset similarity threshold as the negative example samples of the relevant relationships, and resample the negative example samples of the relevant relationships with the positive example samples of the relevant relationships.
[0113] In the embodiments of the present invention, for example, for the entity relationship recognition of all employees in a company. The target relationship is "finance". After finding all relevant employee entities and department entities, corresponding entities can continue to be searched from the negative example samples according to the relevant relationship "sales", so as to further mine knowledge triples.
[0114] In the embodiments of the present invention, the process of resampling the negative example samples of the relevant relationships with the positive example samples of the relevant relationships is similar to the steps of constructing the positive example sample set and the negative example sample set based on the target relationship in the above modules 103-104, and will not be elaborated here.
[0115] In this embodiment, by classifying the service text set, a classified text set is obtained. Entities and relationships are recognized from the texts in the classified text set to obtain an entity set and a relationship set. Based on the entity set and the relationship set, a positive example sample set and a negative example sample set are constructed. Finally, the negative example sample set is resampled based on the positive example sample set to obtain target entity relationship pairs, which can continuously and deeply mine the entity relationships of triples and fully utilize the implicit information of entities, improving the accuracy of entity relationship recognition. Therefore, the entity relationship pair recognition device proposed by the present invention can improve the accuracy of entity relationship recognition.
[0116] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the entity relationship pair recognition method provided by an embodiment of the present invention.
[0117] The electronic device may include a processor 10, a memory 11, a communication interface 12, and a bus 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as an entity relationship pair recognition program.
[0118] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory 11 may also include both the internal storage unit and the external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of the entity relationship pair recognition program, but also to temporarily store data that has been output or will be output.
[0119] In some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, running or executing programs or modules (such as entity relationship pair recognition programs, etc.) stored in the memory 11, and calling data stored in the memory 11 to execute various functions of the electronic device and process data.
[0120] The communication interface 12 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display information processed in the electronic device and to display a visual user interface.
[0121] The bus 13 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 13 may be divided into an address bus, a data bus, a control bus, etc. The bus 13 is configured to implement connection communication between the memory 11 and at least one processor 10, etc.
[0122] Figure 3 Only an electronic device with components is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown, or combine certain components, or have a different component arrangement.
[0123] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to various components. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0124] Furthermore, the electronic device may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the electronic device and other electronic devices.
[0125] Optionally, the electronic device may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device and to display a visual user interface.
[0126] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0127] The entity relationship pair recognition program stored in the memory 11 in the electronic device is a combination of multiple instructions. When running in the processor 10, it can implement:
[0128] Obtain a set of business texts, classify the set of business texts based on the text categories in the set of business texts to obtain a classified text set;
[0129] Perform entity recognition and relationship recognition on the texts in the classified text set to obtain an entity set and a relationship set;
[0130] Construct a positive example sample set and a negative example sample set based on the entity set and the relationship set;
[0131] Perform resampling processing on the negative example sample set based on the positive example sample set to obtain a target entity relationship pair.
[0132] Specifically, for the specific implementation method of the above instructions by the processor 10, reference may be made to the description of the relevant steps in the corresponding embodiments of the accompanying drawings, which will not be elaborated here.
[0133] Furthermore, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).
[0134] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor of an electronic device, can implement:
[0135] Obtain a collection of business texts, perform text classification on the collection of business texts based on the text categories in the collection of business texts to obtain a classified text collection;
[0136] Perform entity recognition and relationship recognition on the texts in the classified text collection to obtain an entity set and a relationship set;
[0137] Construct a positive example sample set and a negative example sample set based on the entity set and the relationship set;
[0138] Perform resampling processing on the negative example sample set based on the positive example sample set to obtain target entity relationship pairs.
[0139] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0140] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, in each embodiment of the present invention, each functional module can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0142] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0143] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0144] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems.
[0145] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0146] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as "second" are used to denote names and do not denote any particular order.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for identifying entity relationship pairs, characterized in that, The method includes: Obtain a collection of business texts, perform text classification on the collection of business texts based on the text categories in the collection of business texts to obtain a collection of classified texts; Perform entity recognition and relationship recognition on the texts in the collection of classified texts to obtain an entity set and a relationship set; Construct a positive example sample set and a negative example sample set based on the entity set and the relationship set; Perform resampling processing on the negative example sample set based on the positive example sample set to obtain target entity-relationship pairs, including: obtaining the relevant relationships of the target relationship in the positive example samples, using the relevant relationships to replace the target relationship in the corresponding negative example samples of the positive example samples to obtain replaced entity-relationship pairs, calculating the similarity scores of the replaced entity-relationship pairs, and taking the replaced entity-relationship pairs with similarity scores greater than a preset similarity threshold as the positive example samples of the relevant relationships, and taking all the positive example samples as the target entity-relationship pairs.
2. The entity relationship pair recognition method according to claim 1, characterized in that The performing text classification on the collection of business texts based on the text categories in the collection of business texts to obtain a collection of classified texts includes: Traverse the business texts in the collection of business texts and calculate the number of words in the traversed business text; When the number of words meets a preset word threshold, determine the traversed business text as a short text, and perform text classification on the short text using a pre-constructed short text classification model; When the number of words does not meet the preset word threshold, determine the traversed business text as a long text, and perform text classification on the long text using a pre-constructed long text classification model; Summarize all the classified business texts to obtain the collection of classified texts.
3. The entity relationship pair recognition method according to claim 1, characterized in that The performing entity recognition and relationship recognition on the texts in the collection of classified texts to obtain an entity set and a relationship set includes: Perform vectorization processing on the classified texts in the collection of classified texts using a pre-constructed BERT model to obtain sentence feature vectors; Perform entity recognition and extraction from the sentence feature vectors using a pre-constructed entity recognition model to obtain an entity set; Determine keywords from the classified texts based on a preset number of entities in the entity set, and determine the relationships between different numbers of entities according to the preset attributes of the keywords, and summarize all the recognized relationships to obtain a relationship set.
4. The entity relationship pair recognition method according to claim 1, characterized in that The constructing a positive example sample set and a negative example sample set based on the entity set and the relationship set includes: Select a target relationship from the relationship set, and select a target entity pair from the entity set based on the target relationship; Calculate the similarity scores of the target entity pair and the target relationship, and take the entity-relationship pairs formed by combining the target entity pairs and the target relationships with similarity scores greater than a preset similarity threshold as the positive example samples of the target relationship, and summarize all the positive example samples to obtain a positive example sample set; Take the entity-relationship pairs formed by combining the target entity pairs and the target relationships with similarity scores less than or equal to the preset similarity threshold as the negative example samples of the target relationship, and summarize all the negative example samples to obtain a negative example sample set.
5. The entity relationship pair recognition method according to claim 1, characterized in that After taking the replaced entity-relationship pairs with similarity scores greater than a preset similarity threshold as the positive example samples of the relevant relationships, the method further includes: Taking the replacement entity relationship pairs with similarity scores less than or equal to a preset similarity threshold as negative example samples of the relevant relationship, and resampling the negative example samples of the relevant relationship using the positive example samples of the relevant relationship.
6. An entity relationship pair recognition device, characterized in that The apparatus includes: A text classification module, configured to obtain a set of business texts, and perform text classification on the set of business texts based on the text categories in the set of business texts to obtain a set of classified texts; An entity relationship recognition module, configured to perform entity recognition and relationship recognition on the texts in the set of classified texts to obtain an entity set and a relationship set; A sample construction module, configured to construct a positive example sample set and a negative example sample set based on the entity set and the relationship set; An entity relationship pair extraction module, configured to resample the negative example sample set based on the positive example sample set to obtain target entity relationship pairs, including: obtaining the relevant relationship of the target relationship in the positive example samples, using the relevant relationship to replace the target relationship in the negative example samples corresponding to the positive example samples to obtain replacement entity relationship pairs, calculating the similarity scores of the replacement entity relationship pairs, taking the replacement entity relationship pairs with similarity scores greater than the preset similarity threshold as positive example samples of the relevant relationship, and taking all the positive example samples as target entity relationship pairs.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the entity relationship pair recognition method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the entity relationship pair recognition method according to any one of claims 1 to 5.
Citation Information
Patent Citations
A field entity attribute relation extraction method based on distance supervision
CN109408642A
Electric power professional knowledge graph automatic completion method based on graph neural network
CN114139709A