Text processing method, system, construction machine, and electronic device

By completing named entity recognition, relationship extraction and entity standardization mapping in one model, the propagation error problem caused by multi-model construction in the prior art is solved, and more efficient and accurate information extraction is achieved.

CN115879471BActive Publication Date: 2025-06-17SANY HEAVY MACHINERY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211449187.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2025-06-17
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

In the prior art, when extracting text information, multiple models need to be built to perform different tasks, resulting in propagation errors and affecting the accuracy of information extraction.

Method used

A text processing method is adopted to complete named entity recognition, relationship extraction and entity standardization mapping through a model, and a text processing model trained based on sample text, standard entities and preset virtual entities is used to realize the standardized mapping of entity relationship triples.

Benefits of technology

It reduces propagation errors, improves the accuracy of information extraction and text processing efficiency, and can process triples and one tuples at the same time, enhancing the scope of application of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115879471B_ABST
    Figure CN115879471B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of text processing, and provides a text processing method, system, construction machinery and electronic device. The method includes: obtaining a text to be processed; inputting the text to be processed into a text processing model to output a standard entity corresponding to the text to be processed. Among them, the text processing model is used to perform a standardized mapping based on the entity relationship triples recognized from the text to be processed to obtain the standard entity corresponding to the text to be processed. The entity relationship triples conform to the semantics of the text to be processed, and are pairwise combinations of potential entities in the text to be processed and / or pairwise combinations of potential entities and preset virtual entities. The present invention is used to solve the defect in the prior art that when extracting information from text, due to the need to construct multiple models to perform different tasks, there is a propagation error, which affects the accuracy of information extraction. It realizes information extraction in one model, thereby avoiding propagation errors and improving the text processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and particularly to a text processing method, system, working machine and electronic device. Background Art

[0002] Currently, the conventional methods for information extraction mainly include two types: pipeline extraction and joint extraction.

[0003] When adopting the pipeline form for information extraction, first, named entity recognition is performed on the input text to obtain all entities in the text, and then the obtained entities are combined pairwise with the original text to process and obtain triples in the form of (entity 1, relationship, entity 2). Finally, standard mapping is respectively performed on entity 1 and entity 2 to obtain standard information. This processing process requires at least three models: a model for named entity recognition, a model for relationship extraction, and a model for entity standard mapping. The prediction results of these models are mutually dependent and there are propagation errors, seriously affecting the accuracy of information extraction.

[0004] For the joint extraction form, named entity recognition and relationship extraction are fused in one model for joint extraction. Compared with pipeline extraction, although this method can reduce a part of the propagation errors, for entity standard mapping, a separate model still needs to be constructed. Therefore, there are still certain propagation errors. Summary of the Invention

[0005] The present invention provides a text processing method, system, working machine and electronic device, which are used to solve the defect in the prior art that when extracting information from text, due to the need to construct multiple models to perform different tasks, there are propagation errors, affecting the accuracy of information extraction. It realizes information extraction in one model, thus avoiding propagation errors and improving the text processing efficiency.

[0006] The present invention provides a text processing method, including:

[0007] Obtain a text to be processed;

[0008] Input the text to be processed into a text processing model to obtain standard entities corresponding to the text to be processed output by the text processing model;

[0009] Among them, the text processing model is trained based on sample texts, the standard entities corresponding to the sample texts, and preset virtual entities; the text processing model is used to perform a standardized mapping on the entity relationship triples recognized from the text to be processed to obtain the standard entities corresponding to the text to be processed, and the entity relationship triples conform to the semantics of the text to be processed, and are pairwise combinations of potential entities in the text to be processed and / or pairwise combinations of the potential entities and the preset virtual entities.

[0010] According to the text processing method of the present invention, the step of inputting the text to be processed into the text processing model to obtain the standard entities corresponding to the text to be processed output by the text processing model includes:

[0011] Obtain text vectors: Input the text to be processed into the text conversion layer of the text processing model to obtain multiple vectors output by the text conversion layer after converting the text to be processed into a vector form;

[0012] Obtain potential entities: Input the vectors into the entity recognition layer of the text processing model to obtain the potential entities in the text to be processed output by the entity recognition layer;

[0013] Obtain entity relationship triples: Input the potential entities into the triple recognition layer of the text processing model to obtain the entity relationship triples output by the triple recognition layer;

[0014] Obtain standard entities: Input the entity relationship triples into the entity alignment layer of the text processing model to obtain the standard entities corresponding to the text to be processed output by the entity alignment layer.

[0015] According to the text processing method of the present invention, the step of inputting the vectors into the entity recognition layer of the text processing model to obtain the potential entities in the text to be processed output by the entity recognition layer includes:

[0016] Obtain phrase fragments: Input the vectors into the phrase combination layer of the entity recognition layer to obtain a phrase combination set output by the phrase combination layer, and the phrase combination set is a set composed of all phrase combination fragments formed by arbitrarily combining the vectors;

[0017] Obtain potential entities: Input the phrase combination set into the classification model layer of the entity recognition layer to obtain the potential entities selected from the phrase combination set output by the classification model layer.

[0018] According to the text processing method of the present invention, the step of inputting the potential entities into the triple recognition layer of the text processing model to obtain the entity relationship triples output by the triple recognition layer includes:

[0019] Obtain entity combinations: Input the potential entities into the entity combination layer of the triple recognition layer to obtain an entity combination set output by the entity combination layer. The entity combination set is a set composed of all entity combinations formed by any pairwise combination of the potential entities, and all entity combinations formed by pairwise combination of the preset virtual entity and any of the potential entities;

[0020] Obtain potential triples: Input the entity combination set into the prediction model layer of the triple recognition layer to obtain the potential triples filtered from the entity combination set and output by the prediction model layer;

[0021] Obtain entity relationship triples: Input the potential triples into the entity relationship confirmation layer of the triple recognition layer to obtain the entity relationship triples filtered from the potential triples and output by the entity relationship confirmation layer.

[0022] According to the text processing method of the present invention, the step of inputting the text to be processed into the text processing model to obtain the standard entities corresponding to the text to be processed and output by the text processing model further includes:

[0023] Obtain text information: Input the text to be processed into the text conversion layer to obtain text information output by the text conversion layer, which is extracted from the text to be processed and represents the semantics of the text to be processed;

[0024] The step of inputting the potential triples into the entity relationship confirmation layer of the triple recognition layer to obtain the entity relationship triples filtered from the potential triples and output by the entity relationship confirmation layer includes:

[0025] Input the potential triples and the text information into the entity relationship confirmation layer to obtain the entity relationship triples filtered from the potential triples based on the text information and output by the entity relationship confirmation layer.

[0026] According to the text processing method of the present invention, the step of inputting the text to be processed into the text processing model to obtain the standard entities corresponding to the text to be processed and output by the text processing model further includes:

[0027] Obtain entity hierarchy: Input the entity relationship triples into the hierarchy attribution layer of the text processing model to obtain the hierarchy attribution information of the potential entities in the entity relationship triples and output by the hierarchy attribution layer.

[0028] The present invention also provides a text processing system, including:

[0029] An acquisition module, configured to acquire a text to be processed;

[0030] A processing module, configured to input the text to be processed into a text processing model, and obtain a standard entity corresponding to the text to be processed output by the text processing model;

[0031] Wherein, the text processing model is trained based on sample texts, standard entities corresponding to the sample texts, and preset virtual entities; the text processing model is configured to perform a standardized mapping on an entity relationship triple recognized from the text to be processed to obtain a standard entity corresponding to the text to be processed, and the entity relationship triple conforms to the semantics of the text to be processed, and is a pairwise combination of potential entities in the text to be processed and / or a pairwise combination of the potential entity and the preset virtual entity.

[0032] The present invention further provides a working machine that uses the text processing method described in any one of the above to process maintenance record texts.

[0033] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the text processing method described in any one of the above is implemented.

[0034] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the text processing method described in any one of the above is implemented.

[0035] A text processing method, system, working machine, and electronic device provided by the present invention obtain a text to be processed, and then input the text to be processed into a text processing model trained based on sample texts, standard entities corresponding to the sample texts, and preset virtual entities, so as to perform a standardized mapping on an entity relationship triple recognized from the text to be processed through the text processing model to obtain a standard entity corresponding to the text to be processed, so that named entity recognition, relationship extraction, and entity standardized mapping are unified in one task, that is, completed by the same model, thereby not only reducing propagation errors, improving the accuracy of information extraction, but also improving text processing efficiency.

[0036] At the same time, through the setting of preset virtual entities, it is possible to identify and distinguish between triples and singletons, avoiding separate modeling for triple and singleton recognition, further reducing cumulative errors, and improving the accuracy of information extraction. Description of the Drawings

[0037] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0038] Figure 1 is a schematic flowchart of a text processing method provided by an embodiment of the present invention;

[0039] Figure 2 is a schematic hierarchical flowchart of processing the text to be processed, "The equipment hydraulic pump leaks hydraulic oil", based on the text processing model provided by an embodiment of the present invention;

[0040] Figure 3 is a schematic structural diagram of a text processing system provided by an embodiment of the present invention;

[0041] Figure 4 is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0043] The following combines Figure 1 and Figure 2 to describe a text processing method of the present invention, which is executed on an electronic device, a component in the electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The present invention does not make specific limitations.

[0044] As Figure 1 shown, the method includes the following steps:

[0045] 101. Obtain the text to be processed;

[0046] Among them, the text to be processed can be any text. In the embodiments of the present invention, the maintenance record text of a work machine is taken as an example to specifically illustrate the text processing method.

[0047] In one embodiment, through the text processing method provided by the embodiments of the present invention, the standard information involved in the maintenance record text can be accurately extracted. For example, the standard information involved, namely "hydraulic pump" and "oil leakage", can be extracted from "the equipment hydraulic pump leaks hydraulic oil", so as to facilitate the analysis of various problems such as the quality, function or performance of the work machine.

[0048] 102. Input the text to be processed into the text processing model to obtain the standard entity corresponding to the text to be processed output by the text processing model;

[0049] Among them, the text processing model is trained based on the sample text, the standard entity corresponding to the sample text, and a preset virtual entity; the text processing model is used to perform a standardized mapping based on the entity relationship triples recognized from the text to be processed to obtain the standard entity corresponding to the text to be processed. The entity relationship triples are in line with the semantics of the text to be processed, and are pairwise combinations of potential entities in the text to be processed and / or pairwise combinations of the potential entities and the preset virtual entity.

[0050] It can be understood that when information extraction is performed in a pipeline form, the output of the model for named entity recognition is used as the input of the model for relationship extraction, and then the output of the model for relationship extraction is used as the input of the model for entity standardized mapping. Therefore, the processing accuracy of the previous model directly affects the processing of the subsequent model, and there are relatively serious propagation errors. At the same time, each model is trained independently, without considering the interaction between named entity recognition, relationship extraction, and entity standardized mapping, which further affects the accuracy of information extraction.

[0051] Furthermore, taking the information extraction in a pipeline form as an example, if a total of three models are used to implement information extraction, assuming that for a text to be processed, one model needs to perform 100,000 calculations, then the three models need to perform 300,000 calculations, which will have a greater impact on the information extraction efficiency.

[0052] It should be noted that in the text processing method provided in the embodiments of the present invention, the text processing model for information extraction of the text to be processed is trained based on the sample text, the standard entities corresponding to the sample text, and the preset virtual entities, and it is a model that can simultaneously implement three tasks: named entity recognition, relationship extraction, and entity standardization mapping. During the training process of this model, the same sample text is used for named entity recognition, relationship extraction, and entity standardization mapping, achieving parameter sharing among tasks, and thus fully considering the dependency relationships among tasks. Therefore, when using the text processing model provided in the embodiments of the present invention to perform information extraction on the text to be processed, the propagation error can be effectively reduced.

[0053] Furthermore, when inputting the text to be processed into the text processing model, the text processing model can directly output the standard entities corresponding to the text to be processed, significantly improving the efficiency of information extraction compared to performing information extraction through multiple models.

[0054] Furthermore, in the existing models for relationship extraction, the relationship is extracted by combining the identified entities pairwise. For example: by combining entity 1 and entity 2, a triple (entity 1, relationship, entity 2) is obtained. However, taking the maintenance record text of construction machinery as an example, there are some default descriptions for the problem description of construction machinery. For example, "the excavator is stalling" is usually described as "stalling". Therefore, when the text to be processed expresses the semantics of "the excavator is stalling", the entity that can be identified from the text to be processed is only "stalling", lacking the default entity "excavator". When using the existing model for relationship extraction, pairwise combination with other entities cannot be achieved, and thus the entity triple cannot be formed, resulting in the failure of information extraction for the text to be processed containing a single entity such as "stalling".

[0055] It should be noted that in the text processing method provided in the embodiments of the present invention, by training the text processing model based on the sample text, the standard entities corresponding to the sample text, and the preset virtual entities, after inputting the text to be processed into the text processing model, when encountering entities such as "stalling", a triple form of (preset virtual entity, relationship, entity) can be formed through the combination of the preset virtual entity and the entity, thereby having the ability to recognize single entities.

[0056] The text processing method provided by the embodiments of the present invention inputs the obtained text to be processed into a text processing model trained based on sample texts, the standard entities corresponding to the sample texts, and preset virtual entities. Through the calculation of the text processing model, the standard entities corresponding to the text to be processed can be output, which not only improves the information extraction efficiency, but also greatly reduces the propagation error and improves the accuracy of information extraction. At the same time, based on the setting of the preset virtual entities in the text processing model, the text processing method provided by the embodiments of the present invention has the ability to process the mixed occurrence of triples and single entities simultaneously, and can distinguish triples and single entities, making it more convenient to use and having a wider scope of application.

[0057] Based on the content of the above embodiments, the step of inputting the text to be processed into the text processing model to obtain the standard entity corresponding to the text to be processed output by the text processing model includes:

[0058] Obtain text vectors: Input the text to be processed into the text conversion layer of the text processing model to obtain multiple vectors output by the text conversion layer after converting the text to be processed into a vector form;

[0059] Obtain potential entities: Input the vectors into the entity recognition layer of the text processing model to obtain the potential entities in the text to be processed output by the entity recognition layer;

[0060] Obtain entity relationship triples: Input the potential entities into the triple recognition layer of the text processing model to obtain the entity relationship triples output by the triple recognition layer;

[0061] Obtain standard entities: Input the entity relationship triples into the entity alignment layer of the text processing model to obtain the standard entities corresponding to the text to be processed output by the entity alignment layer.

[0062] It should be noted that for the text to be processed input into the text processing model, the text to be processed is first encoded into vectors through the text conversion layer, then named entity recognition is performed through the entity recognition layer, relationship extraction is realized through the triple recognition layer, and finally, entity standardization mapping is performed through the entity alignment layer to obtain the standard entities corresponding to the text to be processed.

[0063] Furthermore, the text conversion layer converts the symbolized text into a vectorized representation so that the text to be processed can be recognized by a computer, facilitating subsequent processing.

[0064] In one embodiment, the text conversion layer is a neural network such as a Transformer, LSTM (Long Short-Term Memory), RNN (Recurrent Neural Network), or CNN (Convolutional Neural Network). Specifically, the pre-trained model is used to encode the text to be processed. Here, the pre-trained model can be BERT (Bidirectional Encoder Representations from Transformers), word2vec, ELMo, XLNet, ERNIE (Enhanced Language Representation with Informative Entities), etc., and no specific limitation is made here.

[0065] The text processing method provided by the embodiment of the present invention realizes various tasks such as named entity recognition, relationship extraction, and entity standardization mapping of the text to be processed through a text processing model composed of a text conversion layer, an entity recognition layer, a triple recognition layer, and an entity alignment layer, which not only improves the information extraction efficiency, but also greatly reduces the propagation error and improves the accuracy of information extraction.

[0066] Based on the content of the above embodiment, the step of inputting the vector into the entity recognition layer of the text processing model to obtain the potential entities in the text to be processed output by the entity recognition layer includes:

[0067] Obtain phrase fragments: Input the vector into the phrase combination layer of the entity recognition layer to obtain a phrase combination set output by the phrase combination layer. The phrase combination set is a set composed of all phrase combination fragments arbitrarily combined by the vector.

[0068] Obtain potential entities: Input the phrase combination set into the classification model layer of the entity recognition layer to obtain the potential entities selected from the phrase combination set output by the classification model layer.

[0069] It should be noted that after converting the text to be processed into a vector form and inputting it into the entity recognition layer, the phrase combination layer of the entity recognition layer can enumerate all possible phrase combination fragments, and then determine which type of entity or no entity these phrase combination fragments belong to through the classification model layer of the entity recognition layer.

[0070] In one embodiment, taking the text to be processed as "the hydraulic pump of the device leaks hydraulic oil" as an example, the phrase combination segments output by the phrase combination layer may include: she, shebei, shebeiye, yeya, yeyabeng, etc. Then, the classification model layer uses the classification model to determine what entities "she", "shebei", "shebeiye", etc. belong to respectively, so as to obtain potential entities.

[0071] The text processing method provided by the embodiment of the present invention can obtain all potential entities included in the text to be processed by setting an entity recognition layer composed of a phrase combination layer and a classification model layer, thereby improving the accuracy and comprehensiveness of information extraction.

[0072] Based on the content of the above embodiment, the step of inputting the potential entity into the triple recognition layer of the text processing model to obtain the entity relationship triple output by the triple recognition layer includes:

[0073] Obtain entity combinations: Input the potential entity into the entity combination layer of the triple recognition layer to obtain an entity combination set output by the entity combination layer. The entity combination set is a set composed of all entity combinations formed by any pairwise combination of the potential entities, and all entity combinations formed by pairwise combination of the preset virtual entity and any potential entity;

[0074] Obtain potential triples: Input the entity combination set into the prediction model layer of the triple recognition layer to obtain the potential triples screened out from the entity combination set output by the prediction model layer;

[0075] Obtain entity relationship triples: Input the potential triples into the entity relationship confirmation layer of the triple recognition layer to obtain the entity relationship triples screened out from the potential triples output by the entity relationship confirmation layer.

[0076] It should be noted that after all potential entities included in the text to be processed are obtained by the entity recognition layer, by inputting the potential entity into the entity combination layer of the triple recognition layer, the entity combination layer can obtain all entity combinations that can be formed by the potential entity and the preset virtual entity in a pairwise combination manner. Then, through the prediction model layer of the triple recognition layer, potential entity triples with reasonable relationships are screened out from the entity combinations. Finally, through the entity relationship confirmation layer of the triple recognition layer, entity triples that conform to the semantics of the text to be processed are confirmed from the potential entity triples, that is, entity relationship triples.

[0077] In one embodiment, taking the maintenance record text of a work machine as an example, entity triples generally include a target entity, a view, and the relationship between the target entity and the view. Still taking the text to be processed "The hydraulic pump of the equipment leaks hydraulic oil" as an example, assuming the potential entities include "equipment", "hydraulic pump", "pump", "leaking hydraulic oil", and "hydraulic oil", through the entity combination layer, entity combinations such as "equipment" and "hydraulic pump", "hydraulic pump" and "pump", "hydraulic pump" and "leaking hydraulic oil" can be obtained. Then, through the prediction model layer, potential entity triples such as (hydraulic pump, relationship, leaking hydraulic oil), (equipment, relationship, leaking hydraulic oil) can be screened out. Among them, "hydraulic pump" and "equipment" are the target entities, and "leaking hydraulic oil" is the view. Finally, through the entity relationship confirmation layer, the entity relationship triple that conforms to the semantics of "The hydraulic pump of the equipment leaks hydraulic oil" can be confirmed as (hydraulic pump, relationship, leaking hydraulic oil).

[0078] In another embodiment, taking the text to be processed "engine stalling" as an example, "engine stalling" belongs to the view entity. Obviously, the target entity is missing in this text to be processed. Based on the same processing method as the text to be processed "The hydraulic pump of the equipment leaks hydraulic oil", by combining the preset virtual entity as the target entity with "engine stalling", finally, the entity relationship triple that conforms to the semantics of "engine stalling" can be confirmed as (preset virtual entity, relationship, engine stalling), and then it can be confirmed that "engine stalling" is a single tuple.

[0079] The text processing method provided by the embodiments of the present invention can obtain all triples and single tuples that conform to the semantics of the text to be processed in the text to be processed by setting a triple recognition layer composed of an entity combination layer, a prediction model layer, and an entity relationship confirmation layer, thereby further improving the accuracy and comprehensiveness of information extraction.

[0080] Based on the content of the above embodiments, the step of inputting the text to be processed into the text processing model to obtain the standard entity corresponding to the text to be processed output by the text processing model further includes:

[0081] Obtain text information: Input the text to be processed into the text conversion layer to obtain the text information output by the text conversion layer, which is extracted from the text to be processed and represents the semantics of the text to be processed.

[0082] The step of inputting the potential triples into the entity relationship confirmation layer of the triple recognition layer to obtain the entity relationship triples screened out from the potential triples output by the entity relationship confirmation layer includes:

[0083] Input the potential triples and the text information into the entity relationship confirmation layer to obtain the entity relationship triples screened out from the potential triples based on the text information output by the entity relationship confirmation layer.

[0084] It is understandable that the same character can have different meanings in different texts. For example, "apple" can represent a kind of fruit, or it can represent a mobile phone or a mobile phone brand. Without combining the semantics of the text to be processed, it is easy to cause errors in information extraction.

[0085] It should be noted that taking the text conversion layer as the Transformer network as an example, when encoding the text to be processed, it contains a placeholder CLS. Through the placeholder CLS, the text information of the text to be processed can be represented in vector form. Based on the text information of the text to be processed, the entity relationship confirmation layer can filter out the potential triples that conform to the semantics of the text to be processed from the potential triples, that is, the entity relationship triples.

[0086] The text processing method provided by the embodiments of the present invention obtains the text information representing the semantics of the text to be processed extracted from the text to be processed through the text conversion layer, and then confirms the entity relationship triples through the combination of the text information and the potential triples, so that the obtained entity relationship triples highly conform to the text to be processed, thereby improving the accuracy of information extraction.

[0087] Based on the content of the above embodiments, the step of inputting the text to be processed into the text processing model to obtain the standard entity corresponding to the text to be processed output by the text processing model further includes:

[0088] Obtain the entity hierarchy: Input the entity relationship triples into the hierarchy attribution layer of the text processing model to obtain the hierarchy attribution information of the potential entities in the entity relationship triples output by the hierarchy attribution layer.

[0089] It should be noted that through the setting of the hierarchy attribution layer, after inputting the text to be processed into the text processing model, not only the standard entity corresponding to the text to be processed can be obtained, but also the potential entities of the text to be processed can be classified, thus facilitating the management of text data.

[0090] In one embodiment, the hierarchy attribution layer can be a classifier. Then, by setting multiple different classifications and hierarchically dividing various classifications, the hierarchy attribution marking of the potential entities of the text to be processed is carried out. For example, for construction machinery, classifications can be set including: hydraulic pumps, motors, cylinders, etc. according to equipment targets, classifications can be set including: oil leakage, power failure, liquid leakage, etc. according to demand keywords representing fault phenomena, and classifications can be set including: hydraulic system quality, electrical system quality, lubrication system quality, etc. according to equipment requirements. Then, the equipment requirements are used as the higher level, while the equipment targets and demand keywords are used as the lower levels.

[0091] Taking the text to be processed as "The hydraulic pump of the device leaks hydraulic oil" as an example, the entity relationship triple obtained by the text processing method provided by the embodiment of the present invention is (hydraulic pump, relationship, leaks hydraulic oil). Then, inputting this entity relationship triple into the hierarchical attribution layer, the obtained hierarchical attribution information includes: the device target is the hydraulic pump, the demand keyword is oil leakage, and the device demand is the quality of the hydraulic system.

[0092] The text processing method provided by the embodiment of the present invention realizes the determination of the hierarchical attribution of the entity relationship triple extracted from the text to be processed by setting a hierarchical attribution layer in the text processing model, thus facilitating the management of text data.

[0093] The following Figure 2 illustrates the hierarchical structure of the text processing model applied to the text processing method provided by the embodiment of the present invention. Figure 2 is a schematic diagram of the hierarchical process of processing the text to be processed "The hydraulic pump of the device leaks hydraulic oil" based on the text processing model provided by the embodiment of the present invention. It can be Figure 2 seen that through the text processing method provided by the embodiment of the present invention, the entity relationship triple (hydraulic pump, relationship, leaks hydraulic oil) can be accurately extracted from the text to be processed, and the standard entities (hydraulic pump and oil leakage) corresponding to the text to be processed can be output. At the same time, the hierarchical attribution of the entity relationship triple can also be confirmed, thus facilitating the management of text data.

[0094] The following describes a text processing system provided by the present invention. The text processing system described below can be mutually referred to with the text processing method described above.

[0095] A text processing system provided by the present invention, as Figure 3 shown, includes: an acquisition module 310 and a processing module 320; wherein,

[0096] The acquisition module 310 is used to acquire the text to be processed;

[0097] The processing module 320 is used to input the text to be processed into the text processing model to obtain the standard entity corresponding to the text to be processed output by the text processing model;

[0098] wherein, the text processing model is trained based on sample texts, the standard entities corresponding to the sample texts, and preset virtual entities; the text processing model is used to perform a standardized mapping based on the entity relationship triple recognized from the text to be processed to obtain the standard entity corresponding to the text to be processed. The entity relationship triple is in line with the semantics of the text to be processed, and is a pairwise combination of potential entities in the text to be processed and / or a pairwise combination of the potential entity and the preset virtual entity.

[0099] The text processing system provided by the embodiment of the present invention obtains the text to be processed, and then inputs the text to be processed into a text processing model trained based on a sample text, the standard entity corresponding to the sample text, and a preset virtual entity, so as to perform a standardized mapping on the entity relationship triple identified from the text to be processed through the text processing model, and obtain the standard entity corresponding to the text to be processed, so that named entity recognition, relationship extraction, and entity standardized mapping are unified in one task, that is, completed by the same model, thereby not only reducing the propagation error, improving the accuracy of information extraction, but also improving the text processing efficiency. At the same time, through the setting of the preset virtual entity, the recognition of both triples and single tuples can be realized, avoiding separate modeling for the recognition of triples and single tuples, further reducing the cumulative error, and improving the accuracy of information extraction.

[0100] Optionally, the processing module 320 is used for:

[0101] Obtain text vectors: Input the text to be processed into the text conversion layer of the text processing model, and obtain multiple vectors output by the text conversion layer after converting the text to be processed into a vector form;

[0102] Obtain potential entities: Input the vectors into the entity recognition layer of the text processing model, and obtain the potential entities in the text to be processed output by the entity recognition layer;

[0103] Obtain entity relationship triples: Input the potential entities into the triple recognition layer of the text processing model, and obtain the entity relationship triples output by the triple recognition layer;

[0104] Obtain standard entities: Input the entity relationship triples into the entity alignment layer of the text processing model, and obtain the standard entities corresponding to the text to be processed output by the entity alignment layer.

[0105] Optionally, the step of inputting the vectors into the entity recognition layer of the text processing model and obtaining the potential entities in the text to be processed output by the entity recognition layer includes:

[0106] Obtain phrase fragments: Input the vectors into the phrase combination layer of the entity recognition layer, and obtain a phrase combination set output by the phrase combination layer, where the phrase combination set is a set composed of all phrase combination fragments arbitrarily combined from the vectors;

[0107] Obtain potential entities: Input the phrase combination set into the classification model layer of the entity recognition layer, and obtain the potential entities screened out from the phrase combination set output by the classification model layer.

[0108] Optionally, inputting the potential entity into the triple recognition layer of the text processing model to obtain the entity relationship triple output by the triple recognition layer includes:

[0109] Obtain entity combinations: Input the potential entity into the entity combination layer of the triple recognition layer to obtain a set of entity combinations output by the entity combination layer. The set of entity combinations is a set composed of all entity combinations formed by any pairwise combination of the potential entities, and all entity combinations formed by pairwise combination of the preset virtual entity and any of the potential entities;

[0110] Obtain potential triples: Input the set of entity combinations into the prediction model layer of the triple recognition layer to obtain the potential triples screened out from the set of entity combinations output by the prediction model layer;

[0111] Obtain entity relationship triples: Input the potential triples into the entity relationship confirmation layer of the triple recognition layer to obtain the entity relationship triples screened out from the potential triples output by the entity relationship confirmation layer.

[0112] Optionally, inputting the text to be processed into the text processing model to obtain the standard entity corresponding to the text to be processed output by the text processing model further includes:

[0113] Obtain text information: Input the text to be processed into the text conversion layer to obtain the text information extracted from the text to be processed and representing the semantics of the text to be processed output by the text conversion layer;

[0114] Inputting the potential triples into the entity relationship confirmation layer of the triple recognition layer to obtain the entity relationship triples screened out from the potential triples output by the entity relationship confirmation layer includes:

[0115] Input the potential triples and the text information into the entity relationship confirmation layer to obtain the entity relationship triples screened out from the potential triples based on the text information output by the entity relationship confirmation layer.

[0116] Optionally, inputting the text to be processed into the text processing model to obtain the standard entity corresponding to the text to be processed output by the text processing model further includes:

[0117] Obtain entity hierarchy: Input the entity relationship triples into the hierarchy attribution layer of the text processing model to obtain the hierarchy attribution information of the potential entities in the entity relationship triples output by the hierarchy attribution layer.

[0118] The present invention also provides a working machine that processes maintenance record texts by using any one of the above-described text processing methods.

[0119] It can be understood that the working machine provided in the embodiment of the present invention, which processes maintenance record texts by using the text processing method described in any one of the above embodiments, has all the advantages and technical effects of the text processing method described in any one of the above embodiments, and will not be elaborated here.

[0120] The embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the text processing method provided in any one of the above embodiments.

[0121] Figure 4 An example of the physical structure diagram of an electronic device is shown as Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete mutual communication through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute a text processing method, and the method includes: obtaining a text to be processed; inputting the text to be processed into a text processing model to obtain a standard entity corresponding to the text to be processed output by the text processing model; where the text processing model is trained based on sample texts, the standard entities corresponding to the sample texts, and preset virtual entities; the text processing model is used to perform a standardized mapping on the entity relationship triples identified from the text to be processed to obtain the standard entity corresponding to the text to be processed, and the entity relationship triples are in line with the semantics of the text to be processed, and are pairwise combinations of potential entities in the text to be processed and / or pairwise combinations of the potential entities and the preset virtual entities.

[0122] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0123] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a text processing method provided by the above-mentioned various methods. The method includes: obtaining a text to be processed; inputting the text to be processed into a text processing model to obtain a standard entity corresponding to the text to be processed output by the text processing model; wherein the text processing model is trained based on sample texts, standard entities corresponding to the sample texts, and preset virtual entities; the text processing model is used to perform a standardized mapping on an entity relationship triple recognized from the text to be processed to obtain a standard entity corresponding to the text to be processed. The entity relationship triple is in line with the semantics of the text to be processed, and is a pairwise combination of potential entities in the text to be processed and / or a pairwise combination of the potential entities and the preset virtual entities.

[0124] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, a text processing method is implemented. The method includes: obtaining a text to be processed; inputting the text to be processed into a text processing model to obtain a standard entity corresponding to the text to be processed output by the text processing model; wherein the text processing model is trained based on sample texts, standard entities corresponding to the sample texts, and preset virtual entities; the text processing model is used to perform a standardized mapping on an entity relationship triple recognized from the text to be processed to obtain a standard entity corresponding to the text to be processed. The entity relationship triple is in line with the semantics of the text to be processed, and is a pairwise combination of potential entities in the text to be processed and / or a pairwise combination of the potential entities and the preset virtual entities.

[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A text processing method, characterized in that, Including: Obtain the text to be processed; Input the text to be processed into a text processing model to obtain a standard entity corresponding to the text to be processed output by the text processing model; Wherein, the text processing model is trained based on sample texts, standard entities corresponding to the sample texts, and preset virtual entities; The text processing model is used to perform a standardized mapping on an entity relationship triple recognized from the text to be processed to obtain a standard entity corresponding to the text to be processed, and the entity relationship triple conforms to the semantics of the text to be processed, and is a pairwise combination of potential entities in the text to be processed and / or a pairwise combination of the potential entity and the preset virtual entity.

2. The text processing method according to claim 1, characterized in that, The step of inputting the text to be processed into a text processing model to obtain a standard entity corresponding to the text to be processed output by the text processing model includes: Obtain text vectors: Input the text to be processed into the text conversion layer of the text processing model to obtain multiple vectors output by the text conversion layer after converting the text to be processed into a vector form; Obtain potential entities: Input the vectors into the entity recognition layer of the text processing model to obtain potential entities in the text to be processed output by the entity recognition layer; Obtain entity relationship triples: Input the potential entities into the triple recognition layer of the text processing model to obtain the entity relationship triples output by the triple recognition layer; Obtain standard entities: Input the entity relationship triples into the entity alignment layer of the text processing model to obtain a standard entity corresponding to the text to be processed output by the entity alignment layer.

3. The text processing method according to claim 2, characterized in that, The step of inputting the vectors into the entity recognition layer of the text processing model to obtain potential entities in the text to be processed output by the entity recognition layer includes: Obtain phrase fragments: Input the vectors into the phrase combination layer of the entity recognition layer to obtain a set of phrase combinations output by the phrase combination layer, and the set of phrase combinations is a set composed of all phrase combination fragments arbitrarily combined by the vectors; Obtain potential entities: Input the set of phrase combinations into the classification model layer of the entity recognition layer to obtain potential entities screened out from the set of phrase combinations output by the classification model layer.

4. The text processing method according to claim 2, characterized in that, The step of inputting the potential entities into the triple recognition layer of the text processing model to obtain the entity relationship triples output by the triple recognition layer includes: Obtain entity combinations: Input the potential entities into the entity combination layer of the triple recognition layer to obtain a set of entity combinations output by the entity combination layer, and the set of entity combinations is a set composed of all entity combinations arbitrarily pairwise combined by the potential entities, and all entity combinations pairwise combined by the preset virtual entity and any potential entity; Obtain potential triples: Input the set of entity combinations into the prediction model layer of the triple recognition layer to obtain potential triples screened out from the set of entity combinations output by the prediction model layer; Obtain entity relation triples: Input the potential triples into the entity relation confirmation layer of the triple recognition layer, and obtain the entity relation triples filtered from the potential triples and output by the entity relation confirmation layer.

5. The text processing method according to claim 4, characterized in that, The step of inputting the text to be processed into the text processing model to obtain the standard entities corresponding to the text to be processed and output by the text processing model further includes: Obtain text information: Input the text to be processed into the text conversion layer, and obtain the text information extracted from the text to be processed and output by the text conversion layer, which characterizes the semantics of the text to be processed. The step of inputting the potential triples into the entity relation confirmation layer of the triple recognition layer to obtain the entity relation triples filtered from the potential triples and output by the entity relation confirmation layer includes: Input the potential triples and the text information into the entity relation confirmation layer, and obtain the entity relation triples filtered from the potential triples based on the text information and output by the entity relation confirmation layer.

6. The text processing method according to claim 2, characterized in that, The step of inputting the text to be processed into the text processing model to obtain the standard entities corresponding to the text to be processed and output by the text processing model further includes: Obtain entity hierarchy: Input the entity relation triples into the hierarchy attribution layer of the text processing model, and obtain the hierarchy attribution information of the potential entities in the entity relation triples and output by the hierarchy attribution layer.

7. A text processing system, characterized in that, It includes: An acquisition module for acquiring the text to be processed; A processing module for inputting the text to be processed into the text processing model to obtain the standard entities corresponding to the text to be processed and output by the text processing model; Wherein, the text processing model is trained based on sample texts, the standard entities corresponding to the sample texts, and preset virtual entities; The text processing model is used to perform a standardized mapping on the entity relation triples recognized from the text to be processed to obtain the standard entities corresponding to the text to be processed. The entity relation triples are in line with the semantics of the text to be processed, and are pairwise combinations of the potential entities in the text to be processed and / or pairwise combinations of the potential entities and the preset virtual entities.

8. An operating machine, characterized in that, Use the text processing method according to any one of claims 1 to 6 to process the maintenance record text.

9. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the text processing method according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the text processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Joint extraction method and joint extraction device for entity relationship

    CN113553854A

  • Learning to extract entities from conversations with neural networks

    US20220075944A1