Text processing method and device, electronic equipment and storage medium

By combining text processing methods with contextual semantics and syntactic features, and using a neural network model for text entity recognition in power system information, the problem of low data utilization and low management efficiency in power systems is solved, and the accuracy and efficiency of text entity recognition are improved.

CN116341532BActive Publication Date: 2026-03-31GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-27
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies for power system information management suffer from low data utilization, low management efficiency, high retrieval difficulty, and low recognition efficiency. In particular, the information is described in natural language and the different expression habits of various staff members increase the difficulty for computers to understand it.

Method used

A text processing method combining contextual semantic features and grammatical features of text sequences is adopted. A neural network model is used for text entity recognition, including semantic feature extraction, contextual feature extraction, sequence feature extraction and grammatical feature extraction. Text recognition is performed through a deep learning network with multiple sub-models.

Benefits of technology

It improves the accuracy and efficiency of text entity recognition, enhances the performance and robustness of the model, and solves the problems of low data utilization and low management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116341532B_ABST
    Figure CN116341532B_ABST
Patent Text Reader

Abstract

A text processing method and device, electronic equipment and storage medium are disclosed. The method comprises: obtaining a to-be-processed text comprising a to-be-identified entity; processing the to-be-processed text based on a target text recognition model to determine an identification result corresponding to the to-be-identified entity, wherein the target text recognition model comprises at least one target sub-model, and the at least one target sub-model comprises a semantic feature extraction sub-model, a context feature extraction sub-model, a sequence feature extraction sub-model, a syntax feature extraction sub-model and a relationship classification sub-model; and the identification result comprises an entity identification result and / or an entity relationship identification result of the to-be-identified entity. The technical solution of the embodiment achieves the effect of improving the text entity identification accuracy and efficiency, combines the context semantic features and syntax features of the text sequence, and improves the performance and robustness of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of artificial intelligence algorithms in power systems, and particularly to a text processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the continuous development of artificial intelligence, big data, and Internet of Things technologies, as well as the ongoing construction and upgrading of smart grids, new requirements and directions have been put forward for the intelligent management of power system information. In its daily construction and operation, the smart grid has accumulated massive amounts of information, including data knowledge, mechanistic knowledge, and experiential knowledge of distribution network equipment, as well as heterogeneous knowledge of distribution network operation and maintenance.

[0003] Currently, most data is simply stored in the system, leading to low data utilization, low management efficiency, and high retrieval difficulty. On the other hand, most power grid information is described in natural language, and each staff member has different writing styles, increasing the difficulty for computers to understand the information. Summary of the Invention

[0004] This invention provides a text processing method, apparatus, electronic device, and storage medium to improve the accuracy and efficiency of text entity recognition. It combines the contextual semantic features and syntactic features of text sequences to enhance the performance and robustness of the model.

[0005] According to one aspect of the present invention, a text processing method is provided, the method comprising:

[0006] Obtain the text to be processed, which includes the entity to be identified;

[0007] The text to be processed is identified based on a target text recognition model to determine the recognition result corresponding to the entity to be identified. The target text recognition model includes at least one target sub-model, which includes a semantic feature extraction sub-model, a context feature extraction sub-model, a sequence feature extraction sub-model, a grammatical feature extraction sub-model, and a relation classification sub-model. The recognition result includes the entity recognition result and / or entity relation recognition result of the entity to be identified.

[0008] According to another aspect of the present invention, a text processing apparatus is provided, the apparatus comprising:

[0009] The text to be processed module is used to acquire text to be processed, including entities to be identified.

[0010] The text processing module is used to perform recognition processing on the text to be processed based on the target text recognition model, and determine the recognition result corresponding to the entity to be recognized. The target text recognition model includes at least one target sub-model, which includes a semantic feature extraction sub-model, a context feature extraction sub-model, a sequence feature extraction sub-model, a grammatical feature extraction sub-model, and a relation classification sub-model. The recognition result includes the entity recognition result and / or entity relation recognition result of the entity to be recognized.

[0011] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0012] At least one processor; and

[0013] A memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the text processing method according to any embodiment of the present invention.

[0015] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the text processing method described in any embodiment of the present invention.

[0016] The technical solution of this invention obtains text to be processed, including entities to be identified, and further performs identification processing on the text to be processed based on a target text recognition model to determine the recognition result corresponding to the entities to be identified. This solves the problems of low data utilization, low management efficiency, high retrieval difficulty, and low recognition efficiency in the prior art when performing text recognition. It achieves the effect of improving the accuracy and efficiency of text entity recognition. By combining the contextual semantic features and grammatical features of the text sequence, the performance and robustness of the model are improved.

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a text processing method provided according to Embodiment 1 of the present invention;

[0020] Figure 2 This is a flowchart of a text processing method provided according to Embodiment 2 of the present invention;

[0021] Figure 3 This is a schematic diagram of the structure of a text processing device according to Embodiment 3 of the present invention;

[0022] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the text processing method of the present invention. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0025] Example 1

[0026] Figure 1It is a flowchart of a text processing method provided in the first embodiment of the present invention. This embodiment is applicable to the situation of identifying text entities in the text to be processed based on a neural network model and determining whether each text entity belongs to a preset entity relationship. This method can be executed by a text processing device, which can be implemented in the form of hardware and / or software, and the text processing device can be configured in a terminal and / or a server. As Figure 1 shown, the method includes:

[0027] S110. Obtain the text to be processed including the entity to be identified.

[0028] In this embodiment, the text to be processed can be a text that needs to perform entity extraction and entity relationship analysis. The text to be processed can be a text in any field. Optionally, it can be a text in the power grid field. The specific nouns or pronouns included in the text to be processed that are associated with the corresponding field are the entities to be identified. Exemplarily, if the text to be processed is a text in the power grid field, the entities to be identified can include words of types such as location, institution, status, voltage level, line, substation, bay, disconnect switch, switch, busbar, transformer, and generator. For example, the entities to be identified can be the 5 characters of "变", "压", "器", "漏", and "油".

[0029] It should be noted that the text to be processed can be a pre-stored text retrieved from a relevant database, or text information transmitted received from an external device, or a text obtained by other means, etc. This embodiment does not make specific limitations on this.

[0030] It should also be noted that in a specific application scenario, the text to be processed can be obtained in real time or periodically, or when it is detected that the user uploads text, the text can be obtained and used as the text to be processed. This embodiment does not make specific limitations on this.

[0031] In practical applications, in the original information including the text to be processed, in addition to the text to be processed, there may also be information in other information forms, and there may be a situation where the same text appears repeatedly in the text information included in the original information. Therefore, before obtaining the text to be processed, the obtained original information can be preprocessed first, so as to obtain the text to be processed. It should be noted that another advantage of preprocessing the original information to obtain the text to be processed is that the text information in the original information can be clause-separated and word-separated according to the preprocessing operation, and further, it can help to implement the subsequent entity recognition process and improve the accuracy of entity recognition.

[0032] Based on this, prior to the above technical solutions, the method also includes: acquiring the original information to be processed; preprocessing the original information based on a pre-set information preprocessing method to obtain the text to be processed.

[0033] In this embodiment, the original information can be directly acquired and unprocessed raw corpus. For example, the original information can be heterogeneous knowledge corpus of power grid operation and maintenance. The original information may include at least one form of information representation. Optionally, the form of information representation in the original information may include, but is not limited to, text, images, tables, audio, and video. The information preprocessing method can be a pre-set information processing method that preprocesses the original information including different forms of information representation. Optionally, the information preprocessing method may include at least one of deduplication, noise reduction, sentence segmentation, and word segmentation. Deduplication refers to removing duplicate information; noise reduction refers to deleting information in forms of information other than the target information representation; for example, if the target information representation is text, information in forms of information other than text can be deleted; sentence segmentation refers to dividing text information according to a defined delimiter, where the defined delimiter is a period; word segmentation refers to segmenting words in a sentence.

[0034] It should be noted that the original information can be pre-stored information retrieved from a relevant database, information received from an external device, or information obtained through other means. This embodiment does not specifically limit this.

[0035] In practical applications, the original information to be processed can be obtained first. Then, the original information can be preprocessed according to a pre-set information preprocessing method. Specifically, if the original information contains duplicate information, it can be deduplicated to remove duplicate information and avoid unnecessary processing. If the original information contains information in other forms besides text, it can be denoised to remove other forms of information, such as images and tables. After deduplication and denoising, the original information only contains text without duplicate information. At this point, the original information can be segmented into sentences, dividing the text according to defined delimiters, i.e., dividing the text according to periods, to obtain text containing at least one sentence. Finally, the original information can be segmented into words, dividing the words in the sentences. The segmented text is used as the text to be processed. At this point, the text to be processed includes all non-repeating text sentences from the original information and all words in each sentence.

[0036] In practical applications, due to the difference between Chinese and English, Chinese sentences do not separate words with spaces. Therefore, word segmentation tools can be used to segment the text. In this embodiment, the word segmentation tool can be any tool capable of segmenting Chinese sentences. Optionally, it can be the ICTCLAS (Institute of Computing Technology, Chinese Lexical Analysis System) word segmentation tool. This tool has added a user dictionary feature, meaning that users can first define some words as recognition standards. For example, users can refer to the operating procedures and management regulations of power equipment and the technical standards of power equipment to define words as segmentation standards, thereby solving the problem of segmenting complete words and causing incorrect recognition when using word segmentation tools. For example, if the sentence to be segmented is "The conductor cross-section and splitting type of the transmission line should meet the requirements of corona, radio interference and audible noise, etc.", then the segmentation result can be "Transmission / line / of / conductor / cross-section / and / splitting type / should / meet / corona / radio / interference / and / audible noise / etc. / requirements"; if the sentence to be segmented is "110KV Gushui 470 switch from maintenance to operation", then the segmentation result can be "110KV / Gushui / 470 switch / from / maintenance / to / operation".

[0037] In practical applications, the original information to be processed can be obtained first. Then, the original information can be preprocessed according to the pre-set information preprocessing method to obtain the text to be processed. Then, it can be processed based on the neural network model provided in this embodiment to analyze the entities to be identified in the text to be processed.

[0038] S120. Based on the target text recognition model, the text to be processed is recognized, and the recognition result corresponding to the entity to be recognized is determined.

[0039] In this embodiment, after acquiring the text to be processed, which includes the entities to be identified, the text can be input into a pre-trained target text recognition model for processing. The target text recognition model can be a pre-trained neural network model used to identify entities in text. The target text recognition model can be a deep learning network model including multiple sub-models. The target text recognition model can include at least one target sub-model, which includes a semantic feature extraction sub-model, a context feature extraction sub-model, a sequence feature extraction sub-model, a syntactic feature extraction sub-model, and a relation classification sub-model.

[0040] The semantic feature extraction sub-model can be a neural network model that analyzes and extracts the semantic and syntactic features of each word in the text. Optionally, the semantic feature extraction sub-model may include a word vector module, which is a vector model that maps text content to words in the data space. It can map each word to a fixed-dimensional real-valued vector to describe the meaning and semantic relationships of the words. For example, the word vector module can be the ELMo (Embedding from Language Models) word vector model. The advantage of applying the ELMo word vector model is that it uses a two-layer bidirectional long short-term memory (Bi-LSTM) network as a feature extractor. By stacking two layers of Bi-LSTM, more semantic and syntactic information can be encoded into the word representation vectors, and the word representation vectors can change according to different contexts. Therefore, the word vectors obtained through the ELMo word vector model can better improve the performance of the model.

[0041] The context feature extraction sub-model can be a neural network model that analyzes and extracts the context-dependent features of each sentence in the text. For example, the context feature extraction sub-model can be a two-layer bidirectional long short-term memory network (Bi-LSTM).

[0042] The sequence feature extraction sub-model can be a neural network model that extracts contextual semantic features of a text sequence from different spatial dimensions. For example, the sequence feature extraction sub-model can be a neural network model trained based on a multi-head self-attention algorithm.

[0043] The grammatical feature extraction sub-model can be a neural network model that analyzes and extracts the semantic relationship features between each word in the text. For example, the grammatical feature extraction sub-model can be a Graph Convolution Network (GCN).

[0044] The relationship classification sub-model can be a neural network model that identifies entities in text and classifies entity relationships. Optionally, the relationship classification sub-model may include pooling layers and multi-layer feedforward neural networks.

[0045] For example, the model parameter settings for each target sub-model in the target text recognition model can be shown in the table below:

[0046]

[0047] In practical applications, after obtaining the text to be processed, it can be input into the target text recognition model. The model then processes the text based on its semantic feature extraction sub-model, context feature extraction sub-model, sequence feature extraction sub-model, syntactic feature extraction sub-model, and relation classification sub-model, thereby obtaining the recognition result corresponding to the entity to be identified. This recognition result includes the entity recognition result of the entity to be identified and / or the entity relation recognition result of the entity to be identified.

[0048] In this embodiment, the entity relationship identification result of the entity to be identified can be a binary classification result of classifying the entity to be identified. That is, the entity relationship identification result can include two types of results: the entity to be identified belongs to a preset entity relationship and the entity to be identified does not belong to a preset entity relationship.

[0049] Optionally, the connection relationships between each target sub-model in the target text recognition model can be described by the input-output relationship of the data: the output of the semantic feature extraction sub-model is the input of the context feature extraction sub-model; the output of the context feature extraction sub-model is the input of the sequence feature extraction sub-model; the output of the context feature extraction sub-model is the input of the grammatical feature extraction sub-model; and the outputs of the sequence feature extraction sub-model and the grammatical feature extraction sub-model are the inputs of the relation classification sub-model. In the target text recognition model, the semantic feature extraction sub-model is connected to the context feature extraction sub-model, the context feature extraction sub-model is connected to both the sequence feature extraction sub-model and the grammatical feature extraction sub-model, and both the sequence feature extraction sub-model and the grammatical feature extraction sub-model are connected to the relation classification sub-model.

[0050] It should be noted that before applying the target text recognition model provided in this embodiment, the text recognition model to be trained can be trained first. The training process of the text recognition model to be trained is described in detail below: obtain multiple training samples; input the training samples into the text recognition model to be trained to obtain the actual output results; perform loss processing on the theoretical output results and the actual output results according to the preset loss function; and correct the model parameters in the text processing model to be trained based on the loss value to obtain the target text recognition model.

[0051] The training samples include: training sample text, which contains the entity to be identified, the theoretical output results corresponding to the training sample text, and the theoretical output results, which contain the theoretical entity identification results and theoretical entity relationship identification results corresponding to the entity to be identified.

[0052] The training sample text can be pre-stored in storage or received from an external device. It includes one or more entities, which can be used as entities to be identified. The theoretical entity recognition result can be used to characterize all entities included in the training sample text. The theoretical entity relationship recognition result can be a judgment of whether each entity belongs to a corresponding entity relationship.

[0053] It should be noted that before training the text recognition model, multiple training samples need to be obtained to train the model. To improve the model's accuracy, as many and rich training samples as possible can be obtained, including multiple training sample texts containing the entities to be recognized. Furthermore, these training sample texts are processed to obtain the theoretical entity recognition results and theoretical entity relationship recognition results corresponding to the entities to be recognized, thereby constructing a rich set of training samples based on the above method.

[0054] In practical applications, for each training sample, it can be input into the text recognition model to be trained. The model then processes the training sample to obtain the actual output. The text recognition model to be trained can be a model with default parameters. The model parameters in the text recognition model to be trained are corrected using training samples to obtain the target text recognition model. The actual output includes actual entity recognition results and actual entity relationship recognition results. The actual entity recognition result is the entity recognition result corresponding to the entity to be recognized, output after the training sample text is input into the text recognition model to be trained. The actual entity relationship recognition result is the entity relationship judgment result corresponding to the entity to be recognized, output after the training sample text is input into the text recognition model to be trained.

[0055] Furthermore, after obtaining the actual output results, loss processing can be applied to both the theoretical and actual output results according to a preset loss function. The preset loss function can be any pre-defined loss function; optionally, it can be the cross-entropy loss function.

[0056] In practical applications, the loss value of the model is obtained by applying a loss function to both the theoretical and actual output results. This loss value is then used to correct the model parameters of the text recognition model to be trained. Specifically, when correcting the model parameters using the loss value, the convergence of the preset loss function can be used as a training objective. For example, whether the training error is less than a preset error, whether the error trend is stable, or whether the current iteration count is equal to a preset number. If convergence is detected—for example, the training error of the preset loss function is less than the preset error, or the error trend is stable—it indicates that the text recognition model to be trained has completed training, and iterative training can be stopped. If convergence has not been detected, other training samples can be obtained to continue training the text recognition model until the training error of the preset loss function is within a preset range. When the training error of the preset loss function converges, the trained text recognition model can be used as the target text recognition model. That is, when the text to be processed is input into the target text recognition model, the recognition result of the entity to be recognized in the text can be accurately obtained.

[0057] It should be noted that during model training, in order to prevent the problem of parameter overfitting, the random deactivation (Dropout) technique can be used. Its principle is to randomly disable a portion of the hidden layer nodes in the neural network, so that they do not participate in the weight parameter update process, thereby avoiding excessive dependence between the hidden layer nodes of the neural network and thus alleviating the problem of parameter overfitting.

[0058] The technical solution of this invention obtains text to be processed, including entities to be identified, and further performs identification processing on the text to be processed based on a target text recognition model to determine the recognition result corresponding to the entities to be identified. This solves the problems of low data utilization, low management efficiency, high retrieval difficulty, and low recognition efficiency in the prior art when performing text recognition. It achieves the effect of improving the accuracy and efficiency of text entity recognition. By combining the contextual semantic features and grammatical features of the text sequence, the performance and robustness of the model are improved.

[0059] Example 2

[0060] Figure 2 This is a flowchart of a text processing method provided in Embodiment 2 of the present invention. Based on the foregoing embodiments, S120 has been further refined, and the specific implementation method can be found in the technical solution of this embodiment. Technical terms that are the same as or similar to those in the above embodiments will not be repeated here.

[0061] like Figure 2 As shown, the method includes:

[0062] S210. Obtain the text to be processed, which includes the entity to be identified.

[0063] S220. Extract the semantic features of each word in the text to be processed based on the semantic feature extraction sub-model to obtain the semantic feature matrix.

[0064] In this embodiment, after obtaining the text to be processed, the text to be processed can be input into the target text recognition model. Then, the text to be processed can be processed based on the semantic feature extraction sub-model in the target text recognition model to extract the semantic features of each word in the text to be processed, thereby obtaining the semantic feature matrix.

[0065] In practical applications, the text to be processed is input into the semantic feature extraction sub-model. The word vector features of each word in the text can be extracted by the word vector module in the semantic feature extraction sub-model to obtain the word vector matrix. Furthermore, in order to improve the performance of the model, other feature information can be added, such as relative position features and part-of-speech features. Finally, these features are concatenated together to obtain the semantic feature matrix.

[0066] Optionally, semantic features of each word in the text to be processed are extracted based on the semantic feature extraction sub-model to obtain a semantic feature matrix, including: processing the text to be processed based on the word vector module in the semantic feature extraction sub-model to obtain a word feature matrix corresponding to the text to be processed; processing the text to be processed based on the pre-deployed position vector lookup table in the semantic feature extraction sub-model to obtain a relative position feature matrix corresponding to the text to be processed; and processing the text to be processed based on the pre-deployed part-of-speech vector lookup table in the semantic feature extraction sub-model to obtain a part-of-speech feature matrix corresponding to the text to be processed; and concatenating the word feature matrix, the relative position feature matrix, and the part-of-speech feature matrix to obtain the semantic feature matrix.

[0067] In this embodiment, the word feature matrix can be a matrix representing the word representation vector of each word in the text to be processed. The position vector lookup table can be a table used to determine the distance between each word in the text and two predefined entities. The relative position feature matrix can be a matrix representing the relative distance between each word in the text to be processed and two predefined entities. The part-of-speech vector lookup table can include a part-of-speech dictionary and a corresponding vector lookup table used to represent the part-of-speech vector of each word in the text. The part-of-speech feature matrix can be a matrix representing the part-of-speech features of each word in the text to be processed.

[0068] In practical applications, the text to be processed is input into the semantic feature extraction sub-model. The word vector module processes the text, extracting the word representation vector features of each word. Based on the word representation vectors corresponding to each word, a feature matrix is ​​constructed, which serves as the word feature matrix corresponding to the text. Simultaneously, when the text is input into the feature extraction sub-model, it can also be processed using a pre-deployed position vector lookup table within the semantic feature extraction sub-model. This extracts the relative position features between each word and two pre-defined entities. Specifically, for each word in the text, after determining the vector representations corresponding to two relative positions in the position vector lookup table, the position vectors of the distances to the two entities are concatenated to obtain the final position feature vector. Furthermore, based on the positional feature vectors corresponding to each word, a feature matrix is ​​constructed. This feature matrix serves as the relative positional feature matrix corresponding to the text to be processed. Generally speaking, words closer to entities may provide more useful information for relation classification. Therefore, introducing relative positional information can better improve the performance of the relation extraction model. Simultaneously, when the text to be processed is input into the feature extraction sub-model, it can also be processed according to the pre-deployed part-of-speech vector lookup table to extract the part-of-speech features of each word in the text to be processed. Based on the part-of-speech feature vectors of each word, a feature matrix is ​​constructed. This feature matrix serves as the part-of-speech feature matrix corresponding to the text to be processed. Further, the word feature matrix, the relative positional feature matrix, and the part-of-speech feature matrix can be concatenated together to obtain the semantic feature matrix corresponding to the text to be processed.

[0069] For example, the text to be processed may include N words, and the corresponding text sequence is represented as: {t1,t2,t3,…,t…} N The input is then fed into the pre-trained ELMo word vector model. The word representation vector of each word is extracted from the hidden layer of the top-level LSTM of the ELMo word vector model, denoted as}. A word feature matrix is ​​constructed based on the word representation vector of each word; the position vector lookup table pre-deployed in the semantic feature extraction sub-model can be represented as follows: Where n1 represents the number of positions and d1 represents the dimension of the position vector, for each word in the text to be processed, the vector representations corresponding to two relative positions are looked up in the position vector lookup table, and the position vectors of the distances between the two entities are concatenated to obtain the final position feature vector, represented as follows: The part-of-speech vector lookup table can include a part-of-speech dictionary and the corresponding vector lookup table. Where n2 represents the number of part-of-speech tags in the part-of-speech dictionary, and d2 represents the dimension of the part-of-speech vector. For each word in the text to be processed, the part-of-speech tag can be determined according to the part-of-speech dictionary, and then the corresponding part-of-speech vector can be determined in the vector lookup table based on the part-of-speech tag, represented as... Finally, the word feature matrix, relative position feature matrix, and part-of-speech feature matrix are concatenated to obtain the semantic feature matrix {w1, w2, w3, ..., w N},in,

[0070] S230. Extract the context features of the semantic feature matrix based on the context feature extraction sub-model to obtain the context feature matrix.

[0071] In this embodiment, after obtaining the semantic feature matrix, the semantic feature matrix can be input into the context feature extraction sub-model. The context feature extraction sub-model processes the semantic feature matrix to extract the context features in the semantic feature matrix, thereby obtaining the context feature matrix.

[0072] In practical applications, text entity extraction tasks typically involve long and complex texts. Therefore, after determining the semantic features of each word in the text, the context features of the text can be analyzed based on these semantic features to obtain the context feature matrix corresponding to the text. Specifically, after obtaining the semantic feature matrix representing the semantic features of each word, the semantic feature matrix can be input into the context feature extraction sub-model to extract context features, thereby outputting the context feature matrix.

[0073] For example, in order to simultaneously obtain the contextual information of the text, the context feature extraction sub-model can be a two-layer bidirectional long short-term memory network (Bi-LSTM). The forward LSTM model outputs the hidden state sequence corresponding to each word, and the backward LSTM model concatenates the hidden state sequence according to the text sequence of the text to be processed to obtain the complete hidden state sequence, which can be used as the context feature matrix corresponding to the text to be processed.

[0074] For example, the feature extraction process of the semantic feature matrix by the Bi-LSTM model can be represented by the following formula:

[0075] i t =δ(W i *[h t-1 ,x t ]+b i )

[0076] f t =δ(W f *[h t-1 ,x t ]+b f )

[0077] O t =δ(W o *[h t-1 ,x t ]+b o )

[0078] C t =f t *C t-1 +i t *tan(W c *[h t-1 ,x t ]+b c )

[0079] h t =O t *tanh(C t )

[0080] Among them, i t f t O t This represents the three gated units of each LSTM unit: the input gate, the forget gate, and the output gate, C. t h represents the output state of the output layer at time t. t x represents the output state of the hidden layer at time t. t Let W represent the input at time t, δ() represent the activation function, tan() represent the tangent activation function, tanh() represent the hyperbolic tangent activation function, and W i W f W o Represents the hidden state vector h t and input vector x t The weight matrix, b i b f b o and b c This represents the offset vector.

[0081] S240. The context feature matrix is ​​processed based on the sequence feature extraction sub-model to obtain the first feature matrix.

[0082] In this embodiment, after obtaining the context feature matrix, the context feature matrix can be input into the sequence feature extraction sub-model to process the context feature matrix based on the sequence feature extraction sub-model to obtain the first feature matrix.

[0083] In practical applications, in order to extract the contextual semantic features of the text to be processed under different spatial dimensions, or in order to enhance the features in the context feature matrix, after obtaining the context feature matrix, the context features can be processed based on the sequence feature extraction sub-model, so as to output the first feature matrix.

[0084] For example, the sequence feature extraction sub-model can be a neural network model trained based on a multi-head self-attention mechanism algorithm. Those skilled in the art will understand that the attention mechanism mainly involves three feature spaces: the query space, the key space, and the value space. The first feature matrix can be obtained by projecting the context feature matrix into these three feature spaces respectively, and updating the context feature matrix based on the projection matrices of these three feature spaces.

[0085] For example, the process of multi-head attention mechanism in processing the context feature matrix can be represented by the following formula:

[0086]

[0087] head i =Attention(K) i Q i V i )

[0088] MultiHead(K,Q,V)=Concat(head1,head2,…,head N W O

[0089] Where K, Q, V represent the key space projection matrix, query space projection matrix, and value space projection matrix obtained by performing a nonlinear transformation on the context feature matrix, d represents the update coefficients, and head i Let N represent the output of a certain head, N represent the number of multi-head self-attention heads, and MultiHead(K,Q,V) represent the final feature matrix obtained by concatenating the outputs of all heads and performing a non-linear transformation, i.e., the first feature matrix, denoted as: W O This represents the offset vector.

[0090] S250. The context feature matrix is ​​processed based on the syntax feature extraction sub-model to obtain the second feature matrix.

[0091] In this embodiment, after obtaining the context feature matrix, the context feature matrix can also be input into the syntax feature extraction sub-model to extract the syntax features in the context feature matrix, thereby obtaining the second feature matrix.

[0092] In practical applications, in order to analyze and extract the grammatical features of each word in the text to be processed, the context feature matrix can be input into the sequence feature extraction sub-model at the same time as the grammatical feature extraction sub-model. The grammatical feature extraction sub-model performs multiple convolution processes on the context feature matrix and the target adjacency matrix pre-deployed in the sub-model corresponding to the document dependency graph, so as to aggregate the multi-level neighbor information of the node together as a new representation of the node, and finally output the second feature matrix.

[0093] It should be noted that the document dependency graph can be constructed based on documents with expertise in the relevant domain. For example, a document dependency graph could be constructed based on documents with expertise in the power industry. In practical applications, the document dependency graph is transformed into an adjacency matrix to obtain the corresponding adjacency matrix. Specifically, according to the constructed document dependency graph, each word corresponds to a node in the graph. If there is an edge connecting two nodes i and j, the corresponding position in the adjacency matrix is ​​1, i.e., A. ij =1, otherwise the corresponding position in the adjacency matrix is ​​0, i.e., A ij =0. For reflexive edges in a document dependency graph, we use A ii =1 and A jj =1 indicates that the adjacency matrix corresponding to the final document dependency graph is a symmetric matrix. Typically, each node in the document dependency graph has a different degree. This can easily lead to the model favoring nodes with higher degrees when updating node representations. However, nodes with higher degrees do not necessarily contain more useful information. Normalization removes the influence of node degree, thus yielding the target adjacency matrix corresponding to the document dependency graph.

[0094] For example, the syntax feature extraction sub-model can be a Graph Convolution Network (GCN). The process of processing the context feature matrix based on the Graph Convolution Network can be represented by the following formula:

[0095]

[0096] in, This represents the representation of the i-th node at level l. W represents the representation of the neighboring nodes of the i-th node at level l-1. (l) Let b represent the weight matrix of the l-th layer. (l) This represents the bias vector of the l-th layer. ρ represents the degree of the i-th node, and ρ represents the nonlinear activation function, such as the ReLU activation function.

[0097] For example, the output of GCN, i.e., the second feature matrix, can be represented as H. gcn .

[0098] S260. Based on the relational classification sub-model, the first feature matrix and the second feature matrix are processed to obtain the recognition result corresponding to the entity to be identified.

[0099] In this embodiment, after obtaining the first feature matrix and the second feature matrix, the first feature matrix and the second feature matrix can be input into the relation classification sub-model to process the first feature matrix and the second feature matrix based on the relation classification sub-model, thereby outputting the recognition result corresponding to the entity to be identified.

[0100] In practical applications, the first feature matrix and the second feature matrix are input into the relation classification model. The first feature matrix and the second feature matrix can be processed based on the multiple neural network layers included in the relation classification model, so as to finally output the entity recognition result and entity relation recognition result corresponding to the entity to be identified.

[0101] Optionally, the first feature matrix and the second feature matrix are processed based on the relation classification sub-model to obtain the recognition result corresponding to the entity to be identified, including: performing max pooling operation on the first feature matrix and the second feature matrix respectively based on the pooling layer to obtain the first feature matrix to be concatenated and the second feature matrix to be concatenated; concatenating the first feature matrix to be concatenated and the second feature matrix to be concatenated to obtain the target feature matrix; and processing the target feature matrix based on the multi-layer feedforward neural network module to obtain the recognition result corresponding to the entity to be identified.

[0102] The pooling layer can be a neural network layer that reduces the dimensionality of the feature matrix to obtain higher-level features. Specifically, it can process the feature matrix according to a pre-deployed nonlinear pooling function to reduce the spatial size of the feature matrix, thereby outputting higher-level feature information. In this embodiment, the nonlinear pooling function can be a max pooling function, which performs max pooling on the first and second feature matrices based on the pooling layer to obtain the first and second feature matrices to be concatenated. Those skilled in the art should understand that for the max pooling operation, only the maximum value in the feature matrix is ​​selected to enter the next layer, while other elements are not included. Therefore, the max pooling operation extracts the part with the strongest response from the feature matrix to enter the next layer. This method discards a large amount of redundant information in the network, making the network easier to optimize.

[0103] The Multi-Layer Feed Forward Network (MLB) module can be a neural network consisting of an input layer, one or more hidden layers, and an output layer. Those skilled in the art should understand that a multi-layer feed forward neural network, also known as a multi-layer perceptron, can be used to implement binary or multi-class classification tasks.

[0104] In practical applications, the first feature matrix and the second feature matrix are input into the relation classification sub-model. First, the first feature matrix and the second feature matrix can be max-pooled according to the max-pooling function pre-deployed in the pooling layer to obtain the first feature matrix to be concatenated and the second feature matrix to be concatenated. Then, the first feature matrix and the second feature matrix to be concatenated can be concatenated together to obtain the target feature matrix. Finally, the target feature matrix can be input into the multilayer feedforward neural network module, and the target feature matrix can be processed sequentially based on the input layer, one or more hidden layers and the output layer in the multilayer feedforward neural network module, so as to obtain the recognition result corresponding to the entity to be identified.

[0105] The technical solution of this invention obtains text to be processed, including entities to be identified, and further performs identification processing on the text to be processed based on a target text recognition model to determine the recognition result corresponding to the entities to be identified. This solves the problems of low data utilization, low management efficiency, high retrieval difficulty, and low recognition efficiency in the prior art when performing text recognition. It achieves the effect of improving the accuracy and efficiency of text entity recognition. By combining the contextual semantic features and grammatical features of the text sequence, the performance and robustness of the model are improved.

[0106] Example 3

[0107] Figure 3 This is a schematic diagram of the structure of a text processing device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes a text acquisition module 310 and a text processing module 320.

[0108] The text to be processed module 310 is used to acquire text to be processed that includes entities to be identified.

[0109] The text processing module 320 is used to perform recognition processing on the text to be processed based on the target text recognition model, and determine the recognition result corresponding to the entity to be recognized. The target text recognition model includes at least one target sub-model, which includes a semantic feature extraction sub-model, a context feature extraction sub-model, a sequence feature extraction sub-model, a grammatical feature extraction sub-model, and a relation classification sub-model. The recognition result includes the entity recognition result and / or entity relation recognition result of the entity to be recognized.

[0110] The technical solution of this invention obtains text to be processed, including entities to be identified, and further performs identification processing on the text to be processed based on a target text recognition model to determine the recognition result corresponding to the entities to be identified. This solves the problems of low data utilization, low management efficiency, high retrieval difficulty, and low recognition efficiency in the prior art when performing text recognition. It achieves the effect of improving the accuracy and efficiency of text entity recognition. By combining the contextual semantic features and grammatical features of the text sequence, the performance and robustness of the model are improved.

[0111] Optionally, the device further includes: a raw information acquisition module and a raw information preprocessing module.

[0112] The raw information acquisition module is used to acquire raw information to be processed, wherein the raw information includes at least one form of information representation;

[0113] The original information preprocessing module is used to preprocess the original information based on a pre-set information preprocessing method to obtain the text to be processed, wherein the information preprocessing method includes at least one of deduplication, noise reduction, sentence segmentation, and word segmentation.

[0114] Optionally, the connection relationship of each target sub-model in the target text recognition model includes: the output of the semantic feature extraction sub-model is the input of the context feature extraction sub-model; the output of the context feature extraction sub-model is the input of the sequence feature extraction sub-model; the output of the context feature extraction sub-model is the input of the syntax feature extraction sub-model; and the outputs of the sequence feature extraction sub-model and the syntax feature extraction sub-model are the inputs of the relation classification sub-model.

[0115] Optionally, the text processing module 320 includes: a semantic feature extraction unit, a context feature extraction unit, a first feature matrix determination unit, a second feature matrix determination unit, and a recognition result determination unit.

[0116] The semantic feature extraction unit is used to extract the semantic features of each word in the text to be processed based on the semantic feature extraction sub-model, and obtain a semantic feature matrix;

[0117] The context feature extraction unit is used to obtain a context feature matrix based on the context features of the semantic feature matrix from the context feature extraction sub-model.

[0118] The first feature matrix determination unit is used to process the context feature matrix based on the sequence feature extraction sub-model to obtain a first feature matrix; and

[0119] The second feature matrix determination unit is used to process the context feature matrix based on the syntax feature extraction sub-model to obtain the second feature matrix;

[0120] The identification result determination unit is used to process the first feature matrix and the second feature matrix based on the relationship classification sub-model to obtain the identification result corresponding to the entity to be identified.

[0121] Optionally, the semantic feature extraction sub-model includes a word vector module, and the semantic feature extraction unit includes: a word feature matrix determination sub-unit, a relative position feature matrix determination sub-unit, a part-of-speech vector feature matrix determination sub-unit, and a matrix concatenation sub-unit.

[0122] A word feature matrix determination subunit is used to process the text to be processed based on the word vector module in the semantic feature extraction submodel, to obtain a word feature matrix corresponding to the text to be processed; and...

[0123] A relative position feature matrix determination subunit is used to process the text to be processed based on a pre-deployed position vector lookup table in the semantic feature extraction submodel, to obtain a relative position feature matrix corresponding to the text to be processed; and...

[0124] The part-of-speech vector feature matrix determination subunit is used to process the text to be processed based on the part-of-speech vector lookup table pre-deployed in the semantic feature extraction sub-model, so as to obtain the part-of-speech feature matrix corresponding to the text to be processed;

[0125] The matrix concatenation subunit is used to concatenate the word feature matrix, the relative position feature matrix, and the part-of-speech feature matrix to obtain the semantic feature matrix.

[0126] Optionally, the relationship classification sub-model includes a pooling layer and a multi-layer feedforward neural network module, and the identification result determination unit includes: a sub-unit for determining the feature matrix to be concatenated, a sub-unit for determining the target feature matrix, and a target feature matrix processing unit.

[0127] The feature matrix to be concatenated determination subunit is used to perform max pooling operations on the first feature matrix and the second feature matrix respectively based on the pooling layer pairs in the relation classification submodel to obtain the first feature matrix to be concatenated and the second feature matrix to be concatenated.

[0128] The target feature matrix determination sub-unit is used to concatenate the first feature matrix to be concatenated and the second feature matrix to be concatenated to obtain the target feature matrix;

[0129] The target feature matrix processing unit is used to process the target feature matrix based on the multi-layer feedforward neural network module in the relation classification sub-model to obtain the recognition result corresponding to the entity to be identified.

[0130] Optionally, the device further includes: a training sample acquisition module, an actual output result determination module, a loss processing module, and a model parameter correction module.

[0131] The training sample acquisition module is used to acquire multiple training samples, wherein the training samples include: training sample text, the training sample text includes an entity to be identified, and a theoretical output result corresponding to the training sample text, the theoretical output result includes a theoretical entity identification result and a theoretical entity relationship identification result corresponding to the entity to be identified;

[0132] The actual output result determination module is used to input the training samples into the text recognition model to be trained to obtain the actual output result; wherein, the actual output result includes the actual entity recognition result and the actual entity relationship recognition result;

[0133] The loss processing module is used to perform loss processing on the theoretical output result and the actual output result according to a preset loss function;

[0134] The model parameter correction module is used to correct the model parameters in the text processing model to be trained based on the loss value, so as to obtain the target text recognition model.

[0135] The text processing apparatus provided in the embodiments of the present invention can execute the text processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0136] Example 4

[0137] Figure 4A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0138] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0139] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0140] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as text processing methods.

[0141] In some embodiments, the text processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the text processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the text processing method by any other suitable means (e.g., by means of firmware).

[0142] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0146] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0147] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0148] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0149] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A text processing method characterized by, Comprise: Obtaining a to-be-processed text including a to-be-identified entity; Based on the target text recognition model, the to-be-processed text is identified and processed to determine the recognition result corresponding to the to-be-identified entity, wherein the target text recognition model includes at least one target sub-model, and the at least one target sub-model includes a semantic feature extraction sub-model, a context feature extraction sub-model, a sequence feature extraction sub-model, a syntax feature extraction sub-model and a relationship classification sub-model; the recognition result includes entity recognition result and / or entity relationship recognition result of the to-be-identified entity; The connection relationship of each target sub-model in the target text recognition model comprises: The output of the semantic feature extraction sub-model is the input of the context feature extraction sub-model; The output of the context feature extraction sub-model is the input of the sequence feature extraction sub-model; The output of the context feature extraction sub-model is the input of the syntax feature extraction sub-model; The output of the sequence feature extraction sub-model and the output of the syntax feature extraction sub-model are the input of the relationship classification sub-model; The based on the target text recognition model, the to-be-processed text is identified and processed to determine the recognition result corresponding to the to-be-identified entity, comprising: Based on the semantic feature extraction sub-model, the semantic features of each word in the to-be-processed text are extracted to obtain a semantic feature matrix; Based on the context feature extraction sub-model, the context features of the semantic feature matrix are extracted to obtain a context feature matrix; Based on the sequence feature extraction sub-model, the context feature matrix is processed to obtain a first feature matrix; and, Based on the syntax feature extraction sub-model, the context feature matrix is processed to obtain a second feature matrix; Based on the relationship classification sub-model, the first feature matrix and the second feature matrix are processed to obtain the recognition result corresponding to the to-be-identified entity; The based on the relationship classification sub-model, the first feature matrix and the second feature matrix are processed to obtain the recognition result corresponding to the to-be-identified entity, comprising: Based on the pooling layer in the relationship classification sub-model, the first feature matrix and the second feature matrix are respectively subjected to a max-pooling operation to obtain a first to-be-spliced feature matrix and a second to-be-spliced feature matrix; The first to-be-spliced feature matrix and the second to-be-spliced feature matrix are spliced to obtain a target feature matrix; Based on the multi-layer feedforward neural network module in the relationship classification sub-model, the target feature matrix is processed to obtain the recognition result corresponding to the to-be-identified entity.

2. The method of claim 1, wherein, Also include: Obtaining to-be-processed original information, wherein the original information includes at least one information representation information; Based on the pre-set information preprocessing mode, the original information is preprocessed to obtain a to-be-processed text, wherein the information preprocessing mode includes at least one of de-duplication, denoising, sentence segmentation and word segmentation.

3. The method of claim 1, wherein, The based on the semantic feature extraction sub-model, the semantic features of each word in the to-be-processed text are extracted to obtain a semantic feature matrix, comprising: obtaining a word feature matrix corresponding to the to-be-processed text based on a word vector module in the semantic feature extraction sub-model; and obtaining a relative position feature matrix corresponding to the to-be-processed text based on a pre-deployed position vector lookup table in the semantic feature extraction sub-model; and obtaining a part-of-speech feature matrix corresponding to the to-be-processed text based on a pre-deployed part-of-speech vector lookup table in the semantic feature extraction sub-model; concatenating the word feature matrix, the relative position feature matrix, and the part-of-speech feature matrix to obtain the semantic feature matrix.

4. The method of claim 1, wherein, Further comprising: obtaining a plurality of training samples, wherein the training samples include: training sample texts, the training sample texts including to-be-identified entities, theoretical output results corresponding to the training sample texts, the theoretical output results including theoretical entity identification results and theoretical entity relationship identification results corresponding to the to-be-identified entities; inputting the training samples into a to-be-trained text recognition model to obtain actual output results; wherein the actual output results include actual entity identification results and actual entity relationship identification results; performing loss processing on the theoretical output results and the actual output results according to a pre-set loss function; correcting model parameters in the to-be-trained text recognition model based on a loss value to obtain the target text recognition model.

5. A text processing apparatus characterized by comprising: Comprising: a to-be-processed text acquisition module configured to acquire to-be-processed text including to-be-identified entities; a to-be-processed text processing module configured to perform recognition processing on the to-be-processed text based on a target text recognition model to determine identification results corresponding to the to-be-identified entities, wherein the target text recognition model includes at least one target sub-model, the at least one target sub-model including a semantic feature extraction sub-model, a context feature extraction sub-model, a sequence feature extraction sub-model, a syntax feature extraction sub-model, and a relationship classification sub-model; the identification results include entity identification results and / or entity relationship identification results of the to-be-identified entities; the to-be-processed text processing module is specifically configured to: extract semantic features of each word in the to-be-processed text based on the semantic feature extraction sub-model to obtain a semantic feature matrix; extract context features of the semantic feature matrix based on the context feature extraction sub-model to obtain a context feature matrix; perform processing on the context feature matrix based on the sequence feature extraction sub-model to obtain a first feature matrix; and perform processing on the context feature matrix based on the syntax feature extraction sub-model to obtain a second feature matrix; perform processing on the first feature matrix and the second feature matrix based on the relationship classification sub-model to obtain identification results corresponding to the to-be-identified entities; the to-be-processed text processing module is further specifically configured to: perform maximum pooling operations on the first feature matrix and the second feature matrix based on a pooling layer in the relationship classification sub-model to obtain a first to-be-concatenated feature matrix and a second to-be-concatenated feature matrix; The first to-be-spliced feature matrix and the second to-be-spliced feature matrix are spliced to obtain a target feature matrix; The target feature matrix is processed based on a multi-layer feedforward neural network module in the relationship classification sub-model to obtain an identification result corresponding to the to-be-identified entity; The connection relationship of each target sub-model in the target text recognition model comprises: An output of the semantic feature extraction sub-model is an input of the context feature extraction sub-model; An output of the context feature extraction sub-model is an input of the sequence feature extraction sub-model; An output of the context feature extraction sub-model is an input of the syntax feature extraction sub-model; An output of the sequence feature extraction sub-model and an output of the syntax feature extraction sub-model are inputs of the relationship classification sub-model.

6. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the text processing method of any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the text processing method of any one of claims 1-4 when executed.

Citation Information

Patent Citations

  • Multivariate relation extraction method and extraction system based on multi-model fusion

    CN114925693A