Entity extraction method and device based on adaptive weight fusion
Through the entity extraction method of adaptive weight fusion, combined with the pre-trained language model, BiLSTM network and adaptive multi-head attention mechanism, the semantic loss problem caused by BERT in long text truncation is solved, and the accuracy of entity extraction in industrial quality inspection reports is improved.
Patent Information
- Application Number
- CN202511136738.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-14
AI Technical Summary
When processing long texts, the existing technology, the BERT model, is prone to semantic information loss due to truncation, resulting in low entity extraction accuracy, especially in industrial quality inspection reports where it is difficult to accurately identify key entities.
An entity extraction method with adaptive weight fusion is adopted, combined with a pre-trained language model, a BiLSTM network and an adaptive multi-head attention mechanism. Global and local features are dynamically fused through adaptive weights to enhance the understanding of industrial professional terms and entity recognition capabilities.
It effectively solves the semantic loss problem caused by BERT's truncation of long texts and improves the accuracy of entity extraction, especially the entity recognition effect in complex industrial quality inspection texts.
Smart Images

Figure CN120632111A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text analysis, and in particular to an entity extraction method and device based on adaptive weight fusion. Background Art
[0002] Entity extraction, a core task in natural language processing, faces widespread technical challenges in cross-domain applications (such as industrial quality inspection), including domain terminology and semantic ambiguity, complex relationship modeling, and contextual dependencies. In natural language processing, entity extraction techniques based on pre-trained language models have become mainstream, including BERT (Bidirectional Encoder Representation from Transformer).
[0003] However, while BERT can leverage bidirectional contextual information, it can be prone to semantic loss due to truncation or segmentation when processing long texts, leading to low entity extraction accuracy. For example, when processing industrial quality inspection reports containing complex process descriptions, BERT truncates the long text, resulting in incomplete contextual information for some key entities and ultimately inability to accurately identify them. Summary of the Invention
[0004] (1) Technical issues to be resolved In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides an entity extraction method and device based on adaptive weight fusion, which solves the technical problem of low entity recognition accuracy in the prior art.
[0005] (2) Technical solution In order to achieve the above objectives, the main technical solutions adopted by the present invention include: In the first aspect, an embodiment of the present invention provides an entity extraction method based on adaptive weight fusion, comprising: obtaining an industrial text to be extracted; inputting the industrial text to be extracted into a pre-trained entity extraction model for entity extraction, and obtaining an entity extraction result of the industrial text to be extracted; wherein the entity extraction model comprises an input preprocessing layer, a text encoding layer, a feature extraction layer and a classification layer; the input preprocessing layer is used to segment the industrial text to be extracted, and to mark the part of speech of each segmented word through a preset part-of-speech mapping table to obtain a part-of-speech tag sequence; the text encoding layer is used to generate a word embedding index and a filling mask for the industrial text to be extracted; the feature extraction layer is used to input the word embedding index and the filling mask into In the pre-trained language model, the output features of all hidden layers of the pre-trained language model are obtained, and the outputs of all hidden layers are weightedly fused to obtain weighted semantic features, and the weighted semantic features are input into the BiLSTM network to obtain the forward hidden state and the backward hidden state respectively, and the forward hidden state and the backward hidden state are spliced to obtain the spliced features, and according to the adaptive weights, the spliced features and the part-of-speech tag sequence, the target fusion features for predicting the entity label corresponding to each word in the industrial text to be extracted are obtained; the classification layer is used to input the target fusion features into the fully connected layer, and determine the entity extraction result based on the probability obtained by softmax normalization.
[0006] With the help of the above technical solution, the embodiment of the present application effectively solves the semantic loss problem caused by BERT's long text truncation by integrating the pre-trained language model, BiLSTM network and adaptive multi-head attention mechanism into the entity extraction model, thereby improving the accuracy of entity extraction.
[0007] In a possible embodiment, the parts of speech of the segmented words to be extracted from the industrial text include directional nouns, object nouns, defect type nouns, adjectives, and verbs.
[0008] In one possible embodiment, the adaptive weight includes a first sub-adaptive weight; the feature extraction layer is further used to input the spliced features into the adaptive multi-head attention layer to generate attention features, and the spliced features and the attention features are combined with the first sub-adaptive weight to perform feature fusion to obtain intermediate fusion features, and based on the intermediate fusion features and the part-of-speech tag sequence, the target fusion features are obtained.
[0009] In one possible embodiment, the adaptive weight includes a second sub-adaptive weight; the part-of-speech tag sequence includes a part-of-speech tag for each word segment; the feature extraction layer is further used to input the part-of-speech tag of each word segment into the embedding layer to obtain a part-of-speech vector for each word, and the part-of-speech vector of each word and its corresponding intermediate fusion feature are combined with the second sub-adaptive weight to perform feature fusion to obtain a target fusion feature; wherein, a part-of-speech dictionary is provided in the embedding layer for mapping each part-of-speech tag into a continuous vector representation of a fixed dimension.
[0010] In a possible embodiment, the parts of speech recorded in the part-of-speech dictionary include all known parts of speech recorded in a preset part-of-speech mapping table and an unregistered part of speech added through a reserved position.
[0011] In the second aspect, an embodiment of the present invention provides an entity extraction device based on adaptive weight fusion, comprising: an acquisition module for acquiring industrial text to be extracted; an input module for inputting the industrial text to be extracted into a pre-trained entity extraction model for entity extraction, and obtaining an entity extraction result of the industrial text to be extracted; wherein the entity extraction model comprises an input preprocessing layer, a text encoding layer, a feature extraction layer and a classification layer; the input preprocessing layer is used to segment the industrial text to be extracted, and to mark the part of speech of each segmented word through a preset part-of-speech mapping table to obtain a part-of-speech tag sequence; the text encoding layer is used to generate a word embedding index and a filling mask for the industrial text to be extracted; the feature extraction layer is used to embed the word embedding index and the classification layer. The padding mask is input into the pre-trained language model to obtain the output features of all hidden layers of the pre-trained language model, and the outputs of all hidden layers are weightedly fused to obtain weighted semantic features, and the weighted semantic features are input into the BiLSTM network to obtain the forward hidden state and the backward hidden state respectively, and the forward hidden state and the backward hidden state are spliced to obtain the spliced features, and according to the adaptive weights, the spliced features and the part-of-speech tag sequence, the target fusion features for predicting the entity label corresponding to each word in the industrial text to be extracted are obtained; the classification layer is used to input the target fusion features into the fully connected layer, and determine the entity extraction result based on the probability obtained by softmax normalization.
[0012] In a possible embodiment, the parts of speech of the segmented words to be extracted from the industrial text include directional nouns, object nouns, defect type nouns, adjectives, and verbs.
[0013] In one possible embodiment, the adaptive weight includes a first sub-adaptive weight; the feature extraction layer is further used to input the spliced features into the adaptive multi-head attention layer to generate attention features, and the spliced features and the attention features are combined with the first sub-adaptive weight to perform feature fusion to obtain intermediate fusion features, and based on the intermediate fusion features and the part-of-speech tag sequence, the target fusion features are obtained.
[0014] In one possible embodiment, the adaptive weight includes a second sub-adaptive weight; the part-of-speech tag sequence includes a part-of-speech tag for each word segment; the feature extraction layer is further used to input the part-of-speech tag of each word segment into the embedding layer to obtain a part-of-speech vector for each word, and the part-of-speech vector of each word and its corresponding intermediate fusion feature are combined with the second sub-adaptive weight to perform feature fusion to obtain a target fusion feature; wherein, a part-of-speech dictionary is provided in the embedding layer for mapping each part-of-speech tag into a continuous vector representation of a fixed dimension.
[0015] In a possible embodiment, the parts of speech recorded in the part-of-speech dictionary include all known parts of speech recorded in a preset part-of-speech mapping table and an unregistered part of speech added through a reserved position.
[0016] In a third aspect, an embodiment of the present application also provides an electronic device, comprising: a memory for storing a computer program; a processor for executing the computer program stored in the memory; when the computer program is executed, the processor is used to execute an entity extraction method based on adaptive weight fusion as described above.
[0017] In a fourth aspect, an embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes an entity extraction method based on adaptive weight fusion as described above.
[0018] In order to make the above-mentioned objectives, features and advantages to be achieved by the embodiments of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 A flowchart of an entity extraction method based on adaptive weight fusion provided by an embodiment of the present application is shown; Figure 2 A schematic diagram of the structure of an entity extraction model provided in an embodiment of the present application is shown; Figure 3 A schematic diagram of a feature extraction structure provided by an embodiment of the present application is shown; Figure 4A structural block diagram of an entity extraction device based on adaptive weight fusion provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0021] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0022] While the DeepSeek-coder model currently can extract questions, it still has some biases for more specialized terminology, and contextual logic can still break down. In the field of industrial quality inspection, the DeepSeek-coder model's recognition accuracy for specific industry terms like "surface roughness" and "form and position tolerance" still needs to be improved. Furthermore, understanding and grasping contextual logic within complex text structures remains a challenge.
[0023] Furthermore, existing attention mechanisms (such as multi-head attention) are generally based on static weight allocation and cannot dynamically adjust feature importance based on the input text. For example, in a quality inspection report, the importance of entities such as "material" and "size" varies with the context, but traditional models struggle to adaptively focus on key features, resulting in redundant information interfering with entity boundary judgment.
[0024] Furthermore, the distribution of parts of speech in industrial quality inspection texts exhibits significant domain characteristics (for example, "defect type nouns" account for 30%, far exceeding the 15% in general text). However, existing methods often rely on character or word embeddings and do not explicitly utilize part-of-speech information to enhance entity type recognition. For example, the boundary between "roughness" (a noun) and "rough" (an adjective) is difficult to accurately identify using word embeddings alone.
[0025] Based on this, the embodiment of the present application provides an entity extraction method and device based on adaptive weight fusion. By integrating the DeepSeek-Coder pre-trained language model, the bidirectional long short-term memory network (BiLSTM network) and the adaptive multi-head attention mechanism, it effectively solves the semantic loss problem caused by BERT's truncation of long texts. First, the code generation capability of the DeepSeek-Coder pre-trained language model is used to enhance the deep understanding of industrial professional terms, avoid the fixed-length truncation defect of BERT, and enable it to fully parse the complex semantic logic in long texts (such as the multi-process association in the lithium battery process description); secondly, the full text is end-to-end sequence modeled through the BiLSTM layer, and long-distance dependencies are captured bidirectionally (such as the causal relationship between "electrode material composition" and "coating peeling off after cycling" across paragraphs), eliminating information breaks caused by segmentation processing; further, the adaptive multi-head attention mechanism is introduced to use learnable weights. Dynamically integrate global attention features (directly establish semantic associations between distant terms) with BiLSTM local sequence features to strengthen cross-text dependencies of key entities (such as "diaphragm wrinkles" or "microcracks"); at the same time, combine industry-specific part-of-speech tag systems (such as defect type nouns DTN) and weights , enhance the model's sensitivity to core entities and suppress noise interference such as auxiliary words, and ultimately achieve improved accuracy of entity recognition in long industrial texts through multi-level feature fusion.
[0026] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0027] See Figure 1 , Figure 1 The flowchart of an entity extraction method based on adaptive weight fusion provided by an embodiment of the present application is shown. It should be understood that the entity extraction method can be executed by an entity extraction device based on adaptive weight fusion, and the specific device of the entity extraction device can be set according to actual needs, and the embodiment of the present application is not limited thereto. For example, the entity extraction device can be a computer, or a server, etc. Figure 1 As shown, the entity extraction method includes: Step S110: Obtain the industrial text to be extracted.
[0028] It should be understood that the specific text of the industrial text to be extracted can be set according to actual needs, and the embodiments of the present application are not limited to this.
[0029] Optionally, the industrial text to be extracted may be an industrial quality inspection report, an industrial equipment alarm log, an industrial process parameter record, or the like.
[0030] In step S120 , the industrial text to be extracted is input into a pre-trained entity extraction model to perform entity extraction, thereby obtaining an entity extraction result of the industrial text to be extracted.
[0031] It should be understood that the specific model structure of the entity extraction model can be set according to actual needs, and the embodiments of the present application are not limited thereto.
[0032] Alternatively, as Figure 2As shown, the entity extraction model includes an input preprocessing layer, a text encoding layer, a feature extraction layer and a classification layer; the input preprocessing layer is used to segment the industrial text to be extracted, and mark the part of speech of each segmented word through a preset part-of-speech mapping table to obtain a part-of-speech tag sequence; the text encoding layer is used to generate a word embedding index and a padding mask for the industrial text to be extracted; the feature extraction layer is used to input the word embedding index and the padding mask into the pre-trained language model to obtain the output features of all hidden layers of the pre-trained language model, and perform weighted fusion on the outputs of all hidden layers to obtain weighted semantic features, and input the weighted semantic features into the BiLSTM network to obtain the forward hidden state and the backward hidden state respectively, and splice the forward hidden state and the backward hidden state to obtain the spliced features, and obtain the target fusion features for predicting the entity label corresponding to each word in the industrial text to be extracted based on the adaptive weights, the spliced features and the part-of-speech tag sequence; the classification layer is used to input the target fusion features into the fully connected layer, and determine the entity extraction result based on the probability obtained by softmax normalization.
[0033] To facilitate understanding of the entity extraction model, a specific embodiment is described below.
[0034] Specifically, within the traditional BiLSTM framework, the model primarily consists of three parts: a forward LSTM layer, a backward LSTM layer, and a final merged output layer. Building on this foundation, this application introduces a new entity extraction model that combines the DeepSeek-coder pre-trained language model, employs an adaptive multi-head attention mechanism, employs a feature fusion strategy, and adds a part-of-speech weighting, aiming to enhance the model's performance and feature representation capabilities.
[0035] And, first of all, the input preprocessing layer needs to preprocess the raw data, including reading the industrial text to be extracted, reading the entity label using BIO encoding, and segmenting the text and marking the part of speech through custom part-of-speech tags (for example, "particles" → ON), and then converting the segmented part of speech into integers according to the preset part-of-speech mapping table (for example, ON → 2). Among them, in defining the part-of-speech mapping, it mainly includes the following five categories: directional nouns (for example, DN → 1), object nouns (for example, ON → 2), defect type nouns (for example, DTN → 3), adjectives (for example, JJ → 4) and verbs (for example, VB → 5). And, the above five categories well cover the characteristics of text data for defect detection and appearance inspection of some materials generated in industry.
[0036] Next, the text encoding layer performs data encoding and alignment. Because the DeepSeek-coder pre-trained language model will be used later, the DeepSeek model's tokenizer is used to add two attributes to the text: word embedding indexes and padding masks. The word embedding index breaks the text into processable units (such as words) and maps them to numeric vectors containing semantic information, enabling machine understanding of language. The padding mask indicates whether the corresponding position is an actual word, for example, 1 represents a valid actual word and 0 represents a padding word. For sense labels, the length of the sense label sequence is adjusted to a fixed length through padding or truncation to ensure alignment with the input text. For part-of-speech tags, the length of the part-of-speech tag sequence is adjusted to a fixed length through padding or truncation, and the positional encoding embedding dimension is set to the same as the hidden layer dimension of the BiLSTM network (e.g., 256 dimensions) to ensure compatibility of the concatenated feature dimensions. Finally, for data loading, the data loader is used to load data in batches (i.e., batch_size=1), supporting dynamic padding and random shuffling.
[0037] As well as Figure 3 As shown, Figure 3 FIG. 1 shows a structural diagram of a feature extraction method provided by an embodiment of the present application. Figure 3 As shown, the feature extraction layer is used for word embedding indexing and filling mask input into the pre-trained language model to obtain the output features of all hidden layers of the pre-trained language model. Among them, the specific model of the pre-processing language model can be set according to actual needs, and the embodiments of the present application are not limited to this. For example, the pre-processing language model can be a DeepSeek-Coder model. Subsequently, the pre-trained language model can be loaded by calling the Hugging Face library. The model has been trained on large-scale code and text corpora, has strong semantic understanding capabilities, and can effectively process professional texts in the field of industrial quality inspection. When loading the pre-trained language model, by enabling the full hidden layer output function, the model retains the intermediate results of each layer of Transformer when processing the text, and these results contain the semantic representation of the text at different levels of abstraction. Subsequently, dynamic weights can be assigned to each layer through an adaptive weight mechanism to achieve optimal fusion of multi-granularity features. After that, all hidden layer states of the pre-trained language model can be obtained. Assuming that the pre-trained language model has L hidden layers, the output shape of each hidden layer is (B, T, H), where B represents the batch size, T is the sequence length, and H is the hidden layer dimension. Finally, a list of length L can be obtained, and each element in the list is a tensor of shape (B, T, H). In order to obtain more representative features, it is necessary to perform a weighted fusion of all hidden layer states to obtain weighted semantic features, where the weight of each hidden layer is determined by a learnable parameter vector (i.e., the first sub-adaptive weight) is determined dynamically. For example, the parameter vector can be automatically optimized by backpropagation . Specifically: Assume that the output of l hidden layers is H l And its corresponding weight is , then the weighted semantic feature F is: .
[0038] Furthermore, the weighted semantic features (B, T, H) obtained by weighted fusion of the outputs of each hidden layer can be used as input to subsequent network layers (for example, the BiLSTM network and the adaptive multi-head attention layer). The weighted semantic features can then be used as input to the BiLSTM network. The weighted semantic features are processed by the forward LSTM and backward LSTM in the BiLSTM network to process the forward and reverse order information of the sequence, respectively, to obtain the forward hidden state and backward hidden state. The forward hidden state and backward hidden state are then concatenated to fuse the bidirectional features and obtain the concatenated features.
[0039] Subsequently, the concatenated features can be input into the adaptive multi-head attention layer to generate attention features, and the concatenated features and attention features can be fused to obtain intermediate fusion features. Among them, the adaptive multi-head attention layer performs linear transformations on the input query, key, and value respectively. Among them, for query, key, and value, weighted features are used here, and the shape is (T, B, H), where T is the sequence length, B is the batch size, and H is the feature dimension. These features are then divided into multiple heads, and attention is calculated in parallel. Finally, the results of each head are concatenated and linearly transformed to obtain the final output. And, after obtaining the multi-head attention output, adaptive weights can be used Perform feature fusion, the formula is as follows: Intermediate fusion feature = *Attention features of multi-head attention output + (1- )*Features after splicing.
[0040] The above two features come from the output of multi-head attention and the output of BiLSTM network respectively, and the weight parameters The value range of can be between [0,1]. Close to 1 means that the fusion result is more inclined to the characteristics of multi-head attention, which is suitable for tasks that rely on global semantic associations, and the weight parameter A value close to 0 indicates that the fusion result is more similar to the features of BiLSTM and is suitable for tasks that depend on sequence order.
[0041] In addition, adaptive attention and feature fusion captures the dependencies in the sequence through a multi-head attention mechanism, and then uses learnable adaptive weights to dynamically fuse the attention output and BiLSTM features. This approach can fully utilize feature information from different sources and improve the expressiveness and performance of the model.
[0042] In addition, based on the part-of-speech tags defined in the previous data preprocessing, a part-of-speech feature conversion module is constructed in the entity extraction model. The functions of this part-of-speech feature conversion module include: creating a dictionary containing all known parts of speech based on the set of part-of-speech tags defined in the preprocessing stage, and reserving an additional position to handle unregistered parts of speech (that is, the dictionary size is the number of known parts of speech plus 1. In other words, the parts of speech recorded in the part-of-speech dictionary include all known parts of speech recorded in the preset part-of-speech mapping table and one unregistered part-of-speech added through the reserved position), and mapping each part-of-speech tag to a continuous vector representation of a fixed dimension, which can be specified by the model hyperparameter pos_embedding_dim (position encoding embedding dimension). Among them, unregistered parts of speech refer to part-of-speech categories that are outside the range of part-of-speech tags defined in the preprocessing stage and have never been seen by the model.
[0043] On this basis, this application also introduces a learnable scalar parameter (i.e., the second sub-adaptive weight), and the scalar parameter The function of is to automatically learn the contribution of part-of-speech features to the final representation during training. This parameter is initialized to a random value and continuously optimized through model training to balance the importance of part-of-speech features with other features (such as semantic features). Furthermore, it converts part-of-speech tags into embedding vectors, converting the part-of-speech tag of each word into a continuous vector representation so that the model can better understand and utilize part-of-speech information. In the embedding layer, for each word in the input text, the integer index corresponding to its part-of-speech tag can be passed into the embedding layer to obtain the corresponding vector representation. For example, the input of the embedding layer corresponds to the integer index of the part-of-speech tag (such as 3), and the output corresponds to the vector representation (such as a 256-dimensional vector). The final result is a part-of-speech feature, which is a three-dimensional vector with a shape of [batch size, sequence length, embedding dimension], which represents the part-of-speech vector of each word in each batch.
[0044] After obtaining the part-of-speech features, the model fuses them with other features, concatenating the adaptively fused features and the weighted part-of-speech features along the last dimension. This concatenation operation is performed along the last dimension (i.e., the feature dimension), preserving the sequence length and batch size while only increasing the feature dimension. The concatenated feature tensor integrates semantic, sequence, and part-of-speech information. This multimodal fused feature representation serves as input to the subsequent entity classifier, used to identify various entities in industrial quality inspection texts. The formula is as follows: Target fusion feature = * Part of speech feature output + (1- )*Intermediate fusion features.
[0045] in, The part-of-speech weight controls the fusion ratio between the part-of-speech feature and the previous feature fusion result (i.e., the intermediate fusion feature). Its initial value is randomly generated. The part-of-speech feature comes from the pre-trained language model, and the last hidden state is used as the basic semantic feature.
[0046] also, and They are all independent learnable parameters, which respectively control the fusion ratio of pre-trained features and sequence features, and part-of-speech features and intermediate fusion features.
[0047] In addition, before predicting the entity tag, the model has completed a series of feature extraction and fusion operations, and finally obtained the feature vector for prediction. Specifically, the entity extraction model will fuse the output of the pre-trained language model, the output of the BiLSTM network, and the part-of-speech features to obtain a comprehensive feature representation. Assume that this feature vector is h t , which is the feature at time step t, with the shape of (B, H final ), where B is the batch size, H final is the final fused feature dimension. The fused feature vector is transformed into h t The number of entities is mapped to the entity label, and the output of the fully connected layer can represent the unnormalized score of each sample on each label.
[0048] In order to convert the output of the fully connected layer into a probability distribution, the softmax function is used to normalize the output of the fully connected layer. The softmax function converts each element in the vector into a probability value between 0 and 1, and the sum of the probabilities of all elements is 1.
[0049] After processing the softmax function, we obtain a probability vector, which represents the probability of each sample belonging to each entity label at time step t. Finally, we predict the entity label corresponding to each sample based on the probability vector. The commonly used method is to select the label with the highest probability as the prediction result to obtain the entity extraction result.
[0050] Therefore, with the help of the above technical solution, this application effectively solves the semantic loss problem caused by BERT due to long text truncation by integrating the pre-trained language model, BiLSTM network and adaptive multi-head attention mechanism, thereby improving the accuracy of entity extraction.
[0051] It should be understood that the above-mentioned entity extraction method based on adaptive weight fusion is only exemplary, and those skilled in the art can make various modifications based on the above-mentioned method, and the modified scheme also falls within the scope of protection of this application.
[0052] See Figure 4 , Figure 4 The following is a block diagram of the structure of an entity extraction device 400 based on adaptive weight fusion provided by an embodiment of the present application. It should be understood that the entity extraction device 400 is capable of executing each step in the above-mentioned method embodiment. The specific functions of the entity extraction device 400 can be found in the description above. To avoid repetition, a detailed description is appropriately omitted here. The entity extraction device 400 includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the operating system (OS) of the entity extraction device 400. Specifically, the entity extraction device 400 includes: An acquisition module 410 is used to acquire the industrial text to be extracted; An input module 420 is used to input the industrial text to be extracted into a pre-trained entity extraction model to perform entity extraction and obtain an entity extraction result of the industrial text to be extracted; Among them, the entity extraction model includes an input preprocessing layer, a text encoding layer, a feature extraction layer and a classification layer; the input preprocessing layer is used to segment the industrial text to be extracted, and mark the part of speech of each segmented word through a preset part-of-speech mapping table to obtain a part-of-speech tag sequence; the text encoding layer is used to generate a word embedding index and a padding mask for the industrial text to be extracted; the feature extraction layer is used to input the word embedding index and the padding mask into the pre-trained language model to obtain the output features of all hidden layers of the pre-trained language model, and perform weighted fusion on the outputs of all hidden layers to obtain weighted semantic features, and input the weighted semantic features into the BiLSTM network to obtain the forward hidden state and the backward hidden state respectively, and splice the forward hidden state and the backward hidden state to obtain the spliced features, and according to the spliced features and the part-of-speech tag sequence, obtain the target fusion features for predicting the entity label corresponding to each word in the industrial text to be extracted; the classification layer is used to input the target fusion features into the fully connected layer, and determine the entity extraction result based on the probability obtained by softmax normalization.
[0053] In a possible embodiment, the parts of speech of the segmented words to be extracted from the industrial text include directional nouns, object nouns, defect type nouns, adjectives, and verbs.
[0054] In one possible embodiment, the feature extraction layer is further used to input the spliced features into the adaptive multi-head attention layer to generate attention features, and perform feature fusion on the spliced features and the attention features to obtain intermediate fusion features, and obtain target fusion features based on the intermediate fusion features and the part-of-speech tag sequence.
[0055] In one possible embodiment, the part-of-speech tag sequence includes a part-of-speech tag for each word segment; the feature extraction layer is further used to input the part-of-speech tag of each word segment into an embedding layer to obtain a part-of-speech vector for each word, and perform feature fusion on the part-of-speech vector of each word and its corresponding intermediate fusion feature to obtain a target fusion feature; wherein a part-of-speech dictionary is provided in the embedding layer for mapping each part-of-speech tag into a continuous vector representation of a fixed dimension.
[0056] In a possible embodiment, the parts of speech recorded in the part-of-speech dictionary include all known parts of speech recorded in a preset part-of-speech mapping table and an unregistered part of speech added through a reserved position.
[0057] Since the apparatus described in the above embodiments of the present invention is used to implement the method of the above embodiments of the present invention, those skilled in the art will be able to understand the specific structure and variations of the apparatus based on the method described in the above embodiments of the present invention, and thus will not be described in detail here. All apparatuses used in the method of the above embodiments of the present invention fall within the scope of protection of the present invention.
[0058] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0059] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process flow and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.
[0060] It should be noted that the word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention may be implemented by means of hardware comprising several distinct components and by means of a suitably programmed computer. Among the several devices listed, several of these devices may be embodied by the same hardware. The use of the words first, second, third, etc., is for convenience only and does not imply any order. These words should be understood as part of the component name.
[0061] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0062] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concepts. Therefore, the technical solutions should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0063] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the technical solution of the present invention and its equivalents, the present invention shall also include such modifications and variations.
Claims
1. An entity extraction method based on adaptive weight fusion, characterized in that: include: Obtain the industrial text to be extracted; Inputting the industrial text to be extracted into a pre-trained entity extraction model to perform entity extraction, and obtaining an entity extraction result of the industrial text to be extracted; Wherein, the entity extraction model includes a text encoding layer and a feature extraction layer; The feature extraction layer is used to input the output of the text encoding layer into a pre-trained language model to obtain the output features of all hidden layers of the pre-trained language model, and perform weighted fusion on all the output features to obtain weighted semantic features, and input the weighted semantic features into a BiLSTM network to obtain a forward hidden state and a backward hidden state respectively, and splice the forward hidden state and the backward hidden state to obtain a spliced feature, and obtain a target fusion feature for predicting the entity label corresponding to each word in the industrial text to be extracted based on the adaptive weight, the spliced feature and the part-of-speech tag sequence.
2. The entity extraction method according to claim 1, characterized in that: The parts of speech of the segmented words of the industrial text to be extracted include directional nouns, object nouns, defect type nouns, adjectives and verbs.
3. The entity extraction method according to claim 1, characterized in that: The adaptive weight includes a first sub-adaptive weight; the feature extraction layer is further used to input the spliced features into the adaptive multi-head attention layer to generate attention features, and combine the spliced features and the attention features with the first sub-adaptive weight to perform feature fusion to obtain an intermediate fusion feature, and obtain the target fusion feature based on the intermediate fusion feature and the part-of-speech tag sequence.
4. The entity extraction method according to claim 3, characterized in that: The adaptive weight includes a second sub-adaptive weight; the part-of-speech tag sequence includes a part-of-speech tag for each word segment; the feature extraction layer is further used to input the part-of-speech tag of each word segment into an embedding layer to obtain a part-of-speech vector for each word, and combine the part-of-speech vector of each word and its corresponding intermediate fusion feature with the second sub-adaptive weight to perform feature fusion to obtain the target fusion feature; wherein, a part-of-speech dictionary is provided in the embedding layer for mapping each part-of-speech tag into a continuous vector representation of a fixed dimension.
5. The entity extraction method according to claim 4, characterized in that: The parts of speech recorded in the part-of-speech dictionary include all known parts of speech recorded in a preset part-of-speech mapping table and an unregistered part of speech added through a reserved position.
6. An entity extraction device based on adaptive weight fusion, characterized in that: include: The acquisition module is used to obtain the industrial text to be extracted; An input module, configured to input the industrial text to be extracted into a pre-trained entity extraction model to perform entity extraction, and obtain an entity extraction result of the industrial text to be extracted; Wherein, the entity extraction model includes a text encoding layer and a feature extraction layer; The feature extraction layer is used to input the output of the text encoding layer into a pre-trained language model to obtain the output features of all hidden layers of the pre-trained language model, and perform weighted fusion on all the output features to obtain weighted semantic features, and input the weighted semantic features into a BiLSTM network to obtain a forward hidden state and a backward hidden state respectively, and splice the forward hidden state and the backward hidden state to obtain a spliced feature, and obtain a target fusion feature for predicting the entity label corresponding to each word in the industrial text to be extracted based on the adaptive weight, the spliced feature and the part-of-speech tag sequence.
7. The entity extraction device according to claim 6, characterized in that: The parts of speech of the segmented words of the industrial text to be extracted include directional nouns, object nouns, defect type nouns, adjectives and verbs.
8. The entity extraction device according to claim 6, characterized in that: The adaptive weight includes a first sub-adaptive weight; the feature extraction layer is further used to input the spliced features into the adaptive multi-head attention layer to generate attention features, and combine the spliced features and the attention features with the first sub-adaptive weight to perform feature fusion to obtain an intermediate fusion feature, and obtain the target fusion feature based on the intermediate fusion feature and the part-of-speech tag sequence.
9. The entity extraction device according to claim 8, characterized in that: The adaptive weight includes a second sub-adaptive weight; the part-of-speech tag sequence includes a part-of-speech tag for each word segment; the feature extraction layer is further used to input the part-of-speech tag of each word segment into an embedding layer to obtain a part-of-speech vector for each word, and combine the part-of-speech vector of each word and its corresponding intermediate fusion feature with the second sub-adaptive weight to perform feature fusion to obtain the target fusion feature; wherein, a part-of-speech dictionary is provided in the embedding layer for mapping each part-of-speech tag into a continuous vector representation of a fixed dimension.
10. The entity extraction device according to claim 9, characterized in that: The parts of speech recorded in the part-of-speech dictionary include all known parts of speech recorded in a preset part-of-speech mapping table and an unregistered part of speech added through a reserved position.
Citation Information
Patent Citations
Named entity identification method and device, computer equipment and storage medium
CN112257449A
RoBERTa-BiLSTM-CRF voice dialogue text named entity recognition system fused with attention mechanism
CN117010387A
Relay alarm signal knowledge joint extraction method fusing ALBERT and self-attention mechanism
CN118569377A
Method and apparatus for automatic entity disambiguation
US20070067285A1