Long text instruction extraction method based on lightweight pre-training model and attention mechanism
By using the lightweight pre-trained model MacBERT and an attention mechanism, a parallel extraction network is constructed, which solves the problems of large information content and high entropy in the extraction of long text instructions, and achieves efficient and accurate instruction extraction and improves the model's generality.
Patent Information
- Application Number
- CN202511148422.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-17
- Publication Date
- 2025-11-21
AI Technical Summary
Conventional instruction extraction models suffer from high information content, high information entropy, and severe fragmentation when processing long texts, resulting in complex model structures, a large number of parameters, high requirements for the computing power of the deployment platform, and reduced versatility.
We employ the lightweight pre-trained model MacBERT for encoding and feature mining, construct a multi-level parallel extraction network, utilize the attention mechanism to identify named entities and instruction subjects, and generate structured instructions through information fusion.
While reducing computing power requirements, it improves the accuracy of long text instruction extraction and the versatility of the model, simplifies the model structure, and reduces the computing power requirements of the deployment platform.
Smart Images

Figure CN120996043A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an instruction extraction method, specifically to a long text instruction extraction method based on a lightweight pre-trained model and an attention mechanism. Background Technology
[0002] The emergence of deep learning has brought about a new revolution in Natural Language Processing (NLP). Deep learning references the structure of neurons in the human brain, using multi-level, multi-structured neural network models to infer and fit results. Its excellent reasoning and learning capabilities regarding text data features simplify the process of manually constructing text data features. Deep learning-based NLP can uncover deeper semantic and syntactic information in language, thereby continuously improving the accuracy of various NLP tasks and expanding its application scenarios, laying a solid foundation for many fields such as intelligent interaction.
[0003] Instruction extraction is a crucial task in natural language processing. It aims to automatically extract key information from highly flexible unstructured text data and integrate it into structured data. Based on this structured data, intelligent action sequences are then constructed, ultimately enabling intelligent services. Extracting from long text data containing multiple entities and instructions is particularly challenging. Conventional instruction extraction models are structurally complex with numerous parameters, requiring significant computing power from the deployment platform. This further reduces the versatility of such models, specifically in the following ways:
[0004] 1. Long texts contain a large amount of information; a single text may contain multiple named entities and multiple instructions.
[0005] 2. Long texts have a high degree of information entropy, and there is overlap in textual information between different instructions;
[0006] 3. Long texts contain highly fragmented information, with related named entities in a single instruction distributed over a wide range;
[0007] 4. Currently, the instruction extraction model has a complex structure and a large number of parameters, which places high demands on the computing power of the deployed platform. Summary of the Invention
[0008] The purpose of this invention is to solve the problem that conventional instruction extraction models encounter difficulties when extracting long texts, such as large amounts of information, high information entropy, and severe fragmentation. By using a lightweight pre-trained model to simplify the conventional model structure and reduce model parameters, and by utilizing a multi-level downstream extraction model to fully leverage the deep features of the text, the computational power consumed by repeated calculations is reduced, thereby improving the versatility of the instruction model proposed in this invention and improving the extraction accuracy of this model for complex long texts.
[0009] This invention is achieved through the following technical solutions:
[0010] A long text instruction extraction method based on a lightweight pre-trained model and attention mechanism is proposed. The method employs the lightweight pre-trained encoding model MacBERT to encode long texts and mine textual semantic features. A multi-level extraction model is constructed in the downstream task model. Within this task model, a parallel extraction network structure is designed. One extraction network identifies the text and category information of all named entities in the original text; the other extraction network extracts the text of related named entities based on the identified instruction body. The attention mechanism is improved in this extraction network, using the instruction body identified from the original text as the retrieval basis to enhance the textual information related to the instruction body in the original text, thereby improving the recognition accuracy of the subsequent parallel network. Finally, an information fusion and instruction generation module uses the text of the named entities as the fusion basis to fuse the results of the two extraction networks. The named entities of each fused instruction are then split, and the structured text of each instruction is generated sequentially.
[0011] Furthermore, the method mainly includes four algorithm modules: (1) a named entity recognition module based on a lightweight pre-trained model, (2) a head entity recognition module, (3) a head entity related named entity extraction module based on an attention mechanism, and (4) an information fusion and instruction generation module.
[0012] The input to the named entity recognition module based on the lightweight pre-trained model is the original text to be recognized, and the output is the named entity text and its type contained in the text; the module mainly performs text encoding, feature mining and text type recognition.
[0013] The input to the head entity recognition module is the named entity text and its type contained in the text, and the output is the entity that is the subject of the instruction among the named entities; the module determines whether the entity is the subject of the instruction based on the identified entity type;
[0014] The input of the attention-based head entity related named entity extraction module consists of two parts: the first part is the comprehensive feature tensor set of the instruction body entity, and the second part is the feature tensor of the original text. The module outputs the named entities related to each command body in sequence. The instruction body and the original text in which it is located are used as the reasoning basis of the model. The attention mechanism is introduced to further enhance the weight of the text features related to the instruction body, remove the influence of noisy data on the model, and further improve the accuracy of model recognition.
[0015] The information fusion and instruction generation module has two inputs: one is the entity of the instruction body and its related named entities, and the other is the text and category of each named entity identified from the original text. The module's data is the instruction extracted from the original text. The module uses the text and position information of the entities in the two input data to add categories to the named entities related to the instruction body, and generates instructions based on the fused features.
[0016] Furthermore, in the named entity recognition module based on the lightweight pre-trained model, the encoding layer uses MacBERT's preprocessor Tokenizer to encode the text to be extracted, and uses the MacBERT pre-trained model to perform feature mining on the encoded text. The pre-trained model is composed of three layers of bidirectionally linked Transformers stacked together.
[0017] Furthermore, the downstream recognition model of the model uses only two parallel fully connected layers. One fully connected layer is used to identify the start position of the named entity in the text and its category, and the other fully connected layer is used to identify the end position of the named entity in the text and its category. The result sequence of the inference from the above two fully connected layers is used together with the original text information to determine the text and category of the named entities contained in the original text.
[0018] Furthermore, the head entity recognition module will filter the named entities identified in the original text one by one to determine whether they belong to the main category, and extract the encoded feature tensor corresponding to the entities that match the main category, i.e., the head entities. This tensor has a dimension of m. i ×768 dimensions, where m i 768 represents the length of the encoded character of the head entity, and 768 represents the dimension of the encoded feature of that character. Since the length of the head entities may vary, the module calculates the comprehensive feature tensor of each head entity by summing and averaging. After this operation, all the feature tensors of the head entities are 1×768, and the dimension of each head entity is unified so that subsequent modules can process them uniformly.
[0019] Furthermore, the attention-based head entity related named entity extraction module first takes the comprehensive feature tensor of the head entities as input, and then passes it through a linear fully connected layer to obtain the query value Q of each head entity. i First, a linear fully connected layer is used to increase the fit of the head entity to nonlinear features. Second, using the text feature tensor as input, two parallel linear fully connected layers are constructed to obtain K of the text feature tensor. i Key value and its information value V i Using the retrieval value Q of the head entity i With the feature tensor K of the original text i Perform matrix multiplication to obtain each V iThe corresponding information weight value a i The calculated result is then fed into the Softmax layer to normalize all information weight values, and the normalized a is then... i The value is assigned to the information value V of the corresponding original text feature. i This leads to the enhanced features of the original text.
[0020] Furthermore, the two parallel fully connected layers use the enhanced text features as the basis for reasoning. They identify named entities related to the head entity through the fully connected layers. One fully connected layer is used to identify the start position of the related named entity, and the other fully connected layer is used to identify the end position of the related named entity. Based on this, the named entities related to each head entity are obtained by combining the identification result sequence of the two fully connected layers with the original text information. In this process, the two fully connected layers only identify the position of the named entities related to the head entity and do not identify the type of the related named entity.
[0021] Furthermore, the information fusion and instruction generation module performs information fusion using all named entities identified from the original text and the named entities related to each instruction body. The fusion rules are based on the text and location information of the named entities, and add category attributes to the set of named entities related to each instruction body, thereby obtaining the information of the body of each instruction and its related named entities.
[0022] Furthermore, based on the instruction information after information fusion, elements are split and element type names are added before the corresponding elements. Finally, the split elements are used to generate instruction bodies with a fixed format, thereby obtaining all the instructions contained in the long text.
[0023] The beneficial effects of this invention are:
[0024] This invention proposes a long text instruction extraction method based on a lightweight pre-trained model and an attention mechanism. It employs the lightweight pre-trained encoding model MacBERT to encode long texts and mine textual semantic features. While maintaining extraction performance, the extraction speed and versatility are improved by simplifying the pre-trained model structure and parameters. In the downstream extraction task model of this invention, a parallel extraction network structure is designed. One extraction network identifies the text and category information of all named entities in the original text; the other extraction network extracts the text of related named entities based on the identified instruction body. This extraction network innovatively improves the attention mechanism, using the instruction body identified from the original text as the retrieval basis to enhance the text information related to the instruction body in the original text, thereby improving the recognition accuracy of the subsequent parallel network. Finally, an information fusion and instruction generation module uses the text of the named entities as the fusion basis to fuse the results of the two extraction networks. The named entities of each fused instruction are then split, and the structured text of each instruction is generated sequentially. Attached Figure Description
[0025] Figure 1 This is a flowchart of a long text instruction extraction process based on a lightweight pre-trained model and an attention mechanism.
[0026] Figure 2 This is a flowchart of the named entity recognition module based on a lightweight pre-trained model; Figure 3 This is a flowchart of the head entity recognition module's workflow;
[0027] Figure 4 This is a flowchart of the head entity-related named entity extraction module based on the attention mechanism;
[0028] Figure 5 This is a flowchart of the information fusion and instruction generation module. Detailed Implementation
[0029] This invention is a long text instruction extraction method based on a lightweight pre-trained model and attention mechanism. Addressing the high difficulty of extracting long texts with multiple entities and instructions, a lightweight pre-trained model, MacBERT, is introduced. MacBERT encodes and mines features from the long text, constructing a multi-level extraction model in the downstream task model. It uses this model to analyze the named entity text and its categories in parallel, further identifying the subject and related entities within the named entities. Finally, by identifying overlapping areas of entity positions, the structure of multiple instructions within the long text is extracted.
[0030] It mainly includes four algorithm modules: (1) a named entity recognition module based on a lightweight pre-trained model, (2) a head entity recognition module, (3) a head entity related named entity extraction module based on an attention mechanism, and (4) an information fusion and instruction generation module. The flowchart of this method is shown below. Figure 1 As shown.
[0031] The named entity recognition module based on the lightweight pre-trained model takes the original text to be recognized as input and outputs the named entity texts and their categories contained in the text. This module mainly performs text encoding, feature mining, and text category recognition. The workflow of this module is as follows: Figure 2 As shown.
[0032] The input to the head entity recognition module is the named entity text and its type contained in the text. The output of this module is the entity that is the subject of the instruction among the named entities. This module determines whether the entity is the subject of the instruction based on the identified entity type. The workflow of this module is as follows: Figure 3 As shown.
[0033] The input to the attention-based head entity related named entity extraction module consists of two parts: the first part is the comprehensive feature tensor set of the command body entity, and the second part is the feature tensor of the original text. This module sequentially outputs the named entities related to each command body. This module uses the command body and its corresponding original text as the basis for model inference, and introduces an attention mechanism to further enhance the weight of text features related to the command body, remove the influence of noisy data on the model, and further improve the accuracy of model recognition. The workflow of this module is as follows: Figure 4 As shown.
[0034] The information fusion and instruction generation module receives two inputs: the entity of the instruction body and its associated named entities, and the text and category of each named entity identified from the original text. The data for this module is the instruction extracted from the original text. This module primarily utilizes the text and positional information of entities from the two input data sets to add categories to the named entities related to the instruction body, and generates instructions based on the fused features. The workflow of this module is as follows: Figure 5 As shown.
[0035] In the named entity recognition module based on the lightweight pre-trained model, the encoding layer uses the MacBERT preprocessor (Tokenizer) to encode the text to be extracted, and then uses a MacBERT pre-trained model to perform feature mining on the encoded text. This pre-trained model consists of only three layers of bidirectionally linked Transformers stacked together. This invention uses a pre-trained model that meets the text extraction difficulty requirements and has a small number of stacked layers for feature mining. The number of stacked layers in the pre-trained model affects the number of parameters in the model. By reducing unnecessary model parameters, the computational requirements of the deployment platform are reduced, thereby improving the versatility of the instruction extraction model in this invention.
[0036] The downstream recognition model of this paper uses only two parallel fully connected layers. One fully connected layer is used to identify the start position of named entities in the text and their category, and the other fully connected layer is used to identify the end position of named entities in the text and their category. The resulting sequence of inference from these two fully connected layers is then combined with the original text information to determine the text and category of the named entities contained in the original text.
[0037] The head entity recognition module will filter the named entities identified in the original text one by one to determine whether they belong to the main category, and extract the encoded feature tensor corresponding to the entities that match the main category, i.e., the head entities. The dimension of this tensor is m. i ×768 dimensions, where m i 768 represents the length of the encoded character of the header entity, and 768 represents the dimension of the encoded feature of that character. Since the lengths of header entities may vary, this module calculates the comprehensive feature tensor of each header entity by summing and averaging. After this operation, all header entity character feature tensors are 1×768. This standardization of the dimension of each header entity facilitates unified processing in subsequent modules.
[0038] The attention-based head entity related named entity extraction module first takes the comprehensive feature tensor of the head entities as input, and then passes it through a linear fully connected layer to obtain the query value Q of each head entity. i To enhance the fit of the head entity to nonlinear features, a linear fully connected layer is used. Secondly, the text feature tensor is used as input, and two parallel linear fully connected layers are constructed to obtain K values of the text feature tensor. i Key value and its information value V i Using the retrieval value Q of the head entity. i With the feature tensor K of the original text i Perform matrix multiplication to obtain each V i The corresponding information weight value a i The calculated result is then fed into the Softmax layer to normalize all information weight values. The normalized a... iThe value is assigned to the information value V of the corresponding original text feature. i This leads to the enhanced features of the original text. The improved attention mechanism is then used to retrieve the head entity value Q. i As a basis for retrieval, the original text is enhanced with related named entity information to further achieve the purpose of effective feature enhancement and further improve the accuracy of model recognition.
[0039] An improved attention mechanism links two parallel fully connected layers. These parallel fully connected layers use enhanced text features as the basis for inference, identifying named entities related to the head entity. One fully connected layer identifies the start position of the relevant named entity, and the other identifies the end position. Based on this, the named entities related to each head entity are obtained by combining the identification results sequence of the two fully connected layers with the original text information. During this process, the two fully connected layers only identify the position of the named entities related to the head entity, without identifying the type of the related named entities, thus avoiding redundant text recognition operations and unnecessary computational consumption.
[0040] The information fusion and instruction generation module uses all named entities identified from the original text and the named entities related to each instruction body to perform information fusion. The fusion rules are based on the text and position information of the named entities, and add category attributes to the set of named entities related to each instruction body, thereby obtaining the information of the body of each instruction and its related named entities.
[0041] Based on the fused instruction information, elements are split, and element type names are added before the corresponding elements, such as adding "Main" before "Vehicle No. 1", "Action" before "Motor", and "Position" before "8690". Finally, the split elements are used to generate a fixed-format instruction body, thus obtaining all the instructions contained in the long text.
[0042] This invention has the following characteristics:
[0043] 1. This invention proposes a model for instruction extraction based on a lightweight pre-trained model. While ensuring that the instruction extraction effect is not affected, the encoding and feature mining layers of instruction extraction are simplified, reducing the computing power required for deployment of the proposed model and thus improving the model's versatility.
[0044] 2. This invention proposes an improved attention mechanism that uses the instruction subject as the retrieval value to enhance the information value weight of the text related to the instruction subject in the original text by retrieving the key value of the relevant information in the original text. This improves the utilization rate of effective information in the original text and thus improves the extraction accuracy of this model.
[0045] 3. This invention proposes a parallel extraction and information fusion instruction extraction model. Parallel extraction includes two parts of information: one is the complete named entity text and category information in the original text, and the other is the identification of named entities related to each instruction entity. Using the identified named entity categories and relevance as the fusion basis, category attributes are added to the relevant entities in the instruction body of the original text, and each instruction is element-wise split to finally generate the structure of each instruction.
Claims
1. A method for extracting long text instructions based on a lightweight pre-trained model and an attention mechanism, characterized by: A lightweight pre-trained encoding model is used to encode long texts and mine textual semantic features. A multi-level extraction model is constructed in the downstream task model. In the task model, a parallel extraction network structure is designed. One extraction network is used to identify the text and category information of all named entities in the original text. The other extraction network extracts the text of the named entities related to the identified instruction subject. The attention mechanism is improved in this extraction network. The instruction subject identified from the original text is used as the retrieval basis to enhance the text information related to the instruction subject in the original text and improve the recognition accuracy of the subsequent parallel network. Finally, the information fusion and instruction generation module uses the text of the named entities as the fusion basis to fuse the results of the above two extraction networks. The named entities of each instruction after fusion are split and the structure text of each instruction is generated in sequence.
2. The method for extracting long text instructions based on a lightweight pre-trained model and attention mechanism according to claim 1, characterized in that: The pre-trained encoding model uses MacBERT.
3. A long text instruction extraction method based on a lightweight pre-trained model and attention mechanism according to claim 1 or 2, characterized in that: The method mainly includes four algorithm modules: (1) a named entity recognition module based on a lightweight pre-trained model, (2) a head entity recognition module, (3) a head entity related named entity extraction module based on an attention mechanism, and (4) an information fusion and instruction generation module. The input to the named entity recognition module based on the lightweight pre-trained model is the original text to be recognized, and the output is the named entity text and its type contained in the text; the module mainly performs text encoding, feature mining and text type recognition. The input to the head entity recognition module is the named entity text and its type contained in the text, and the output is the entity that is the subject of the instruction among the named entities; the module determines whether the entity is the subject of the instruction based on the identified entity type; The input of the attention-based head entity related named entity extraction module consists of two parts: the first part is the comprehensive feature tensor set of the instruction body entity, and the second part is the feature tensor of the original text. The module outputs the named entities related to each command body in sequence. The instruction body and the original text in which it is located are used as the reasoning basis of the model. The attention mechanism is introduced to further enhance the weight of the text features related to the instruction body, remove the influence of noisy data on the model, and further improve the accuracy of model recognition. The input to the information fusion and instruction generation module consists of two parts: one is the entity of the instruction body and its related named entities, and the other is the text and category of each named entity identified from the original text. The data of the module is the instruction extracted from the original text. The module uses the text and position information of the entities in the two input data to add categories to the named entities related to the instruction body, and generates instructions based on the fused features.
4. The method for extracting long text instructions based on a lightweight pre-trained model and attention mechanism according to claim 3, characterized in that: In the named entity recognition module based on the lightweight pre-trained model, the encoding layer uses MacBERT's preprocessor Tokenizer to encode the text to be extracted, and uses the MacBERT pre-trained model to perform feature mining on the encoded text. The pre-trained model is composed of three layers of bidirectionally linked Transformers stacked together.
5. The method for extracting long text instructions based on a lightweight pre-trained model and attention mechanism according to claim 4, characterized in that: The downstream recognition model of the aforementioned model uses only two parallel fully connected layers. One fully connected layer is used to identify the start position of the named entity in the text and its category, and the other fully connected layer is used to identify the end position of the named entity in the text and its category. The result sequence of the inference from the above two fully connected layers is used together with the original text information to determine the text and category of the named entities contained in the original text.
6. The method for extracting long text instructions based on a lightweight pre-trained model and attention mechanism according to claim 3, characterized in that: The head entity recognition module will filter the named entities identified in the original text one by one to determine whether they belong to the main category, and extract the encoded feature tensor corresponding to the entities that match the main category, i.e., the head entities. The dimension of this tensor is m. i ×768 dimensions, where m i 768 represents the length of the encoded character of the head entity, and 768 represents the dimension of the encoded feature of that character. Since the length of the head entities may vary, the module calculates the comprehensive feature tensor of each head entity by summing and averaging. After this operation, all the feature tensors of the head entities are 1×768, and the dimension of each head entity is unified so that subsequent modules can process them uniformly.
7. The method for extracting long text instructions based on a lightweight pre-trained model and attention mechanism according to claim 3, characterized in that: The attention-based head entity related named entity extraction module first takes the comprehensive feature tensor of the head entities as input, and then passes it through a linear fully connected layer to obtain the query value Q of each head entity. i This increases the head entity's fit to nonlinear features by using a linear fully connected layer; Secondly, using the text feature tensor as input, two parallel linear fully connected layers are constructed to obtain K of the text feature tensor. i Key value and its information value V i Using the retrieval value Q of the head entity i With the feature tensor K of the original text i Perform matrix multiplication to obtain each V i The corresponding information weight value a i The calculated result is then fed into the Softmax layer to normalize all information weight values, and the normalized a is then... i The value is assigned to the information value V of the corresponding original text feature. i This leads to the enhanced features of the original text.
8. The method for extracting long text instructions based on a lightweight pre-trained model and attention mechanism according to claim 7, characterized in that: The two parallel fully connected layers use the enhanced text features as the basis for reasoning. They identify named entities related to the head entity through the fully connected layers. One fully connected layer is used to identify the start position of the related named entity, and the other fully connected layer is used to identify the end position of the related named entity. Based on this, the named entities related to each head entity are obtained by combining the identification result sequence of the two fully connected layers with the original text information. In this process, the two fully connected layers only identify the position of the named entities related to the head entity and do not identify the type of the related named entity.
9. The method for extracting long text instructions based on a lightweight pre-trained model and attention mechanism according to claim 3, characterized in that: The information fusion and instruction generation module uses all named entities identified from the original text and the named entities related to each instruction body to perform information fusion. The fusion rules are based on the text and position information of the named entities, and add category attributes to the set of named entities related to each instruction body, thereby obtaining the information of the body of each instruction and its related named entities.
10. The method for extracting long text instructions based on a lightweight pre-trained model and attention mechanism according to claim 9, characterized in that: The elements are split based on the fused instruction information, and the element type name is added before the corresponding element. Finally, the split elements are generated into a fixed-format instruction body, thus obtaining all the instructions contained in the long text.