Text extraction method, system and medium based on fusion pre-training
Through the text extraction method based on fusion pre-training, the problem of low accuracy of text extraction in the financial field is solved, and higher text extraction accuracy and boundary recognition capabilities are achieved.
Patent Information
- Application Number
- CN202210038607.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-13
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-01-13
AI Technical Summary
The existing text extraction methods have problems with low accuracy in the financial field, especially in the current securities field, such as the extraction of partial texts that lead to blurred boundaries.
The text extraction method based on fusion pre-training is adopted, and the text is pre-trained and encoded by the pre-training model, character vectors are obtained, and effective word feature vectors are obtained through semantic extraction and feature selection fusion, and finally shunt decoding is performed to obtain word segmentation results and entity recognition results.
It improves the accuracy of text extraction, effectively avoids the problem of blurred boundaries, and makes the text extraction results more accurate.
Smart Images

Figure CN114398855B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a text extraction method, system and medium based on fusion pre-training. Background Art
[0002] Text information extraction is a relatively mature algorithm technology in the field of deep learning, and has been successfully applied in various business scenarios. However, in the financial field, especially in the field of current securities, the existing text extraction methods still have certain boundary problems. For example, for the extraction of the digital text "3.0975", only "3.09" will be extracted, or for the extraction of the digital text "4000", only "400" will be extracted, making the accuracy of text extraction not high enough.
[0003] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the object of the present invention is to provide a text extraction method, system and medium based on fusion pre-training, aiming to improve the accuracy of text extraction.
[0005] The technical solution of the present invention is as follows:
[0006] A text extraction method based on fusion pre-training, comprising:
[0007] Get the text to be extracted;
[0008] Performing pre-training encoding on the text to be extracted through a pre-training model to obtain a corresponding character vector;
[0009] Select at least part of the character vectors to perform semantic extraction on adjacent texts, and concatenate them to obtain a semantic feature vector;
[0010] Performing feature selection on the semantic feature vector and fusing it to obtain a valid word feature vector;
[0011] The effective word feature vector is subjected to shunt decoding to obtain word segmentation results and entity recognition results respectively.
[0012] In one embodiment, before pre-training and encoding the text to be extracted by using a pre-trained model to obtain a corresponding character vector, the method further includes:
[0013] The pre-trained model is subjected to adversarial training.
[0014] In one embodiment, performing adversarial training on the pre-trained model includes:
[0015] Constructing an adversarial sample, and adding the adversarial sample to the input embedding layer of the pre-trained model for perturbation;
[0016] The pre-trained model is subjected to adversarial training according to the adversarial sample to update model parameters, and the adversarial training ends when the number of updates reaches a preset number.
[0017] In one embodiment, constructing the adversarial sample specifically includes:
[0018] The adversarial sample is calculated according to the following formula:
[0019]
[0020]
[0021] Among them, g adv represents the gradient of the pre-trained model during adversarial training, X represents the input information, y represents the label information, and δ t-1 represents the disturbance size at time t-1, f θ represents the output of the pre-trained model, L represents the loss function, represents the gradient of the perturbation in the loss function, α represents the learning rate, ‖ ‖ F is the Frobenius norm, g t It represents the gradient of the pre-trained model at time t, and ∏ is the multiplication symbol.
[0022] In one embodiment, the adversarial training of the pre-trained model according to the adversarial sample to update the model parameters until the number of updates reaches a preset number, then the adversarial training ends, specifically includes:
[0023] After the pre-trained model is perturbed according to the adversarial sample, according to the formula Accumulate the gradient of the parameter θ, where K represents the number of times the gradient is increased, E represents the mathematical expectation, and g t-1 is the gradient of the pre-trained model at time t-1, Indicates the gradient of the parameters in the loss function;
[0024] The parameters of the pre-trained model are updated according to the accumulated gradients until the adversarial training ends when the number of updates reaches a preset number.
[0025] In one embodiment, selecting at least part of the character vectors to perform semantic extraction on adjacent texts and concatenating them to obtain a semantic feature vector includes:
[0026] Selecting encoding layers at several preset positions in the pre-trained model as target encoding layers;
[0027] The output results of the target coding layer are respectively input into text classification models connected in a one-to-one correspondence to perform semantic extraction of adjacent texts, the number of the text classification models is the same as the target coding layer, and the kernel sizes of the text classification models are different;
[0028] The extraction results of each text classification model are fused and spliced to obtain the semantic feature vector.
[0029] In one embodiment, the feature selection and fusion of the semantic feature vector to obtain the effective word feature vector specifically includes:
[0030] The semantic feature vector is selected and fused through a fully connected layer to obtain a valid word feature vector, where the input of the fully connected layer is F input , the output is F output ,
[0031] F input =concat(E1,E2,E i …,E n ),
[0032] F output =softmax(F input )=softmax(concat(E1,E2,E i …,E n )), where E i is the output result of the i-th target coding layer, and n is the number of target coding layers.
[0033] In one embodiment, the kernel size of the text classification model is 3-7.
[0034] A text extraction system based on fusion pre-training, the system comprising at least one processor; and
[0035] a memory communicatively connected to the at least one processor; wherein,
[0036] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned text extraction method based on fusion pre-training.
[0037] A non-volatile computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the one or more processors can execute the above-mentioned text extraction method based on fusion pre-training.
[0038] Beneficial effect: The present invention discloses a text extraction method, system and medium based on fusion pre-training. Compared with the prior art, the embodiments of the present invention obtain character vectors by encoding based on a pre-training model framework, and fuse at least part of the character vectors to perform semantic extraction of adjacent texts to learn text semantic information, thereby enhancing the semantic learning ability, so that the final word segmentation result can effectively avoid the problem of blurred boundaries and improve the accuracy of text extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0040] Figure 1 A flow chart of a text extraction method based on fusion pre-training provided in an embodiment of the present invention;
[0041] Figure 2 A schematic diagram of a model framework of a text extraction method based on fusion pre-training provided in an embodiment of the present invention;
[0042] Figure 3 A schematic diagram of functional modules of a text extraction device based on fusion pre-training provided in an embodiment of the present invention;
[0043] Figure 4 A schematic diagram of the hardware structure of a text extraction system based on fusion pre-training provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. The embodiments of the present invention are described below in conjunction with the accompanying drawings.
[0045] See also Figure 1 , Figure 1 The flowchart of an embodiment of the text extraction method based on fusion pre-training provided by the present invention is applicable to the case of automatically identifying the counterparty in the transaction process. Figure 1 As shown, the method specifically comprises the following steps:
[0046] S100: Obtain the text to be extracted.
[0047] In this embodiment, the text to be extracted can be the text information of the transaction dialogue in the current trading process, such as order information, consultation information, etc. sent between different trading institutions. By obtaining the text information in the transaction dialogue as the text to be extracted and performing automatic text extraction processing, the efficiency of financial information recognition processing is improved. Of course, in other embodiments, the text to be extracted is not limited to the text information in the current trading, but can also be text information in other transactions, or text information obtained by recognizing and converting transaction voice information, etc. This embodiment does not limit this.
[0048] S200, pre-training encoding is performed on the text to be extracted through a pre-training model to obtain a corresponding character vector.
[0049] The pre-trained model is trained through large-scale corpus information. It can achieve good results in downstream tasks by training and fine-tuning through downstream tasks. Therefore, in this embodiment, the pre-trained model is used to pre-train and encode the extracted text to obtain the corresponding character vector. Specifically, this embodiment preferably uses the Bert pre-training model for character encoding. Bert is a pre-trained language representation model, which emphasizes that it is no longer pre-trained by using the traditional unidirectional language model or the shallow splicing method of two unidirectional language models as in the past, but a new MLM (masked language model) is used to generate a deep bidirectional language representation, that is, for the input text, some words in the text are randomly masked with a certain probability, and then the Bert model is used to predict these masked words for pre-training to obtain the vector encoding of each character. Of course, in other embodiments, pre-training models such as Albert or RoBerta can also be used for pre-training encoding, which is not limited in this embodiment.
[0050] In one embodiment, before step S200, the method further includes:
[0051] The pre-trained model is subjected to adversarial training.
[0052] In this embodiment, before character encoding is performed on the extracted text, the pre-trained model is first combined with an adversarial training learning method to maximize the robustness and accuracy of the model. Specifically, adversarial training algorithms such as FreeLB, FGM, and PGD may be selected, but this embodiment does not limit this.
[0053] In one embodiment, adversarial training is performed on the pre-trained model, including:
[0054] Constructing an adversarial sample, and adding the adversarial sample to the input embedding layer of the pre-trained model for perturbation;
[0055] The pre-trained model is subjected to adversarial training according to the adversarial sample to update model parameters, and the adversarial training ends when the number of updates reaches a preset number.
[0056] In this embodiment, adversarial training is an important way to enhance the robustness of the model. During the adversarial training process, adversarial samples are constructed and added to the input embedding layer of the pre-trained model for perturbation. The input samples of the pre-trained model will be mixed with some small perturbations. The model is attacked by these perturbed adversarial samples, so that the model can recognize the true labels of these adversarial samples. That is, the pre-trained model is adversarially trained according to the adversarial samples during training, so that the model adapts to this change to update the model parameters until the adversarial training is completed, thereby improving the robustness of the model when encountering adversarial samples, and at the same time, it can also improve the performance and generalization ability of the model to a certain extent.
[0057] In the specific implementation, FreeLB adversarial training is adopted. The perturbation is first calculated by the following formula to attack the weights of the pre-trained model:
[0058]
[0059]
[0060] Among them, g adv represents the gradient of the pre-trained model during adversarial training, X represents the input information, y represents the label information, and δ t-1 represents the disturbance size at time t-1, f θ represents the output of the pre-trained model, L represents the loss function, represents the gradient of the perturbation in the loss function, α represents the learning rate, ‖ ‖ F is the Frobenius norm, g t It represents the gradient of the pre-trained model at time t, and ∏ is the multiplication symbol.
[0061] After the pre-trained model is perturbed according to the adversarial sample, according to the formula Accumulate the gradient of the parameter θ, where K represents the number of times the gradient is increased, E represents the mathematical expectation, and g t-1 is the gradient of the pre-trained model at time t-1, It means to find the gradient of the parameters in the loss function.
[0062] After obtaining the accumulated gradient, the parameters of the pre-trained model are updated. When the number of updates reaches the preset number, the adversarial training ends. The model parameters are regularized through adversarial training, a training method that introduces noise, thereby improving the model's robustness and generalization ability.
[0063] S300, selecting at least part of the character vectors to perform semantic extraction on adjacent texts, and concatenating them to obtain a semantic feature vector.
[0064] In this embodiment, after encoding to obtain the corresponding character vector, since the Bert series pre-trained model uses a single-word mode when constructing embeddings, this mode will lead to the loss of lexical semantic information in the Chinese context, and will also lead to boundary problems in text extraction, such as the digital text "3.0975" only extracting "3.09". To avoid this boundary problem, this embodiment selects at least part of the encoding result of the pre-trained model to perform semantic extraction on adjacent texts, that is, learns the semantic relationship between adjacent characters to better capture local correlations, thereby avoiding boundary problems caused by single characters and improving the accuracy of text extraction.
[0065] In one embodiment, step S300 includes:
[0066] Selecting encoding layers at several preset positions in the pre-trained model as target encoding layers;
[0067] The output results of the target coding layer are respectively input into text classification models connected in a one-to-one correspondence to perform semantic extraction of adjacent texts, the number of the text classification models is the same as the target coding layer, and the kernel sizes of the text classification models are different;
[0068] The extraction results of each text classification model are fused and spliced to obtain the semantic feature vector.
[0069] In this embodiment, the pre-trained model usually includes multiple coding layers, that is, a hidden layer structure including multiple Transformers. Since the higher the number of coding layers of the pre-trained model, the more detailed the data features obtained by the output latent vector are, the existing Bert and other pre-trained models will only output the coding results of the last layer (the highest layer). In this embodiment, in order to learn short-distance semantic features, coding layers at several preset positions are selected as target coding layers. Specifically, the last 25%-50% of all coding layers can be selected. For example, when the number of coding layers, that is, Transformers layers, in the pre-trained model is 12, the last 3 to 6 layers (that is, 3 to 6 layers from the last layer) are selected. When it is 18 layers, the last 5 to 9 layers are selected.
[0070] A text classification model is connected behind each selected target coding layer. In this embodiment, a text classification model TextCNN based on a convolutional neural network is used. For example, when 12 Transformers layers are used, the last 6 Transformers layers are selected as the target coding layers. A TextCNN module is connected after each of the 6 Transformers layers to perform semantic extraction on the adjacent text. In addition, in order to better capture local correlation, the kernel size of each TextCNN in this embodiment is different, and the kernel size is preferably set to 3-7. Since the TextCNN module can learn the semantic relationship between the word and the word whose distance is the size of the kernel, the setting of the kernel size is equivalent to setting the learning range of the TextCNN model. When learning the relationship between words, the distance size cannot exceed the size of the kernel. Therefore, in this embodiment, by setting kernels of different sizes, the model can learn text semantic information from multiple angles, increase the generalization ability and semantic understanding ability of the model, and improve the boundary recognition ability. In this embodiment, TextCNN can well solve the problem that Bert is fine-grained as words and cannot understand the semantics of the entire word.
[0071] After semantic extraction of neighboring texts through TextCNN, vector fusion is used to fuse and splice the extraction results output by each TextCNN to obtain a semantic feature vector. Through fusion and splicing, n hidden_size-dimensional vectors are converted into 1 hidden_size-dimensional feature vector, where n is the number of target encoding layers. The fusion method can enable the model to retain the semantic information learned by different TextCNNs from different angles, further enhancing the model's ability to learn semantics.
[0072] S400: performing feature selection on the semantic feature vector and fusing the feature vector to obtain a valid word feature vector.
[0073] In this embodiment, after the text classification model is integrated with some coding layers to implement short-distance semantic extraction, the semantic feature vector obtained by fusion and splicing is used to perform feature selection in a fully connected layer, and the effective word feature vector is selected and integrated.
[0074] In one embodiment, step S400 includes:
[0075] The semantic feature vector is selected and fused through a fully connected layer to obtain a valid word feature vector, where the input of the fully connected layer is F input , the output is F output ,
[0076] F input=concat(E1,E2,E i …,E n ),
[0077] F output =softmax(F input )=softmax(concat(E1,E2,E i …,E n )), where E i is the output result of the i-th target coding layer, and n is the number of target coding layers.
[0078] In this embodiment, the output result of the target programming layer fusion text classification model concat (E1, E2, E i …,E n ) to perform feature selection, specifically through the softmax function for classification, and select the most effective word features.
[0079] S500, performing split decoding on the effective word feature vector to obtain word segmentation results and entity recognition results respectively.
[0080] In this embodiment, based on the output of the fully connected layer, the downstream tasks are shunted and decoded, so that text extraction can be performed efficiently while realizing entity recognition, and word segmentation results and entity recognition results are obtained. Specifically, the effective word feature vectors are respectively input into the trained entity recognition task layer and word segmentation task layer. For the entity recognition task, the output of the fully connected layer is used again through the LSTM (Long Short-Term Memory) network structure to extract long-distance semantic features, and its output is used as the input of the decoding layer in the entity recognition task. The decoding layer adopts CRF (conditional random fields, conditional random fields) to predict entity labels, and finally output the corresponding entity annotations; for the word segmentation task, the output of the fully connected layer is decoded by a CRF decoder, and the character tags in the effective word feature vector are output to obtain the word segmentation result, and the character tags include entity start tags, entity remaining tags, and non-entity tags. For example, the text "A debt B institution issued to C institution" is finally segmented as "BI0BIIBIBII", where "B" is the entity start tag, "I" is the entity remaining tag, that is, the other positions in the entity except the starting position, and "O" is a non-entity tag, which is the parsing result of the space. The words in a sentence can be well segmented in the form of B, I, and O, so that the model can learn how to cut a sentence well and achieve accurate text segmentation extraction.
[0081] In order to better understand the implementation process of the text extraction method based on fusion pre-training provided by the present invention, the following is combined with Figure 2 The specific model structure in the text extraction process based on fusion pre-training provided by the present invention is introduced:
[0082] like Figure 2 As shown in the figure, the text to be extracted "A Debt B Mechanism..." is obtained. First, the input text is vectorized through the Bert pre-training model to obtain a fixed-dimensional character or word vector, and FreeLB adversarial training is added to the input embedding layer of the pre-training model to perturb the input embedding to increase the robustness of the model. The Bert pre-training model adopts a hidden layer structure of 12 Transformers. In order to learn short-distance semantic features, the last 6 layers of Transformers are selected through the semantic feature selection module to fuse the TextCNN module to perform semantic extraction on the adjacent text, that is, a TextCNN with kernels of different sizes is connected after the last 6 layers of Transformers to extract the key information in the sentence, so that the model can learn text semantic information from multiple angles, improve the generalization and boundary recognition ability of the model, and then use the vector fusion module to fuse each TextCNN output, and convert the 6 hidden_size-dimensional vectors into 1 hidden_size-dimensional feature vector; after passing through the semantic feature selection module, the spliced vector will pass through a fully connected layer (Fully Connected The output of the fully connected layer is decoded and annotated by the CRF decoder in the word segmentation task to obtain the character annotation results in the form of B, I, and O to accurately segment the sentence. The annotation result of "A debt B mechanism... institution" is "B, I, B, I, O". In the entity recognition task, the output of the fully connected layer is decoded by LSTM and CRF in turn to obtain the entity annotation result. For example, the annotation result of "A debt B mechanism... institution" is "B-BN, I-BN, B-ORG, I-ORG O". One character corresponds to one mark. BN and ORG are different entity annotations. BN represents the bond entity, and ORG represents the institution entity. Therefore, while realizing entity recognition extraction, accurate word segmentation is also achieved, thereby improving the accuracy of extraction.
[0083] Another embodiment of the present invention provides a text extraction device based on fusion pre-training, such as Figure 3 As shown, the device comprises:
[0084] An acquisition module 11 acquires the text to be extracted;
[0085] A pre-training module 12 performs pre-training encoding on the text to be extracted through a pre-training model to obtain a corresponding character vector;
[0086] A semantic extraction module 13 selects at least part of the character vectors to perform semantic extraction on adjacent texts, and concatenates them to obtain a semantic feature vector;
[0087] A fusion module 14 performs feature selection on the semantic feature vectors and fuses them to obtain a valid word feature vector;
[0088] The segmentation and recognition module 15 performs shunt decoding on the effective word feature vector to obtain word segmentation results and entity recognition results respectively.
[0089] The acquisition module 11, the pre-training module 12, the semantic extraction module 13, the fusion module 14 and the segmentation and recognition module 15 are connected in sequence. The module referred to in the present invention refers to a series of computer program instruction segments that can complete specific functions, which is more suitable for describing the execution process of text extraction based on fusion pre-training than a program. For the specific implementation method of each module, please refer to the corresponding method embodiment above, which will not be repeated here.
[0090] Another embodiment of the present invention provides a text extraction system based on fusion pre-training, such as Figure 4 As shown, the system 10 includes:
[0091] One or more processors 110 and memory 120, Figure 4 A processor 110 is used as an example for description. The processor 110 and the memory 120 may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.
[0092] The processor 110 is used to complete various control logics of the system 10, and it can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component, or any combination of these components. In addition, the processor 110 can also be any traditional processor, microprocessor or state machine. The processor 110 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a DSP and / or any other such configuration.
[0093] The memory 120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions corresponding to the text extraction method based on fusion pre-training in the embodiment of the present invention. The processor 110 executes various functional applications and data processing of the system 10 by running the non-volatile software programs, instructions and units stored in the memory 120, that is, implementing the text extraction method based on fusion pre-training in the above method embodiment.
[0094] The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required by at least one function; the data storage area may store data created according to the use of the system 10, etc. In addition, the memory 120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 120 may optionally include a memory remotely arranged relative to the processor 110, and these remote memories may be connected to the system 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0095] One or more units are stored in the memory 120, and when executed by one or more processors 110, the text extraction method based on fusion pre-training in any of the above method embodiments is executed, for example, the above described Figure 1 The method comprises steps S100 to S500.
[0096] An embodiment of the present invention provides a non-volatile computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors, for example, to execute the above-described Figure 1 The method comprises steps S100 to S500.
[0097] As an example, non-volatile storage media can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) as external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms such as synchronous RAM (SRAM), dynamic RAM, (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The disclosed memory components or memories of the operating environment described herein are intended to include one or more of these and / or any other suitable types of memory.
[0098] In summary, in the text extraction method, system and medium based on fusion pre-training disclosed by the present invention, the method obtains the text to be extracted; performs pre-training encoding on the text to be extracted through a pre-training model to obtain a corresponding character vector; selects at least part of the character vector to perform semantic extraction on adjacent texts, and splices to obtain a semantic feature vector; performs feature selection on the semantic feature vector and fuses it to obtain a valid word feature vector; performs shunt decoding on the valid word feature vector to obtain word segmentation results and entity recognition results respectively. Character vectors are obtained by encoding based on the pre-training model framework, and at least part of the character vectors are fused to perform semantic extraction on adjacent texts to learn text semantic information, thereby enhancing the semantic learning ability, so that the final word segmentation result can effectively avoid the problem of blurred boundaries and improve the accuracy of text extraction.
[0099] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium, and the computer program can include the processes of the above-mentioned method embodiments when executed. The storage medium can be a memory, a disk, a floppy disk, a flash memory, an optical storage device, etc.
[0100] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A text extraction method based on fusion pre-training, characterized in that: include: Get the text to be extracted; Performing pre-training encoding on the text to be extracted through a pre-training model to obtain a corresponding character vector; Select at least part of the character vectors to perform semantic extraction on adjacent texts, and concatenate them to obtain a semantic feature vector; Performing feature selection on the semantic feature vector and fusing it to obtain a valid word feature vector; Performing split decoding on the effective word feature vector to obtain word segmentation results and entity recognition results respectively; Before performing pre-training encoding on the text to be extracted by using a pre-training model to obtain a corresponding character vector, the method further includes: Performing adversarial training on the pre-trained model; The selecting at least part of the character vectors to perform semantic extraction on adjacent texts and concatenating them to obtain a semantic feature vector includes: Selecting encoding layers at several preset positions in the pre-trained model as target encoding layers; The output results of the target coding layer are respectively input into text classification models connected in a one-to-one correspondence to perform semantic extraction of adjacent texts, the number of the text classification models is the same as the target coding layer, and the kernel sizes of the text classification models are different; The extraction results of each text classification model are fused and spliced to obtain the semantic feature vector; Inputting the effective word feature vectors into the entity recognition task layer and word segmentation layer that have completed training respectively; The output of the fully connected layer is used to extract long-distance semantic features through the LSTM network structure, and the output result of the LSTM network structure is used as the input of the decoding layer in the entity recognition task. The decoding layer uses CRF to predict the entity label and finally outputs the corresponding entity annotation; The output of the fully connected layer is decoded by a CRF decoder, and the character tags in the effective word feature vector are output to obtain a word segmentation result, where the character tags include an entity start tag, an entity remainder tag, and a non-entity tag.
2. The text extraction method based on fusion pre-training according to claim 1 is characterized in that: The performing adversarial training on the pre-trained model comprises: Constructing an adversarial sample, and adding the adversarial sample to the input embedding layer of the pre-trained model for perturbation; The pre-trained model is subjected to adversarial training according to the adversarial sample to update model parameters, and the adversarial training ends when the number of updates reaches a preset number.
3. The text extraction method based on fusion pre-training according to claim 2 is characterized in that: The constructing of the adversarial sample specifically includes: The adversarial sample is calculated according to the following formula: Among them, g adv represents the gradient of the pre-trained model during adversarial training, X represents the input information, y represents the label information, and δ t-1 represents the disturbance size at time t-1, f θ represents the output of the pre-trained model, L represents the loss function, represents the gradient of the perturbation in the loss function, α represents the learning rate, ‖‖ F is the Frobenius norm, g t It represents the gradient of the pre-trained model at time t, and ∏ is the multiplication symbol.
4. The text extraction method based on fusion pre-training according to claim 3 is characterized in that: The adversarial training is performed on the pre-trained model according to the adversarial sample to update the model parameters until the number of updates reaches a preset number, and then the adversarial training ends, specifically including: After the pre-trained model is perturbed according to the adversarial sample, according to the formula Accumulate the gradient of the parameter θ, where K represents the number of times the gradient is increased, E represents the mathematical expectation, and g t-1 is the gradient of the pre-trained model at time t-1, Indicates the gradient of the parameters in the loss function; The parameters of the pre-trained model are updated according to the accumulated gradients until the adversarial training ends when the number of updates reaches a preset number.
5. The text extraction method based on fusion pre-training according to claim 1 is characterized in that: The step of performing feature selection and fusion on the semantic feature vector to obtain a valid word feature vector specifically includes: The semantic feature vector is selected and fused through a fully connected layer to obtain a valid word feature vector, where the input of the fully connected layer is F input , the output is F output , F input =concat(E1,E2,E i …,E n ), F output =softmax(F input )=softmax(concat(E1,E2,E i …,AND n )), Among them, E i is the output result of the i-th target coding layer, and n is the number of target coding layers.
6. The text extraction method based on fusion pre-training according to claim 1 is characterized in that: The kernel size of the text classification model is 3-7.
7. A text extraction system based on fusion pre-training, characterized in that: The system includes at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the text extraction method based on fusion pre-training described in any one of claims 1-6.
8. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, enable the one or more processors to execute the text extraction method based on fusion pre-training as described in any one of claims 1-6.
Citation Information
Patent Citations
Event joint extraction method fusing word features and deep learning
CN113190602A
Deep learning-based department semantic information extraction method and device
CN113268576A