Information parsing method, device and electronic device based on vocabulary enhancement

By constructing an input sequence containing character vectors and vocabulary vectors, and using the BERT model to generate semantic representations, the problem of combining Chinese optimization model and pre-trained model is solved, and the accuracy and efficiency of information parsing are improved.

CN114692635BActive Publication Date: 2025-05-06BEIJING KUAQUO INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210168591.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-23
Publication Date
2025-05-06
Estimated Expiration
2042-02-23

AI Technical Summary

Technical Problem

In the prior art, Chinese optimization models cannot be combined with pre-trained models, and the training stage cannot be parallelized, resulting in the rich semantic representation of pre-trained models being wasted and the model training efficiency is low.

Method used

By obtaining bond information and constructing an input sequence containing character vectors and vocabulary vectors, the input sequence is processed using a pre-trained model based on the BERT model, semantic representations are generated, and target vectors are output by discarding layers and normalizing layers to achieve information parsing.

Benefits of technology

It fully utilizes the vocabulary information in the text, improves the ability to identify entities' boundaries, and improves the accuracy of extraction of transaction elements in the secondary transaction business of financial bonds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114692635B_ABST
    Figure CN114692635B_ABST
Patent Text Reader

Abstract

The present invention discloses an information parsing method, device and electronic device based on vocabulary enhancement, the method comprising: obtaining bond information to be parsed, constructing an input sequence according to the bond information, the input sequence comprising a character vector in the bond information and a vocabulary vector corresponding to the character; constructing a pre-training model, processing the input sequence through the pre-training model, generating a semantic representation of the input sequence; after passing the semantic representation through a discarding layer and a normalization layer, outputting a target vector, the target vector being the parsed structured bond information. The embodiment of the present invention makes full use of the vocabulary information in the text, so that the device can better identify entity boundaries and improve the accuracy of transaction element extraction in the secondary transaction business of financial bonds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an information parsing method, device and electronic equipment based on vocabulary enhancement. Background Art

[0002] In text processing, a common requirement is to extract valuable information from a text. For example, when booking a hotel, it is necessary to extract key information such as location and time from unstructured text information. This requirement also exists in the field of financial bonds, which is to extract valuable information from unstructured text information.

[0003] In text processing, a common requirement is to extract valuable information from a text. For example, when booking a hotel, it is necessary to extract key information such as location and time from unstructured text information. This requirement also exists in the field of financial bonds, which is to extract valuable information from unstructured text information.

[0004] When constructing Chinese language models, the embedding layer of existing pre-trained models uses character-level input. The disadvantage is that it ignores the rich vocabulary information in the text and has limited improvement in the extraction of text information, especially boundary information.

[0005] Optimized models for Chinese, such as Lattice-LSTM, cannot be combined with pre-trained models. At the same time, the training phase cannot be parallelized, which results in the waste of rich semantic representations of pre-trained models and low training efficiency.

[0006] Therefore, the prior art still needs to be improved and developed. Summary of the invention

[0007] In view of the above-mentioned deficiencies in the prior art, the present invention provides an information parsing method, device and electronic device based on vocabulary enhancement, aiming to solve the problem that the Chinese optimization model in the information parsing method in the prior art cannot be combined with the pre-training model, and the training stage cannot be parallelized, which leads to the waste of rich semantic representation of the pre-training model and low model training efficiency.

[0008] The technical solution of the present invention is as follows:

[0009] The first embodiment of the present invention provides an information parsing method based on vocabulary enhancement, the method comprising:

[0010] Obtaining bond information to be parsed, and constructing an input sequence according to the bond information, wherein the input sequence includes character vectors in the bond information and vocabulary vectors corresponding to the characters;

[0011] Construct a pre-trained model, process the input sequence through the pre-trained model, and generate a semantic representation of the input sequence;

[0012] After the semantic representation passes through the discard layer and the normalization layer, the target vector is output, which is the parsed structured bond information.

[0013] Furthermore, the step of obtaining the bond information to be parsed and constructing an input sequence according to the bond information includes:

[0014] Get the bond information to be parsed, and process the bond information into single characters;

[0015] Constructs input sequences from individual characters.

[0016] Further, constructing an input sequence according to a single character includes:

[0017] Get the Chinese sentences in the bond information to be parsed;

[0018] Match the Chinese sentences according to the preset dictionary to obtain the vocabulary in the Chinese sentences;

[0019] Each character is paired with a word containing the character to generate an input sequence.

[0020] Furthermore, the constructing of the pre-training model, processing the input sequence through the pre-training model to generate a semantic representation of the input sequence, includes:

[0021] Build a pre-trained model based on the BERT model;

[0022] The input sequence is processed through the BERT model to generate the semantic representation corresponding to the input sequence.

[0023] Furthermore, the processing of the input sequence by the BERT model to generate a semantic representation corresponding to the input sequence includes:

[0024] Performing a nonlinear transformation on the vocabulary vectors in the input sequence through the BERT model to generate a nonlinearly transformed vocabulary vector, wherein the transformed vocabulary vector is aligned with the dimension of the character vector;

[0025] Calculate the correlation between the character vector and the vocabulary vector, calculate the weights of all vocabulary vectors based on the correlation, and calculate the target vocabulary vector based on the weights;

[0026] The target vocabulary vector is fused into the character vector to generate a semantic representation of the input sequence.

[0027] Further, the calculating the correlation between the character vector and the vocabulary vector, calculating the weights of all vocabulary vectors according to the correlation, and calculating the target vocabulary vector according to the weights, includes:

[0028] Compute the correlation between character vectors and vocabulary vectors based on a bilinear attention layer;

[0029] Calculate the weights of all vocabulary vectors based on their relevance;

[0030] Calculate the target vocabulary vector based on the weights.

[0031] Furthermore, the step of fusing the target vocabulary vector into the character vector to generate a semantic representation of the input sequence includes:

[0032] The target vocabulary vector is added to the character vector to generate a fused representation, which is a semantic representation of the input sequence.

[0033] Another embodiment of the present invention provides an information parsing device based on vocabulary enhancement, the device comprising:

[0034] A sequence construction module, used to obtain bond information to be parsed, and construct an input sequence according to the bond information, wherein the input sequence includes a character vector in the bond information and a vocabulary vector corresponding to the character;

[0035] Input sequence processing is used to build a pre-trained model, process the input sequence through the pre-trained model, and generate a semantic representation of the input sequence;

[0036] The target vector output module is used to output a target vector after passing the semantic representation through a discard layer and a normalization layer. The target vector is the parsed structured bond information.

[0037] Another embodiment of the present invention provides an electronic device, the electronic device comprising at least one processor; and,

[0038] a memory communicatively connected to the at least one processor; wherein,

[0039] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the above-mentioned information parsing method based on vocabulary enhancement.

[0040] Another embodiment of the present invention further provides a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, enable the one or more processors to execute the above-mentioned information parsing method based on vocabulary enhancement.

[0041] Beneficial effects: The embodiment of the present invention makes full use of the vocabulary information in the text, so that the device can better identify the entity boundary and improve the accuracy of transaction element extraction in the secondary transaction business of financial bonds. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0043] Figure 1 A flowchart of a preferred embodiment of an information parsing method based on vocabulary enhancement of the present invention;

[0044] Figure 2 A schematic diagram of a network structure of a preferred embodiment of an information parsing method based on vocabulary enhancement of the present invention;

[0045] Figure 3 A schematic diagram of character and vocabulary pairs in a preferred embodiment of an information parsing method based on vocabulary enhancement of the present invention;

[0046] Figure 4 It is a flowchart of a specific application embodiment of a preferred embodiment of an information parsing method based on vocabulary enhancement of the present invention;

[0047] Figure 5 A network diagram of a vocabulary adapter device added to a specific application embodiment of a preferred embodiment of an information parsing method based on vocabulary enhancement of the present invention;

[0048] Figure 6 A schematic diagram of functional modules of a preferred embodiment of an information parsing device based on vocabulary enhancement according to the present invention;

[0049] Figure 7 The figure is a schematic diagram of the hardware structure of a preferred embodiment of an electronic device of the present invention. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0051] The embodiments of the present invention are described below with reference to the accompanying drawings.

[0052] In view of the above problems, the present invention provides an information parsing method based on vocabulary enhancement. Figure 1 , Figure 1 FIG. 1 is a flow chart of a preferred embodiment of an information parsing method based on vocabulary enhancement according to the present invention. Figure 1 As shown, it includes:

[0053] Step S100, obtaining bond information to be parsed, and constructing an input sequence according to the bond information, wherein the input sequence includes character vectors in the bond information and vocabulary vectors corresponding to the characters;

[0054] Step S200: construct a pre-training model, process the input sequence through the pre-training model, and generate a semantic representation of the input sequence;

[0055] Step S300: After passing the semantic representation through a discard layer and a normalization layer, a target vector is output, and the target vector is the parsed structured bond information.

[0056] In specific implementation, the embodiment of the present invention is used to parse bond information, and its main task is named entity recognition. Named entity recognition is an encoder-decoder model from the model point of view. Specifically, the encoder learns semantic representation, and the decoder learns downstream tasks such as entity extraction. Based on the basic end-to-end deep learning framework, the embodiment of the present invention constructs model input and vocabulary adapters to align with the original character-based text representation, so that vocabulary information can be integrated into the encoder, guiding the neural network to learn deeper vocabulary relationships.

[0057] like Figure 2 As shown, the network structure diagram of the named entity recognition model of the present invention is as follows: Figure 2 InputEmbedder: input embedding layer; Multi-Head Attention: multi-head attention; Add&LN: residual connection and layer normalization; Feed Forward: forward feedback neural network; Lexicon Adapter: vocabulary adapter device. i : the i-th input character; e i : the i-th character vector; The i-th output of the l-th layer of the model.

[0058] Construct an input sequence of character and word pairs; construct a pre-trained model to perform nonlinear transformation on the word vectors in the character and word sequence; obtain the weights and sum of all words through the bilinear attention layer to generate word information by analyzing the correlation between the transformed characters and the number of words; fuse the word information into the character vector to obtain the semantic representation of the input layer; obtain the fused representation by adding the word vector to the character vector; generate the target vector after passing the fused representation through the discard layer and the layer normalization layer to complete the parsing of the bond information.

[0059] In one embodiment, obtaining the bond information to be parsed and constructing an input sequence according to the bond information include:

[0060] Get the bond information to be parsed, and process the bond information into single characters;

[0061] Constructs input sequences from individual characters.

[0062] In specific implementation, because word segmentation is difficult to transfer to the pre-trained model of the Transformer structure and has uncertainty, the traditional pre-trained models for Chinese are all based on character-level input. The present invention proposes to construct an input sequence of "character-word" pairs, which can integrate the vocabulary related to the characters into the input layer of the model, so that the model can learn the semantics implied in the vocabulary.

[0063] In one embodiment, constructing an input sequence based on a single character includes:

[0064] Get the Chinese sentences in the bond information to be parsed;

[0065] Match the Chinese sentences according to the preset dictionary to obtain the vocabulary in the Chinese sentences;

[0066] Each character is paired with a word containing the character to generate an input sequence.

[0067] In the specific implementation, for a given Chinese sentence s c = {c1, c2, ..., c3}, and use the dictionary D to automatically match the potential words contained in the sentence. Then, among the matched words, each character and the word containing the character form a word pair, represented by s cw ={c1, ws1), (c2, ws2),..., (c n , ws n )}, where c i Indicates the i-th character in the sentence, ws i Indicates that it contains c i A vocabulary collection. Figure 3 The following is a schematic diagram of character and vocabulary pairs.

[0068] In one embodiment, a pre-training model is constructed, and an input sequence is processed by the pre-training model to generate a semantic representation of the input sequence, including:

[0069] Build a pre-trained model based on the BERT model;

[0070] The input sequence is processed through the BERT model to generate the semantic representation corresponding to the input sequence.

[0071] In specific implementation, the embodiment of the present invention constructs a pre-training model based on the BERT model, performs pre-training based on the BERT model, and generates semantic representations corresponding to the input sequence. The BERT model mainly utilizes the Encoder structure of the Transformer, and adopts the most primitive Transformer. In general, BERT has the following characteristics:

[0072] Structure: The Transformer Encoder structure is adopted, but the model structure is deeper than the Transformer. The Transformer Encoder contains 6 Encoder blocks, the BERT-base model contains 12 Encoder blocks, and the BERT-large contains 24 Encoder blocks. Training: The training is mainly divided into two stages: the pre-training stage and the Fine-tuning stage. The pre-training stage is similar to Word2Vec, ELMo, etc., and is trained on a large data set based on some pre-training tasks. The Fine-tuning stage is for fine-tuning when it is used for some downstream tasks later, such as text classification, part-of-speech tagging, question-answering systems, etc. BERT can be fine-tuned on different tasks without adjusting the structure.

[0073] In one embodiment, the input sequence is processed by the BERT model to generate a semantic representation corresponding to the input sequence, including:

[0074] Performing a nonlinear transformation on the vocabulary vectors in the input sequence through the BERT model to generate a nonlinearly transformed vocabulary vector, wherein the transformed vocabulary vector is aligned with the dimension of the character vector;

[0075] Calculate the correlation between the character vector and the vocabulary vector, calculate the weights of all vocabulary vectors based on the correlation, and calculate the target vocabulary vector based on the weights;

[0076] The target vocabulary vector is fused into the character vector to generate a semantic representation of the input sequence.

[0077] In one embodiment, calculating the correlation between the character vector and the vocabulary vector, calculating the weights of all vocabulary vectors according to the correlation, and calculating the target vocabulary vector according to the weights includes:

[0078] Compute the correlation between character vectors and vocabulary vectors based on a bilinear attention layer;

[0079] Calculate the weights of all vocabulary vectors based on their relevance;

[0080] Calculate the target vocabulary vector based on the weights.

[0081] In one embodiment, the target vocabulary vector is fused into the character vector to generate a semantic representation of the input sequence, including:

[0082] The target vocabulary vector is added to the character vector to generate a fused representation, which is a semantic representation of the input sequence.

[0083] In specific implementation, there are two major difficulties in integrating vocabulary information into character information, and existing technologies have no effective solutions: aligning the word vector and character vector dimensions to construct the input layer of the model; allowing the model to learn contextual semantically related words, that is, focusing on semantically related words.

[0084] In order to integrate the transformed input, i.e., the "character-vocabulary" pair sequence, into a pre-trained model such as BERT, the present invention constructs a vocabulary adapter device at the input layer to solve the above shortcomings.

[0085] like Figure 4 As shown, the word vectors in the "character-word" pair sequence are nonlinearly transformed to align the dimensions with the character vectors; the input of the i-th position of the "character-word" pair sequence is defined as: in:

[0086] Character vector, i.e. the output of a transformer layer in BERT.

[0087] The collection of vocabulary word vectors corresponding to the characters.

[0088] The representation of the jth word corresponding to the character. w is the pre-trained word vector table, w ij It is ws i The jth word in the vocabulary set.

[0089] Then, the dimensionality of vocabulary word vectors and character vectors is aligned through the following nonlinear transformation:

[0090]

[0091] Where W1 is d c ×d w The matrix of W2 is d c ×d c d is a matrix of , b1 and b2 are bias. c ,d w They represent the hidden layer vector dimension and word vector dimension of BERT respectively.

[0092] In order to give greater weight to the most relevant words matched by the dictionary, the following bilinear attention layer is used:

[0093] Using the obtained word vector Construct the vocabulary word vector corresponding to the i-th character V i is m×d c , where m is the number of words that each character matches. Then calculate the correlation a between the character and the word i :

[0094]

[0095] W attn is the weight matrix of the bilinear attention layer, and then the weights of all words are calculated:

[0096]

[0097] The purpose of the output layer is to obtain the vocabulary information Fuse into the character vector to obtain the semantic representation of the input layer. First, add the word vector to the character vector to get the fused representation

[0098]

[0099] Then, the obtained fusion representation The robustness is further improved by using the Dropout layer and the LayerNorm layer respectively.

[0100] The traditional BERT-based pre-training model lacks a device to integrate vocabulary information into each layer of Transformer modules, and the data input of each layer of Transformer modules is character vectors.

[0101] The pre-training model of the embodiment of the present invention is as follows Figure 5 As shown in Figure 1, by applying the vocabulary adapter device to the BERT pre-trained model, the powerful semantic representation ability of BERT and the contextual semantic information contained in the vocabulary are combined. The resulting vocabulary adapter is integrated into any Transformer layer in the BERT model.

[0102] The embodiment of the present invention adopts the BERT pre-training model to obtain text encoding. In some other embodiments, other Transformer-based pre-training models can be used according to business needs.

[0103] In some other embodiments, the vocabulary adapter device can be extended to be integrated between all Transformer layers.

[0104] In some other embodiments, different word vectors may be selected according to different businesses to obtain more accurate semantic representation.

[0105] The method of fusing vocabulary information and character information in the vocabulary adapter of the embodiment of the present invention utilizes nonlinear changes and bidirectional linear attention mechanism; integrates the vocabulary adapter into any layer of BERT transformer; fully utilizes the vocabulary information in the text, so that the device can better identify entity boundaries; and improves the accuracy of transaction element extraction in the secondary transaction business of financial bonds by 2%-4%.

[0106] It should be noted that there is not necessarily a certain order between the above-mentioned steps. A person skilled in the art can understand, based on the description of the embodiments of the present invention, that in different embodiments, the above-mentioned steps may have different execution orders, that is, they may be executed in parallel, may be executed interchangeably, and so on.

[0107] Another embodiment of the present invention provides an information parsing device based on vocabulary enhancement, such as Figure 6 As shown, the device 1 comprises:

[0108] A sequence construction module 11 is used to obtain bond information to be parsed and construct an input sequence according to the bond information, wherein the input sequence includes a character vector in the bond information and a vocabulary vector corresponding to the character;

[0109] Input sequence processing 12, used to build a pre-training model, process the input sequence through the pre-training model, and generate a semantic representation of the input sequence;

[0110] The target vector output module 13 is used to output a target vector after passing the semantic representation through a discard layer and a normalization layer, and the target vector is the parsed structured bond information.

[0111] The specific implementation method is shown in the method embodiment, which will not be described in detail here.

[0112] Another embodiment of the present invention provides an electronic device, such as Figure 7 As shown, the electronic device 10 includes:

[0113] One or more processors 110 and memory 120, Figure 7 A processor 110 is used as an example for description. The processor 110 and the memory 120 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.

[0114] The processor 110 is used to complete various control logics of the electronic device 10, and it can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISCMachine) or other programmable logic device, a discrete gate or transistor logic, a discrete hardware control, or any combination of these components. In addition, the processor 110 can also be any traditional processor, microprocessor or state machine. The processor 110 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.

[0115] The memory 120 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions corresponding to the information parsing method based on vocabulary enhancement in the embodiment of the present invention. The processor 110 executes various functional applications and data processing of the device 10 by running the non-volatile software programs, instructions and units stored in the memory 120, that is, implementing the information parsing method based on vocabulary enhancement in the above method embodiment.

[0116] The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store applications required for operating the device and at least one function; the data storage area may store data created according to the use of the device 10, etc. In addition, the memory 120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 120 may optionally include a memory remotely arranged relative to the processor 110, and these remote memories may be connected to the device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0117] One or more units are stored in the memory 120, and when executed by one or more processors 110, perform the information parsing method based on vocabulary enhancement in any of the above method embodiments, for example, perform the above described Figure 1 The method comprises steps S100 to S300.

[0118] An embodiment of the present invention provides a non-volatile computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors, for example, to execute the above-described Figure 1 The method comprises steps S100 to S300.

[0119] As an example, non-volatile storage media can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) as external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms such as synchronous RAM (SRAM), dynamic RAM, (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The disclosed memory controls or memories of the operating environment described herein are intended to include one or more of these and / or any other suitable types of memory.

[0120] Another embodiment of the present invention provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the information parsing method based on vocabulary enhancement of the above method embodiment. For example, the above described Figure 1 The method comprises steps S100 to S300.

[0121] The embodiments described above are merely illustrative, in which the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0122] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can exist in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.

[0123] Conditional language such as "can," "can," "might," or "may," among other things, unless specifically stated otherwise or otherwise understood within the context as used, is generally intended to convey that particular embodiments can include (while other embodiments do not) particular features, elements, and / or operations. Thus, such conditional language is also generally intended to imply that features, elements, and / or operations are anyway required for one or more embodiments or that one or more embodiments must include logic for determining, with or without input or prompting, whether such features, elements, and / or operations are included or will be performed in any particular embodiment.

[0124] What has been described herein in this specification and the accompanying drawings includes examples of information parsing methods and devices that can provide vocabulary-based enhancement. Of course, it is not possible to describe every conceivable combination of elements and / or methods for the purpose of describing the various features of the present disclosure, but it is recognized that many other combinations and permutations of the disclosed features are possible. Therefore, it is obvious that various modifications can be made to the present disclosure without departing from the scope or spirit of the present disclosure. In addition, or in an alternative, other embodiments of the present disclosure may be apparent from consideration of this specification and the accompanying drawings and from the practice of the present disclosure as presented herein. It is intended that the examples set forth in this specification and the accompanying drawings are considered to be illustrative and not restrictive in all respects. Although specific terms are employed herein, they are used in a general and descriptive sense and are not used for limiting purposes.

Claims

1. An information parsing method based on vocabulary enhancement, characterized in that: The method comprises: Obtaining bond information to be parsed, and constructing an input sequence according to the bond information, wherein the input sequence includes character vectors in the bond information and vocabulary vectors corresponding to the characters; Construct a pre-trained model, process the input sequence through the pre-trained model, and generate a semantic representation of the input sequence; After passing the semantic representation through the discard layer and the normalization layer, a target vector is output, which is the parsed structured bond information; The step of constructing a pre-trained model and processing the input sequence through the pre-trained model to generate a semantic representation of the input sequence includes: Build a pre-trained model based on the BERT model; The input sequence is processed through the BERT model to generate the semantic representation corresponding to the input sequence; The input sequence is processed by the BERT model to generate a semantic representation corresponding to the input sequence, including: Performing a nonlinear transformation on the vocabulary vectors in the input sequence through the BERT model to generate a nonlinearly transformed vocabulary vector, wherein the transformed vocabulary vector is aligned with the dimension of the character vector; Calculate the correlation between the character vector and the transformed vocabulary vector according to the bilinear attention layer, calculate the weights of all transformed vocabulary vectors according to the correlation, and calculate the target vocabulary vector according to the weights; The target vocabulary vector is merged into the character vector to generate a semantic representation of the input sequence; The step of fusing the target vocabulary vector into the character vector to generate a semantic representation of the input sequence includes: The target vocabulary vector is added to the character vector to generate a fused representation, which is a semantic representation of the input sequence.

2. The method according to claim 1, characterized in that The step of obtaining the bond information to be parsed and constructing an input sequence according to the bond information includes: Get the bond information to be parsed, and process the bond information into single characters; Constructs input sequences from individual characters.

3. The method according to claim 2, characterized in that The step of constructing an input sequence according to a single character comprises: Get the Chinese sentences in the bond information to be parsed; Match the Chinese sentences according to the preset dictionary to obtain the vocabulary in the Chinese sentences; Each character is paired with a word containing the character to generate an input sequence.

4. An information parsing device based on vocabulary enhancement, characterized in that: The device comprises: A sequence construction module, used to obtain bond information to be parsed, and construct an input sequence according to the bond information, wherein the input sequence includes a character vector in the bond information and a vocabulary vector corresponding to the character; Input sequence processing is used to build a pre-trained model, process the input sequence through the pre-trained model, and generate a semantic representation of the input sequence; A target vector output module is used to output a target vector after passing the semantic representation through a discard layer and a normalization layer, wherein the target vector is the parsed structured bond information; The step of constructing a pre-trained model and processing the input sequence through the pre-trained model to generate a semantic representation of the input sequence includes: Build a pre-trained model based on the BERT model; The input sequence is processed through the BERT model to generate the semantic representation corresponding to the input sequence; The input sequence is processed by the BERT model to generate a semantic representation corresponding to the input sequence, including: Performing a nonlinear transformation on the vocabulary vectors in the input sequence through the BERT model to generate a nonlinearly transformed vocabulary vector, wherein the transformed vocabulary vector is aligned with the dimension of the character vector; Calculate the correlation between the character vector and the transformed vocabulary vector according to the bilinear attention layer, calculate the weights of all transformed vocabulary vectors according to the correlation, and calculate the target vocabulary vector according to the weights; The target vocabulary vector is merged into the character vector to generate a semantic representation of the input sequence; The step of fusing the target vocabulary vector into the character vector to generate a semantic representation of the input sequence includes: The target vocabulary vector is added to the character vector to generate a fused representation, which is a semantic representation of the input sequence.

5. An electronic device, characterized in that: The electronic device comprises at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the information parsing method based on vocabulary enhancement as described in any one of claims 1-3.

6. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, enable the one or more processors to execute the information parsing method based on vocabulary enhancement as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Chinese named entity recognition method and device based on vocabulary enhancement and multiple features

    CN114021549A