Multimodal Feature Fusion Method, Apparatus, Device, Medium and Product
Through the multimodal feature fusion method and BERT model, the difficulty in extracting key information caused by document layout changes is solved, and efficient entity relationship extraction is achieved in financial audit scenarios, improving audit efficiency.
Patent Information
- Application Number
- CN202210416064.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-20
AI Technical Summary
In the case where the document layout is not fixed, it is difficult for the prior art to effectively extract key information, resulting in inefficient auditing. Especially in financial audit scenarios, existing optical character recognition technology cannot adapt to layout adjustments or new layout documents.
The multimodal feature fusion method is used to identify document images through OCR to obtain text features and position features, divide the image into multiple regions according to preset rules, extract regional image features, and use the BERT model to perform deep feature fusion, and extract entity relationships based on the table sequence relationship extraction model.
It realizes accurate extraction of entities and entity relationships in documents of different layouts, improves audit efficiency, reduces dependence on human resources, and is highly adaptable.
Smart Images

Figure CN114821255B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to technologies such as natural language processing and optical character recognition, and can be applied to intelligent financial scenarios. Background Art
[0002] In some scenarios, it is necessary to review key information in documents. For example, in a reimbursement form, it is necessary to review information such as the name of the reimburser, the reimbursement amount, and the consumption date. And the review of this information often requires a large amount of manpower. In order to improve the review efficiency, a neural network is used to process the image of the document to automatically extract entities and entity relationships that users are interested in from the document. The related technology writes specific rules for documents with specific formats, and this method has great limitations. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, medium, and product for fusing multi-modal features.
[0004] According to one aspect of the present disclosure, there is provided a method for fusing multi-modal features, including: obtaining an image including text; performing feature recognition on the image to obtain text features and position features of the text; dividing the image into multiple regions according to a preset rule, and extracting image features of at least one region among the multiple regions; encoding the text features to obtain a text vector; and encoding the image features of the at least one region to obtain an image vector of the at least one region; and encoding the position features to obtain a position vector; fusing the text vector, the image vector of the at least one region, and the position vector to obtain a fused target vector.
[0005] According to another aspect of the present disclosure, there is provided a multi-modal feature fusion apparatus, including: an obtaining unit for obtaining an image including text; a recognition unit for performing feature recognition on the image to obtain text features and position features of the text; a dividing and extracting unit for dividing the image into multiple regions according to a preset rule and extracting image features of at least one region among the multiple regions; a determining vector unit for encoding the text features to obtain a text vector; and encoding the image features of the at least one region to obtain an image vector of the at least one region; and encoding the position features to obtain a position vector; a fusion unit for fusing the text vector, the image vector of the at least one region, and the position vector to obtain a fused target vector.
[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described above.
[0008] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements the method described above when executed by a processor.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Description of the Drawings
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a flowchart of a method for fusing multi-modal features provided according to an embodiment of the present disclosure;
[0012] Figure 2 is a flowchart of a method for obtaining a text vector provided according to an embodiment of the present disclosure;
[0013] Figure 3 is a flowchart of a method for obtaining an image vector provided according to an embodiment of the present disclosure;
[0014] Figure 4 is a flowchart of a method for obtaining a two-dimensional position vector provided according to an embodiment of the present disclosure;
[0015] Figure 5 is a flowchart of a method for obtaining input features provided according to an embodiment of the present disclosure;
[0016] Figure 6 is a flowchart of a method for obtaining a fused target vector provided according to an embodiment of the present disclosure;
[0017] Figure 7 is a filling representation intention provided according to an embodiment of the present disclosure;
[0018] Figure 8 is a schematic diagram of an application form review scenario provided according to an embodiment of the present disclosure;
[0019] Figure 9 It is a block diagram of a multi-modal feature fusion device shown according to an exemplary embodiment;
[0020] Figure 10 It is a block diagram of an electronic device for implementing the multi-modal feature fusion method of the embodiments of the present disclosure. Specific embodiments
[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0022] The present disclosure is applied to extracting entities and entity relationships in a document to be reviewed in a document review scenario. In related technologies, the Optical Character Recognition (OCR) technology is used to recognize the text in a document image. For example, in a financial review scenario, specific rule codes are usually written according to business requirements to extract corresponding key fields. The method of writing specific rules to extract key fields is applicable to business scenarios where the document layout is fixed or does not change much. However, in scenarios where the document layout is not fixed, or after the business is launched, the document layout is adjusted or new layouts are added, the originally written rules may not be applicable to the documents with adjusted or newly added layouts. In addition, in the case of a relatively complex document layout, the pre-written rules cannot cover all situations, resulting in failure to achieve the expected effect.
[0023] In view of this, the present disclosure provides a method for fusing multi-modal features. For a document image to be recognized, OCR is used to recognize the document image to obtain the text features and location features in the document. The document image is divided into multiple regions according to a preset rule, and the image features of at least one region among the multiple regions are extracted. The text features, location features, and image features are encoded and then input into BERT (Bidirectional Encoder Representation from Transformers) for deep feature fusion, and the output features are used as the overall features of the document. When the present disclosure obtains the image features, it does not obtain the image features by taking the document image as a whole, nor does it obtain the image features by taking each character in the document image as a region. Instead, the document image is divided into multiple regions according to a preset rule, and the image features of at least one region among the multiple regions are extracted. By using the above method for obtaining image features, during the process of multi-modal feature fusion, different attentions are assigned according to the location features of different image features, which can make the multi-modal feature fusion more sufficient.
[0024] Using the method for fusing multi-modal features provided by the present disclosure, the overall features of the document image can be applied to entity relationship extraction. When applying, a relationship extraction model based on Table Sequence is used to extract entity relationships from the overall features of the document image.
[0025] In the following embodiments of the present disclosure, the method for fusing multi-modal features provided by the present disclosure will be described with reference to the accompanying drawings.
[0026] Figure 1 is a flowchart of the method for fusing multi-modal features provided by the embodiments of the present disclosure; as Figure 1 shown, the method for fusing multi-modal features provided by the present disclosure includes the following steps S101-S105.
[0027] In step S101, an image including text is obtained.
[0028] In the present disclosure, the image including text can be a document image. This image can be an image local to the device running the multi-modal feature fusion method. This image can also be an image sent by other devices that have a communication connection with the device running the multi-modal feature fusion method to this device. This image can also be an image obtained in real time through a user instruction.
[0029] In step S102, the image is feature-recognized to obtain the text features and location features of the text.
[0030] In the present disclosure, OCR technology can be used to perform feature recognition on images. The text features and position features of the images are extracted through OCR technology. The position features in the present disclosure include one-dimensional position features and two-dimensional position features. Among them, the one-dimensional position feature refers to the relative positions between the text features. The two-dimensional position feature refers to the position coordinates of the text features in a coordinate system established with a specified origin. The specified origin in the present disclosure can be the upper left corner of the image.
[0031] In step S103, the image is divided into multiple regions according to a preset rule, and the image features of at least one region among the multiple regions are extracted.
[0032] In the present disclosure, in order to be able to extract entity relationships more comprehensively in the entity relationship extraction stage, during the process of obtaining the overall features after image fusion, the image is divided into multiple regions. The image features of at least one region among the multiple regions are extracted. In some examples, in order to improve the result accuracy, the image features of each region among the multiple regions can be extracted. The present disclosure does not make a specific limitation on how many regions of image features are selected for extraction. Compared with extracting the image features with the image as a whole, more image features of the image can be included in the overall features of the image, which is beneficial to referring to the position features of the image features during the fusion process to obtain more attention.
[0033] The rule for dividing the image into multiple regions in the present disclosure can be set according to actual needs. For example, a rule can be set to divide the image into four regions along the horizontal center line and the vertical center line. Or, it can also be set to equally divide the image into M regions along the horizontal or vertical direction, where M is a positive integer.
[0034] In one implementation, the image is divided into four regions: the upper left corner, the upper right corner, the lower right corner, and the lower left corner, denoted as V1, V2, V3, and V4 according to the set rule. In this embodiment, the ResNeXt-FPN module is used as the backbone network for image feature extraction. The FPN in this embodiment refers to the intermediate network of the feature pyramid, which is the abbreviation of Feature Pyramid Network. The image features of V1, V2, V3, and V4 are respectively extracted through the backbone network.
[0035] In step S104, the text features are encoded to obtain text vectors; and, the image features of at least one region are encoded to obtain at least one region's image vectors; and, the position features are encoded to obtain position vectors.
[0036] After performing feature recognition on the image, extracting text features, image features of at least one region, and position features, the text features are encoded to obtain text vectors. The image features of at least one region among multiple regions are encoded to obtain image vectors of at least one region among multiple regions. The position features are encoded to obtain position vectors. Of course, in some examples, the image features of each region among the regions can be encoded to obtain image vectors of each region among the regions.
[0037] In step S105, the text vector, the image vectors of at least one region, and the position vector are fused to obtain a fused target vector.
[0038] In the present disclosure, a Bert model can be used to fuse the text vector, the image vectors of at least one region, and the position vector to obtain a fused target vector. The fused target vector is used as the overall feature of the image for subsequent applications.
[0039] In the present disclosure, by performing feature recognition on the image, text features and position features are obtained, and after dividing the image into multiple regions, image features of at least one region among the multiple regions are extracted. After encoding and fusing the text features, position features, and image features of at least one region, compared with the image features extracted by taking the image as a whole, the degree of fusion of multi-modal features can be improved.
[0040] The following embodiments of the present disclosure will specifically describe the process of obtaining the text vector, the image vectors of at least one region, and the position vector.
[0041] Figure 2 is a flowchart of a method for obtaining a text vector according to an embodiment of the present disclosure; as Figure 2 shown, the process of encoding the text features to obtain text vectors provided by the present disclosure includes the following steps S201-S204.
[0042] In step S201, the text is tokenized, and the tokenization results are serialized to obtain multiple sequences.
[0043] After using OCR technology to recognize the text in the image, in the present disclosure, existing tokenization methods (such as WordPiece) can be used to tokenize the text to obtain tokenization results. To serialize the tokenization results, [CLS] can be used to mark the start of the sequence, and [SEP] is used as the end of the sequence. In the present disclosure, the tokenization results are marked with Tokens.
[0044] Word embedding (Token Embedding) is used to represent the semantic information of Tokens.
[0045] In step S202, according to the relative position information among the sequences in multiple sequences, the one-dimensional position encoding of each sequence is determined.
[0046] In the present disclosure, one-dimensional position encoding (Position Embedding) is adopted to represent the encoding of the relative position information of each Token in each sequence.
[0047] In step S203, based on the word embedding representing the semantic information of the sequence, the one-dimensional position encoding of the sequence, and the segment embedding different from other sequences, the sequence vector of the sequence is determined.
[0048] In order to distinguish different sequences and distinguish sequences from image features in the present disclosure, segment embedding (Segment Embedding) is added to each sequence vector. Therefore, each sequence vector of a single text sequence is composed of the following three vectors, denoted as:
[0049] t = Token Embedding + Position Embedding + Segment Embedding
[0050] In step S204, based on the sequence vectors corresponding to the sequences in the text, a text vector is generated.
[0051] The present disclosure obtains a text vector by segmenting the text into multiple sequences, according to the word embedding, one-dimensional position encoding, and segment embedding representing the semantic information of each sequence, so as to prepare for multi-modal feature fusion.
[0052] The above embodiments are combined Figure 2 to illustrate the process of determining the text vector. The following embodiments will be combined Figure 3 to illustrate the process of determining the image vector.
[0053] Figure 3 is a flowchart of a method for obtaining an image vector according to an embodiment of the present disclosure; as Figure 3 shown, the process of encoding the image features of at least one region provided by the present disclosure to obtain the image vectors of at least one region includes the following steps S301-S304.
[0054] In step S301, the image features of at least one region are respectively subjected to pooling processing to obtain the initial image vectors of at least one region.
[0055] The image features of at least one region are transformed into initial image vectors of a fixed size through the pooling operation of a convolutional neural network. Of course, in some examples, the image features of each region can be transformed into initial image vectors of a fixed size through the pooling operation of a convolutional neural network. Suppose there are image features of 4 regions, namely the first region image feature, the second region feature, the third region image feature, and the fourth region feature. The initial image vector of the first region image is obtained by performing the pooling operation of the convolutional neural network on the first region image feature. The same operation as that on the first region image feature is performed on the second region feature, the third region image feature, and the fourth region feature to obtain their respective corresponding initial image vectors. Of course, it is also possible to select some of the first region image feature, the second region feature, the third region image feature, and the fourth region feature for pooling processing to obtain the initial image vectors corresponding to the corresponding region image features, which is not limited in this disclosure.
[0056] In step S302, linear transformations are respectively performed on the initial image vectors of at least one region.
[0057] In order to make the length of the image vector consistent with that of the text vector, this disclosure adds a projection layer (ProjectLayer) to perform a linear transformation on the image vector to make its length consistent with that of the text vector. In some examples, linear transformations can be respectively performed on the initial image vectors of each region.
[0058] In step S303, according to the positional relationship of at least one region, one-dimensional positional encodings corresponding to the initial image vectors of at least one region are determined.
[0059] Since the convolutional neural network does not have information about the order of image vectors, this disclosure adds a one-dimensional positional encoding (Position Embedding) representing the order of image vectors. The relative order of image vectors in this disclosure can be from left to right and from top to bottom. In some examples, one-dimensional positional encodings corresponding to the initial image vectors of each region can be determined according to the positional relationship of each region.
[0060] In step S304, based on the initial image vectors of at least one region after linear transformation, the one-dimensional positional encodings, and the segment embeddings different from other initial image vectors, the image vectors of at least one region are determined.
[0061] In some examples, based on the initial image vectors of each region after linear transformation, the one-dimensional positional encodings, and the segment embeddings different from other initial image vectors, the image vectors of each region can be determined.
[0062] In the present disclosure, in order to distinguish the initial image vectors corresponding to different regions and to distinguish image vectors from text vectors, a Segment Embedding different from them can be added to the initial image vectors. In the present disclosure, the image vectors can be denoted as:
[0063] v = Proj(Vis Embedding)+Position Embedding+Segment Embedding
[0064] In the present disclosure, the initial image vectors are obtained by performing pooling processing on the image features, and then the initial image vectors are linearly transformed to make the lengths of the initial image vectors consistent with those of the text vectors. According to the linearly transformed initial image vectors, one-dimensional position encodings, and segment embeddings that are different from other initial image vectors, the final image vectors are determined to prepare for multi-modal feature fusion.
[0065] It can be understood that the present disclosure can perform corresponding operations on each of multiple regions, making the results more accurate.
[0066] Next, the present disclosure will combine the attached Figure 4 to illustrate the process of determining the two-dimensional position vectors.
[0067] Figure 4 is a flowchart of a method for obtaining two-dimensional position vectors according to an embodiment of the present disclosure; as Figure 4 shown, the position features can be two-dimensional position features, and the position vectors can be two-dimensional position vectors. The process of encoding the two-dimensional position features to obtain two-dimensional position vectors provided by the present disclosure includes the following steps S401-S404.
[0068] In step S401, the first coordinate and the second coordinate of the text box characterized by the two-dimensional position features, as well as the height and width of the text box, are encoded.
[0069] After the present disclosure uses OCR technology to recognize an image, it can obtain the position information of the text box therein. In the present disclosure, the position of the text box is represented by two-dimensional position features. The two-dimensional position features include the first coordinate and the second coordinate, as well as the height and width of the text box. The first coordinate and the second coordinate are respectively the coordinates at the diagonal positions of the text box.
[0070] Exemplarily, the two-dimensional position features can be (x0, x1, y0, y1, w, h), where (x0, y0) are the coordinates of the upper left corner of the text box, (x1, y1) are the coordinates of the lower right corner of the text box, w is the width of the text box, and h is the height of the text box.
[0071] The present disclosure encodes the x - coordinate and y - coordinate in the first coordinate and the second coordinate using two - dimensional Position Embedding, and encodes the height of the text box and the width of the text box using two - dimensional Position Embedding.
[0072] In step S402, the x - coordinate in the encoded first coordinate and the x - coordinate in the second coordinate are concatenated with the width of the encoded text box to obtain a position vector in the x - axis direction.
[0073] In step S403, the y - coordinate in the encoded first coordinate and the y - coordinate in the second coordinate are concatenated with the height of the encoded text box to obtain a position vector in the y - axis direction.
[0074] In step S404, the position vector in the x - axis direction and the position vector in the y - axis direction are used as the two - dimensional position vector of the text box.
[0075] In the present disclosure, the two - dimensional position vector obtained through steps S401 - S404 is denoted as:
[0076] i = Concat(Pos Embedding(x0, x1, w), Pos Embedding(y0, y1, h))
[0077] It should be noted that when the present disclosure encodes two - dimensional position information representing an image region, it encodes the upper - left coordinate and the lower - right coordinate of the corresponding image region to obtain the two - dimensional position vector of the image region. For the special vectors [CLS] and [SEP], (0, 0, 0, 0, 0, 0) is used.
[0078] After the present disclosure encodes the two - dimensional position features, it concatenates the coordinates in the x - axis direction with the width of the text box, and concatenates the coordinates in the y - axis direction with the height of the text box, finally obtaining a two - dimensional position vector to prepare for multi - modal feature fusion.
[0079] To more clearly illustrate the process of the present disclosure for obtaining text vectors, image vectors of at least one region, and position vectors, an appendix Figure 5 is used for illustration. Figure 5 is a flowchart of obtaining input features according to the method provided by the embodiment of the present disclosure; as Figure 5As shown, the present disclosure performs OCR recognition on a document image to obtain text features and location features. In some examples, the location features can be two-dimensional location features. Of course, under other specific conditions, they can also be location features with more dimensions or fewer dimensions, etc. The present disclosure does not make any limitations. The document image is divided into 4 regions according to a preset rule. Image features matching the image of the corresponding region are obtained by performing image feature extraction on the document image of at least one region. The image features and text features are encoded to obtain an image vector and a text vector, and the image vector and the text vector are concatenated. A position encoding is superimposed on the concatenated image vector and text vector. For example, a two-dimensional position encoding and a one-dimensional position encoding can be superimposed. The vector obtained by superimposing the two-dimensional position encoding and the one-dimensional position encoding after concatenation is used as an input vector and input into the BERT model. The full fusion of multi-modal features is achieved in the BERT model.
[0080] Through Figure 5 It can be seen that in the present disclosure, the text vector, the image vectors of at least one region, and the position vector are fused to obtain a fused target vector, including concatenating the text vector and the image vectors of at least one region. In some examples, it can be concatenating the text vector and the image vectors of each region. Then, a position vector, such as a two-dimensional position vector, is superimposed on the concatenated vector to obtain an input vector. The input vector is input into the Bert model for fusion to obtain a fused target vector.
[0081] The present disclosure concatenates the text vector and the image vectors of at least one region, and superimposes a position vector on the concatenated vector to obtain an input vector for the Bert model. In the Bert model, different attentions are respectively given to the text vectors or image vectors at different positions according to the position vector, which can enable the Bert model to focus on the parts that are more helpful for the application and ignore the interfering parts.
[0082] The Bert model includes multiple encoders. For example, the Bert model can include multiple sequentially connected encoders. In the following embodiments of the present disclosure, Figure 6 the fusion process of one of the encoders in the Bert model will be described.
[0083] Figure 6 is a flowchart of obtaining a fused target vector according to the method provided by the embodiments of the present disclosure; as Figure 6 shown, in the present disclosure, inputting the input vector into the Bert model for fusion includes the following steps S601 - S605.
[0084] In step S601, the input vector is input into the first encoder.
[0085] In step S602, similarity attention scores are determined in the first encoder based on the similarities between the text vectors and the image vector in the input vector.
[0086] For the convenience of description in the following embodiments of the present disclosure, each text vector and image vector included in the input vector are represented by tokens. There are multiple tokens in the input vector.
[0087] Three matrices W are randomly initialized and generated Q 、W K 、W V . The tokens in the input vector are multiplied by W Q 、W K 、W V respectively according to the following formula to obtain three vectors: Query, Key, and Value.
[0088] Query = XW Q
[0089] Key = XW K
[0090] Value = XW V
[0091] For each token, the similarity between the Query and Key vectors is calculated
[0092]
[0093] In the above formula, Query i is the Query vector of the i-th token, key j is the Key vector of the j-th token, and Similarity(Query i , Key j ) is the similarity between the i-th token and the j-th token. The value ranges of both i and j are N, where N represents the number of tokens. After determining i in the above formula, the value of j is taken from 1 to N. In this way, the similarity between the i-th token and N tokens can be obtained, that is, the similarity between the current token and itself, and the similarity between the current token and N - 1 other tokens.
[0094] According to the similarity between the i-th token and N tokens, the similarity attention score is calculated according to the following formula.
[0095]
[0096] In the formula, a i is the similarity attention score of the i-th token. Represents the sum of the similarities between the i-th Token and the j-th Token. Represents the similarity of the current Token with itself.
[0097] In step S603, based on the similarity attention scores, and the position vectors corresponding to each text vector and the position vectors corresponding to the image vectors, determine the spatial attention scores.
[0098] In some examples, for instance, it can be based on the similarity attention scores, and the two-dimensional position vectors corresponding to each text vector and the two-dimensional position vectors corresponding to the image vectors, to determine the spatial attention scores.
[0099] The present disclosure assigns different attentions to Tokens with different position information. In order to distinguish the similarity attention scores, the attention assigned according to the two-dimensional position vectors is represented by the spatial attention scores.
[0100]
[0101] a′ i is the spatial attention score of the i-th Token, W is the spatial distance between the i-th and the j-th Tokens in the x-axis direction, and the calculation method is the absolute value of the difference between the x coordinate of the i-th Token and the x coordinate of the j-th Token. H is the spatial distance between the i-th Token and the j-th Token in the y-axis direction, and the calculation method is the absolute value of the difference between the y coordinate of the i-th text vector and the y coordinate of the j-th text vector. is the bias of the i-th Token and the j-th Token in the two-dimensional space x-axis direction. is the bias of the i-th Token and the j-th Token in the two-dimensional space y-axis direction. Where 2D represents the two-dimensional space.
[0102] Based on the spatial attention score of the i-th Token, determine the output vector of the i-th Token according to the following formula.
[0103]
[0104] By calculating a′ i The formula can obtain N spatial attention scores of the i-th Token. For the i-th Token, calculate the product of the spatial attention score of each Token and the corresponding Value vector, obtain N products, and take the sum of the N products as the output vector of the i-th Token.
[0105] In the present disclosure, the value range of i is N. According to the above process, the output vectors of N Tokens can be obtained.
[0106] In step S604, based on the spatial attention scores, the output of the first encoder is obtained.
[0107] The present disclosure concatenates the outputs of N Tokens as the output of the first encoder.
[0108] The present disclosure performs a linear transformation on the output of the first encoder through a Project Layer to keep its dimension consistent with the input.
[0109] In step S605, the output of the first encoder is used as the input of the second encoder. After passing through all the encoders, a fused target vector is obtained.
[0110] Assume that the Bert model has 12 stacked encoders (Attention Encoder). The output of the first encoder is used as the input of the second encoder. After such calculations for multiple times, deep feature interactions are performed among the image vector, the two-dimensional position vector, and the text vector.
[0111] The present disclosure assigns different spatial attention scores to the vectors at different positions according to the two-dimensional position vectors corresponding to the respective text vectors and the two-dimensional position vectors corresponding to the image vectors, which can increase the in-depth fusion among the image vector, the text vector, and the two-dimensional position vector and avoid focusing on local image vectors or text vectors.
[0112] The multi-modal feature fusion method provided by the present disclosure can be used in the scenario of extracting entities and entity relationships from the target vector.
[0113] Extracting entities and entity relationships from the target vector obtained by the multi-modal feature fusion method of the present disclosure can be not limited by the layout limitation and more accurately extract entities and entity relationships.
[0114] In the present disclosure, entities and entity relationships are extracted from the target vector based on the Table Sequence. Generally, the relationship extraction task needs to involve two steps. The first step is named entity recognition (NER), and the second step is relationship recognition (RE). Generally, there are two ways to implement relationship extraction. The first is the serial way, that is, entities are extracted first and then relationships are recognized. The second is the joint way, that is, extraction and recognition are performed simultaneously. The present disclosure can use the joint extraction way.
[0115] For NER and RE, the algorithm learns different sequence representations and table representations respectively, and these two representations can capture task-related information respectively. For the NER task, it is assumed to be a sequence tagging problem. For the RE task, given a sentence x = [x i 1≤i≤N , if there is a relationship r, it is represented by and , and ⊥ is used where there is no relationship. Figure 7 is the padding representation intention provided according to the embodiments of the present disclosure. For the sentence "Xiaoming is from Beijing", the table shown in Figure 7 can be obtained by padding.
[0116] The table representation is an N*N vector table. The MD-RNN (multi-dimensional RNN) based on the GRU structure is used as the Text Encoder. When updating the information of the current cell in the table, the structural characteristics of the table are utilized, and the information in the four directions of up, down, left, and right is fused through the MD-RNN. At the same time, the representations of the two words corresponding to the current cell under the Sequence Encoder are introduced to enable feature interaction between the Table Encoder and the Sequence Encoder. The Sequence Encoder adopts a structure similar to Transformers. The models in the present disclosure all adopt the cross-entropy Loss for the selection of the Loss function.
[0117] The present disclosure shows the application scenarios of the present disclosure through Figure 8 . Figure 8 is a schematic diagram of the application form review scenario provided according to the embodiments of the present disclosure. An image including text is obtained, and the target vector in the image is obtained by using the multi-modal feature fusion method provided by the present disclosure. Entities and entity relationships are extracted from the target vector based on the Table Sequence. In the application form review scenario shown in Figure 8 , the basic information of the applicant needs to be reviewed. It can be seen that the key entities can be extracted through the present disclosure and the relationships can be corresponding one by one.
[0118] Based on the same concept, the embodiments of the present disclosure also provide a multi-modal feature fusion device.
[0119] It can be understood that, in order to implement the above functions, the multi-modal feature fusion device provided by the embodiments of the present disclosure includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of the various examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.
[0120] Figure 9 is a block diagram of a multi-modal feature fusion device shown according to an exemplary embodiment. Referring to Figure 9 , the device includes an acquisition unit 901, an identification unit 902, a division and extraction unit 903, a determination vector unit 904, a fusion unit 905, and a relationship extraction unit 906.
[0121] The acquisition unit 901 is configured to acquire an image including text; the identification unit 902 is configured to perform feature identification on the image to obtain the text features and position features of the text; the division and extraction unit 903 is configured to divide the image into multiple regions according to a preset rule, and extract the image features of at least one region among the multiple regions; the determination vector unit 904 is configured to encode the text features to obtain a text vector; and encode the image features of at least one region to obtain at least one region image vector; and encode the position features to obtain a position vector; the fusion unit 905 is configured to fuse the text vector, at least one region image vector, and the position vector to obtain a fused target vector.
[0122] In one implementation, the determination vector unit 904 is configured to: perform word segmentation on the text, and serialize the word segmentation result to obtain multiple sequences; determine the one-dimensional position encoding of each sequence according to the relative position information between the sequences in the multiple sequences; determine the sequence vector of the sequence based on the word embedding representing the semantic information of the sequence, the one-dimensional position encoding of the sequence, and the segment embedding different from other sequences; generate a text vector based on the sequence vectors corresponding to the sequences in the text.
[0123] In one embodiment, the determining vector unit 904 is further configured to: perform pooling processing on the image features of at least one region to obtain initial image vectors of at least one region; perform linear transformation on the initial image vectors of at least one region respectively; determine one-dimensional position encodings of the initial image vectors corresponding to at least one region according to the positional relationship of at least one region; and determine the image vectors of at least one region based on the initial image vectors after linear transformation of at least one region, the one-dimensional position encodings, and a segment embedding different from other initial image vectors.
[0124] In one embodiment, the position feature is a two-dimensional position feature and the position vector is a two-dimensional position vector; the determining vector unit 904 is further configured to: encode the first coordinate and the second coordinate of the text box characterized by the two-dimensional position feature, as well as the height and width of the text box, where the first coordinate and the second coordinate are respectively the coordinates at the diagonal positions of the text box; splice the x coordinate in the encoded first coordinate and the x coordinate in the encoded second coordinate with the encoded width of the text box to obtain a position vector in the x-axis direction; splice the y coordinate in the encoded first coordinate and the y coordinate in the encoded second coordinate with the encoded height of the text box to obtain a position vector in the y-axis direction; and use the position vector in the x-axis direction and the position vector in the y-axis direction as the two-dimensional position vector of the text box.
[0125] In one embodiment, the fusion unit 905 is configured to: splice the text vector and the image vectors of at least one region; superimpose the position vector on the spliced vector to obtain an input vector; and input the input vector into the Bert model for fusion to obtain a fused target vector.
[0126] In one embodiment, the Bert model includes multiple encoders; the fusion unit 905 is further configured to: input the input vector into the first encoder; determine similarity attention scores in the first encoder based on the similarity between each text vector and the image vector in the input vector; determine spatial attention scores based on the similarity attention scores, the position vectors corresponding to each text vector, and the position vectors corresponding to the image vectors; obtain the output of the first encoder based on the spatial attention scores; and use the output of the first encoder as the input of the second encoder until after passing through all the encoders, to obtain the fused target vector.
[0127] In one embodiment, the apparatus 900 further includes a relationship extraction unit 906, configured to extract entities and entity relationships from the target vector.
[0128] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0129] …
[0130] In the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0131] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0132] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0133] As Figure 10 shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0134] A plurality of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0135] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the multi-modal feature fusion method. For example, in some embodiments, the multi-modal feature fusion method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the multi-modal feature fusion method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the multi-modal feature fusion method in any other suitable way (e.g., by means of firmware).
[0136] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0137] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0138] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0139] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0140] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0141] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0142] It should be understood that the various forms of the process shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0143] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for fusing multi-modal features, comprising: Obtaining an image including text; Performing feature recognition on the image to obtain the text feature and the position feature of the text; Dividing the image into multiple regions according to a preset rule, and extracting the image features of at least one region among the multiple regions; Encoding the text feature to obtain a text vector; And encoding the image features of the at least one region to obtain the image vectors of the at least one region; and encoding the position feature to obtain a position vector; Fusing the text vector, the image vectors of the at least one region, and the position vector to obtain a fused target vector; Wherein, the encoding the image features of the at least one region to obtain the image vectors of the at least one region includes: Performing pooling processing on the image features of the at least one region to obtain the initial image vectors of the at least one region; Performing linear transformation on the initial image vectors of the at least one region respectively; Determining the one-dimensional position encoding of the initial image vectors corresponding to the at least one region according to the position relationship of the at least one region; Determining the image vectors of the at least one region based on the linearly transformed initial image vectors of the at least one region, the one-dimensional position encoding, and a segment embedding different from other initial image vectors.
2. The method according to claim 1, wherein The encoding the text feature to obtain a text vector includes: Performing word segmentation on the text and serializing the word segmentation results to obtain multiple sequences; Determining the one-dimensional position encoding of each sequence according to the relative position information between the sequences in the multiple sequences; Determining the sequence vectors of the sequences based on the word embedding representing the semantic information of the sequences, the one-dimensional position encoding of the sequences, and a segment embedding different from other sequences; Generating a text vector based on the sequence vectors corresponding to the sequences in the text.
3. The method according to claim 1, wherein, The position feature is a two-dimensional position feature, and the position vector is a two-dimensional position vector; The encoding the position feature to obtain a position vector includes: Encoding the first coordinate and the second coordinate of the text box represented by the two-dimensional position feature, and the height and the width of the text box, wherein the first coordinate and the second coordinate are the coordinates at the diagonal positions of the text box respectively; Concatenating the x coordinate in the encoded first coordinate and the x coordinate in the encoded second coordinate with the encoded width of the text box to obtain a position vector in the x-axis direction; Concatenating the y coordinate in the encoded first coordinate and the y coordinate in the encoded second coordinate with the encoded height of the text box to obtain a position vector in the y-axis direction; Taking the position vector in the x-axis direction and the position vector in the y-axis direction as the two-dimensional position vector of the text box.
4. The method according to claim 1, wherein The fusing the text vector, the image vectors of the at least one region, and the position vector to obtain a fused target vector includes: Concatenating the text vector and the image vectors of the at least one region; Overlaying the position vector on the concatenated vector to obtain an input vector; Input the input vector into the Bert model for fusion to obtain the fused target vector.
5. The method according to claim 4, wherein The Bert model includes multiple encoders; The step of inputting the input vector into the Bert model for fusion to obtain the fused target vector includes: Input the input vector into the first encoder; In the first encoder, determine the similarity attention scores based on the similarity between each text vector and the image vector in the input vector; Based on the similarity attention scores, and the position vectors corresponding to each text vector and the position vector corresponding to the image vector, determine the spatial attention scores; Based on the spatial attention scores, obtain the output of the first encoder; Use the output of the first encoder as the input of the second encoder, and until after passing through all the encoders, obtain the fused target vector.
6. The method according to any one of claims 1-5, further comprising: Extract entities and entity relationships from the target vector.
7. A multi-modal feature fusion device, comprising: An acquisition unit for acquiring an image including text; An identification unit for performing feature identification on the image to obtain the text features and position features of the text; A division and extraction unit for dividing the image into multiple regions according to a preset rule and extracting the image features of at least one of the multiple regions; A vector determination unit for encoding the text features to obtain text vectors; And encoding the image features of the at least one region to obtain the image vectors of the at least one region; and encoding the position features to obtain position vectors; A fusion unit for fusing the text vectors, the image vectors of the at least one region, and the position vectors to obtain the fused target vector; Wherein, the vector determination unit is further used for: Performing pooling processing on the image features of the at least one region to obtain the initial image vectors of the at least one region; Performing linear transformation on the initial image vectors of the at least one region respectively; According to the position relationship of at least one region, determine the one-dimensional position encoding of the initial image vector corresponding to the at least one region; Based on the linearly transformed initial image vectors of at least one region, the one-dimensional position encoding, and the segment embedding different from other initial image vectors, determine the image vectors of at least one region.
8. The apparatus according to claim 7, wherein, The vector determination unit is used for: Segment the text and serialize the segmentation results to obtain multiple sequences; According to the relative position information between each sequence in the multiple sequences, determine the one-dimensional position encoding of each sequence; Based on the word embedding representing the semantic information of the sequence, the one-dimensional position encoding of the sequence, and the segment embedding different from other sequences, determine the sequence vector of the sequence; Generate text vectors based on the sequence vectors corresponding to each sequence in the text.
9. The device according to claim 7, wherein The position feature is a two-dimensional position feature, and the position vector is a two-dimensional position vector; The vector determination unit is further used for: Encode the first coordinate and the second coordinate of the text box representing the two-dimensional position feature, as well as the height and the width of the text box, where the first coordinate and the second coordinate are the coordinates at the diagonal positions of the text box respectively; Concatenate the x coordinate in the encoded first coordinate and the x coordinate in the encoded second coordinate with the encoded width of the text box to obtain a position vector in the x-axis direction; Concatenate the y coordinate in the encoded first coordinate and the y coordinate in the encoded second coordinate with the encoded height of the text box to obtain a position vector in the y-axis direction; Use the position vector in the x-axis direction and the position vector in the y-axis direction as the two-dimensional position vector of the text box.
10. The apparatus according to claim 7, wherein the fusion unit is configured to: Concatenate the text vector and the image vectors of the at least one region; Overlay the position vector on the concatenated vectors to obtain an input vector; Input the input vector into a Bert model for fusion to obtain a fused target vector.
11. The apparatus according to claim 10, wherein, The Bert model includes a plurality of encoders; the fusion unit is further configured to: Input the input vector into the first encoder; Determine similarity attention scores in the first encoder based on the similarities between the text vectors and the image vectors in the input vector; Determine spatial attention scores based on the similarity attention scores, the position vectors corresponding to the respective text vectors, and the position vectors corresponding to the image vectors; Obtain the output of the first encoder based on the spatial attention scores; Use the output of the first encoder as the input of the second encoder, and after passing through all the encoders, obtain a fused target vector.
12. The apparatus according to any one of claims 7-11, further comprising: A relation extraction unit, configured to extract entities and entity relations from the target vector.
13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
15. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image-text retrieval method and system based on attention mechanism and gating mechanism
CN112966135A
Document classification method and device, electronic equipment and storage medium
CN113742483A