Construction method of multi-modal information extraction instruction data set, extraction model and extraction method
By constructing a multimodal information extraction instruction data set and adopting a stacked attention network fusion strategy, the problem of insufficient training data of multimodal large model is solved, and the accuracy and efficiency of multimodal information extraction is improved.
Patent Information
- Application Number
- CN202510357903.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing multimodal large models have problems in information extraction with insufficient training data, large resource usage, high time cost and poor extraction effect, and lack of multimodal information extraction instruction data sets, resulting in the model not following instructions, errors and omissions.
A multimodal information extraction instruction data set is constructed, and a stacked attention network fusion strategy is adopted. By designing modal-specific prompt word templates, initial screening and verification steps, combining a stacked attention network to enhance modal fusion, and improving the multimodal large model architecture.
It improves the accuracy and completeness of multimodal information extraction, reduces the illusion problem of the model not following instructions, and improves the efficiency and effectiveness of multimodal information extraction.
Smart Images

Figure CN120296384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a method for constructing a multi-modal information extraction instruction data set, an extraction model, and an extraction method. Background Art
[0002] Multi-modal information extraction technology refers to extracting information such as entities, relationships, attributes, and events from modalities such as text, images, and videos generated in the industrial field. Multi-modal information extraction generally uses different information extraction technologies according to different modal data, which easily leads to problems such as long training time consumption, high development cost, large resource occupation, long research and development cycle, and poor extraction effect. With the development of large language models based on the Decoder-Only Transformer architecture such as GPT4, research and applications focusing on multi-modal large models have also made great progress, and technologies for information extraction using multi-modal large models have also received extensive attention, research, and applications.
[0003] Traditional information extraction technologies for entities, relationships, attributes, events, etc. in multi-modal data need to develop a corresponding extraction algorithm according to different modalities and extraction tasks. Dozens of different models are required to complete the multi-modal information extraction task for the combination of different modalities and different tasks, resulting in high resource and time costs. At the same time, the pipeline information extraction method will transmit and accumulate errors, leading to poor extraction effects. In the structural framework of multi-modal large models such as LLaVA, at the stage of fusing image features and text features, feature vectors of different modalities are only integrated through simple splicing, and the fused features do not fully utilize text and image information. At the same time, due to the lack of an instruction fine-tuning data set, the information extraction effect based on multi-modal large models is poor.
[0004] In summary, there are two main problems in extracting multi-modal information based on multi-modal large models. One is the lack of a multi-modal information extraction instruction data set. Due to the particularity of the field, it is difficult to collect training data, and the lack of information extraction instruction data sets for text, images, and videos results in poor multi-modal information extraction effects. The other is that due to the Transformer architecture belonging to the sequential prediction of single words and errors in pre-training data, there are problems such as non-standardization, non-compliance with instructions, errors, and omissions in the extraction results of multi-modal data. Summary of the Invention
[0005] The object of the present invention is to provide a construction method, an extraction model and an extraction method for a multi-modal information extraction instruction data set. Based on the extraction technology of a multi-modal large model with a stacked attention network fusion strategy, firstly, a multi-modal instruction data set of real-time and historical texts, images, videos, etc. is constructed for fine-tuning the multi-modal large model information extraction instructions. Secondly, a stacked attention network is adopted to enhance the modal fusion strategy to reduce hallucination problems such as the multi-modal large model not following instructions, fabricating out of thin air, errors and omissions.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a construction method for a multi-modal information extraction instruction data set, including the following steps:
[0008] S1. Design different prompt templates according to the extraction tasks for different modal data, and each prompt template can meet the extraction functions of entity, relationship, attribute and event information of the corresponding modal data;
[0009] S2. Send an information extraction request and call an open-source multi-modal large model to obtain entity, relationship, attribute and event information of the multi-modal data, that is, insert the corresponding modal data at the preset position of the prompt template;
[0010] S3. Run the execution program of the prompt template and output the initial Json result;
[0011] S4. Screen out the Json results that meet the criteria from the initial Json results based on a preset preliminary screening criterion to obtain the preliminary screening Json results;
[0012] As a possible implementation, S4 includes:
[0013] S40. Judge whether the initial Json result meets the standard Json format. If so, execute S41, otherwise execute S3;
[0014] As a possible implementation, S40 includes the following steps:
[0015] S400. Judge whether the string generated by the open-source multi-modal large model is a number. If so, return False, otherwise, execute S401;
[0016] S401. Use the built-in Json library in Python to convert the string into Json. If the conversion is successful, return True, otherwise return False.
[0017] S41. Determine whether all the Json results that meet the standard format output by S40 conform to the Json format defined in the prompt template. At the same time, for the entity, relationship, and attribute extraction results, determine whether the keys of the Json exist "entities", "relations", and "properties", and for the event extraction results, determine whether the keys of the Json exist "events". In addition, determine whether the values corresponding to the keys conform to the specified data types.
[0018] S42. Save all the conforming Json results output by S41, which are defined as the initially screened Json results.
[0019] S5. Based on the preset verification criteria, screen out the Json results that meet the criteria from the initially screened Json results to obtain the verified Json results. The multi-modal information extraction instruction dataset is composed of multiple verified Json results corresponding to the multi-modal data.
[0020] In a second aspect, the present invention provides a multi-modal information extraction model. The multi-modal information extraction model applies the multi-modal information extraction instruction dataset and is obtained by fine-tuning training based on supervised instructions.
[0021] As a possible implementation, the multi-modal information extraction model includes:
[0022] A visual encoder, which is used to receive the original image or video key frame sequence and output visual block features. The visual encoder can embed each visual block and extract image features to obtain feature vectors.
[0023] A text encoder, which embeds the prompt template into a vector space with the same dimension as the feature vector, obtains a query vector with the same dimension as the feature vector, and inputs it to the stacked attention layer.
[0024] A visual language connector, which embeds the visual block features input by the visual encoder into the vector space of the large language model for aligning the visual and text modalities.
[0025] A stacked attention layer, which receives the visual block features aligned with the text modality input by the visual language connector and the query vector input by the text encoder, and calculates the attention scores between the query vector and the visual blocks to search for regions strongly related to the prompt.
[0026] And a large language model, which is the basic layer of the multi-modal information extraction model. It uses the tokenizer and embedding layer to tokenize the text and map it into the input vector space of the large language model.
[0027] As a possible implementation, the attention scores are calculated in the following way:
[0028] Map the visual block features to a k-dimensional vector space through a fully connected layer. In the k-dimensional vector space, add the query vector and the visual block features to calculate the hidden layer feature vector;
[0029] Map the hidden layer feature vector to a distributed space through a fully connected layer and calculate the attention probability of the visual block;
[0030] Use multiple attention layers, calculate and update the attention scores of each attention layer. The last attention layer only needs to assign weights to the feature vectors of the visual blocks.
[0031] As a possible implementation, calculate the hidden layer feature vector through the following method:
[0032]
[0033] where h A is the hidden layer feature vector, v I ∈R d×m is the visual block feature vector, d is the dimension of the visual feature, m is the number of visual blocks, v Q ∈R d is the query vector; W I,A , W I,A ∈R k×d is the mapping matrix, b A ∈R k is the bias, p I ∈R m is an m-dimensional vector representing the attention probability of each visual block v Q .
[0034] As a possible implementation, calculate the attention probability p of the visual block through the following method I :
[0035] p I = softmax(W P h A + b P )
[0036] where W P ∈R 1×k is the mapping matrix, b P ∈R 1×m is the bias.
[0037] As a possible implementation, the calculation formula for the k-th attention layer is as follows:
[0038]
[0039] where u 0 is initialized to v Q, where k is the number of attention layers, and u k (k > 0) is initialized as shown in the following formula:
[0040]
[0041] Update u using the above iterative formula at each layer k .
[0042] Thirdly, the present invention provides an information extraction method applying the multimodal information extraction model provided in the second aspect, including the following steps:
[0043] Send an information extraction request, and call the multimodal information extraction model provided in the second aspect to obtain entity, relationship, attribute, and event information of multimodal data, that is, insert corresponding modal data at the preset positions of the prompt template;
[0044] Run the execution program of the prompt template and output the initial Json result;
[0045] Based on the preset preliminary screening criteria, screen out the Json results that meet the criteria from the initial Json results to obtain the preliminarily screened Json results;
[0046] Based on the preset verification criteria, screen out the Json results that meet the criteria from the preliminarily screened Json results to obtain the verified Json results as the information extraction results.
[0047] Compared with the prior art, the present invention has the following effects:
[0048] 1. The construction method of the multimodal information extraction instruction dataset proposed by the present invention constructs a relatively accurate and complete multimodal information extraction instruction dataset through steps such as the design of information extraction prompt templates for multimodal data entities, relationships, attributes, events, etc., the generation of results by open-source multimodal large models, the automatic preliminary screening of results, and verification, providing data support for training dedicated large models for multimodal information extraction.
[0049] 2. The multimodal information extraction model proposed by the present invention, including the construction of the multimodal information extraction instruction dataset and the improvement of the multimodal large model architecture based on the stacked attention network fusion strategy, can relieve problems such as insufficient multimodal information extraction instruction datasets and hallucinations of multimodal large models.
[0050] 3. The multimodal information extraction model architecture proposed by the present invention replaces simple feature vector concatenation operations with stacked attention modules, enhances the fusion of text and image information, and improves the accuracy of multimodal information extraction. Description of the Drawings
[0051] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0052] Figure 1 It is a flowchart of a method for constructing a multi-modal information extraction instruction data set proposed in an embodiment of the present invention;
[0053] Figure 2 It is a schematic diagram of a multi-modal information extraction model adopting a stacked attention network strategy in an embodiment of the present invention;
[0054] Figure 3 It is a schematic diagram of the structure of a multi-modal information extraction model in an embodiment of the present invention. Detailed implementation manners
[0055] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0056] It should be noted that when an element is referred to as being "fixedly disposed on" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.
[0057] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0058] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "upper" and "lower" is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention.
[0059] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the term "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0060] For the information extraction technologies of entities, relationships, attributes, events, etc. of traditional multi-modal data, a corresponding extraction algorithm needs to be developed according to different modalities and extraction tasks. Dozens of different models are required to complete the multi-modal information extraction task for the combination of different modalities and different tasks, resulting in high resource and time costs. At the same time, the pipeline information extraction method will transmit and accumulate errors, leading to poor extraction effects. In the structural framework of multi-modal large models such as LLaVA, at the stage of fusing image features and text features, the feature vectors of different modalities are only integrated through simple splicing, and the fused features do not fully utilize text and image information. At the same time, due to the lack of an instruction fine-tuning data set, the information extraction effect in the industrial field based on multi-modal large models is poor. In this embodiment, taking the industrial field as an example, a construction method, an extraction model, and an extraction method for a multi-modal information extraction instruction data set are proposed. For the extraction technology of multi-modal large models based on the stacked attention network fusion strategy, first, a multi-modal instruction data set of real-time and historical text, images, videos, etc. is constructed for the information extraction instruction fine-tuning of multi-modal large models. Secondly, the stacked attention network is used to enhance the modal fusion strategy to reduce hallucination problems such as multi-modal large models not following instructions, fabricating out of thin air, errors, and omissions.
[0061] In a first aspect, an embodiment of the present invention proposes a construction method for a multi-modal information extraction instruction data set. Refer to Figure 1 , including the following steps:
[0062] S1. Design different prompt templates according to the extraction tasks for different modal data, and each prompt template can meet the extraction functions of entity, relationship, attribute, and event information of the corresponding modal data;
[0063] As an example, S1 includes the following steps:
[0064] S10. Construct an instruction data set for different modal document data extraction tasks, including: the construction of entity, relationship, and attribute instruction data sets and the construction of event information extraction instruction data sets.
[0065] As an example, the templates of entity, relationship, and attribute instruction data sets in different modal document data extraction tasks are as follows:
[0066]
[0067]
[0068]
[0069] As an example, the template of the instruction dataset for event information extraction in different modality document data extraction tasks is as follows:
[0070]
[0071]
[0072] S11. Construct an instruction dataset for different modality image and video data extraction tasks, including: construction of entity, relationship, and attribute instruction datasets and construction of event information extraction instructions.
[0073] As an example, the templates of entity, relationship, and attribute instruction datasets in different modality image and video data extraction tasks are as follows:
[0074]
[0075]
[0076]
[0077] As an example, the template of the instruction dataset for event information extraction in different modality image and video data extraction tasks is as follows:
[0078]
[0079]
[0080] S2. Send an information extraction request and call an open-source multimodal large model to obtain entity, relationship, attribute, and event information of multimodal data, that is, insert the corresponding modality data at the preset position in the prompt template;
[0081] As an example, using the instruction dataset constructed in S1, call the local open-source multimodal large model through the API to obtain entity, relationship, attribute, and event information of multimodal data. Exemplarily, the document data extraction request method is as follows:
[0082]
[0083]
[0084] The image and video data extraction request method is as follows:
[0085]
[0086] S3. Run the execution program of the prompt template and output the initial Json result;
[0087] Exemplarily, for each extraction request, first fill the documents or key frame sequences of images and videos to be extracted in the sample into the preset <begin_of_doc><end_of_doc> or <begin_of_images><end_of_images> positions. After calling the interface, obtain the result returned by the multi-modal large model through completion.choices[0].message.
[0088] S4. Screen out the Json results that meet the criteria from the initial Json result based on the preset initial screening criteria to obtain the initially screened Json result;
[0089] As a possible implementation, S4 includes:
[0090] S40. Determine whether the initial Json result meets the standard Json format. If so, execute S41; otherwise, execute S3;
[0091] As a possible implementation, S40 includes the following steps:
[0092] S400. Determine whether the string generated by the open-source multi-modal large model is a number. If so, return False; otherwise, execute S401;
[0093] S401. Use the built-in Json library in Python to convert the string into Json. If the conversion is successful, return True; otherwise, return False.
[0094] Specifically, the following code is used to determine whether the initial Json result meets the standard Json format:
[0095]
[0096] S41. Determine whether all the Json results that meet the standard format output by S40 conform to the Json format defined in the prompt template. It should be noted that the Json in the standard format refers to JavaScript Object Notation defined in the computer field, that is, JavaScript object notation. The Json format defined in the prompt template refers to conforming to JavaScript object notation and the keys and values defined in the prompt template we designed. At the same time, for the entity, relationship, and attribute extraction results, judge whether the keys of the Json exist "entities", "relations", and "properties", and for the event extraction results, judge whether the keys of the Json exist "events". In addition, judge the value corresponding to the key, that is, whether the data following the key such as "events" conforms to the specified data type, that is, the type of value defined in the prompt template. For example, if the data type of the value corresponding to the "entities" key is "list", then judge whether the value of "entities" in the generated Json result is a list.
[0097] S42. Save all the conforming Json results output by S41, which are defined as the preliminary screening Json results.
[0098] S5. Based on the preset verification criteria, screen out the Json results that meet the standards from the preliminary screening Json results to obtain the verified Json results. The multi-modal information extraction instruction dataset is composed of multiple verified Json results corresponding to the multi-modal data. Exemplarily, the preset verification criteria are to judge whether the content is correct to ensure the accuracy of the finally obtained extraction instruction dataset, mainly examining the accuracy and completeness of the generated content, that is, determining whether the generated content is correct and whether there are omissions or redundancies in the extraction results.
[0099] In a second aspect, an embodiment of the present invention provides a multi-modal information extraction model. The multi-modal information extraction model applies the multi-modal information extraction instruction dataset and is obtained by supervised instruction fine-tuning training. Existing multi-modal large models generally use simple splicing means to fuse image and text features, and do not make full use of the graphic and text information during the generation process, resulting in hallucinations in the model generation results, that is, situations such as errors and omissions. To avoid this situation, this embodiment proposes a new fusion strategy, that is, using a stacked attention network to achieve the fusion of graphic and text features.
[0100] See Figures 2 to 3 , as a possible implementation, the multi-modal information extraction model includes:
[0101] A visual encoder, which is used to receive a sequence of original images or video key frames and output visual block features. The visual encoder can embed each visual block and extract image features to obtain a feature vector. Exemplarily, CLIP-ViT-L / 336px is selected as the visual encoder. Its image processor scales images of different sizes to [336, 336] pixels, and the last hidden feature of CLIP is selected as the feature matrix of the image or video key frame sequence.
[0102] A text encoder, which embeds a prompt template into a vector space of the same dimension as the feature vector to obtain a query vector of the same dimension as the feature vector. That is, Query Vector, which is a special term in the field of deep learning. To prevent ambiguity, "Query Vector" can be used as a note and input into a stacked attention layer. Exemplarily, LSTM is selected as the text encoder. After the information extraction prompt words are tokenized, feature vectors are obtained through look-up, and then a fixed-dimensional feature vector is obtained through the LSTM network for input into the stacked attention layer.
[0103] A vision-language connector, which embeds the visual block features input by the visual encoder into the vector space of a large speech model for aligning the visual and text modalities. Exemplarily, a two-layer perceptron is used as the vision-language connector. After the visual features are mapped to the vector space of the large language model, the dimension becomes [576, 4096] dimensions.
[0104] A stacked attention layer, which receives the visual block features aligned with the text modality input by the vision-language connector and the query vector input by the text encoder, and calculates the attention scores between the query vector and the visual blocks to search for regions strongly related to the prompt words. As a possible implementation, the attention scores are calculated in the following way:
[0105] The visual block features are mapped to a k-dimensional vector space through a fully connected layer. In the k-dimensional vector space, the query vector and the visual block features are added together to calculate and obtain a hidden layer feature vector. As a possible implementation, the hidden layer feature vector is calculated by the following method:
[0106]
[0107] where, h A is the hidden layer feature vector, v I ∈R d×m is the visual block feature vector, d is the dimension of the visual feature, m is the number of visual blocks, v Q ∈R d is the query vector; W I,A , W I,A ∈R k×d is the mapping matrix, b A ∈Rk is the offset, p I ∈R m is an m-dimensional vector representing the attention probability of each visual block v Q .
[0108] The hidden layer feature vector is mapped to a distributed space through a fully connected layer, and the attention probability of the visual block is calculated;
[0109] As a possible implementation, the attention probability p of the visual block is calculated in the following way I :
[0110] p I = softmax(W P h A + b P )
[0111] where W P ∈R 1×k is the mapping matrix, b P ∈R 1×m is the offset.
[0112] Compared with the simple concatenation of text and image feature vectors, the stacked attention module can better represent and allocate visual blocks related to text instructions. For complex instructions, multiple attention layers are used, and the attention scores of each attention layer are calculated and updated. The last attention layer only needs to assign weights to the feature vectors of the visual blocks.
[0113] As a possible implementation, the calculation formula for the k-th layer attention layer is as follows:[[]]
[0114]
[0115] where u 0 is initialized as v Q in the first layer, k is the number of attention layers, and the initialization method of u k (k > 0) is as shown in the following formula:[[]]
[0116]
[0117] In each layer, u k is updated using the above iterative formula, and the last layer only needs to assign weights to the feature vectors of the image blocks.
[0118] In a third aspect, an information extraction method applying the multi-modal information extraction model provided in the second aspect is provided in an embodiment of the present invention, including the following steps:
[0119] Send an information extraction request, and call the multi-modal information extraction model provided by the second party to obtain entity, relationship, attribute, and event information of the multi-modal data, that is, insert the corresponding modal data at the preset position of the prompt template;
[0120] Run the execution program of the prompt template and output the initial Json result;
[0121] Based on the preset initial screening criteria, screen out the Json results that meet the criteria from the initial Json results to obtain the initially screened Json results;
[0122] Based on the preset verification criteria, screen out the Json results that meet the criteria from the initially screened Json results to obtain the verified Json results as the information extraction results.
[0123] And the large language model, which is the basic layer of the multi-modal information extraction model. It uses the tokenizer and the embedding layer to tokenize the text and map it to the input vector space of the large language model. At each step, it predicts the probability of the (N + 1)-th word based on the previous N words until the end-of-sequence token is generated.
[0124] In the description of the above embodiments, the specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0125] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A construction method for a multi-modal information extraction instruction dataset, characterized in that, It includes the following steps: S1. Design different prompt templates according to the extraction tasks for different modal data, and each of the prompt templates can meet the extraction functions of entity, relationship, attribute, and event information of the corresponding modal data; S2. Send an information extraction request and call an open-source multi-modal large model to obtain entity, relationship, attribute, and event information of the multi-modal data, that is, insert the corresponding modal data at the preset position of the prompt template; S3. Run the execution program of the prompt template and output the initial Json result; S4. Screen out the Json results that meet the criteria from the initial Json result based on the preset initial screening criteria to obtain the initially screened Json result; S5. Screen out the Json results that meet the criteria from the initially screened Json result based on the preset verification criteria to obtain the verified Json result, and the multiple verified Json results corresponding to the multi-modal data constitute the multi-modal information extraction instruction dataset.
2. The construction method of the multi-modal information extraction instruction data set according to claim 1, characterized in that The S4 includes: S40. Judge whether the initial Json result meets the standard Json format. If so, execute S41; otherwise, execute S3; S41. Judge whether all the Json results output by S40 that meet the standard format conform to the Json format defined in the prompt template; at the same time, for the entity, relationship, and attribute extraction results, judge whether the keys of the Json exist "entities", "relations", and "properties", and for the event extraction results, judge whether the keys of the Json exist "events"; in addition, judge whether the values corresponding to the keys conform to the specified data types; S42. Save all the conforming Json results output by S41 and define them as the initially screened Json result.
3. The construction method of the multi-modal information extraction instruction data set according to claim 2, wherein, The S40 includes the following steps: S400. Judge whether the string generated by the open-source multi-modal large model is a number. If so, return False; otherwise, execute S401; S401. Use the built-in Json library in Python to convert the string into Json. If the conversion is successful, return True; otherwise, return False.
4. A multi-modal information extraction model, characterized in that, The multi-modal information extraction model applies the multi-modal information extraction instruction dataset and is obtained based on supervised instruction fine-tuning training.
5. The multimodal information extraction model according to claim 4, wherein The multi-modal information extraction model includes: A visual encoder, which is used to receive the original image or video key frame sequence and output visual block features. The visual encoder can embed each visual block and extract image features to obtain feature vectors; A text encoder, which embeds the prompt template into a vector space with the same dimension as the feature vector, obtains a query vector with the same dimension as the feature vector, and inputs it into the stacked attention layer; A visual language connector, which embeds the visual block features input by the visual encoder into the vector space of the large speech model for aligning the visual and text modalities; A stacked attention layer, which receives the visual block features aligned with the text modality input by the visual language connector and the query vector input by the text encoder, and calculates the attention scores of the query vector and the visual block to search for regions strongly related to the prompt; And the large language model, which is the basic layer of the multi-modal information extraction model, uses the tokenizer and the embedding layer to tokenize the text and map it to the input vector space of the large language model.
6. The multimodal information extraction model according to claim 5, wherein Calculate the attention scores in the following way: Map the visual block features to a k-dimensional vector space through a fully connected layer. In the k-dimensional vector space, add the query vector and the visual block features to calculate the hidden layer feature vector; Map the hidden layer feature vector to a distributed space through a fully connected layer and calculate the attention probability of the visual blocks; Use multiple attention layers, calculate and update the attention scores of each attention layer. The last attention layer only needs to assign weights to the feature vectors of the visual blocks.
7. The multi-modal information extraction model according to claim 6, wherein Calculate the hidden layer feature vector through the following method: Among them, h A is the hidden layer feature vector, v I ∈R d×m is the visual block feature vector, d is the dimension of the visual feature, m is the number of visual blocks, v Q ∈R d is the query vector; W I,A , W I,A ∈R k×d are the mapping matrices, b A ∈R k is the bias, p I ∈R m is an m-dimensional vector representing the attention probability of each visual block v Q .
8. The multimodal information extraction model according to claim 7, wherein Calculate the attention probability p of the visual block in the following way I : p I = softmax(W P h A + b P ) where, W P ∈ R 1×k is the mapping matrix, and b P ∈ R 1×m is the bias term.
9. The multimodal information extraction model according to claim 8, wherein For the k-th attention layer, the calculation formula is as follows: where u 0 is initialized to v in the first layer Q , k is the number of attention layers, and u k (k > 0) is initialized as shown in the following formula: Update u using the above iterative formula at each layer k .
10. An information extraction method using the multimodal information extraction model according to any one of claims 4 to 9, characterized in that, Include the following steps: Send an information extraction request, call the multi-modal information extraction model according to any one of claims 4 to 9 to obtain the entity, relationship, attribute, and event information of the multi-modal data, that is, insert the corresponding modal data at the preset position of the prompt template; Run the execution program of the prompt template and output the initial Json result; Based on the preset initial screening criteria, screen out the Json results that meet the criteria from the initial Json results to obtain the initially screened Json results; Based on the preset verification criteria, screen out the Json results that meet the criteria from the initially screened Json results to obtain the verified Json results as the information extraction results.
Citation Information
Patent Citations
Multi-modal evaluation object extraction method based on regional perception alignment network
CN114693949A
Link prediction method based on entity and relation representation fused with multi-modal information
CN116680343A
Entity alignment method and system based on image generation algorithm and multi-modal large model
CN117725230A
Scientific and technical literature flow chart entity and relation extraction method based on retrieval enhancement
CN119003788A
Multi-source and multi-mode fused knowledge reasoning method, system and device and medium
CN119005340A
Cited By
Method for processing multi-modal data and electronic device
CN122596048A