Multi-modal image data entity relationship extraction method based on large model

By constructing a multimodal image data entity relationship extraction model and using a large model to convert image data into text and perform entity relationship extraction, the gap in image data entity relationship extraction is solved and the accuracy and efficiency of entity relationship extraction are improved.

CN120599448APending Publication Date: 2025-09-05NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510753950.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In the existing technology, the entity relationship extraction method of image data is little known, and it is difficult to effectively use large models to convert image-text mixed data and extract entity relationships.

Method used

Build an entity relationship extraction model for multimodal image data, including a text conversion module, a connection module and an entity relationship extraction module. Use a large model to convert multimodal image data into text data and perform entity relationship extraction. Use transformer and convolutional structures to improve data understanding and processing capabilities, and combine prompt word engineering to clarify instructions and output formats.

Benefits of technology

It improves the feature extraction effect of multimodal data, reduces the loss in the feature extraction and fusion process, and improves the accuracy and efficiency of entity relationship extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599448A_ABST
    Figure CN120599448A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image data entity relationship extraction method based on a large model, and the method comprises the following steps: obtaining multi-modal image data, and taking the multi-modal image data as a data set; training a pre-constructed large model based on the data set to obtain a corresponding multi-modal image data entity relationship extraction model; wherein the multi-modal image data entity relationship extraction model comprises a text conversion module, a connection module and an entity relationship extraction module which are connected in sequence; and inputting the multi-modal image data needing to be subjected to entity relationship extraction into the trained multi-modal image data entity relationship extraction model to obtain a corresponding entity relationship triple. According to the method, entity relationship extraction is carried out on image input based on the large model, the data understanding capability and processing capability of the large model under the multi-modal condition are improved, and the entity relationship extraction capability under the multi-modal condition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of entity relationship extraction, and more particularly to a method for extracting entity relationships from multimodal image data based on a large model. Background Art

[0002] With the rapid development of artificial intelligence (AI), various AI-specific tasks require not only continuously updated models but also increasing data requirements. With the rapid generation of massive amounts of data, obtaining effective data from this vast amount of open-domain, unstructured data while minimizing noise interference has become a major challenge. Therefore, information extraction has become a crucial task in natural language processing (NLP). Entity relationship extraction, as a core task of information extraction, is fundamental to downstream NLP tasks such as knowledge graph construction, machine reading, text summarization, question-answering systems, artificial intelligence, autonomous driving, machine translation, and semantic web annotation.

[0003] With the emergence and continuous development of large models in recent years, their powerful text understanding and processing capabilities have made them the preferred method in natural language processing. In entity relationship extraction tasks, large models can often rival existing fully supervised methods, achieving near-state-of-the-art (SOTA) results. While they may not achieve optimal results in low-sample scenarios, they can achieve optimal results through supervision and fine-tuning using Chain of Thought (CoT).

[0004] However, existing methods basically focus on entity relationship extraction from text data, but pay little attention to entity relationship extraction from image data.

[0005] Therefore, how to use large models to convert image-text mixed data into text and extract text entity relationships is a problem that technical personnel in this field urgently need to solve. Summary of the Invention

[0006] In view of this, the present invention provides a method for extracting entity relationships from multimodal image data based on a large model, constructs a multimodal image data entity relationship extraction model based on the large model, converts the multimodal data into text data and performs entity relationship extraction.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] The present invention provides a method for extracting entity relationships from multimodal image data based on a large model, comprising the following steps:

[0009] S1. Obtain multimodal image data as a dataset;

[0010] S2. Training a pre-built large model based on the data set to obtain a corresponding multimodal image data entity relationship extraction model;

[0011] The multimodal image data entity relationship extraction model includes a text conversion module, a connection module, and an entity relationship extraction module connected in sequence; the text conversion module is used to convert the multimodal image data into text; the connection module is used to adjust the instructions and format of the output text of the text conversion module and input it into the entity relationship extraction module; the entity relationship extraction module is used to extract entity relationships from the input text;

[0012] S3. Input the multimodal image data that needs to be subjected to entity relationship extraction into the trained multimodal image data entity relationship extraction model to obtain the corresponding entity relationship triples.

[0013] Furthermore, the multimodal image data includes: an image and corresponding text data.

[0014] Furthermore, the text conversion module includes an Embedding layer, a plurality of consecutive encoder layers, a feature fusion layer and a decoder layer connected in sequence;

[0015] The Embedding layer converts the input multimodal image data into multiple sub-images and sub-texts through embedding and position encoding;

[0016] The plurality of consecutive encoder layers respectively encode the input sub-image and sub-text to obtain corresponding image features and text features;

[0017] The feature fusion layer performs feature fusion on the image features and text features, and converts them into corresponding text through the decoder layer.

[0018] Furthermore, the plurality of consecutive encoder layers include a plurality of consecutive image encoder layers and a plurality of consecutive text encoder layers.

[0019] Furthermore, the image encoder layer includes a convolutional layer, an image multi-head attention layer and an image feedforward neural network;

[0020] After the convolution layer reorganizes the input sub-image, it performs three different convolution and flattening operations to obtain three different mappings respectively;

[0021] The three different mappings are used as K, Q and V of the multi-head attention layer of the image, and self-attention calculation is performed to obtain the multi-head attention features of the image;

[0022] The multi-head attention features of the image are passed through the image feedforward neural network to obtain corresponding image features as input to the next image encoder layer.

[0023] Furthermore, the text encoder layer includes a text multi-head attention layer and a text feedforward neural network;

[0024] The subtext is subjected to self-attention calculation by the text multi-head attention layer to obtain the multi-head attention features of the text;

[0025] The multi-head attention features of the text are passed through the text feedforward neural network to obtain corresponding text features as input to the next text encoder layer.

[0026] Furthermore, the feature fusion layer includes a multi-head attention fusion layer, a forward propagation layer, and a summation and normalization layer;

[0027] The multi-head attention fusion layer independently calculates attention scores for the image features and text features output by the encoder layer and concatenates the results.

[0028] The splicing result passes through the forward propagation layer to further extract and fuse features;

[0029] After the forward propagation layer, the concatenation result is added to the output of the forward propagation layer through the summation and normalization layer; and layer normalization is performed.

[0030] Furthermore, the processing of the connection module includes:

[0031] The text output by the text conversion module is optimized through prompt word engineering, and a prompt word triple containing instructions, target output results and content is output as input text for the entity relationship extraction module.

[0032] Furthermore, the processing of the entity relationship extraction module includes the following steps:

[0033] The input text of the entity relationship extraction module is divided into blocks, and each block of input text is embedded and mapped into a vector matrix;

[0034] Based on the vector matrix, calculate the corresponding K, Q and V, and then calculate the self-attention score through the softmax function;

[0035] The corresponding K and V are calculated through residual connection and feedforward neural network and used as the input of the next decoder until the last decoder completes the output;

[0036] The output of the last decoder is passed through the softmax function to calculate the probability of the output text, and the text with the highest probability is output as the corresponding entity relationship triplet.

[0037] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a method for extracting entity relationships from multimodal image data based on a large model, which has the following beneficial effects:

[0038] This paper combines the structural features of transformers and convolution to construct a text conversion module, improving the understanding and processing capabilities of multimodal data and reducing losses during feature extraction. By extracting graphic and text information through the text conversion module, the effectiveness of feature extraction is improved, losses during the extraction and fusion processes are reduced, and text that integrates multimodal information is generated, ultimately improving the effectiveness of entity relationship extraction.

[0039] By building a connection module and combining it with the prompt word project, we can achieve the characteristics of clear instructions and clear output, and solve the problem of entity relationship extraction failure caused by unclear instructions or unclear output format during the text data extraction process, so as to improve the effect of entity relationship extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0041] Figure 1 The present invention provides a multimodal image data entity relationship extraction model architecture.

[0042] Figure 2 A structural diagram of multiple consecutive encoder layers of a text conversion module provided by an embodiment of the present invention.

[0043] Figure 3 A structural diagram of a convolutional layer provided in an embodiment of the present invention.

[0044] Figure 4 This is a structural diagram of the feature fusion layer of the text conversion module provided by an embodiment of the present invention.

[0045] Figure 5 This is a structural diagram of the entity relationship extraction module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0047] The embodiment of the present invention discloses a method for extracting entity relationships from multimodal image data based on a large model, comprising the following steps:

[0048] S1. Obtain multimodal image data as a dataset;

[0049] S2. Train the pre-built large model based on the dataset to obtain the corresponding multimodal image data entity relationship extraction model;

[0050] The multimodal image data entity relationship extraction model includes a text conversion module, a connection module, and an entity relationship extraction module connected in sequence; the text conversion module is used to convert the multimodal image data into text; the connection module is used to adjust the instructions and format of the output text of the text conversion module and input it into the entity relationship extraction module; the entity relationship extraction module is used to extract entity relationships from the input text;

[0051] S3. Input the multimodal image data that needs to be subjected to entity relationship extraction into the trained multimodal image data entity relationship extraction model to obtain the corresponding entity relationship triples.

[0052] This embodiment is applied to the field of autonomous driving, using the large-scale model-based multimodal image data entity relationship extraction method provided by the present invention. First, a large number of road scene images must be collected as a training dataset. These images include not only standard visual information but also data from sensors such as LiDAR and radar, providing additional dimensions such as depth and speed. This data is crucial for accurately identifying entities such as vehicles, pedestrians, and traffic lights.

[0053] Secondly, the model is built and trained. A text conversion module is constructed to convert multimodal image data into text form. This process involves encoding features such as the position, color, shape of objects and their relative positions. A connection module is constructed to adjust the instructions and format of the text output by the text conversion module for subsequent processing. This process involves standardizing the text format to ensure that all extracted information is represented in a consistent manner to facilitate the understanding and processing of the entity relationship extraction module. An entity relationship extraction module is constructed to extract entities and their mutual relationships from the adjusted text. This embodiment uses the structure of a generative pre-trained converter to identify key entities in the image (such as vehicles, pedestrians, traffic lights) and the relationships between them (such as "approaching", "far away", "waiting", etc.).

[0054] Finally, the model is applied to entity relationship extraction. New road scene images are fed into the trained multimodal image data entity relationship extraction model. The model generates corresponding entity relationship triples (e.g., "Vehicle A is approaching Pedestrian B," "Traffic light C is red"). This information helps the autonomous driving system understand its surroundings and make appropriate decisions, such as slowing down, stopping, or taking a detour.

[0055] The method of the present invention enables autonomous vehicles to more intelligently perceive their surroundings, improving driving safety and efficiency. This multimodal data-based approach provides strong support for solving complex driving scenarios.

[0056] The embodiment of the present invention uses the pipeline concept to split the model building task into two subtasks. A large model is trained for each subtask, which is used to convert multimodal data into text data and extract entity relationships from text data. A connection strategy is designed to reduce error propagation and try to consider the correlation between subtasks to reduce information loss. Figure 1 As shown, the input image data is first preprocessed, and then model 1 is used to convert the data into text, which is then processed through a pre-designed connection module, and finally entity relationships are extracted from the text.

[0057] Each step of the present invention is described in detail below:

[0058] Step S1, obtaining multimodal image data as a data set; specifically comprising:

[0059] In order to build a high-quality multimodal image dataset, it is necessary to deploy a variety of sensor devices to capture different dimensions of road scenes. These devices include but are not limited to:

[0060] Camera: Captures visual information within the visible light range and records road environment, vehicles, pedestrians, traffic lights, road signs, etc.

[0061] LiDAR (laser radar): generates point cloud data by emitting laser pulses and receiving reflected signals. It is used to accurately measure the distance, position, and shape of objects, especially at night or in low-light conditions.

[0062] Millimeter-wave radar: Uses radio waves to detect the speed, distance, and direction of a target. It can detect dynamic objects (such as moving vehicles and pedestrians) and targets in adverse weather conditions.

[0063] Infrared sensor: Captures thermal radiation information to detect life forms. It can identify pedestrians and animals hidden in the shadows or moving at night.

[0064] IMU (Inertial Measurement Unit): records the vehicle's own acceleration, angular velocity, and attitude information. It helps understand the vehicle's motion state and its relationship with the surrounding environment.

[0065] To ensure the diversity and comprehensiveness of the dataset, it is necessary to cover a variety of typical road scenarios and driving conditions:

[0066] Urban roads: including busy intersections, one-way lanes, two-way lanes, crosswalks, etc. Capture complex traffic conditions such as dense traffic flow, pedestrians crossing the road, cyclists, etc.

[0067] Highways: including straight sections, curves, uphill and downhill sections, toll booths, etc. Capture high-speed vehicles, overtaking behavior, emergency stops, etc.

[0068] Rural roads: including narrow roads, irregular road surfaces, livestock crossings, etc. Capturing special scenes that rarely occur but may pose risks.

[0069] Extreme weather conditions: including rainy, snowy, foggy, and nighttime conditions. Simulate sensor performance in harsh environments to improve model robustness.

[0070] Static scenes: including open roads, parking lots, construction areas, etc. Used to test the system's perception capabilities in non-dynamic environments.

[0071] In this embodiment, the annotation of multimodal image data is an important step in building a high-quality dataset, which specifically includes the following:

[0072] Entity annotation: This involves selecting and classifying key objects in an image, such as vehicles, pedestrians, traffic lights, and road signs. The objects are labeled with their location (e.g., bounding box coordinates), category (e.g., "car," "truck," "pedestrian"), and attributes (e.g., color, size).

[0073] Relationship annotation: describes the relationship between entities, such as "Vehicle A is following Vehicle B" and "Pedestrian C is waiting at the traffic light." These relationships are represented as triples (subject, relationship, object), providing supervisory signals for subsequent entity relationship extraction.

[0074] Time series annotation: If the dataset contains video streams or multiple frames of continuous data, changes in the time dimension need to be annotated, such as "the vehicle moves from left to right."

[0075] Sensor data alignment: Synchronizes time and space between data from different sensors. For example, it fuses camera images with LiDAR point cloud data to ensure temporal and spatial consistency between the two.

[0076] Through the embodiments of the present invention, a high-quality, diverse, and well-annotated multimodal image dataset can be obtained, laying a solid foundation for subsequent model training and application.

[0077] In step S2, the multimodal image data entity relationship extraction model includes a text conversion module, a connection module and an entity relationship extraction module;

[0078] 1. The processing of the text conversion module includes:

[0079] The text conversion module of this embodiment is improved based on the MIEFormer model and includes an embedding layer, multiple consecutive encoder layers, a feature fusion layer, and a decoder layer;

[0080] The Embedding layer first digitizes the input data and converts it into a vector as the input of the encoder through word embedding and positional encoding. The input data is divided into n parts (Tokens). For the i-th part, the formula is expressed as:

[0081]

[0082] E i =H i W

[0083]

[0084] V i =E i +PE i

[0085] Among them, H i [j] represents the value of the jth bit of the one-hot encoding, H irepresents one-hot encoding, W is a trainable parameter matrix, E i Indicates that the one-hot encoding is converted into word embedding, PE pos,i Represents position coding, where pos represents the position coding of the posth part, i represents the i-th component of the coding, and d model Represents word embedding E i The final input vector V i Embedding E for words i and position encoding PE pos,i Add together.

[0086] In the embedding process of this embodiment, for an A*B input image, the image is first divided into n 114*114 sub-images, then convolution is used to calculate the token of each sub-image, and finally embedding is used to convert the entire image into a mapping (tokenmap) of length n.

[0087] After obtaining the tokenmap, the next step is the encoding process of the image data, refer to Figure 2 As shown, this embodiment uses 12 consecutive image encoders. For each layer of image encoder, its input is the output of the previous layer, that is, the tokenmap. First, the reshape function is used to transform the tokenmap of length n into h*w, where h = A%114 and w = B%114. Then, three convolutional layers are used to generate three different tokenmaps, which are recorded as T. K 、T Q 、T V ,

[0088] Then perform multi-head attention calculation: K = T K *W K , Q=T Q *W Q 、V=T V *W V Among them, W Q , W K , W V is the trainable parameter matrix.

[0089] Among them, the specific structure of the convolution layer refers to Figure 3 As shown in the figure, after the input token map is reorganized, three different convolution and flattening operations are performed to obtain three different maps respectively; K, Q, and V are obtained by calculation. Finally, the output of the encoder of this layer is the image feature.

[0090] In the encoding process of text data in this embodiment, features are obtained through the Bert structure. Specifically, this embodiment includes multiple consecutive text encoders. Each layer of text encoder first performs a layered operation on the input text, and then embeds each text token part, including token embedding and position embedding, and then inputs the Bert encoder to obtain the output of each layer. The output of the last layer is the feature of the input text.

[0091] After obtaining the features of the image and text respectively, the embodiment of the present invention designs a feature fusion layer based on the characteristics of the image and text features to obtain a better feature fusion effect and reduce the loss in the feature process. Figure 4 shown.

[0092] Specifically, in the encoding process of images and texts, the parameters are passed in the last three layers of image and text encoders respectively, so that the model can fuse K and V of the text encoder when calculating K and V of the image features, that is, K = K V +K T , V=V V +V T , where K and V are the fused features, K V 、V V K is the K and V value of the image encoder, K T 、V T are the K and V values ​​of the text encoder, thereby obtaining the image and text fusion features.

[0093] After obtaining the image-text fusion features, the decoder generates text. The specific process involves first generating a special start marker, then calculating attention using the output of the previous step and the image-text fusion features. The softmax function then calculates the probability distribution of the next token, selecting the token with the highest probability as the output. Ultimately, the decoder can generate text that incorporates the combined image and text information. For example, if an image of a cat sitting on a blanket is input along with the text "A cat is resting," the model will output text that reflects the image and text, i.e., "A black cat is resting on a white mat."

[0094] In the autonomous driving scenario of this embodiment, for example, if an image of a blue car waiting at a traffic light is input, and a text message "A vehicle is stopped and waiting ahead" is also input, the model will output text containing the image and text, namely, "A blue car is stopped and waiting at a red traffic light. Pedestrians are crossing the road at the crosswalk ahead."

[0095] In this embodiment, self-attention is calculated, and the specific formula is as follows:

[0096] Q=X×W Q

[0097] K=X×W K

[0098] V=X×W V

[0099]

[0100] Where X is the input vector, W Q , W K , W V is a trainable parameter matrix, softmax(·) is the activation function, d k is the vector dimension.

[0101] 2. The processing of the connection module includes:

[0102] The connection module of this embodiment mainly optimizes the prompt words to be input into the large model for entity relationship extraction through prompt engineering technology. The embodiment of the present invention adopts a small sample prompt method, that is, by providing a small amount of example text for the prompt words, the large model can be more adapted to the tasks to be completed by this invention.

[0103] The input of the connection module in this embodiment is the text data output by the text conversion module. A Python script is written to add fields to the input text data. Specifically, the following text is added to the header of the text data: "Output in JSON format and extract entity relationships for the following content."

[0104] In the autonomous driving scenario of this embodiment, for example:

[0105] Text: "The traffic light ahead is red, a car is waiting, and pedestrians are crossing the road."

[0106] Entity relationship: <car, waiting, red light>;

[0107] Entity relationship: <pedestrian, crossing, road>.

[0108] Finally, the output of this module includes clear instructions, target output results and content prompts; this module can significantly improve the extraction success rate of large models.

[0109] 3. The processing of the entity relationship extraction module includes:

[0110] The entity relationship extraction module of this embodiment is a large model built using the transformer architecture, which can support a variety of natural language processing tasks. This embodiment adopts the generative pre-trained transformer model 3.0 (GPT3.0). Compared with the standard transformer model, it only uses the decoder part of the standard transformer model. It is a model with only a decoder, which has great advantages in few samples and zero samples, and has greater generalization performance.

[0111] This embodiment refers to Figure 5 As shown in the figure, for the input text, each token is first embedded, and then input into the decoding layer of the GPT model to calculate K, Q, and V of each layer. The input of each decoder in the decoding layer is the K, Q, and V calculated by the decoder of the previous layer. When the last layer of decoder completes the output KQV, the softmax function is used to calculate the probability distribution of the output token, and the text with the largest probability is selected as the output text decoded by the token. The output text is the entity relationship triplet of the input text.

[0112] The specific process of this embodiment is as follows:

[0113] Assume that the input text is x, which contains m tokens. Each token is mapped into a vector through embedding, denoted as x1…x m , so the input can be viewed as a matrix x = [x1…x m ], then calculate K'=x*W K 、Q'=x*W Q 、V'=x*W V , and then calculate the self-attention value through the softmax function, After obtaining the attention, further features are obtained through residual connections and feedforward neural networks to obtain the output matrix. At the end of the decoding layer, a softmax function is included to calculate the probability of the output text corresponding to each row of the output matrix, and the text with the highest probability is selected as the output, finally obtaining the entity relationship triplet of the input text.

[0114] In this embodiment, according to step S2, the pre-built large model is trained based on the data set to obtain the corresponding multimodal image data entity relationship extraction model; the overall form is expressed by the formula:

[0115] y=GPT(C(MIEFormer(IT)))

[0116] Among them, IT represents the input image-text mixed data, and y represents the output entity relationship triplet.

[0117] In this embodiment, image and text data are first input. In the text conversion module, preliminary features are first obtained through embedding and position encoding. Then, the encoder in the converter layer is used for encoding. The information of the image and text is integrated during the encoding process. Finally, the decoder is used to generate text content containing text and image data. Then, prompt words are added to the text through the connection module, thereby improving the effect of entity relationship extraction in the generative pre-trained converter model 3.0. The text with the added prompt words is input into the generative pre-trained converter model 3.0, and preliminary features are generated through text and position embedding. Then, the decoder module in the converter architecture is used for encoding. All entities in the text are obtained through an entity classifier. Then, entity pairs are formed by entities, and each entity pair is classified by the relationship classifier to finally obtain the entity relationship in the text. This embodiment is a multimodal entity relationship extraction method based on a large model.

[0118] The present invention proposes a process method for entity relationship extraction of multimodal data based on a large model. By converting multimodal data into text data that is easy to process, the large model can extract information from more data. In view of the characteristics of graphic and text data, a text conversion model is constructed to extract graphic and text information, improve the effect of feature extraction, reduce the loss in the extraction process and the fusion process, generate text that integrates multimodal information, and ultimately improve the effect of entity relationship extraction. Using a pipeline method, a transition strategy for connection between modules is designed, and combined with a multimodal method, the image-text mixed data is converted into text data to assist in the entity relationship extraction of image information, thereby improving the extraction efficiency and capability in multimodal situations. If there is no design of a connection module, the failure of entity relationship extraction will occur during the text data extraction process due to unclear instructions and unclear output formats. The designed connection module can effectively avoid this situation.

[0119] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0120] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for extracting entity relationships from multimodal image data based on a large model, characterized in that: The following steps are involved: S1. Obtain multimodal image data as a dataset; S2. Training a pre-built large model based on the data set to obtain a corresponding multimodal image data entity relationship extraction model; The multimodal image data entity relationship extraction model includes a text conversion module, a connection module, and an entity relationship extraction module connected in sequence; the text conversion module is used to convert the multimodal image data into text; the connection module is used to adjust the instructions and format of the output text of the text conversion module and input it into the entity relationship extraction module; the entity relationship extraction module is used to extract entity relationships from the input text; S3. Input the multimodal image data that needs to be subjected to entity relationship extraction into the trained multimodal image data entity relationship extraction model to obtain the corresponding entity relationship triples.

2. The method for extracting entity relationships from multimodal image data based on a large model according to claim 1, wherein: The multimodal image data includes: an image and corresponding text data.

3. The method for extracting entity relationships from multimodal image data based on a large model according to claim 1, wherein: The text conversion module includes an embedding layer, a plurality of consecutive encoder layers, a feature fusion layer and a decoder layer connected in sequence; The Embedding layer converts the input multimodal image data into multiple sub-images and sub-texts through embedding and position encoding; The plurality of consecutive encoder layers respectively encode the input sub-image and sub-text to obtain corresponding image features and text features; The feature fusion layer performs feature fusion on the image features and text features, and converts them into corresponding text through the decoder layer.

4. The method for extracting entity relationships from multimodal image data based on a large model according to claim 3, wherein: The plurality of consecutive encoder layers include a plurality of consecutive image encoder layers and a plurality of consecutive text encoder layers.

5. The method for extracting entity relationships from multimodal image data based on a large model according to claim 4, wherein: The image encoder layer includes a convolutional layer, an image multi-head attention layer and an image feedforward neural network; After the convolution layer reorganizes the input sub-image, it performs three different convolution and flattening operations to obtain three different mappings respectively; The three different mappings are used as K, Q and V of the multi-head attention layer of the image, and self-attention calculation is performed to obtain the multi-head attention features of the image; The multi-head attention features of the image are passed through the image feedforward neural network to obtain corresponding image features as input to the next image encoder layer.

6. The method for extracting entity relationships from multimodal image data based on a large model according to claim 4, wherein: The text encoder layer includes a text multi-head attention layer and a text feedforward neural network; The subtext is subjected to self-attention calculation by the text multi-head attention layer to obtain the multi-head attention features of the text; The multi-head attention features of the text are passed through the text feedforward neural network to obtain corresponding text features as input to the next text encoder layer.

7. The method for extracting entity relationships from multimodal image data based on a large model according to claim 4, wherein: The feature fusion layer includes a multi-head attention fusion layer, a forward propagation layer, and a summation and normalization layer; The multi-head attention fusion layer independently calculates attention scores for the image features and text features output by the encoder layer and concatenates the results. The splicing result passes through the forward propagation layer to further extract and fuse features; After the forward propagation layer, the concatenation result is added to the output of the forward propagation layer through the summation and normalization layer; and layer normalization is performed.

8. The method for extracting entity relationships from multimodal image data based on a large model according to claim 1, wherein: The processing process of the connection module includes: The text output by the text conversion module is optimized through prompt word engineering, and a prompt word triple containing instructions, target output results and content is output as input text for the entity relationship extraction module.

9. The method for extracting entity relationships from multimodal image data based on a large model according to claim 1, wherein: The processing process of the entity relationship extraction module includes the following steps: The input text of the entity relationship extraction module is divided into blocks, and each block of input text is embedded and mapped into a vector matrix; Based on the vector matrix, calculate the corresponding K, Q and V, and then calculate the self-attention score through the softmax function; The corresponding K and V are calculated through residual connection and feedforward neural network and used as the input of the next decoder until the last decoder completes the output; The output of the last decoder is passed through the softmax function to calculate the probability of the output text, and the text with the highest probability is output as the corresponding entity relationship triplet.