Method for carrying out relation extraction on specific industry based on multi-modal large language model

By introducing knowledge fusion module and reinforcement learning reward function module into the multimodal large language model, the problem of incomplete multimodal information fusion in the marketing industry is solved, and a more accurate and stable relationship extraction effect is achieved.

CN120011577AActive Publication Date: 2025-05-16HUNAN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510089552.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The existing multimodal large language model is difficult to effectively integrate modal information such as text, images and audio in the marketing industry, resulting in an incomplete understanding of complex multimodal relationships.

Method used

The knowledge fusion module is introduced, and the information of the knowledge graph dynamically infers related entities and relationships, combined with the cross-modal attention module and the reward function module for reinforcement learning, optimize the model parameters to improve the accuracy of relationship extraction.

Benefits of technology

Through the combination of knowledge fusion and reinforcement learning, the model can understand and analyze multimodal information more deeply, significantly improving the accuracy and robustness of relationship extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011577A_ABST
    Figure CN120011577A_ABST
Patent Text Reader

Abstract

The invention discloses a method for carrying out relation extraction on a specific industry based on a multi-modal large language model, which comprises the following steps of: inputting sample information into a Transformer encoder, and carrying out multi-modal information fusion by utilizing a cross-modal attention module according to a hierarchical fusion mode to capture a semantic relation in a text; a knowledge fusion technology is introduced and utilized to construct a knowledge graph from crawled specific industry data to serve as supplementary information during modal fusion, so that when a model processes texts, images and audios, contents in the knowledge graph can be compared to better understand potential relationships among different modals; and inputting the fusion features into a reward function module based on reinforcement learning, and outputting the relationship between the two entities in the original text with relatively high accuracy based on the matching degree of a prediction result and a real entity relationship according to the setting of a reward function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for extracting relations for a specific industry based on a multimodal large language model. Background Art

[0002] With the rapid progress of artificial intelligence and natural language processing technology, large language models have demonstrated excellent performance in a wide range of application scenarios. Among them, multimodal large models are designed to fuse data of different modalities, but in practical applications, how to effectively fuse information of modalities such as text, images, and audio remains a challenge. Taking the marketing industry as an example, the marketing field involves many sub-industries, each of which has its own unique terminology and expressions. The current fusion method cannot fully explore the deep relationship between different modalities, resulting in the model's incomplete understanding of the complex multimodal relationships in marketing scenarios. For example, when analyzing advertising videos, it is impossible to accurately and effectively associate and integrate the visual content in the video with the key information in the advertising copy. Summary of the invention

[0003] The embodiment of the present application provides a method for extracting relationships for a specific industry based on a multimodal large language model, introduces a knowledge fusion module in the modal fusion process, utilizes the information of the knowledge graph, dynamically infers related entities and relationships according to task requirements, enhances the semantic understanding ability of the model, and enables the model to use prior knowledge for deeper reasoning and analysis.

[0004] The embodiment of the present application provides a method for extracting relations for a specific industry based on a multimodal large language model, including the following process:

[0005] S1: Input sample information (text content, image content or video content) into the Transformer encoder of the multimodal large language model, and the Transformer encoder extracts deep features from the text;

[0006] S2: Use the cross-modal attention module to fuse the text modality with the image modality and audio modality in a multi-level fusion manner;

[0007] S3: Introducing the knowledge fusion module to build the crawled industry-specific data into a knowledge graph. When performing modal fusion, the knowledge graph is used as supplementary information. When the model processes text, images, and audio, it compares the features of these modalities with the content in the knowledge graph.

[0008] S4: The fused features are input into the reward function module based on reinforcement learning. According to the setting of the reward function, the parameters of the model are adjusted based on the degree of match between the predicted results and the real entity relationship. By continuously adjusting the model according to the feedback of the reward function, the model can eventually output the relationship between the two entities in the original sample with a higher accuracy.

[0009] The further cross-modal attention module fuses the text modality with the image modality and the audio modality in accordance with the basic layer fusion, the middle layer feature fusion, and the high-level semantic fusion. The specific operation process of introducing the knowledge fusion module to build the knowledge graph as supplementary information includes:

[0010] First, the base layer performs feature fusion, segmenting the text, and converting each word into a vector through the word embedding model to form a text feature matrix.

[0011] Extract image frames from the advertising video and use convolutional neural network to extract image features to obtain feature maps where n i is the number of image regions, d i is the characteristic dimension of each region;

[0012] The audio in the advertising video is converted into a spectrogram, and then the convolutional neural network is used to extract the audio features to obtain the audio feature vector where n a is the number of audio clips, d a is the characteristic dimension of each fragment;

[0013] Crawl specific industry data to build a knowledge graph, use graph embedding methods to extract entity and relationship features in the knowledge graph, and represent the knowledge graph in vector form where n k is the number of knowledge entities, d k is the characteristic dimension of each entity;

[0014] The basic layer features F after fusion using the fusion function base for:

[0015] F base =Concat(T,I,A,K);

[0016] Concat represents a vector concatenation operation, which connects the feature vectors of the four modes end to end to form a longer fused feature vector. The fused base layer feature vector is:

[0017]

[0018] In the middle-level feature fusion, we focus on the interaction and fusion between modalities, calculate the correlation between text and image modalities, and audio modalities, generate attention weights, and use the softmax function to obtain the attention weights of the text-image modality, text-audio modality, and text-knowledge graph modality respectively:

[0019] For the image modality, the similarity between the text feature and each image region feature is calculated to obtain the attention weight.

[0020] For the audio modality, the similarity between the text features and the features of each audio clip is calculated to obtain the attention weight.

[0021] For the knowledge graph modality, the similarity between the text feature and each knowledge entity feature is calculated to obtain the attention weight.

[0022] Where W t is the weight matrix between text features and image features, W a is the weight matrix between text features and audio features, W k It is the weight matrix between text features and knowledge graph features;

[0023] The features of the modalities are weighted and fused to obtain the fused middle-level feature vector F mid :

[0024]

[0025] According to the task requirements, the knowledge graph fusion module obtains the entities and relationships inferred from the knowledge graph and obtains the inference result R. Concatenate the middle-level fusion features F mid Calculate weights:

[0026] γ=sigmoid(W g ×Concat(F mid , R)+b g )

[0027] Finally, the weight γ is used to dynamically balance the middle-level features and the knowledge graph reasoning results to obtain the fused high-level semantic feature vector F high :

[0028] F high =γ⊙F mid +(1-γ)⊙R

[0029] The high-level semantic feature vector output after fusion

[0030] The fused features are input into the reward function module based on reinforcement learning. Based on the immediate reward and long-term reward, the weight factor α is introduced to dynamically adjust the importance of the immediate reward and long-term reward, so that the model can focus on different goals at different stages, thereby improving learning efficiency. The specific operation process of the reward function module based on reinforcement learning includes:

[0031] First, the high-level semantic feature vector F is output high As the state representation of the reinforcement learning environment, the immediate reward is used to reflect the matching degree between the currently predicted entity relationship and the real relationship as the weighted composition of the total reward. The calculation of the immediate reward is:

[0032]

[0033] in, is the predicted entity relationship, r is the real entity relationship, is some distance measure between the two (such as edit distance or semantic distance), d max is the maximum possible value of this distance metric;

[0034] Adding long-term rewards considers the overall performance of the model on a series of actions to encourage the model to learn a more stable strategy as a weighted component of the total reward. The calculation of long-term rewards is:

[0035]

[0036] Among them, γ' is the discount factor, Q(S, c) is the state-action value function of taking action c in state S, S' is the new state after taking action c, A is the definition of the action space as all possible entity relationship labels, and S is the high-level semantic feature vector F after output high As a state representation of a reinforcement learning environment;

[0037] The weight factor α is introduced to weight the immediate reward and the long-term reward. The weight factor α is used to dynamically adjust the importance of the immediate reward and the long-term reward, so that the model can focus on different goals at different stages, thereby improving learning efficiency. The weighted sum of the immediate reward and the long-term reward is defined as the total reward

[0038] R 总奖励 =α·R 即时 +(1-α)·R 长期

[0039] According to R 总奖励 The reinforcement learning algorithm Q-learning is used to continuously update the state-action value function Q(S, c) during the learning process to maximize the cumulative reward:

[0040]

[0041] After training, for a given input state S, select Q(S, c) to maximize the action c as the predicted entity relationship label, and output the label as the relationship between the two entities.

[0042] The relationship between the two output entities is fed back to the knowledge fusion module to update and improve the entities and relationships in the knowledge graph. The updated knowledge graph contains the latest and most accurate information, reducing the impact of outdated or erroneous data, helping the model to obtain more accurate information support during modal fusion and improving the overall analysis and prediction accuracy.

[0043] In the knowledge graph update process, in addition to relying on the matching degree between the prediction results and the real entity relationships, the semantic similarity between the newly extracted entities / relationships and the corresponding items in the knowledge graph is calculated by cosine similarity. The newly extracted entities are defined as E new , the existing entity in the knowledge graph is defined as E old First, the entity E new and E old Input into the cross-modal attention module to get vector representation Then use the cosine similarity formula SSS (Semantic Similarity Score) to calculate the cosine similarity between the new relationship and the existing relationship in the knowledge graph:

[0044]

[0045] The knowledge graph fusion module obtains the entities and relationships inferred from the knowledge graph and obtains the inference result R. Mining potential entity relationships by designing specific semantic reasoning;

[0046] First, we collect professional knowledge in the marketing field, such as industry reports and market research papers, to build a semantic knowledge base, which contains the semantic associations between various marketing concepts, such as the association between "brand positioning" and "target audience characteristics", the connection between "promotion activity types" and "sales growth expectations", etc.

[0047] Each entity pair in the knowledge graph is defined as (e i ,e j ), based on the cosine similarity, a correction term based on the marketing semantic knowledge base is introduced to calculate the semantic similarity of each entity pair in the embedding space: Definition and Entity e i and e j The embedding vector of is calculated based on the constructed marketing semantic knowledge base:

[0048]

[0049] Among them, β is a weight parameter used to balance the contribution of the two similarities; K(e i ,e j ) represents the entity e calculated based on the marketing semantic knowledge base i and e j If the semantic knowledge base clearly states that “high-end brands” and “high-income target audiences” are strongly associated, then when e i For "high-end brands", j When the target audience is high income, K(e i ,e j ) will have a higher value;

[0050] At the same time, the graph neural network is introduced to effectively capture the complex relationships and features in the graph structure data, and the entities in the knowledge graph are used as the nodes of the graph, and the relationship between the entities is used as the edge of the graph; for example, "brand positioning" and "target audience characteristics" are two nodes, and the relationship between them is the edge connecting the two nodes. A feature vector is initialized for each node. This feature vector can contain some basic information about the entity, such as the type of entity, the frequency of occurrence, etc.

[0051] Using the message passing mechanism, in each iteration, each node receives information from its neighbor nodes and updates its own feature vector based on this information. i The eigenvector in round t is Its neighbor node set is N(e i ), then node e i The update formula of the feature vector in the t+1th round can be:

[0052]

[0053] Where W' is the learnable weight matrix and σ is the activation function;

[0054] After several rounds of iterations, the feature vector of each node has fully integrated the information of its neighboring nodes. At this time, the similarity between nodes is calculated to mine potential entity relationships. i and e j , calculate the cosine similarity of their final feature vectors as the strength of the relationship between them R ij :

[0055]

[0056] Among them, T ij is the total number of iterations;

[0057] According to the set threshold θ', when R ij>θ', infer entity e i and e j There is a potential relationship between them. At the same time, we can refer to the relationship types in the marketing semantic knowledge base to assign specific semantic types to the inferred relationships, such as "association", "influence", etc.

[0058] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: introducing a knowledge fusion module in the modal fusion process, utilizing the information of the knowledge graph, dynamically inferring related entities and relationships according to task requirements to enhance the semantic understanding ability of the model, and enabling the model to use prior knowledge for deeper reasoning and analysis; the fused features are dynamically adjusted and optimized through a reward function module based on reinforcement learning to predict entity results, and the model is continuously optimized through trial and error and feedback mechanisms to improve the accuracy and robustness of entity relationship prediction. The relationship between the two entities after output is fed back to the knowledge fusion module to update and improve the entities and relationships in the knowledge graph. The updated knowledge graph contains the latest and most accurate information, reduces the impact of outdated or erroneous data, and helps the model obtain more accurate information support during modal fusion, thereby improving the overall analysis and prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 Extract the model diagram for this application relationship;

[0060] Figure 2 Extract the flow chart for this application relationship;

[0061] Figure 3 Schematic diagram of multi-level feature fusion of the cross-modal attention module in this application. DETAILED DESCRIPTION

[0062] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. The preferred embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thoroughly understood.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.

[0064] Embodiment 1

[0065] See also Figure 1-3A method for extracting relations for a specific industry based on a multimodal large language model includes: inputting text, image and video content into a Transformer encoder for deep feature extraction; using a cross-modal attention module to perform multi-level fusion of text modality with image modality and audio modality, introducing a knowledge fusion module to construct a knowledge graph, and using it as supplementary information during modal fusion. The fused features are used through a reward function module based on reinforcement learning to adjust model parameters to improve the accuracy of relation extraction.

[0066] The cross-modal attention module performs multi-level feature fusion according to the basic layer feature fusion, middle layer feature fusion and high-level semantic fusion. When the basic layer feature is fused,

[0067] After the text is segmented, the word embedding model is used to form a text feature matrix T. Image frames are extracted from the video. The convolutional neural network is used to extract image features to obtain the feature map I. The audio in the video is converted into a spectrogram, and the audio features are extracted to obtain the audio feature vector A. Specific industry data is crawled to build a knowledge graph, and entity and relationship features are extracted and represented as a vector K. The fusion function is used to concatenate T, I, A, and K to obtain the base layer feature vector F. base ;

[0068] When mid-level features are fused,

[0069] Calculate the similarity between text and image, audio, and knowledge graph modalities, generate attention weights through the softmax function, perform weighted fusion on the features of each modality, and obtain the middle-level feature vector F mid ;

[0070] When high-level semantic fusion is performed,

[0071] The knowledge graph fusion module obtains the inference results R of entities and relationships, and concatenates the middle-level features F mid The gating mechanism is used to calculate the weights of R and R, and the gating weight γ is used to dynamically balance the middle-level features and the knowledge graph inference results to obtain the high-level semantic feature vector F high .

[0072] The fused features are input into the reward function module based on reinforcement learning, and F high The state representation of the reinforcement learning environment is defined as S, and the action space is defined as A for all possible entity relationship labels. The immediate reward and long-term reward are calculated, and the weight factor α is introduced for weighting. The reinforcement learning algorithm Q-learning is used to continuously update the state-action value function Q(S, c) during the learning process to maximize the cumulative reward. After the training is completed, for a given input state S, the action c that maximizes Q(S, c) is selected as the predicted entity relationship label.

[0073] The matching degree between the output results and the real entity relationships is used in combination with the cosine similarity to calculate the semantic similarity between the newly extracted entities / relationships and the corresponding items in the knowledge graph, so as to update and improve the entities and relationships in the original knowledge graph, reduce the impact of outdated or erroneous data, and help the model obtain more accurate information support during modal fusion, thereby improving the overall analysis and prediction accuracy.

[0074] The knowledge graph fusion module obtains the entities and relationships inferred from the knowledge graph and obtains the inference result R. Mining potential entity relationships by designing specific semantic reasoning;

[0075] First, we collect a wide range of professional knowledge in the field of marketing, including but not limited to industry reports, market research papers, classic marketing case studies, professional marketing books, etc., to build a marketing semantic knowledge base with clear hierarchy and rich content, which contains the semantic associations between various marketing concepts, such as the association between "brand positioning" and "target audience characteristics", the connection between "promotion activity types" and "sales growth expectations", etc.;

[0076] Each entity pair in the knowledge graph is introduced into a correction term based on the marketing semantic knowledge base using cosine similarity to calculate the semantic similarity of each entity pair in the embedding space, wherein a weight parameter is introduced to balance the contribution of the two similarities;

[0077] According to the semantic similarity S, a threshold θ is set to infer whether there is a potential relationship between the entity pairs; at the same time, referring to the relationship types in the marketing semantic knowledge base, the inferred relationship is given a specific semantic type, such as "association" and "influence".

[0078] Experimental Example 1

[0079] In the construction of knowledge graphs in the marketing industry, it is crucial to accurately identify entity relationships in advertising videos. The accuracy of entity relationship reasoning is improved through multimodal data fusion and reinforcement learning methods.

[0080] Experimental data

[0081] At least 1,000 marketing industry advertising videos are selected as experimental data. Each piece of data contains a video, the corresponding advertising copy, and annotated entity relationships. Preprocess the data:

[0082] Extract frames from the video to obtain image information of each frame for subsequent extraction of visual features;

[0083] Perform spectrum conversion on the audio, converting the time domain signal into the frequency domain signal, in preparation for audio feature extraction;

[0084] Perform word segmentation and word embedding on the advertising copy to convert the text into a vector form that can be processed by computers, providing a basis for subsequent feature extraction.

[0085] The sample content (a certain brand of mobile phone advertising video, with the text content "XX mobile phone, clearer photos, longer battery life", the image content "mobile phone camera interface, battery icon", and the audio content "background music and narration introduction") is input into the Transformer encoder of the multimodal large model for feature extraction.

[0086] The text content is converted into a text feature matrix T (with a dimension of 512) through word embedding technology, the image frame is processed by convolutional neural network to obtain an image feature map (with a dimension of 1024), the audio is converted into a spectrum map and the audio feature vector A (with a dimension of 512) is obtained through convolutional neural network, and the mobile phone industry data is crawled to build a knowledge graph to obtain a knowledge feature vector K (with a dimension of 512);

[0087] The text feature matrix T, image feature map I, audio feature vector A and knowledge feature vector K are concatenated to obtain a basic fusion feature with a dimension of 2560, providing rich data for subsequent in-depth fusion;

[0088] The softmax function is used to calculate the attention weights of text, image, audio, and knowledge graph for weighted fusion to obtain the middle-level fusion feature F. mid (Dimension is 512); concatenate middle-level features F mid The inference results R of entities and relationships are obtained by the knowledge graph fusion module. The gating mechanism is used to calculate the weights to obtain the high-level semantic feature vector F. high (The dimension is 512);

[0089] Using the high-level semantic feature vector F high As the state representation of the reinforcement learning environment, we define α as 0.8, R 即时 is 1, R 长期 is 0.5, R 总奖励 The calculation is as follows:

[0090] R 总奖励 =0.8·1+(1-0.8)·0.5=0.9

[0091] Reinforcement learning algorithm Q-learning based on R 总奖励 Update the state-action value function, and the final output relationship is "XX mobile phone-photo effect", and the actual labeled relationship is "XX mobile phone-photo effect".

[0092] The model output results were evaluated using precision, recall, and F1 values. The number of relationships correctly predicted by the model was 184, the number of negative examples incorrectly predicted as positive was 16, and the number of positive examples incorrectly predicted as negative was 20:.

[0093]

[0094] Among them, TP represents true positive examples (positive examples correctly predicted by the model), FP represents false positive examples (negative examples incorrectly predicted as positive examples by the model), and FN represents false negative examples (positive examples incorrectly predicted as negative examples by the model).

[0095] Most traditional multimodal models simply concatenate multimodal features without fully considering the interaction between different modalities. For example, when processing advertising videos, they may simply connect text, image, and audio features directly, ignoring the differences in the importance of each modality in expressing entity relationships in different scenarios. This experiment uses the softmax function to calculate attention weights for weighted fusion, which can dynamically adjust the contribution of different modal features and capture entity relationships more accurately;

[0096] Some models lack further optimization mechanisms after feature fusion and inference. This experiment introduces reinforcement learning and sets a reasonable reward mechanism to enable the model to continuously adjust its strategy based on environmental feedback, thereby improving the accuracy and stability of predictions. Models without reinforcement learning optimization may fall into a local optimal solution and cannot be dynamically adjusted according to actual conditions.

[0097] Experimental Results

[0098]

[0099]

[0100] According to the experimental results, this experimental model is superior to the existing traditional multimodal model and non-reinforcement learning optimization model in terms of accuracy, recall rate and F1 value. When using the cross-modal attention module for multi-level feature fusion, the knowledge fusion module is introduced to obtain the entities and relationships inferred from the knowledge graph as supplementary information for feature fusion, which significantly improves the accuracy of relationship extraction of the multimodal large language model, can more accurately identify entity relationships in advertising videos, and provide more reliable data support for the construction of knowledge graphs in the marketing industry.

[0101] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for extracting relations for a specific industry based on a multimodal large language model, characterized in that: Including the following process, S1: Input the sample information into the Transformer encoder of the multimodal large language model, and the Transformer encoder performs deep feature extraction on the text; S2: Use the cross-modal attention module to fuse the text modality with the image modality and audio modality in a multi-level fusion manner; S3: Introduce the knowledge fusion module to construct the crawled specific industry data into a knowledge graph, and use the knowledge graph as supplementary information when performing modal fusion; S4: The fused features are input into the reward function module based on reinforcement learning. By continuously adjusting the model according to the feedback of the reward function, the model can eventually output the relationship between the two entities in the original sample with high accuracy. The cross-modal attention module performs multi-level feature fusion according to the basic layer fusion, middle layer feature fusion, and high-level semantic fusion methods, and introduces the knowledge fusion module to construct a knowledge graph as supplementary information. The operations include: The base layer performs feature fusion and forms a text feature matrix through the word embedding model Use convolutional neural network to extract image features and obtain feature maps Then use the convolutional neural network to extract audio features and obtain the audio feature vector Crawl specific industry data to build a knowledge graph, use graph embedding methods to extract entity and relationship features in the knowledge graph, and represent the knowledge graph in vector form The basic layer features after fusion using the fusion function The middle-level feature fusion uses the softmax function to obtain the attention weights α of the text-image modality. t,i =softmax Attention weight α for text-audio modality t,a =softmax Attention weight α for text-knowledge graph modality t,k =softmax The features of the modalities are weighted and fused to obtain the fused middle-level feature vector F mid : The knowledge graph fusion module obtains the entities and relationships inferred from the knowledge graph and obtains the inference results. Concatenate the middle-level fusion features F mid Calculate weights: γ=sigmoid(W g ×Concat(F mid ,R)+b g ) Finally, the weight γ is used to dynamically balance the middle-level features and the knowledge graph reasoning results to obtain the fused high-level semantic feature vector F high =γ⊙F mid +(1-γ)⊙R。 The knowledge graph fusion module obtains the entities and relations inferred from the knowledge graph and obtains the inference result R. Each entity pair in the knowledge graph is defined as (e i ,e j ), a correction term based on the marketing semantic knowledge base is introduced based on cosine similarity to calculate the semantic similarity of each entity pair in the embedding space, and a graph neural network is introduced to effectively capture the complex relationships and features in graph structured data.

2. The method for extracting relations for a specific industry based on a multimodal large language model as claimed in claim 1, characterized in that: The output high-level semantic feature vector F high The state representation of the reinforcement learning environment is defined as S, the action space is defined as A for all possible entity relationship labels, and the immediate reward is used to reflect the matching degree between the current predicted entity relationship and the real relationship as the weighted composition of the total reward. The calculation of the immediate reward is: in, is the predicted entity relationship, r is the real entity relationship, is some distance measure between the two, d max is the maximum possible value of this distance metric.

3. The method for extracting relations for a specific industry based on a multimodal large language model as claimed in claim 2, characterized in that: Adding long-term rewards Considering the overall performance of the model on a series of actions, the calculation of long-term rewards is: Where γ' is the discount factor, Q(S,c) is the state-action value function of taking action c in state S, and S' is the new state after taking action c.

4. The method for extracting relations for a specific industry based on a multimodal large language model as claimed in claim 2, characterized in that: The immediate reward and the long-term reward are weighted, and the weight factor α is used to dynamically adjust the importance of the immediate reward and the long-term reward. The weighted sum of the immediate reward and the long-term reward is defined as the total reward: R 总奖励 =α·R 即时 +(1-a)·R 长期 According to R 总奖励 The reinforcement learning algorithm Q-learning is used to continuously update the state-action value function Q(S, c) during the learning process to maximize the cumulative reward: After training, for a given input state S, select Q(S, c) to maximize the action c as the predicted entity relationship label, and output the label as the relationship between the two entities.

5. The method for extracting relations for a specific industry based on a multimodal large language model as claimed in claim 4, characterized in that: The relationship between the two output entities is fed back to the knowledge fusion module to update and improve the entities and relationships in the knowledge graph. The semantic similarity between the newly extracted entities / relationships and the corresponding items in the knowledge graph is calculated by cosine similarity. The newly extracted entity is defined as E new , the existing entity in the knowledge graph is defined as E old First, the entity E new and E old Input into the cross-modal attention module to get vector representation Then use the cosine similarity formula SSS to calculate the cosine similarity between the new relationship and the existing relationship in the knowledge graph:

6. The method for extracting relations for a specific industry based on a multimodal large language model as claimed in claim 1, characterized in that: The knowledge graph fusion module mines potential entity relationships through semantic reasoning. The specific process is as follows: First, we collect professional knowledge in the marketing field to build a semantic knowledge base, which contains the semantic associations between various marketing concepts. Each entity pair in the knowledge graph is defined as (e i ,e j ), based on the cosine similarity, a correction term based on the marketing semantic knowledge base is introduced to calculate the semantic similarity of each entity pair in the embedding space, and the definition and Entity e i and e j The embedding vector of is calculated based on the constructed marketing semantic knowledge base: Among them, β is a weight parameter used to balance the contribution of the two similarities; K(e i ,e j ) represents the entity e calculated based on the marketing semantic knowledge base i and e j The semantic similarity between According to the semantic similarity S, a threshold θ is set. i ,e j )>θ when inferring entity e i and e j There is a certain potential relationship between them; at the same time, refer to the relationship type in the marketing semantic knowledge base to assign specific semantic types to the inferred relationship.

7. The method for extracting relations for a specific industry based on a multimodal large language model as claimed in claim 1, characterized in that: The graph neural network uses entities in the knowledge graph as nodes of the graph and the relationship between entities as edges of the graph; for example, "brand positioning" and "target audience characteristics" are two nodes, and the relationship between them is the edge connecting the two nodes. A feature vector is initialized for each node. This feature vector can contain some basic information of the entity. Using the message passing mechanism, in each iteration, each node receives information from its neighbor nodes and updates its own feature vector based on this information. i The eigenvector in round t is Its neighbor node set is N(e i ), then node e i The update formula of the feature vector in the t+1th round can be: Where W' is the learnable weight matrix and σ is the activation function; After several rounds of iterations, the feature vector of each node has fully integrated the information of its neighboring nodes. At this time, the similarity between nodes is calculated to mine potential entity relationships. i and e j , calculate the cosine similarity of their final feature vectors as the strength of the relationship between them R ij : Among them, T ij is the total number of iterations; According to the set threshold θ', when R ij >θ', infer entity e i and e j There is a potential relationship between them. At the same time, the relationship types in the marketing semantic knowledge base can be referred to to assign specific semantic types to the inferred relationships.

Citation Information

Patent Citations

  • Social platform multi-modal unified information extraction method

    CN117149916A

  • Relationship extraction method based on multi-modal large language model

    CN117972121A

  • Multi-modal document structured processing and knowledge extraction method based on large language model

    CN119227794A

  • Extracting document hierarchy using a multimodal, layer-wise link prediction neural network

    US20240161529A1

Cited By

  • Cross-border industry knowledge graph construction method and system based on prompt project

    CN120893533A

  • Data tamper-proofing method, system and device based on block chain evidence storage and storage medium

    CN121188845A

  • Multi-modal reasoning method and device based on knowledge graph enhancement and medium

    CN121457627A