A Multimodal Entity Recognition Method Based on Large Language Models
Through the large language model and adaptive modal interaction framework to process multimodal data, the problems of insufficient fusion of cross-modal data and lagging knowledge graph update in the existing technology are solved, and the entity recognition and knowledge graph update are achieved with high accuracy and high efficiency.
Patent Information
- Application Number
- CN202411393490.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-10-08
AI Technical Summary
The existing multimodal entity recognition technology has insufficient accuracy in cross-modal data fusion and knowledge graph update, and cannot fully utilize the complementarity of multimodal data, and the traditional methods are inefficient and cannot meet the needs of real-time dynamic updates.
The large language model is used to combine the adaptive modal interaction framework, and the text, image and audio data are preprocessed and feature fusion through the modal perception mechanism to generate a comprehensive semantic representation, and the knowledge graph is updated through a dynamic mapping algorithm to automatically generate new entity nodes and relationships.
It significantly improves the accuracy and robustness of entity recognition in complex scenarios, can update the knowledge graph in a timely manner, and improves the dynamic adaptability and efficiency of the system.
Smart Images

Figure CN119167937B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal entity recognition, and in particular to a multimodal entity recognition method based on a large language model. Background Art
[0002] In the prior art, multimodal entity recognition mainly relies on the independent processing of single-modal data, and entity recognition is carried out separately in each modality through traditional machine learning or deep learning methods. The single-modal data method can achieve certain recognition effects in the single modality. However, in the face of complex scenarios and the presence of multiple modality information, the prior art has significant limitations in data fusion and the comprehensive processing ability of cross-modal information. Since data in different modalities are usually processed independently in their respective ways, this fragmented approach leads to the dispersion of information, unable to make full use of the complementarity between multimodals, and ultimately resulting in a low accuracy of entity recognition.
[0003] In addition, existing multimodal processing technologies also show deficiencies in dealing with complex contexts and ambiguous information. When processing the fusion of multimodal data such as text, images, and audio, they are unable to effectively integrate data from different modalities, resulting in inaccurate recognized entities and imprecise relationship reasoning in the case of information loss or misalignment between modalities, seriously affecting the actual application effect.
[0004] There are also certain bottlenecks in the update and maintenance of the existing knowledge graph. Traditional knowledge graph construction and update mostly rely on manual or semi-automated methods, which are not only inefficient but also unable to timely reflect the latest entity information and their relationships. In the scenario of continuously growing multimodal data, the complexity of updating the knowledge graph manually increases significantly, making it difficult to meet the requirements of real-time expansion and dynamic update.
[0005] In summary, the prior art mainly has the following defects: First, the accuracy of cross-modal entity recognition is not high, and the complementarity of multimodal data cannot be fully utilized; second, the existing multimodal data fusion methods perform poorly in dealing with complex contexts and ambiguous information, resulting in poor entity recognition and relationship reasoning effects; finally, the existing knowledge graph update method is inefficient and cannot meet the real-time dynamic update requirements in the multimodal data scenario. Summary of the Invention
[0006] An object of the present invention is to propose a multimodal entity recognition method based on a large language model, and the present invention improves the accuracy of complex entity recognition.
[0007] A multimodal entity recognition method based on a large language model according to an embodiment of the present invention includes the following steps:
[0008] S1. Receive text data, image data, and audio data, and construct a multi-modal input data set;
[0009] S2. Determine and apply preprocessing methods for different modal input data through an adaptive modal interaction framework, and preprocess the multi-modal input data set;
[0010] S3. Input the preprocessed multi-modal input data set into a large language model with a modal perception mechanism. The large language model is used to generate a comprehensive semantic representation including multi-modal context understanding, combining the semantic information of text, images, and audio;
[0011] S4. Use the adaptive modal interaction framework to adaptively fuse the comprehensive semantic representation generated by the large language model and other modal features, and generate an optimized fused feature representation by dynamically adjusting the interaction weights of different modalities;
[0012] S5. Perform cross-modal entity recognition based on the fused feature representation. Cross-modal entity recognition includes using the modal perception mechanism to identify complex entities and relationships in text and other modal information, and generating corresponding classification labels and relationship mappings for the entities;
[0013] S6. Automatically match the identified entities with the existing knowledge graph through a dynamic mapping algorithm. If the entity is a new entity, then according to the results of the adaptive modal interaction, automatically generate new nodes and update the relationship of the existing knowledge graph nodes;
[0014] S7. By continuously inputting new multi-modal data, the system automatically adjusts and optimizes the modal perception mechanism of the large language model and the adaptive modal interaction framework, continuously learns new knowledge, and dynamically expands the knowledge graph.
[0015] Optionally, the S1 step includes:
[0016] S11. Receive a text data set, an image data set, and an audio data set;
[0017] S12. For the text data set D t , each text sample x t ∈ D t is composed of a set of sentences {s1, s2, …, s n}, where s n is the nth sentence in the text sample;
[0018] S13. For the image data set D i , each image sample x i ∈ D i is represented by a set of pixel matrices M = {p 11 , p 12 , …, p mn}, where pmn represents the pixel value at the m-th row and n-th column in the image;
[0019] S14. For the audio dataset D a , each audio sample x a ∈ D a is represented by a set of frequency-domain feature vectors F = {f1, f2, …, f k}, where f k is the k-th frequency component of the audio signal;
[0020] S15. Combine the text dataset D t , the image dataset D i and the audio dataset D a into a multi-modal input dataset:
[0021] D = {x t , x i , x a}.
[0022] Optionally, the S2 step includes:
[0023] S21. Perform modal recognition on the multi-modal input dataset D = {x t , x i , x a} through an adaptive modal interaction framework to determine the type and data structure of each modality;
[0024] S22. Apply different preprocessing methods according to the characteristics of each modality, where perform word segmentation and denoising processing on the text dataset D t , perform standardization and size normalization processing on the image dataset D i , and perform denoising and frequency-domain conversion processing on the audio dataset D a ;
[0025] S23. Combine the preprocessed text dataset P t (x t ), the image dataset P i (x i ) and the audio dataset P a (x a ) into a preprocessed multi-modal input dataset:
[0026] P D = {P t (x t ), P i (x i ), P a (x a )}.
[0027] Optionally, the step S3 includes:
[0028] S31. Input the preprocessed multi-modal input data set P D into a large language model with a modality perception mechanism;
[0029] S32. For the text data P t (x t ) generate a text semantic representation through the text embedding layer of the large language model:
[0030]
[0031] where E t (x t ) represents the semantic representation vector of the text data x t ; w j is the j-th word of the text; φ(w j ) is the word embedding obtained through the pre-trained embedding matrix, α j is the attention weight, and γ j is the context vector;
[0032] S33. For the image data P i (x i ) generate an image semantic representation through the image embedding layer:
[0033]
[0034] where E i (x i ) represents the semantic representation vector of the image data x i ; p mn is the pixel point of the image, and f(p mn ) represents the image features extracted through the convolutional neural network; β mn is the weight of the image features, representing the visual importance of a specific area, and δ mn is the position encoding;
[0035] S34. For the audio data P a (x a ) generate an audio semantic representation through the audio embedding layer:
[0036]
[0037] where E a (x a ) represents the semantic representation vector of the audio data x a ; f k is the k-th frequency component in the audio signal, and ψ(f k ) is the frequency domain feature extracted through the Fourier transform; ηk is the feature weight of the audio signal, σ k is the time step encoding;
[0038] S35. Integrate the generated text semantic representation E t (x t ), the image semantic representation E i (x i ) and the audio semantic representation E a (x a ) through a modality-aware mechanism to generate a comprehensive semantic representation that includes multi-modal context understanding:
[0039] E multi = λ1·g(E t (x t )) + λ2·g(E i (x i )) + λ3·g(E a (x a )) + θ(C);
[0040] where, E multi is the multi-modal comprehensive semantic representation, g(E) represents the non-linear transformation function of each modality representation, λ1, λ2, λ3 are the importance weights of the features of each modality, and θ(C) is the multi-modal context fusion function.
[0041] Optionally, the S4 step includes:
[0042] S41. Use an adaptive modality interaction framework to adaptively fuse the comprehensive semantic representation E multi generated by the large language model and the features of each modality, and the adaptive modality interaction framework dynamically adjusts the interaction weights of different modalities;
[0043] S42. Through the interaction function h(E multi , E m ) in the adaptive modality interaction framework, perform feature fusion on the text semantic representation E t (x t ), the image semantic representation E i (x i ) and the audio semantic representation E a (x a ):
[0044]
[0045] where, is the feature interaction function at the p-th row and q-th column during the fusion process, combining the local interaction of the comprehensive semantic representation E multi and the modality feature E m , α pqis the local attention weight, β pq is the weight of the feature position, combined with the position encoding of the modal features to optimize the fused spatial feature representation;
[0046] S43. Through the cross-modal interaction mechanism, dynamically adjust the weights ω of different modalities according to the importance of each modality m :
[0047]
[0048] where, μ m is the basic weight of modality m, is the multiple interaction function between modalities, ζ i is the dynamic adjustment parameter related to the i-th feature dimension;
[0049] S44. Generate the optimized fused feature representation F opt , and the optimization process of the fused feature representation is carried out through the non-linear transformation function g(F fusion ):
[0050]
[0051] where, F opt represents the finally optimized fused feature representation, γ r is the weight of the r-th non-linear transformation channel, is the non-linear activation function of the fused feature representation, δ r is the dynamic weight of each channel.
[0052] Optionally, the S5 step includes:
[0053] S51. Based on the fused feature representation F opt Identify complex entities and relationships in text, image, and audio data through the modality-aware mechanism;
[0054] S52. Identify explicit and implicit entities in the text modality through the modality-aware mechanism, and the identification of implicit entities is expressed as:
[0055]
[0056] where, E implicit is the implicit entity representation; w j is the j-th word in the text; δ j represents the implicit weight of the current word; c j is the context information of this word; ξ(w j , ∑λ k ·g(φ k (c j , h j))) is a comprehensive relationship function between words and context, where λ k is an implicit weight adjustment parameter, φ k (c j ,h j ) represents the interaction function of context c j and the hidden layer state h j ; g is an activation function for non-linear transformation, ∈ k is random noise;
[0057] S53. For image and audio modalities, the relationships between entities are identified through a modality perception mechanism:
[0058]
[0059] Among them, R relation is the relationship mapping between entities; κ pq is the relationship weight between modalities; is the mutual influence between the entities of modality p and modality q; α pq is the local interaction weight of different modality features; is the local feature interaction function in modality p and modality q, z ∈ Ω is the integration domain, representing the feature spaces of modality p and q;
[0060] S54. Generate corresponding classification labels C entity :
[0061]
[0062] Among them, C entity is the classification label of the entity; W c and b c are the weight matrix and bias term of the classifier respectively; γ l and ζ r are the weight coefficients of entity features and relationship mappings respectively; tanh represents the activation function, which is used to generate the final classification result after fusing entity features and relationship information; the softmax function is used to map the result to a probability distribution.
[0063] Optionally, the S6 step includes:
[0064] S61. Automatically match the identified entity E entity with the existing knowledge graph G = (V, R) through a dynamic mapping algorithm. The multi-modal features of entity E entity are matched with the knowledge graph node v i ∈ V through modal similarity calculation;
[0065] S62. If the identified entity Eentity No existing node v is matched i , then a new node v is automatically generated according to the result of the adaptive modal interaction new , and the new node is represented as:
[0066]
[0067] where v new is the newly generated node, γ l is the weight of the feature dimension, φ l represents the interaction of the non - linear activation function for generating the new node combined with the entity feature E entity and the classification label C entity of α k is the weight of the classification label, is the transformation function of the multi - modal feature, and y ∈ Θ is the offset in the feature space;
[0068] S63. Update the relationship R between entities, and generate a new entity v through the relationship mapping new and the relationship between the existing entity node v i :
[0069]
[0070] where r(v new , v i ) is the newly generated relationship mapping, β j is the weight of the relationship mapping, ρ j represents the feature interaction function between the entity and the node, and γ j (z) represents the position offset function of the local feature, and z ∈ Ω is the feature space.
[0071] S64. Automatically update the knowledge graph to generate the extended knowledge graph G'=(V', R'):
[0072]
[0073] where G' is the extended knowledge graph, η n is the weight of the new relationship, ζ n represents the non - linear transformation function of the relationship, updates the nodes and relationships of the knowledge graph by combining the relationship mapping between the new node and the existing node, and dz is the integral operation in the feature space.
[0074] The beneficial effects of the present invention are:
[0075] (1) By dynamically adjusting the information interaction weights between different modalities, the present invention makes full use of the complementarity of text, image, and audio multi-modal data to achieve more precise feature fusion. Under the adaptive modality interaction framework, the weights between modalities will be adaptively adjusted according to different tasks and data characteristics, thereby improving the understanding and recognition ability of cross-modal data. Compared with traditional single-modal or simple fusion methods, the present invention significantly improves the accuracy and robustness of entity recognition in complex scenarios, especially showing stronger adaptability when dealing with multi-modal data containing noise or incomplete information.
[0076] (2) By introducing the modality perception mechanism of the large language model, the present invention can simultaneously process and understand data from different modalities to achieve semantic fusion of multi-modal contexts. While processing text data, the large language model can combine the information in images and audio to generate cross-modal comprehensive semantic representations, which can not only identify explicit entity information but also infer implicit entities and their relationships through context, improving the accuracy of complex entity recognition.
[0077] (3) The present invention combines the multi-modal entity recognition results with a dynamic knowledge graph, automatically generates new entity nodes and updates the relationships between entities. The new entities identified through the dynamic mapping algorithm can be added to the existing knowledge graph in real time, and the relationship mappings between nodes can be automatically generated according to the relationship information between modalities, effectively coping with the constantly changing entities and relationships in multi-modal data, ensuring that the knowledge graph is always in the latest state, significantly improving the dynamic adaptability of the system and the integrity of the knowledge graph. Compared with traditional manual or semi-automated update methods, the dynamic update mechanism of the present invention greatly improves the efficiency and can timely reflect the latest entity and relationship structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0079] Figure 1 is a flowchart of a multi-modal entity recognition method based on a large language model proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0080] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0081] Refer to Figure 1 , a multi-modal entity recognition method based on a large language model, comprising the following steps:
[0082] S1. Receive text data, image data, and audio data, and construct a multi-modal input data set;
[0083] S2. Determine and apply preprocessing methods for different modal input data through an adaptive modal interaction framework, and preprocess the multi-modal input data set;
[0084] S3. Input the preprocessed multi-modal input data set into a large language model with a modal perception mechanism. The large language model is used to generate a comprehensive semantic representation containing multi-modal context understanding, combining the semantic information of text, images, and audio;
[0085] S4. Use the adaptive modal interaction framework to adaptively fuse the comprehensive semantic representation generated by the large language model and other modal features, and generate an optimized fused feature representation by dynamically adjusting the interaction weights of different modalities;
[0086] S5. Perform cross-modal entity recognition based on the fused feature representation. Cross-modal entity recognition includes using the modal perception mechanism to identify complex entities and relationships in text and other modal information, and generating corresponding classification labels and relationship mappings for the entities;
[0087] S6. Automatically match the identified entities with the existing knowledge graph through a dynamic mapping algorithm. If the entity is a new entity, then according to the results of the adaptive modal interaction, automatically generate new nodes and update the relationship of the existing knowledge graph nodes;
[0088] S7. By continuously inputting new multi-modal data, the system automatically adjusts and optimizes the modal perception mechanism of the large language model and the adaptive modal interaction framework, continuously learns new knowledge, and dynamically expands the knowledge graph.
[0089] In this embodiment, step S1 includes:
[0090] S11. Receive a text data set, an image data set, and an audio data set;
[0091] S12. For the text data set D t , each text sample x t ∈ D t is composed of a set of sentences {s1, s2,..., s n}, where s n is the nth sentence in the text sample;
[0092] S13. For the image data set D i , each image sample x i ∈ D i is represented by a set of pixel matrices M = {p 11 , p 12 ,..., p mn}, where pmn represents the pixel value at the m-th row and n-th column in the image;
[0093] S14. For the audio dataset D a , each audio sample x a ∈D a is represented by a set of frequency-domain feature vectors F = {f1, f2, …, f k}, where f k is the k-th frequency component of the audio signal;
[0094] S15. Combine the text dataset D t , the image dataset D i and the audio dataset D a into a multi-modal input dataset:
[0095] D = {x t , x i , x a}.
[0096] In this embodiment, step S2 includes:
[0097] S21. Perform modal recognition on the multi-modal input dataset D = {x t , x i , x a} through the adaptive modal interaction framework to determine the type and data structure of each modality;
[0098] S22. Apply different preprocessing methods according to the characteristics of each modality, where perform word segmentation and denoising processing on the text dataset D t , perform normalization and size normalization processing on the image dataset D i , and perform denoising and frequency-domain conversion processing on the audio dataset D a ;
[0099] S23. Combine the preprocessed text dataset P t (x t ), the image dataset P i (x i ) and the audio dataset P a (x a ) into a preprocessed multi-modal input dataset:
[0100] P D = {P t (x t ), P i (x i ), P a (x a )}.
[0101] In this embodiment, step S3 includes:
[0102] S31. Input the preprocessed multi-modal input data set P D into a large language model with a modality perception mechanism;
[0103] S32. For the text data P t (x t ) generate a text semantic representation through the text embedding layer of the large language model:
[0104]
[0105] where E t (x t ) represents the semantic representation vector of the text data x t ; w j is the j-th word of the text; φ(w j ) is the word embedding obtained through the pre-trained embedding matrix, α j is the attention weight, and γ j is the context vector;
[0106] S33. For the image data P i (x i ) generate an image semantic representation through the image embedding layer:
[0107]
[0108] where E i (x i ) represents the semantic representation vector of the image data x i ; p mn is the pixel point of the image, and f(p mn ) represents the image features extracted through the convolutional neural network; β mn is the weight of the image features, representing the visual importance of a specific region, and δ mn is the position encoding;
[0109] S34. For the audio data P a (x a ) generate an audio semantic representation through the audio embedding layer:
[0110]
[0111] where E a (x a ) represents the semantic representation vector of the audio data x a ; f k is the k-th frequency component in the audio signal, and ψ(f k ) is the frequency domain feature extracted through the Fourier transform; ηk is the feature weight of the audio signal, σ k is the time step encoding;
[0112] S35. Integrate the generated text semantic representation E t (x t ), the image semantic representation E i (x i ) and the audio semantic representation E a (x a ) through a modality-aware mechanism to generate a comprehensive semantic representation that includes multi-modal context understanding:
[0113] E multi = λ1·g(E t (x t )) + λ2·g(E i (x i )) + λ3·g(E a (x a )) + θ(C);
[0114] where E multi is the multi-modal comprehensive semantic representation, g(E) represents the non-linear transformation function of each modality representation, λ1, λ2, λ3 are the importance weights of each modality feature, and θ(C) is the multi-modal context fusion function.
[0115] In this embodiment, step S4 includes:
[0116] S41. Use an adaptive modality interaction framework to adaptively fuse the comprehensive semantic representation E multi generated by the large language model and each modality feature, and the adaptive modality interaction framework dynamically adjusts the interaction weights of different modalities;
[0117] S42. Through the interaction function h(E multi , E m ) in the adaptive modality interaction framework, perform feature fusion on the text semantic representation E t (x t ), the image semantic representation E i (x i ) and the audio semantic representation E a (x a ):
[0118]
[0119] where is the feature interaction function at the p-th row and q-th column during the fusion process, combining the local interaction of the comprehensive semantic representation E multi and the modality feature E m , α pqis the local attention weight, β pq is the weight of the feature position, which combines the position encoding of the modal features to optimize the fused spatial feature representation;
[0120] S43. Through the interaction mechanism between modalities, dynamically adjust the weights ω of different modalities according to the importance of each modality m :
[0121]
[0122] where μ m is the basic weight of modality m, is the multiple interaction function between modalities, ζ i is the dynamic adjustment parameter related to the i-th feature dimension;
[0123] S44. Generate the optimized fused feature representation F opt , and the optimization process of the fused feature representation is carried out through the non-linear transformation function g(F fusion ):
[0124]
[0125] where F opt represents the finally optimized fused feature representation, γ r is the weight of the r-th non-linear transformation channel, is the non-linear activation function of the fused feature representation, δ r is the dynamic weight of each channel.
[0126] In this embodiment, step S5 includes:
[0127] S51. Based on the fused feature representation F opt Identify complex entities and relationships in text, image, and audio data through the modality perception mechanism;
[0128] S52. Identify explicit and implicit entities in the text modality through the modality perception mechanism. The identification of implicit entities is expressed as:
[0129]
[0130] where E implicit is the implicit entity representation; w j is the j-th word in the text; δ j represents the implicit weight of the current word; c j is the context information of this word; ξ(w j , ∑λ k ·g(φ k (c j , h j))) is a comprehensive relationship function of words and context, where λ k is an implicit weight adjustment parameter, φ k (c j ,h j ) represents the interaction function of context c j and the hidden layer state h j ; g is an activation function for non-linear transformation, ∈ k is random noise;
[0131] S53. For image and audio modalities, identify the relationships between entities through the modality perception mechanism:
[0132]
[0133] Among them, R relation is the relationship mapping between entities; κ pq is the relationship weight between modalities; is the mutual influence between entities of modality p and modality q; α pq is the local interaction weight of different modality features; is the local feature interaction function in modality p and modality q, z ∈ Ω is the integration domain, representing the feature spaces of modality p and q;
[0134] S54. Generate corresponding classification labels C entity for the implicit entities in the identified text modality and the entities in the image and audio modalities:
[0135]
[0136] Among them, C entity is the classification label of the entity; W c and b c are the weight matrix and bias term of the classifier respectively; γ l and ζ r are the weight coefficients of entity features and relationship mapping respectively; tanh represents the activation function, which is used to generate the final classification result after fusing entity features and relationship information; the softmax function is used to map the result to a probability distribution.
[0137] In this embodiment, the S6 step includes:
[0138] S61. Automatically match the identified entity E entity with the existing knowledge graph G = (V, R) through the dynamic mapping algorithm. The multi-modal features of entity E entity are matched with the knowledge graph node v i ∈ V through modal similarity calculation;
[0139] S62. If the identified entity Eentity No existing node v was matched i , then a new node v is automatically generated according to the result of the adaptive modal interaction new , and the new node is represented as:
[0140]
[0141] where v new is the newly generated node, γ l is the weight of the feature dimension, φ l represents the interaction of the non - linear activation function for generating the new node combined with the entity feature E entity and the classification label C entity is the weight of the classification label, k is the weight of the classification label, is the transformation function of the multi - modal feature, y ∈ Θ is the offset in the feature space;
[0142] S63. Update the relationship R between entities, and generate a new entity v through the relationship mapping new and the relationship between the existing entity node v i :
[0143]
[0144] where r(v new , v i ) is the newly generated relationship mapping, β j is the weight of the relationship mapping, ρ j represents the feature interaction function between the entity and the node, γ j (z) represents the position offset function of the local feature, z ∈ Ω is the feature space.
[0145] S64. Automatically update the knowledge graph to generate the extended knowledge graph G'=(V', R'):
[0146]
[0147] where G' is the extended knowledge graph, η n is the weight of the new relationship, ζ n represents the non - linear transformation function of the relationship, updates the nodes and relationships of the knowledge graph by combining the relationship mapping between the new node and the existing node, and dz is the integral operation in the feature space.
[0148] Example 1:
[0149] In this Embodiment 1, taking the traffic management system of a smart city as an application scenario, it is demonstrated how the present invention effectively identifies and processes entity information in traffic conditions, and automatically updates relevant entities and their relationships through a knowledge graph, proving the feasibility and effectiveness of the present invention. The traffic management system of a smart city needs to process data from different modalities in real time, including text reports, images captured by traffic cameras, and audio information collected in the monitoring system. The above data together constitute the multi-modal input of the system, aiming to timely identify traffic events and related entities (vehicles, pedestrians, and traffic lights), and dynamically update the road status and traffic flow conditions through a knowledge graph.
[0150] In City A on August 15, 2024, a bus rear-ended a private car, causing traffic congestion on that road. After the accident, the traffic management system received three types of data through intelligent devices:
[0151] The first type is audio data sent by the alarm operator through voice, describing the occurrence of the accident;
[0152] The second type is image data captured by traffic cameras, including on-site photos of the bus and the private car;
[0153] The third type is a road condition monitoring text report, describing the affected area of the accident and the traffic flow conditions.
[0154] The system processes the data through the method of the present invention to automatically identify the key entities (vehicles, road sections, and accident types) in the accident, and updates the relevant knowledge graph to optimize road control measures.
[0155] First, the system receives data from three modalities: text, image, and audio. The text data includes a road condition report, specifically describing the location of the accident (at about 500 meters from the intersection in the east-west direction on the main road of the Second Ring Road in City A) and the congestion level (the blocked length is about 3 kilometers and the vehicle passage is slow); the image data comes from traffic cameras and captures the specific scene of the collision between the bus and the private car; the audio data comes from the description of the alarm operator, with the content of "The bus rear-ended the private car, and the traffic flow was blocked on the main road of the Second Ring Road in City A, and urgent handling is required."
[0156] The data are used as multimodal inputs in the method of the present invention and preprocessed through an adaptive modal interaction framework. In the text data, the system first performs word segmentation and denoising on the accident report, extracts key location, time, and event information, and identifies entities such as "bus", "private car", "rear-end accident", and "Second Ring Road of City A". In the image data processing, the system uses a convolutional neural network to extract visual features in the image, identifies the collision location and shape features of the bus and the private car, and through the modal perception mechanism, matches the vehicles in the image with the information of "bus" and "private car" in the text. For audio data, the system converts the description information in the audio into text through frequency domain feature extraction and speech recognition, and associates it with the entities in the text and image.
[0157] Next, the system generates a comprehensive semantic representation through the modal perception mechanism of the large language model, combining semantic information from different modalities. For the "bus" entity in the text, the system combines the visual features in the image, determines its vehicle type as "city bus", and identifies that the accident occurred on the Second Ring Road of City A. At the same time, the "rear-end" event in the audio data is also matched with the collision scene in the image through comprehensive semantic analysis. Through the adaptive interaction mechanism between modalities, the system dynamically adjusts the weights of different modalities to ensure that data information from different sources is fully utilized.
[0158] The system then uses the dynamic mapping algorithm in the present invention to match the identified entities of "bus", "private car", and "rear-end accident" with the existing traffic management knowledge graph. Since the vehicles and events involved in this accident are new entities, the system automatically generates a new node "bus rear-end accident" and establishes new relationships with the entities of "Main Road of Second Ring Road of City A", "bus", and "private car". The knowledge graph is automatically expanded, generating new nodes and relationship mappings, and the updated knowledge graph can timely reflect the latest road traffic conditions.
[0159] To prove the superiority of the method of the present invention, a large number of experiments were conducted to compare the performance of the method of the present invention with existing single-modal entity recognition methods in different traffic scenarios. The following Table 1 shows the experimental results based on a real traffic data set, including multimodal data of text, image, and audio modalities, and evaluates the performance of the two methods in terms of entity recognition accuracy and knowledge graph update speed:
[0160] Table 1 Experimental Results Based on Real Traffic Data Set
[0161]
[0162] As can be seen from Table 1, the method of the present invention achieves an entity recognition accuracy rate of 95.8%, significantly higher than the existing single-modal methods (78.5%, 80.2% and 76.4% for text, image and audio respectively). Especially in the fusion and utilization rate of multi-modal data, the method of the present invention reaches 92%, far exceeding the utilization rate of the existing single-modal methods. In addition, the method of the present invention can complete the automatic update of the knowledge graph within 1.2 seconds, with a speed significantly faster than the 3.5 to 4.0 seconds of the existing technology. The data in Table 1 proves the significant advantages of the present invention in improving recognition accuracy, reducing time delay and enhancing the efficiency of knowledge graph update.
[0163] Through the above application scenario of smart city traffic management, Example 1 details the feasibility and effectiveness of the multi-modal entity recognition method based on the large language model of the present invention in practice. By processing multi-modal data and combining the semantic understanding ability of the large language model and the adaptive modal interaction framework, the system can achieve efficient and accurate entity recognition and automatically update the knowledge graph through the dynamic mapping algorithm. Through the comparison of experimental data, the method of the present invention is superior to the existing technology in terms of entity recognition accuracy rate, knowledge graph update speed and multi-modal data utilization rate, and can effectively solve the problems of insufficient cross-modal data fusion and lagging knowledge graph update in the existing technology, and has a wide range of application prospects.
[0164] The present invention realizes more precise feature fusion by dynamically adjusting the information interaction weights between different modalities and making full use of the complementarity of multi-modal data of text, image and audio. Under the adaptive modal interaction framework, the weights between modalities will be adaptively adjusted according to different task and data characteristics, thus improving the understanding and recognition ability of cross-modal data. Compared with the traditional single-modal or simple fusion methods, the present invention significantly improves the accuracy rate and robustness of entity recognition in complex scenarios, especially showing stronger adaptability when dealing with multi-modal data containing noise or incomplete information.
[0165] The present invention can simultaneously process and understand data from different modalities by introducing the modal perception mechanism of the large language model, and realize the semantic fusion of multi-modal contexts. While processing text data, the large language model can combine the information in images and audio to generate cross-modal comprehensive semantic representations, which can not only identify explicit entity information, but also infer implicit entities and their relationships through context, improving the accuracy of complex entity recognition.
[0166] The present invention combines the results of multi-modal entity recognition with a dynamic knowledge graph, automatically generates new entity nodes and updates the relationships between entities. The new entities identified through the dynamic mapping algorithm can be added to the existing knowledge graph in real time, and the relationship mappings between nodes are automatically generated based on the relationship information between modalities, effectively coping with the constantly changing entities and relationships in multi-modal data, ensuring that the knowledge graph is always in the latest state, significantly enhancing the dynamic adaptability of the system and the integrity of the knowledge graph. Compared with traditional manual or semi-automated update methods, the dynamic update mechanism of the present invention greatly improves the efficiency and can timely reflect the latest entity and relationship structures.
[0167] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A multi-modal entity recognition method based on a large language model, characterized in that, It includes the following steps: S1. Receive text data, image data, and audio data, and construct a multi-modal input dataset; S2. Determine and apply preprocessing methods for different modal input data through an adaptive modal interaction framework, and preprocess the multi-modal input dataset; S3. Input the preprocessed multi-modal input dataset into a large language model with a modal perception mechanism. The large language model is used to generate a comprehensive semantic representation including multi-modal context understanding, combining the semantic information of text, images, and audio; S4. Use the adaptive modal interaction framework to adaptively fuse the comprehensive semantic representation generated by the large language model and other modal features, and generate an optimized fused feature representation by dynamically adjusting the interaction weights of different modalities; S5. Perform cross-modal entity recognition based on the fused feature representation. Cross-modal entity recognition includes using the modal perception mechanism to identify complex entities and relationships in text and other modal information, and generating corresponding classification labels and relationship mappings for the entities; S6. Automatically match the identified entities with the existing knowledge graph through a dynamic mapping algorithm. If the entity is a new entity, then according to the results of the adaptive modal interaction, automatically generate new nodes and update the relationship of the existing knowledge graph nodes; S7. By continuously inputting new multi-modal data, the system automatically adjusts and optimizes the modal perception mechanism of the large language model and the adaptive modal interaction framework, continuously learns new knowledge, and dynamically expands the knowledge graph; The S2 step includes: S21. Perform modality recognition on the multi-modal input data set D = {x t , x i , x a} through the adaptive modality interaction framework to determine the type and data structure of each modality; S22. Apply different preprocessing methods according to the characteristics of each modality, where for the text dataset D t perform word segmentation and denoising processing, and for the image dataset D i perform normalization and size normalization processing, and for the audio dataset D a perform denoising and frequency domain conversion processing; S23. Combine the preprocessed text dataset P t (x t ), the image dataset P i (x i ) and the audio dataset P a (x a ) into the preprocessed multi-modal input dataset P D = {P t (x t ), P i (x i ), P a (x a )}; The S4 step includes: S41. Use the adaptive modal interaction framework to adaptively fuse the comprehensive semantic representation E generated by the large language model multi and each modal feature, and the adaptive modal interaction framework dynamically adjusts the interaction weights of different modalities; S42. Through the interaction function h(E multi , E m ) in the adaptive modality interaction framework, perform feature fusion on the text semantic representation E t (x t ), the image semantic representation E i (x i ), and the audio semantic representation E a (x a ); S43. Dynamically adjust the weights ω of different modalities according to the importance of each modality through the interaction mechanism between modalities m ; The S5 step includes: S51. Based on the fused feature representation F opt Identify complex entities and relationships in text, image, and audio data through a modality perception mechanism; S52. Identify explicit and implicit entities in the text modality through the modal perception mechanism; S53. For the image and audio modalities, identify the relationships between entities through the modal perception mechanism; Generate corresponding classification label C for the implicit entities in the identified text modality and the entities in the image and audio modalities entity ; The S6 step includes: S61. Map the identified entity E entity automatically to the existing knowledge graph G=(V, R) through a dynamic mapping algorithm. The multi-modal features of entity E entity are matched with the knowledge graph node v i ∈V through modal similarity calculation; S62. If the recognized entity E entity fails to match the existing node v i , a new node v is automatically generated according to the result of the adaptive modal interaction new ; S63. Update the relationship R between entities, and generate a new entity v through relationship mapping new and the existing entity node v i between the relationships; S64. Automatically update the knowledge graph to generate an extended knowledge graph G'=(V',R'); 2. The multimodal entity recognition method based on a large language model according to claim 1, wherein The S1 step includes: S11. Receive a text dataset, an image dataset, and an audio dataset; S12. For the text dataset D t , each text sample x t ∈D t is composed of a set of sentences {s1, s2, …, s n}, where s n is the nth sentence in the text sample; S13. For the image dataset D i , each image sample x i ∈ D i is represented by a set of pixel matrices M = {p 11 , p 12 , …, p mn}, where p mn represents the pixel value at the m-th row and n-th column in the image; S14. For the audio dataset D a , each audio sample x a ∈D a is represented by a set of frequency-domain feature vectors F = {f1, f2, …, f k}, where f k is the k-th frequency component of the audio signal; S15. Combine the text dataset D t , the image dataset D i and the audio dataset D a into a multi-modal input dataset: D = {x t , x i , x a}.
3. A multimodal entity recognition method based on a large language model according to claim 1, characterized in that The S3 step includes: S31. Input the preprocessed multi-modal input data set P D into a large language model with a modality perception mechanism; S32. For the text data P t (x t ) Generate a text semantic representation through the text embedding layer of the large language model: Among them, E t (x t ) represents the semantic representation vector of the text data x t ; w j is the j-th word of the text; φ(w j ) is the word embedding obtained through the pre-trained embedding matrix, α j is the attention weight, and γ j is the context vector; S33. For the image data P i (x i ), generate an image semantic representation through the image embedding layer: Among them, E i (x i ) represents the semantic representation vector of the image data x i ; p mn is the pixel point of the image, and f(p mn ) represents the image features extracted by the convolutional neural network; β mn is the weight of the image features, representing the visual importance of a specific region, and δ mn is the position encoding; S34. For audio data P a (x a ), generate an audio semantic representation through the audio embedding layer: Among them, E a (x a ) represents the semantic representation vector of the audio data x a ; f k is the k-th frequency component in the audio signal, and ψ(f k ) is the frequency-domain feature extracted by Fourier transform; η k is the feature weight of the audio signal, and σ k is the time step encoding; S35. The generated text semantic representation E t (x t ), the image semantic representation E i (x i ) and the audio semantic representation E a (x a ) are fused through a modality perception mechanism to generate a comprehensive semantic representation containing multi-modal context understanding: E multi = λ1·g(E t (x t )) + λ2·g(E i (x i )) + λ3·g(E a (x a )) + θ(C); Among them, E multi is a multi-modal comprehensive semantic representation, g(E) represents the non-linear transformation function of each modal representation, λ1, λ2, and λ3 are the importance weights of each modal feature, and θ(C) is the multi-modal context fusion function.
4. A multimodal entity recognition method based on a large language model according to claim 1, characterized in that The S4 step includes: S41. Use the adaptive modal interaction framework to adaptively fuse the comprehensive semantic representation E generated by the large language model multi and the features of each modality, and the adaptive modal interaction framework dynamically adjusts the interaction weights of different modalities; S42. Through the interaction function h(E multi , E m ) in the adaptive modality interaction framework, perform feature fusion on the text semantic representation E t (x t ), the image semantic representation E i (x i ), and the audio semantic representation E a (x a ): Among them, is the feature interaction function of the p-th row and q-th column in the fusion process, which combines the comprehensive semantic representation E multi with the local interaction of the modality feature E m , α pq is the local attention weight, and β pq is the weight of the feature position, which combines the position encoding of the modality feature to optimize the spatial feature representation of the fusion; S43. Dynamically adjust the weights ω of different modalities according to the importance of each modality through the interaction mechanism between modalities m :[[]]END]] Among them, μ m is the basic weight of mode m, is the multiple interaction function between modes, ζ i is the dynamic adjustment parameter related to the i-th feature dimension; S44. Generate the optimized fused feature representation F opt , and the optimization process of the fused feature representation is performed by the non - linear transformation function g(F fusion ): Among them, F opt represents the finally optimized fused feature representation, γ r is the weight of the r-th non-linear transformation channel, is the non-linear activation function of the fused feature representation, δ r is the dynamic weight of each channel.
5. The multimodal entity recognition method based on a large language model according to claim 1, wherein The S5 step includes: S51. Based on the fused feature representation F opt Identify complex entities and relationships in text, image, and audio data through a modality perception mechanism; S52. Identify explicit and implicit entities in the text modality through the modal perception mechanism. The identification of implicit entities is expressed as: Among them, E implicit is the implicit entity representation; w j is the j-th word in the text; δ j represents the implicit weight of the current word; c j is the context information of the word; ξ(w j , ∑λ k ·g(φ k (c j , h j ))) is the comprehensive relationship function between the word and the context, where λ k is the implicit weight adjustment parameter, φ k (c j , h j ) represents the interaction function between the context c j and the hidden layer state h j ; g is the activation function for non-linear transformation, ∈ k is the random noise; S53. For the image and audio modalities, identify the relationships between entities through the modal perception mechanism: where, R relation is the relationship mapping between entities; κ pq is the relationship weight between modalities; is the mutual influence between the entities of modality p and modality q; α pq is the local interaction weight of different modality features; is the local feature interaction function in modality p and modality q, z ∈ Ω is the integration domain, representing the feature spaces of modality p and q; Generate corresponding classification labels C for the implicit entities in the identified text modality and the entities in the image and audio modalities entity : Among them, C entity is the classification label of the entity; W c and b c are the weight matrix and bias term of the classifier respectively; γ l and ζ r are the weight coefficients of entity features and relationship mapping respectively; tanh represents the activation function, which is used to generate the final classification result after fusing entity features and relationship information; the softmax function is used to map the result to a probability distribution.
6. A multimodal entity recognition method based on a large language model according to claim 1, characterized in that The S6 step includes: S61. Identify the entity E entity Automatically match the identified entity E with the existing knowledge graph G = (V, R) through a dynamic mapping algorithm. The multi-modal features of entity E entity are matched with the knowledge graph node v i ∈V through modal similarity calculation; S62. If the recognized entity E entity fails to match the existing node v i , then a new node v is automatically generated according to the result of the adaptive modal interaction new . The new node is represented as: Among them, v new is the newly generated node, γ l is the weight of the feature dimension, φ l represents the interaction of the non - linear activation function for generating new nodes combined with the entity feature E entity and the classification label C entity where α k is the weight of the classification label, is the transformation function of the multi - modal feature, and y ∈ Θ is the offset in the feature space; S63. Update the relationship R between entities and generate a new entity v through relationship mapping new and the existing entity node v i The relationship between: Among them, r(v new , v i ) is the newly generated relationship mapping, β j is the weight of the relationship mapping, ρ j represents the feature interaction function between the entity and the node, γ j (z) represents the position offset function of the local feature, z ∈ Ω is the feature space; S64. Automatically update the knowledge graph to generate an extended knowledge graph G'=(V',R'): Among them, G' is the extended knowledge graph, η n is the weight of the new relationship, ζ n represents the non-linear transformation function of the relationship, updates the nodes and relationships of the knowledge graph by combining the relationship mapping between the new nodes and the existing nodes, and dz is the integral operation in the feature space.
Citation Information
Patent Citations
Multi-modal semantic collaborative interaction image-text joint named entity recognition method
CN115455970A
Multi-relation perception heterogeneous graph visual question and answer method fusing syntax tree
CN117034185A