Key information extraction method and system based on multi-modal large model
Through the application of data preprocessing, cross-modal attention mechanism and feature distillation loss function of multimodal large models, the information extraction problem of multimodal large models in noise interference and low-resolution image scenarios is solved, and the high recall rate and low error propagation rate of key information are achieved, which improves the accuracy and resolution of data.
Patent Information
- Application Number
- CN202510339826.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
Existing multimodal large models have low recall and high error propagation rates in information extraction, making it difficult to work effectively in noise interference and low-resolution image scenarios.
Through standardized preprocessing of input data, cross-modal attention mechanism, hierarchical feature extraction architecture and adaptive output modules for reinforcement learning, modal feature weights are dynamically adjusted, and combined with feature distillation loss function and regularization strategy, intelligent selection of features between modals and elimination of redundant features are achieved.
In noise interference and low-resolution image scenarios, the recall rate of key information is significantly improved, the error propagation rate is reduced, and the data accuracy and resolution are improved.
Smart Images

Figure CN120296182A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information processing, and particularly relates to a key information extraction method and system based on a multimodal large model. Background Art
[0002] A multimodal large model is an artificial intelligence model that can simultaneously process multiple different types of data (such as text, images, voice, video, sensor data, etc.). Through deep neural network technology, it fuses and correlates information of different modalities, thereby achieving a more comprehensive and intelligent understanding and generation ability. By leveraging the powerful capabilities of the multimodal large model, valuable information can be accurately extracted from diverse input data.
[0003] However, existing information extraction methods have problems such as low recall rate of key information and high error propagation rate. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a key information extraction method and system based on a multimodal large model.
[0005] The technical solution adopted to solve the above technical problem is: A key information extraction method based on a multimodal large model, comprising the following specific steps:
[0006] Step 1: Perform standardized preprocessing on the input heterogeneous modal data, including text vectorization, image feature encoding, and speech spectrum analysis;
[0007] Step 2: Establish an inter-modal correlation matrix through a cross-modal attention mechanism to dynamically adjust the feature weights of each modality;
[0008] Step 3: Adopt a hierarchical feature extraction architecture to sequentially perform context semantic understanding, entity relationship modeling, and core information localization;
[0009] Step 4: An adaptive output module based on reinforcement learning dynamically optimizes the output form according to the application scenario.
[0010] Through the above technical solution, intelligent selection of features between modalities is achieved through dynamic weight allocation, greatly improving the recall rate of key information in noise interference scenarios. By introducing a feature distillation loss function and eliminating redundant features, the error propagation rate is effectively reduced in scenarios where noisy speech and low-resolution images coexist.
[0011] Furthermore, the second step includes: constructing a cross-modal contrastive learning loss function; implementing a regularization strategy for inter-modal feature distillation; dynamically discarding feature channels below the threshold.
[0012] Through the above technical solution, the accuracy of the data is greatly improved.
[0013] Further, Step 1 adopts the following specific formula:
[0014] Suppose the input data contains M modalities: {X (1) , …, X (m)}, and an encoder is constructed for each modality:
[0015]
[0016] Where: is the modality-specific encoder, T m is the modality time step, and d m is the feature dimension;
[0017] Spatio-temporal alignment mechanism:
[0018] Define the alignment function to handle the temporal differences of different modalities:
[0019]
[0020] Where τ (m) is the timestamp vector of each modality, and the dynamic time warping algorithm is adopted:
[0021]
[0022] A is the alignment path matrix.
[0023] Through the above technical solution, the problem of temporal misalignment of information from different sources in common scenarios is solved, and the cross-modal correlation accuracy is greatly improved.
[0024] Further, the cross-modal attention mechanism in Step 2 adopts the following formula:
[0025] Construct the inter-modal attention weight matrix:
[0026]
[0027] Where, are the Query and Key projection matrices.
[0028] Through the above technical solution, the intelligent selection of features between modalities is realized, and the recall rate of key information is significantly improved.
[0029] Further, Step 3 adopts the following formula:
[0030] Coarse-grained screening, graph neural network:
[0031] Construct the feature graph G = (V, E), and update the node features:
[0032]
[0033] Edge weight calculation:
[0034] e uv = MLP([h u , h v )
[0035] Fine-grained parsing, knowledge graph enhancement:
[0036] Entity connection uses joint embedding:
[0037]
[0038] Where φ(e') is the embedding vector of entity e in the knowledge base.
[0039] Through the above technical solutions, the parsing ability of information is improved, and the data is made more accurate.
[0040] Furthermore, it includes a multi-source data access interface, a feature fusion engine, a knowledge enhancement module, and an adaptive output unit;
[0041] The multi-source data access interface supports API, file upload, and real-time streaming input. The feature fusion engine integrates a dual-channel processing architecture of Transformer and graph convolutional network. The knowledge enhancement module has a built-in extensible domain knowledge base and semantic rule base. The adaptive output unit has the ability to output in multiple modes such as structured data generation, visual display, and voice broadcast;
[0042] The knowledge enhancement module includes: a self-updating domain term dictionary, an ontology-based relationship reasoning engine, and a visual knowledge editing interface.
[0043] Through the above technical solutions, the hardware cost can be reduced, and the manual substitution rate can be significantly improved.
[0044] The beneficial effects of the present invention are as follows: The present invention realizes the intelligent selection of features between modalities through dynamic weight allocation, greatly improving the recall rate of key information in the noise interference scenario. By introducing a feature distillation loss function and eliminating redundant features, the error propagation rate is effectively reduced in the scenario where noisy speech and low-resolution images coexist. Brief Description of the Drawings
[0045] Figure 1 is the system architecture diagram of the present invention. Detailed Embodiments
[0046] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] As Figure 1 shown, the key information extraction method and system based on the multimodal large model in this embodiment include the following specific steps:
[0048] Step 1: Perform standardized preprocessing on the input heterogeneous modal data, including text vectorization, image feature encoding, and speech spectrum analysis;
[0049] Step 2: Establish an inter-modal correlation matrix through a cross-modal attention mechanism to dynamically adjust the feature weights of each modality;
[0050] Step 3: Adopt a hierarchical feature extraction architecture to sequentially perform context semantic understanding, entity relationship modeling, and core information localization;
[0051] Step 4: The adaptive output module based on reinforcement learning dynamically optimizes the output form according to the application scenario, realizes intelligent selection of features between modalities through dynamic weight allocation, greatly improves the recall rate of key information in the noise interference scenario, introduces a feature distillation loss function, and effectively reduces the error propagation rate in the scenario where noisy speech and low-resolution images coexist after eliminating redundant features.
[0052] Furthermore, the second step includes: constructing a cross-modal contrastive learning loss function; implementing a regularization strategy for inter-modal feature distillation; dynamically discarding feature channels below the threshold, which greatly improves the accuracy of the data.
[0053] The first step adopts the following specific formula:
[0054] Assume that the input data contains M modalities: {X (1) ,…,X (m)}, construct an encoder for each modality:
[0055]
[0056] where: is a modality-specific encoder, T m is the modality time step, and d m is the feature dimension;
[0057] Spatio-temporal alignment mechanism:
[0058] Define the alignment function to handle the temporal differences of different modalities:
[0059]
[0060] where τ (m) is the timestamp vector of each modality, and the dynamic time warping algorithm is adopted:
[0061]
[0062] A is an alignment path matrix, which solves the problem of temporal misalignment of information from different sources in common scenarios and greatly improves the cross-modal correlation accuracy.
[0063] The cross-modal attention mechanism in the second step adopts the following formula:
[0064] Construct an inter-modal attention weight matrix:
[0065]
[0066] where are Query and Key projection matrices, which realize the intelligent selection of inter-modal features and significantly improve the recall rate of key information.
[0067] The third step adopts the following formula:
[0068] Coarse-grained screening, graph neural network:
[0069] Construct a feature graph G=(V, E), and update the node features:
[0070]
[0071] Calculate the edge weights:
[0072] e uv = MLP([h u , h v )
[0073] Fine-grained parsing, knowledge graph enhancement:
[0074] Entity connection adopts joint embedding:
[0075]
[0076] where φ(e') is the embedding vector of entity e in the knowledge base, which improves the parsing ability of information and makes the data more accurate.
[0077] It includes a multi-source data access interface, a feature fusion engine, a knowledge enhancement module, and an adaptive output unit;
[0078] The multi-source data access interface supports API, file upload, and real-time streaming input. The feature fusion engine integrates a dual-channel processing architecture of Transformer and graph convolutional network. The knowledge enhancement module builds an extensible domain knowledge base and semantic rule base. The adaptive output unit has the multi-mode output capabilities of structured data generation, visual display, and voice broadcast;
[0079] The knowledge enhancement module includes: a self-updating domain term dictionary, an ontology-based relationship reasoning engine, and a visual knowledge editing interface, which can reduce hardware costs and significantly increase the manual substitution rate.
[0080] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.
Claims
1. A key information extraction method based on a multimodal large model, characterized in that It includes the following specific steps: Step 1: Perform standardized preprocessing on the input heterogeneous modal data, including text vectorization, image feature encoding, and speech spectrum analysis; Step 2: Establish an inter-modal correlation matrix through a cross-modal attention mechanism to dynamically adjust the feature weights of each modality; Step 3: Adopt a hierarchical feature extraction architecture to sequentially perform context semantic understanding, entity relationship modeling, and core information localization; Step 4: An adaptive output module based on reinforcement learning dynamically optimizes the output form according to the application scenario.
2. The key information extraction method based on a multimodal large model according to claim 1, wherein The said Step 2 includes: constructing a cross-modal contrastive learning loss function; implementing a regularization strategy for inter-modal feature distillation; dynamically discarding feature channels below the threshold.
3. The key information extraction method based on a multimodal large model according to claim 2, characterized in that The said Step 1 adopts the following specific formula: Suppose the input data contains M modalities: {X (1) , …, X (m)}, and an encoder is constructed for each modality: Wherein: is a modality-specific encoder, T m is the modality time step, d m is the feature dimension; Spatio-temporal alignment mechanism: Define the alignment function Handle the timing differences of different modalities: where τ (m) is the timestamp vector of each modality, and the dynamic time warping algorithm is adopted: A is the alignment path matrix.
4. The key information extraction method based on a multi-modal large model according to claim 3, wherein, The cross-modal attention mechanism in the said Step 2 adopts the following formula: Construct an inter-modal attention weight matrix: Among them, is the Query and Key projection matrix.
5. The key information extraction method and system based on a multi-modal large model according to claim 4, wherein The said Step 3 adopts the following formula: Coarse-grained screening, graph neural network: Construct a feature graph G=(V, E), and update the node features: Calculate the edge weights: e uv = MLP([h u , h v ) Fine-grained parsing, knowledge graph enhancement: Entity connection adopts joint embedding: Where φ(e') is the embedding vector of entity e in the knowledge base.
6. The key information extraction system based on the multi-modal large model according to claim 5, wherein, It includes a multi-source data access interface, a feature fusion engine, a knowledge enhancement module, and an adaptive output unit; The said multi-source data access interface supports API, file upload, and real-time streaming input. The feature fusion engine integrates a dual-channel processing architecture of Transformer and graph convolutional network. The knowledge enhancement module has a built-in extensible domain knowledge base and semantic rule base. The adaptive output unit has the multi-mode output capabilities of structured data generation, visual display, and voice broadcast; The said knowledge enhancement module includes: a self-updating domain term dictionary, an ontology-based relationship reasoning engine, and a visual knowledge editing interface.
Citation Information
Cited By
Power market information extraction and pushing method and system based on multi-modal semantic fusion
CN120508992A
Power market information extraction and pushing method and system based on multi-modal semantic fusion
CN120508992B