Multi-modal explicit memory system and device, storage medium and program product

By introducing multimodal coding, unified memory representation, adaptive sparseness and hierarchical retrieval and fusion modules into the explicit memory system, the fusion difficulties and insufficient efficiency in multimodal data processing are solved, and more efficient multimodal information understanding and reasoning effects are achieved.

CN120218237APending Publication Date: 2025-06-27HUA DATA TECH (SHANGHAI) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510272444.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing explicit memory technology faces problems such as multimodal fusion difficulties, insufficient data processing efficiency, limited context correlation and insufficient system scalability when processing multimodal data, resulting in insufficient processing of visual and auditory information in scenarios such as medical diagnosis, autonomous driving and interactive agents.

Method used

A multimodal explicit memory system is proposed. The text, visual and auditory data are converted into vectorized features through the multimodal encoding module, the unified memory representation module realizes cross-modal alignment, the adaptive sparse module reduces storage and calculation costs, and the layered search and fusion module performs dynamic search and fusion, and finally generates the inference results of comprehensive multimodal information through the cross-modal inference and integration module.

Benefits of technology

It significantly improves the model's understanding and inference effect of multimodal information, improves data processing efficiency and multimodal correlation, and enhances the accuracy and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218237A_ABST
    Figure CN120218237A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal explicit memory system and device, a storage medium and a program product, and relates to the technical field of multi-modal data retrieval. The multi-modal explicit memory system comprises a multi-modal coding module used for converting multi-modal data such as text, visual and auditory data into vectorized multi-modal features; the unified memory representation module is used for realizing cross-modal alignment and constructing a multi-modal layer explicit memory bank; the self-adaptive rarefaction module is used for pruning or quantifying the multi-modal features; the hierarchical retrieval and fusion module is used for performing hierarchical retrieval and cross-modal fusion in the memory bank; and the cross-modal reasoning and integration module is used for integrating hierarchical retrieval results by using a large language model to obtain a reasoning result of comprehensive multi-modal information. According to the method, by processing text, image and audio information, a reasoning result integrating visual clues, auditory features and text contexts is output, and the data processing efficiency and the multi-modal correlation degree are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data retrieval, and particularly to a multimodal explicit memory system, device, storage medium, and program product. Background Art

[0002] Traditional large language models are mostly based on implicit memory (knowledge stored in model parameters) and working memory (key-value pairs in the context window). As the model scale continues to expand, the knowledge stored in model parameters shows obvious limitations in terms of flexibility, update cost, and interpretability. To improve the knowledge capacity without significantly increasing the model parameter scale, researchers have proposed frameworks based on explicit memory mechanisms, which store specific knowledge, context, or skills in an external system in the form of retrievable memory blocks, and retrieve and integrate these memory blocks during reasoning to enhance the reasoning ability. Such methods are superior to implicit memory-based methods in terms of updatability, scalability, and dependence on model parameters.

[0003] Currently, most explicit memory technologies mainly focus on the text modality, and often encounter technical bottlenecks such as difficulties in multimodal fusion, insufficient data processing efficiency, limited context relevance, and insufficient system scalability when dealing with non-text modalities such as vision and audition. Such bottlenecks lead to insufficient processing of visual and auditory information in many specific scenarios. For example, in medical diagnosis, it is necessary to combine text medical records and medical images for multimodal analysis and diagnosis; in autonomous driving technology, it is necessary to integrate visual and audio information collected by various sensors such as cameras and radars while executing text-based navigation instructions; in the information processing technology of interactive agents, it involves multimodal interactions between agents and the environment or humans, such as speech, visual cues, and ambient sounds.

[0004] It can be seen that explicit memory technology urgently needs to be improved to make full use of multimodal context information for collaborative reasoning. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the present invention proposes a multimodal explicit memory system, device, storage medium, and program product, which incorporates text, visual, and auditory data into the same explicit memory framework to support cross-modal collaborative reasoning; relying on an adaptive sparsification and hierarchical retrieval mechanism, it reduces storage and computational costs, provides flexible support for knowledge update and expansion, dynamically retrieves and integrates the most relevant multimodal memory blocks, and significantly improves the model's understanding and reasoning effects on multimodal information.

[0006] In a first aspect, the present invention provides a multimodal explicit memory system, including the following modules:

[0007] A multimodal encoding module, configured to convert multimodal data such as text, vision, and audition into vectorized multimodal features;

[0008] A unified memory representation module, which is used to project multi-modal features into a shared vector space to achieve cross-modal alignment; and construct an explicit memory bank for multi-modal layers in the shared vector space;

[0009] An adaptive sparsification module, which is used to prune or quantize multi-modal features;

[0010] A hierarchical retrieval and fusion module, which is used to perform hierarchical retrieval and cross-modal fusion in the memory bank;

[0011] A cross-modal reasoning and integration module, which uses a large language model to integrate the hierarchical retrieval results to obtain an inference result with comprehensive multi-modal information.

[0012] As a further improvement of the present invention, the multi-modal encoding module includes:

[0013] A text encoder, which is used to encode text input into a vector representation;

[0014] A visual encoder, which is used to extract image features and retain key spatial and semantic information;

[0015] An auditory encoder, which is used to generate audio vector embeddings.

[0016] As a further improvement of the present invention, for specific task scenarios, the visual encoder or the auditory encoder is customized and pre-trained.

[0017] As a further improvement of the present invention, the unified memory representation module includes: achieving cross-modal alignment through an additional projection layer or an alignment network.

[0018] As a further improvement of the present invention, the memory bank is composed of memory blocks of different modal layers, and each memory block includes:

[0019] Modal type identifier;

[0020] Vectorized features;

[0021] Relevant context information, including data source, time sequence or location information.

[0022] As a further improvement of the present invention, the adaptive sparsification module, based on a sparsification mechanism of attention or statistical metrics, prunes or quantizes features with lower importance by setting a dynamic threshold.

[0023] As a further improvement of the present invention, the hierarchical retrieval and fusion module includes:

[0024] Modal layer retrieval unit: perform preliminary retrieval in the memory bank, where the modal layer corresponding to the task requirements is preferentially retrieved, and then other modal layers are retrieved;

[0025] Cross-modal fusion unit: Calculate the correlation degree of the retrieval results of different modal layers to screen the most relevant memory blocks; Assign weights to different modalities through a hierarchical attention mechanism to dynamically control the retrieval depth and range.

[0026] As a further improvement of the present invention, the cross-modal reasoning and integration module uses the self-attention layer of the large language model or a dedicated cross-modal attention layer to explicitly model the associations between different modalities through a joint attention mechanism, and obtains an inference result that synthesizes multi-modal information.

[0027] In a second aspect, the present invention provides a multi-modal explicit memory integration method, including the following steps:

[0028] Multi-modal encoding: Used to convert multi-modal data such as text, vision, and audition into vectorized multi-modal features;

[0029] Unified memory representation: Used to project multi-modal features into a shared vector space to achieve cross-modal alignment; And in the shared vector space, construct an explicit memory bank for multi-modal layers;

[0030] Adaptive sparsification: Used to prune or quantize multi-modal features;

[0031] Hierarchical retrieval and fusion: Used to perform hierarchical retrieval and cross-modal fusion in the memory bank;

[0032] Cross-modal reasoning and integration: Use the large language model to integrate the hierarchical retrieval results to obtain an inference result that synthesizes multi-modal information.

[0033] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method described in the first aspect.

[0034] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.

[0035] In a fifth aspect, the present invention provides a computer program product, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.

[0036] Compared with the prior art, the present invention utilizes text, image, and audio information simultaneously during the inference stage, realizes multi-modal parallel retrieval and weighted fusion through a hierarchical retrieval mechanism, explicitly models the associations between different modalities through a joint attention mechanism, enables the model output to synthesize visual cues, auditory features, and text context, and generates more interpretable and accurate inference results. Compared with the retrieval method mainly targeting text, the present invention significantly improves the data processing efficiency and multi-modal correlation, while enhancing the accuracy and robustness of the operation of the retrieval system. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flowchart of the operation of a multi-modal explicit memory system disclosed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the present invention will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Among them, the steps S1, S2... and modules M1, M2 described in the embodiments of the present invention do not limit the only execution manner of the present invention; the various models, simulation environments, and software described in the present invention are not the only limiting manners of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0039] In the present invention, a computer device / equipment / system refers to a related entity applied to a computer, such as hardware, a combination of hardware and software, software, or software in execution. Specifically, for example, software includes, but is not limited to, a process running on a processor, a processor, an object, executable software, an execution thread, a program, and / or a computer. Also, an application program or a script program running on a server, and the server can both be software. One or more software can be in the process of execution and / or thread, and the software can be localized on one computer and / or distributed between two or more computers, and can be run by various computer-readable media.

[0040] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0041] In a first aspect, the present invention provides a multi-modal explicit memory system, as Figure 1 shown, including the following modules:

[0042] M1: A multi-modal encoding module, including:

[0043] Text Encoder: Using pre-trained models such as Transformer and BERT, encodes the text input into a vector representation.

[0044] Visual Encoder: Uses a Convolutional Neural Network (CNN) or Vision Transformer (ViT) to extract image features, retaining key spatial and semantic information.

[0045] Auditory Encoder: Preprocesses audio data (such as converting to Mel spectrogram) and uses a convolutional network or self-attention structure to generate audio vector embeddings.

[0046] M2: Unified Memory Representation Module,

[0047] Projects features of different modalities into a shared vector space compatible with the language model, and cross-modal alignment can be achieved through an additional projection layer or alignment network.

[0048] Builds "Memory Slots" in the shared vector space through this module. Each memory slot includes:

[0049] Modality type identifier;

[0050] Vectorized features (key-value or embedding vectors);

[0051] Related context information (such as data source, temporal or location information, etc.).

[0052] M3: Adaptive Sparsification Module,

[0053] Aiming at the redundancy problem of high-dimensional multi-modal features, introduces a sparsification mechanism based on attention or statistical metrics to prune or quantize less important features, reducing the overall storage requirement and improving the retrieval efficiency.

[0054] Retains high-resolution features for key information such as medical images; quantizes or aggregates general or unimportant features, ensuring data integrity while guaranteeing system performance.

[0055] Furthermore, quantization techniques (such as floating-point compression) can be optionally used in this module to further optimize the storage performance.

[0056] M4: Hierarchical Retrieval and Fusion Module, including:

[0057] Modality Layer Retrieval: According to the context and requirements of the current task, performs a preliminary retrieval in the explicit memory bank of the corresponding modality. For example, if visual information is needed, the visual memory bank is preferentially retrieved; if speech information is needed, the retrieval is performed in the auditory memory bank.

[0058] Cross-modal fusion: Calculate the semantic relevance or similarity of the retrieval results, and assign weights to different modalities through a hierarchical attention mechanism to dynamically control the retrieval depth and scope.

[0059] Calculate the correlation of the retrieval results of different modalities, and screen the most relevant memory blocks based on the attention scores to inject them into the attention layer of the large language model.

[0060] In addition, cross-modal fusion can also be performed through direct weighted averaging or other non-linear fusion algorithms.

[0061] M5: Cross-modal inference and integration module,

[0062] Inject the retrieved multi-modal memory blocks into the self-attention layer or the dedicated cross-modal attention layer of the large language model.

[0063] Explicitly model the associations between different modalities through a joint attention mechanism, enabling the model output to synthesize visual cues, auditory features, and text context, and generating more interpretable and accurate inference results.

[0064] In a second aspect, the present invention provides an embodiment of a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method in the first aspect.

[0065] In a third aspect, the present invention provides an embodiment of a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in the first aspect are implemented.

[0066] In a fourth aspect, the present invention provides an embodiment of a computer program product, and when the computer program is executed by a processor, the steps of the method in the first aspect are implemented.

[0067] Embodiment 1: Application in the medical diagnosis scenario

[0068] S1. Data input:

[0069] Text: Medical record information and doctor's written instructions;

[0070] Vision: Medical images (such as X-rays, CTs, MRIs), after extracting image features through a visual encoder, are stored in an explicit memory block;

[0071] Audition: Patient voice consultation records, after generating embedding vectors through an auditory encoder, are stored in the corresponding memory bank.

[0072] S2. Unified memory representation: Align the input data across modalities and construct a multi-modal layer explicit memory bank; perform adaptive pruning or quantization on the multi-modal data features;

[0073] S3. Hierarchical Retrieval and Fusion:

[0074] When a doctor queries or existing medical record information is input into the system, the required modality (such as X-ray, MRI, text, and voice) is first determined.

[0075] The hierarchical retrieval module retrieves the most relevant image and audio features from the explicit memory bank of the corresponding modality and fuses them in the cross-modal attention layer.

[0076] S4. Cross-modal Reasoning and Integration:

[0077] After obtaining multi-modal evidence, the large language model combines the medical record text with the imaging and voice features to comprehensively generate a diagnosis recommendation with an explanation (such as indicating the location of the lesion in the image or the key symptoms in the patient's voice).

[0078] Compared with the prior art, this method significantly improves the accuracy and interpretability of the diagnosis.

[0079] Example 2: Application in the Autonomous Driving Scenario

[0080] S1. Data Input:

[0081] Text: Vehicle navigation instructions, traffic rules, and road condition information;

[0082] Vision: Images or video frames captured by vehicle cameras, which are encoded by the Vision Transformer and stored in the explicit memory bank;

[0083] Audition: Ambient sounds or voice instructions, which generate vector representations through the auditory encoder.

[0084] S2. Unified Memory Representation: Align the input data across modalities and construct a multi-modal layer explicit memory bank; perform adaptive pruning or quantization on the multi-modal data features;

[0085] S3. Hierarchical Retrieval and Fusion:

[0086] The system receives multi-modal inputs in real time and calls the corresponding memory blocks in combination with the vehicle position and the external environment context (such as weather and traffic conditions), such as the visual features of road signs.

[0087] When identifying a congestion scenario, the hierarchical retrieval module preferentially retrieves the visual information of surrounding vehicles and obstacles and conducts a comprehensive analysis in combination with the ambient sound alarm.

[0088] S4. Cross-modal Reasoning and Integration:

[0089] The large language model generates driving decisions through cross-modal reasoning, thereby enhancing the driving safety and precision of the vehicle.

[0090] Compared with the existing technologies that only support text retrieval, the present invention can effectively integrate visual and auditory perception information, greatly improving the decision-making efficiency.

[0091] Embodiment Three: Interactive Agent Application

[0092] S1. Data Input:

[0093] Text: Conversation content, knowledge base Q&A, etc.;

[0094] Vision: Key frames of the video image of the user or the environment, generating visual features through a visual encoder;

[0095] Audition: User voice commands or environmental sounds, extracting features through an auditory encoder and storing them in an explicit memory bank.

[0096] S2. Unified Memory Representation: Aligning the input data across modalities and constructing a multi-modal layer explicit memory bank; adaptively pruning or quantifying the multi-modal data features;

[0097] S3. Hierarchical Retrieval and Fusion:

[0098] When the user asks a question containing visual cues, the system automatically retrieves visual memory and performs cross-modal fusion in combination with the conversation context to generate a natural interactive answer.

[0099] S4. Cross-modal Reasoning and Integration:

[0100] If there is ambiguity in speech recognition, the system can correct it in combination with visual cues, thereby improving the interaction accuracy and user experience.

[0101] Due to supporting multi-modal memory retrieval and fusion, the performance of this system in complex interaction scenarios is better than that of the RAG technology that only relies on text retrieval.

Claims

1. A multimodal explicit memory system, characterized in that: Includes the following modules: Multimodal encoding module, which is used to convert multimodal data such as text, vision and hearing into vectorized multimodal features; A unified memory representation module is used to project multimodal features into a shared vector space to achieve cross-modal alignment; and to construct a multimodal layer explicit memory library in the shared vector space; Adaptive sparsification module for pruning or quantizing multimodal features; A hierarchical retrieval and fusion module, used for performing hierarchical retrieval and cross-modal fusion in the memory bank; The cross-modal reasoning and integration module uses a large language model to integrate hierarchical retrieval results to obtain reasoning results that integrate multimodal information.

2. The system according to claim 1, characterized in that The multimodal encoding module comprises: A text encoder, which encodes text input into a vector representation; Visual encoder, which is used to extract image features and preserve key spatial and semantic information; Auditory encoder for generating audio vector embeddings.

3. The system according to claim 2, characterized in that Customized pre-training of visual encoders or auditory encoders for specific task scenarios.

4. The system according to claim 1, characterized in that The unified memory representation module includes: achieving cross-modal alignment through an additional projection layer or an alignment network.

5. The system according to claim 1, characterized in that The memory bank is composed of memory blocks of different modal layers, each memory block including: Modality type identifier; Vectorized features; Relevant contextual information, including data source, timing, or location information.

6. The system according to claim 1, characterized in that The adaptive sparseness module, based on the sparseness mechanism of attention or statistical indicators, prunes or quantizes features with lower importance by setting dynamic thresholds.

7. The system according to claim 1, characterized in that The hierarchical retrieval and fusion module includes: Modal layer retrieval unit: performs preliminary retrieval in the memory bank, where the modal layer corresponding to the task requirements is retrieved first, and then other modal layers are retrieved; Cross-modal fusion unit: Filter the most relevant memory blocks by calculating the correlation between retrieval results of different modal layers; assign weights to different modalities through the hierarchical attention mechanism to dynamically control the retrieval depth and scope.

8. The system according to claim 1, characterized in that The cross-modal reasoning and integration module uses the self-attention layer or the dedicated cross-modal attention layer of the large language model to explicitly model the association between different modalities through a joint attention mechanism to obtain an inference result that integrates multimodal information.

9. A multimodal explicit memory integration method, characterized in that: The following steps are involved: Multimodal encoding: used to convert multimodal data such as text, vision and hearing into vectorized multimodal features; Unified memory representation: used to project multimodal features into a shared vector space to achieve cross-modal alignment; and to build a multimodal layer explicit memory library in the shared vector space; Adaptive sparsification: used to prune or quantize multimodal features; Hierarchical retrieval and fusion: used to perform hierarchical retrieval and cross-modal fusion in the memory bank; Cross-modal reasoning and integration: Use a large language model to integrate hierarchical retrieval results to obtain reasoning results that integrate multimodal information.

10. A computer device embodiment, comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the system functions or method steps described in any one of claims 1-9.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the system functions or method steps described in any one of claims 1 to 9 are implemented.

12. A computer program product, characterized in that When the computer program is executed by a processor, the system functions or method steps described in any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Language model reasoning method and device based on hierarchical state memory bank and combined initialization, terminal, medium and product

    CN121638477A

  • Language model inference method and device based on hierarchical state memory bank and combined initialization, terminal, medium and product

    CN121638477B

  • Multimodal explicit memory system, apparatus, storage medium, and program product

    WO2026184075A1