A method, apparatus, device and medium for processing multi-modal data

By constructing a graph-structured memory graph with entities as nodes and a multimodal large model, visual, audio, and textual information are integrated, solving the problem of insufficient response of the multimodal large model in long-term interaction scenarios, and realizing efficient visual long-term memory and cross-time reasoning capabilities.

CN122114172APending Publication Date: 2026-05-29MALANSHAN AUDIO & VIDEO LABORATORY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MALANSHAN AUDIO & VIDEO LABORATORY
Filing Date
2026-02-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing multimodal large models struggle to achieve long-term interaction and cross-time reasoning in scenarios such as home assistants, educational companionship, industrial inspection, and medical rehabilitation. They lack sustained visual memory and semantic association capabilities, leading to response failures or insufficient accuracy.

Method used

Construct a graph-structured memory graph with entities as nodes, integrate semantic understanding of visual text, audio text, and text commands, perform semantic understanding and storage through a multimodal large model, and combine it with a language large model for language reasoning to achieve long-term visual memory and cross-modal retrieval.

Benefits of technology

It improves the effectiveness and accuracy of the system's response in long-term interactions, and has the capabilities of continuous perception, persistent storage and intelligent retrieval, realizing a leap from instantaneous visual recognition to long-term visual cognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114172A_ABST
    Figure CN122114172A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence, in particular to a multi-modal data processing method and device, equipment and medium. The method comprises the following steps: after a user request is converted into a retrieval vector, the retrieval vector is matched and retrieved in combination with a memory graph constructed by multi-modal historical data to obtain memory information as context semantics; the memory graph integrates multi-source semantic information from visual target detection, audio speech and voiceprint recognition, and text instructions, so that the system has a cross-modal long-term memory capability; and a language large model is used to fuse the memory information and the original request for reasoning, and the historical context associated with the semantics is dynamically called for reasoning, so that the effectiveness and accuracy of the response are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method, apparatus, device, and medium for processing multimodal data. Background Technology

[0002] With the development of artificial intelligence, question-and-answer interaction based on language models has become possible. These models can generate answers by combining images and question text, but they generally remain at the "instant recognition" stage, leading to limitations in scenarios such as home assistants, educational companionship, industrial inspection, and medical rehabilitation, resulting in ineffective or inaccurate responses. Summary of the Invention

[0003] The purpose of this application is to provide a method, apparatus, device, and medium for processing multimodal data, which can improve the accuracy of the response.

[0004] In a first aspect, a method for processing multimodal data is provided, comprising: acquiring a user's request and generating a retrieval vector based on the request; determining memory information matching the retrieval vector from a memory graph; performing language inference based on the memory information and the request using a large language model, and outputting the request result; wherein the memory graph is constructed based on historical memory information obtained through semantic understanding of historical visual text information, audio text information, and text instructions, the visual text information is obtained by target detection of video frames in multimodal input data, and the audio text information is obtained by speech recognition and voiceprint recognition of audio streams in multimodal input data.

[0005] In a preferred embodiment, this application may be further configured to include: acquiring historical multimodal input data, the multimodal input data including video frames, audio streams, and text instructions; performing object detection on the video frames to obtain visual text information, the visual text information including entity information and attribute information corresponding to the entity information; performing speech recognition and voiceprint recognition on the audio stream to obtain audio text information, the audio text information including speaker identity and speech content; and performing semantic understanding using a multimodal large model based on the visual text information, the audio text information, and the text instructions to obtain historical memory information, and storing the historical memory information in a memory graph.

[0006] In a preferred embodiment, this application can be further configured such that: the multimodal large model includes: an encoder and a decoder; based on the visual text information, the audio text information, and the text instructions, the multimodal large model is used to perform semantic understanding to obtain memory information, including: based on the visual text information, the audio text information, and the text instructions, the encoder performs unified position encoding to obtain encoded information; and the decoder is used to decode the encoded information to obtain historical memory information.

[0007] In a preferred embodiment, this application may be further configured as follows: after obtaining historical memory information by performing semantic understanding using a multimodal large model based on the visual text information, the audio text information, and the text instructions, the application further includes: determining whether the historical memory information is key information; if so, performing the step of storing the historical memory information to a memory graph; if not, discarding the historical memory information.

[0008] In a preferred embodiment, this application may be further configured to include: updating the memory graph based on the multimodal large model using update instructions, wherein the update instructions include at least one of the following: key information storage instructions and node update instructions.

[0009] In a preferred embodiment, this application can be further configured such that the memory graph is a graph structure with entities as nodes, and temporal, spatial, and semantic relationships are established between the nodes.

[0010] In a preferred embodiment, this application can be further configured to: determine memory information matching the retrieval vector from the memory graph, including: determining whether a second retrieval is needed based on a large language model; if so, determining supplementary memory information matching the retrieval vector from the memory graph; and determining the final memory information matching the retrieval vector based on the large language model according to the memory information and the supplementary memory information.

[0011] Secondly, a multimodal data processing apparatus is provided, comprising: an acquisition module for acquiring a user's request; a generation module for generating a retrieval vector based on the request; a retrieval module for determining memory information matching the retrieval vector from a memory graph; wherein the memory graph is constructed based on historical memory information obtained through semantic understanding of historical visual text information, audio text information, and text instructions, the visual text information being obtained by target detection from video frames in multimodal input data, and the audio text information being obtained by speech recognition and voiceprint recognition from audio streams in multimodal input data; and a request processing module for performing language reasoning based on a large language model, according to the memory information and the request, and outputting the request result.

[0012] Thirdly, an electronic device is provided, the electronic device including a memory and a processor, the memory storing a computer program, the processor executing the multimodal data processing method according to any one of the first aspects when running the computer program.

[0013] Fourthly, a computer-readable storage medium is provided, wherein at least one piece of program code is stored therein, the program code being loaded and executed by a processor to implement the multimodal data processing method as described in any of the first aspects.

[0014] Fifthly, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the multimodal data processing method as described in any of the first aspects.

[0015] In summary, the multimodal data processing method provided in this application has the following beneficial technical effects:

[0016] After converting user requests into retrieval vectors, matching and retrieval are performed using a memory graph constructed from multimodal historical data to obtain memory information, which serves as contextual semantics. The memory graph integrates multi-source semantic information from visual object detection, audio speech and voiceprint recognition, and text commands, enabling the system to have cross-modal long-term memory capabilities. Furthermore, it utilizes a large language model to fuse memory information with the original request for reasoning, dynamically invoking the historical context semantically associated with it for reasoning, thereby improving the effectiveness and accuracy of the response.

[0017] In addition, this application also provides a multimodal data processing apparatus, device, and medium, all of which have the aforementioned beneficial technical effects. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic flowchart of a multimodal data processing method provided in an embodiment of this application;

[0020] Figure 2 This is a schematic diagram of the multimodal data processing flow provided in the embodiments of this application;

[0021] Figure 3 This is a schematic diagram of a memory storage process provided in an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of a multimodal large model processing method provided in an embodiment of this application;

[0023] Figure 5 This is a schematic diagram of a request processing flow provided in an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of the structure of a multimodal data processing device provided in an embodiment of this application;

[0025] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] This specific embodiment is merely an explanation of this application and is not intended to limit it. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of this application.

[0027] It should be noted that, in the optional embodiments of this application, the data related to object information, when applied to specific products or technologies, requires the permission or consent of the object. Furthermore, the collection, use, and processing of this data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if the embodiments of this application involve data related to an object, it must be obtained with the permission and consent of the object, the permission and consent of relevant departments, and in accordance with the relevant laws, regulations, and standards of the country and region. If the embodiments involve personal information, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required. The embodiments also need to be implemented with the permission and consent of the object.

[0028] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0030] While existing multimodal large models possess capabilities such as image understanding and object recognition, they generally remain at the "instant recognition" stage, unable to form sustained visual memory. Consequently, they struggle to meet the demands of long-term interaction, cross-temporal reasoning, and complex task scenarios. First, current models process visual information in a one-off manner; the input is lost once the context is removed, failing to achieve "seeing and remembering," and lacking time-series management and historical state backtracking mechanisms.

[0031] The motivation behind this design is to overcome the limitations of large-scale models that "forget after one viewing" and "lack continuous understanding," enabling the system to possess long-term visual memory capabilities similar to humans. In most current visual models, image or video input is often only used in the current task, unable to accumulate experience or actively retrieve past information for future tasks. This makes the system ill-suited for scenarios with continuous, context-dependent, and long-term interaction requirements, such as home assistants, educational companionship, industrial inspection, and medical rehabilitation.

[0032] To address this pain point, this application embodiment enables the model to understand and abstract visual content, forming a structured "visual memory," which can then be accurately retrieved through semantic retrieval in subsequent interactions, thereby achieving a closed loop of "seeing—remembering—recalling—reasoning."

[0033] In summary, this invention aims to solve the above problems by constructing a structured visual memory generation and retrieval mechanism, enabling the model to have the ability of continuous memory, semantic association, key recall and cross-time reasoning, thereby enabling AI to achieve a leap from instantaneous visual recognition to long-term visual cognition.

[0034] This application provides a method for processing multimodal data, such as... Figure 1 As shown, the method provided in this application embodiment can be executed by an electronic device, which can be a server or a terminal device. The electronic device is a server, which can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smartphone, tablet computer, laptop computer, desktop computer, etc., but is not limited to these. The terminal device and the electronic device can be directly or indirectly connected through wired or wireless communication methods. This application embodiment does not impose any limitations on this connection. The method includes:

[0035] S1. Obtain the user's request and generate a retrieval vector based on the request.

[0036] A user's request represents a user's expressed need or instruction, such as where something was previously placed. Of course, a user's request can also be in the form of text, audio, video, etc., which are not limited in the embodiments of this application.

[0037] A retrieval vector is a numerical vector that represents the features and semantic information of a request after it has been converted into a text request. This vector is used for subsequent similarity calculations in the vector space to facilitate matching and retrieval of relevant content.

[0038] In some embodiments, optionally, a text input box is set in the user input interface, where the user can enter a text request. The electronic device acquires the user's input text in real time by listening to changes in the content of the input box; the text is preprocessed, including removing special characters and standardizing capitalization, to normalize the text format; the preprocessed text is input into a pre-trained language model, such as BERT, for deep understanding and feature extraction; the feature vector of the text is output and used as a retrieval vector. Optionally, speech recognition technology is used. When the user enters a request via voice, the speech signal is first converted into text; semantic analysis is performed on the converted text to identify key information and intent; based on the results of the semantic analysis, a suitable template is selected from a pre-set semantic template library to structure the text; the structured text is input into a vector generation algorithm, such as a term frequency-inverse document frequency (TF-IDF) algorithm, to generate the corresponding retrieval vector. It is understood that other methods can also be used to acquire the user's request and generate the retrieval vector, which is not limited here.

[0039] S2. Determine the memory information that matches the retrieval vector from the memory graph.

[0040] Among them, the memory map is constructed based on historical memory information obtained through semantic understanding of historical visual text information, audio text information and text instructions. The visual text information is obtained by target detection of video frames in multimodal input data, and the audio text information is obtained by speech recognition and voiceprint recognition of audio streams in multimodal input data.

[0041] Conventional memory graphs are mostly in simple vector form, lacking structured semantic representations of "person-object-scene-time". In this application, the memory graph is a graph structure with entities as nodes, establishing temporal, spatial, and semantic relationships between nodes. This application constructs an entity relationship and event semantic graph, which can improve the accuracy of retrieval and positioning.

[0042] Among them, based on the vector model, memory information that matches the retrieval vector is determined from the memory graph.

[0043] Vector models can measure the relevance between retrieval vectors and memory information in a memory graph. For large text models, which are conventional large models with retrieval capabilities, their structure is not limited in this application's embodiments. A memory graph is a graph structure used to organize and store information, representing information as nodes and edges. Nodes represent various entities (such as people, objects, etc.), and edges represent relationships between entities (events, time, relationships between people, etc.), clearly presenting the connections and hierarchical structure between information. For example, in a memory graph, different people are nodes, and events, chronological relationships, and relationships between people are edges. Memory information refers to the various data contents stored in the memory graph.

[0044] After a user submits a request and a corresponding retrieval vector has been generated, the vector model is used to search for matching memory information from a pre-built memory graph (containing a large amount of knowledge information).

[0045] Specifically, first, the generated retrieval vector is input into the vector model; all memory information in the memory graph is vectorized, and each memory information is also converted into a numerical vector with the same dimension as the retrieval vector; the similarity between the retrieval vector and each memory information vector is calculated; all memory information is sorted according to the calculated similarity, arranged in descending order of similarity; a similarity threshold is set, and only memory information with a similarity higher than the threshold is considered to be memory information that matches the retrieval vector.

[0046] In this embodiment of the application, based on the vector model, multimodal retrieval is performed to find the most relevant entities, events and context fragments from the memory graph to obtain memory information, and then return the candidate memory set.

[0047] S3. Based on the large language model, perform language reasoning based on memory information and requests, and output the request results.

[0048] Construct a recall context based on memory information; output the request result based on the recall context and the request, realizing reasoning and response driven by visual long-term memory.

[0049] For the large language model, which is a conventional language model, the embodiments in this application are not limited.

[0050] In this embodiment, a multimodal large model is endowed with visual memory generation capabilities. In simpler terms, this means enabling AI not only to understand images but also to "remember what it sees," just like a human, and to retrieve these memories through a language large model for comprehension and response when needed. This solves the problem of traditional AI forgetting what it has seen and being unable to utilize visual information long-term.

[0051] Specifically, it understands scenes and objects from images or videos, extracts key visual text information, transforms it into storable structured visual memories, performs speech recognition and voiceprint recognition on audio streams to obtain audio text information, and quickly retrieves corresponding memories based on user semantic input in dialogues or tasks, achieving cross-modal association and semantic understanding.

[0052] For example, in a home setting, a user can ask, "Where are my keys?" and the AI ​​can retrieve yesterday's visual recordings to answer the location; in education and healthcare settings, the AI ​​can remember students' practice work or rehabilitation movements, compare progress, and provide feedback.

[0053] By constructing a "multimodal long-term memory intelligent agent," a complete closed loop is achieved from continuous visual / audio perception to memory generation, storage, retrieval, and reasoning. This builds a "visual long-term memory system" for AI, enabling it to have continuous perception, persistent storage, and intelligent retrieval capabilities, thus evolving it from a "one-time recognition tool" into an intelligent assistant with cognitive continuity and experience accumulation.

[0054] As can be seen, in this embodiment, after the user request is converted into a retrieval vector, it is matched and retrieved by combining a memory graph constructed from multimodal historical data to obtain memory information as contextual semantics. The memory graph integrates multi-source semantic information from visual object detection, audio speech and voiceprint recognition, and text commands, enabling the system to have cross-modal long-term memory capabilities. Furthermore, it utilizes a large language model to fuse memory information with the original request for reasoning and dynamically calls historical contexts semantically associated with it for reasoning, thereby improving the effectiveness and accuracy of the response.

[0055] In this embodiment, during the "perception-memory generation" stage, images, videos, and audio streams are continuously received to identify people, objects, scenes, and events, and these are converted into structured representations. Memory storage adopts an "entity-centric" graph structure, with visual targets as nodes. Temporal, spatial, event, and semantic relationships are established between nodes, and the memory is divided into "episodic memory" (recording specific events) and "semantic memory" (extracting patterns from multiple observations). Each memory is accompanied by meta-information such as a timestamp and credibility, and can be associated across modalities, such as binding the face and voice of the same person as a unified entity. Furthermore, when a user asks a question or triggers a task, the system enters the "retrieval-reasoning" stage, using vector indexing and graph structure retrieval mechanisms to find relevant memory fragments, i.e., memory information, and further filters and associates historical information through multiple rounds of reasoning. The system also has the ability to filter redundant information, retain key memories, and dynamically update nodes, allowing the memory to continuously evolve. Redundant information includes casual conversation, weather greetings, etc., while only important events are recorded in the memory graph. Finally, the model combines the retrieved visual memories with a language reasoning model to output a decision or answer.

[0056] Overall, this solution upgrades AI from a "short-term recognition tool" to an intelligent agent with long-term visual experience accumulation and cross-time reasoning ability, realizing a human-like cognitive process of "seeing-remembering-recalling-understanding-decision".

[0057] Specifically, one possible implementation of the embodiments of this application further includes:

[0058] S201. Obtain historical multimodal input data, including video frames, audio streams, and text commands;

[0059] Understandably, in the historical question-and-answer process, the input data may involve video and text commands. The video is then broken down into video frames and audio streams to obtain multimodal input data, resulting in video frames, audio streams, and text commands.

[0060] S202. Perform target detection on the video frame to obtain visual text information, which includes entity information and the attribute information corresponding to the entity information.

[0061] In this embodiment, target detection, face recognition, scene recognition, and behavior recognition are performed on video frames to extract preliminary entity information and its attribute information (such as people, objects, locations, and actions). Entity information refers to the name or category of the specific target object detected in the video frame, and is the core part of visual text information. For example, when detecting an animal image, the entity information might be dog, cat, bird, etc. Attribute information is associated with entity information and is used to describe the entity's characteristics, state, relationships, etc. Taking the detected "dog" as an example, its attribute information might include color (e.g., black), breed (husky), actions (running, sleeping), etc.

[0062] In this embodiment, conventional object detection and face recognition models can be used for preliminary analysis to obtain visual text information. This recognized information is then fed into a multimodal large model for analysis. While the multimodal large model cannot recognize faces, it can summarize who is performing a certain action and what they are saying from the recognized visual and audio text information.

[0063] S203. Perform speech recognition and voiceprint recognition on the audio stream to obtain audio text information, which includes the speaker's identity and speech content.

[0064] In this embodiment of the application, speech recognition (ASR) and voice fingerprint recognition (VAD+SpeakerID) are performed on the audio stream to extract the speaker's identity and voice content.

[0065] See Figure 2 , Figure 2This is a schematic diagram of the multimodal data processing flow provided in the embodiments of this application. After multimodal input, video frames and audio streams are obtained, and video processing (object detection, face recognition, scene recognition, behavior recognition) and audio processing (ASR, VAD, SpeakerID) are performed respectively to obtain visual features (person / object / position / action) and audio features (speaker / content / voiceprint).

[0066] S204. Based on visual text information, audio text information, and text instructions, use a multimodal large model to perform semantic understanding, obtain historical memory information, and store the historical memory information in a memory map.

[0067] In this embodiment, visual text information, audio text information, and multimodal input data are input into a multimodal large model (image + audio + ID information joint encoding) to perform deep semantic understanding and enhance the accuracy of entity recognition, action understanding, and scene relationship parsing.

[0068] One possible implementation of this application embodiment is a multimodal large model, which includes an encoder and a decoder. Based on visual text information, audio text information, and text instructions, the multimodal large model is used to perform semantic understanding to obtain memory information. This includes: performing unified position encoding based on the encoder to obtain encoded information based on the visual text information, audio text information, and text instructions; and decoding the encoded information using the decoder to obtain historical memory information.

[0069] Specifically, the multimodal large model aligns and integrates image embeddings, audio embeddings, and text embeddings to generate a unified semantic representation vector. Based on time series and scene events, it semantically structures perceptual information (basic information that the multimodal large model can understand, used to generate semantically structured memories) to generate visual event descriptions, including: entities involved in the event (people / objects), the time and location of the event, and the relationships and behaviors between entities. Meta-information is generated for each memory, including a timestamp (referring to the specific timestamp), a source fragment index, a credibility score, and episodic memory or semantic classification tags.

[0070] Among them, see Figure 3 The multimodal large model receives visual text information, audio text information, and text commands, performs joint encoding, and achieves deep semantic understanding. It then integrates the encoded data, aligning and fusing the image text, audio text, and text commands to unify the semantic vector and obtain encoded information. Decoding yields memory information, which represents event descriptions, including entities, time, location, and relationships. This information is then stored in a memory graph, creating / updating unified cross-scene identifiers for entity nodes and generating metadata, including timestamps, indexes, credibility, and classification tags, to complete memory storage.

[0071] Furthermore, in the embodiments of this application, the multimodal large model is further described.

[0072] In one feasible approach, see [link to relevant documentation] Figure 4 The video, audio, and text are encoded into embeddings using VisionEncoder, AudioEncoder, and Tokenizer respectively. After concatenation, a unified 3DRoPE positional encoding is added, and the LLMDecoder is input to generate a memory text description through autoregression.

[0073] For prompt words, such as:

[0074] You will receive a video and a set of character traits. Each trait (some of which may belong to the same character) can be in one of two forms:

[0075] Facial features: represented by a single frame in the video and its corresponding bounding box;

[0076] Speech features: Composed of several speech segments, each segment includes a start time, an end time (both in MM:SS format), and the corresponding speech content.

[0077] Each facial or voice feature is identified by a unique ID, which is enclosed in angle brackets (<>).

[0078] Your task:

[0079] Based on the provided feature ID, generate a detailed and coherent description of the current video segment. The description should cover all observable or reasonably inferred events in the video. The generated description should include (but is not limited to) the following aspects:

[0080] 1. Character appearance. 2. Character's actions and movements. 3. Character's dialogue. 4. Character's contextualized behavior: character positioning, interaction methods, with a focus on behavior, emotional state, and relationships.

[0081] Strict requirements:

[0082] If a character has an associated feature ID in the input context (whether it's a face or a voice), that feature ID must and can only be used (e.g., ...).<face_1> ,<voice_2> () is used to refer to the character.

[0083] If a role does not have an associated feature ID in the input context, a short descriptive phrase is used to refer to the role.

[0084] Each description must represent only one atomic-level event or detail.

[0085] When the context allows for inference, incorporate natural temporal expressions and spatial location information.

[0086] The final output must be a list of strings, where each string corresponds to exactly one atomic event or description.

[0087] Please return only a list of valid strings, without any additional explanations or formatting instructions.

[0088] The discussion further elaborates on multimodal input information.

[0089] <input_video> ,

[0090] "<face_1> ": ,

[0091] "<face_2> ": ,

[0092] "<face_3> ": ,

[0093] "<voice_1> [{"start_time":"00:05","end_time":"00:08","asr":"Hello"},{"start_time":"00:09","end_time":"00:12","asr":"Let's begin today's lesson."}],

[0094] "<voice_2> [{"start_time":"00:15","end_time":"00:18","asr":"Thank you for giving me the opportunity"}]

[0095] Note: This input is used for a multimodal speaker recognition task and includes video, face images, and timestamped speech transcription information.

[0096] <input_video> : Input video, including video footage and audio.

[0097] <face1> / <face2> / <face3>Face ID and its corresponding face image ( (This is used to provide visual references for known figures.)

[0098] <voice1> / <voice2>Speaker (voiceprint) ID, each ID contains multiple voice segments.

[0099] Each audio segment includes a start time, an end time, and the corresponding ASR text, and is aligned with the video timeline.

[0100] The model aligns spoken content with video footage using voice timestamps and combines this with facial reference images to identify the person in the video corresponding to each voice segment.

[0101] Output example:

[0102] In the bright conference room,<face_1> Walk in confidently, projecting a professional image, and head towards...<face_2> Shake hands with him.

[0103] <face_1> He was wearing a black suit, a white shirt, and a tie.<face_1> He has short black hair and wears glasses.

[0104] <face_2> She wore a striking red dress and had long brown hair.

[0105] <face_2> With a warm smile<face_1> Say hello. Then...<face_2> He sat down at the table, briefly checked his phone, and occasionally looked up to glance around.

[0106] <voice_1> He addressed the crowd, "Good afternoon, everyone. Let's begin the meeting." His commanding voice silenced the room, drawing everyone's attention to him.

[0107] <face_2> Listen attentively<voice_1> He nodded in agreement as he spoke, occasionally checking his phone. The overall atmosphere was professional, and the participants gradually settled into their respective roles.

[0108] <face_1> Adjust your tie and begin discussing the meeting agenda, engaging in productive exchanges with other participants.

[0109] Specifically, the visual text information ImageEmbedding, audio text information AudioEmbedding, and text instruction TokenEmbedding are concatenated; then, unified position encoding 3DRoPE is performed based on the encoder to obtain encoded information. Through encoding, the model can perceive the relative positional relationships of elements in different modalities.

[0110] The encoded information is decoded using the LLMDecoder*24 decoder to obtain the historical memory information.

[0111] Among them, LLMDecoder*24 is a Transformer decoder structure, which repeats 24 times.

[0112] Each layer contains the following components (from top to bottom):

[0113] LayerNorm: Normalization. Normalizes the concatenated multimodal embeddings and 3DRoPE position codes, outputting h_norm1.

[0114] Multi-HeadSelf-Attention: Calculates attention weights among all tokens to capture global dependencies. Using a causal mask, it computes the attention of h_norm1, outputting attn_output = MultiHeadAttention(Q = Wh_q·h_norm1, K = Wh_k·h_norm1, V = Wh_v·h_norm1, mask = causal_mask). Then, it adds a residual connection, outputting h_mid = h_in + attn_output.

[0115] LayerNorm: Normalization. Normalizes h_mid to obtain h_norm2.

[0116] FFNSwiGLU: A feedforward network using the Swish-GatedLinearUnit activation function (enhancing nonlinear expressiveness). The standard FFN is replaced with the SwishGLU structure, resulting in FFN(x) = Swish(W_1x + b_1) ⊗ (W_2x + b_2). With residual connections, the output h_out = h_mid + FFN(h_norm2).

[0117] The h_out is used as the input to the next layer, stacked 24 times, and then H is obtained;

[0118] Then, it also includes: LMHead: the language model head, which maps to the vocabulary space through LM to predict the next token and obtain historical memory information.

[0119] Furthermore, the inability to differentiate the importance of stored information and filter key visual memories may lead to redundancy or information loss. For long videos and continuous observation tasks, it is generally difficult to extract and compress key events, making it difficult to accumulate "visual experience." Therefore, one possible implementation of this application, after obtaining historical memory information by using a multimodal large model for semantic understanding based on visual text information, audio text information, and text instructions, further includes: determining whether the historical memory information is key information; if so, performing the step of storing the historical memory information to a memory map; if not, discarding the historical memory information.

[0120] Each historical memory corresponds to one event. Non-critical information, such as casual conversation or weather inquiries, is pre-set. Furthermore, the system determines whether a memory is critical; only critical information is stored. This allows the system to filter redundant information, retain key memories, and dynamically update nodes, enabling continuous evolution of the memory.

[0121] As can be seen, in this embodiment of the application, a key judgment mechanism is introduced after generating historical memory information, and only information that is determined to be key is stored in the memory graph, avoiding the accumulation of redundant or low-value data; maintaining the high signal-to-noise ratio and high correlation of the memory graph, significantly enhancing the reasoning stability and response quality of the system in long-term interaction.

[0122] One possible implementation of this application further includes: updating the memory graph based on a multimodal large model using update instructions, wherein the update instructions include at least one of the following: key information storage instructions and node update instructions.

[0123] The update instruction is an input prompt for the multimodal large model. This update instruction is used to indicate that you want to focus more on valuable information rather than casual chatter. Furthermore, the update instruction can also instruct the continuous updating of the memory graph based on time-series data. For example, if person A and person B are classmates at one time and spouses at another time, the relationship between nodes can be updated.

[0124] As can be seen, in this embodiment of the application, the dynamic evolution of the memory graph is driven by update instructions (such as key information storage instructions and node update instructions), which enables the system to correct, expand or eliminate existing memory nodes in real time according to new inputs or external feedback, thus ensuring the timeliness and accuracy of the memory graph.

[0125] One possible implementation of this application embodiment is to determine memory information matching the retrieval vector from the memory graph, including: determining whether a second retrieval is needed based on a large language model; if so, determining supplementary memory information matching the retrieval vector from the memory graph; and determining the final memory information matching the retrieval vector based on the large language model according to the memory information and supplementary memory information.

[0126] In this embodiment, the language big data model performs multiple rounds of semantic reasoning and filtering on the retrieval results to construct the final recall context. Multiple rounds of reasoning refer to the fact that a single dialogue may involve multiple retrievals. Because the user's dialogue or question involves multiple rounds of memory retrieval, the language big data model, as the judgment subject, will determine whether the currently retrieved content can resolve the user's doubts. If the retrieved content is insufficient, it will modify the content to be retrieved and perform a second retrieval.

[0127] As can be seen, in this embodiment, a secondary judgment mechanism based on a large language model is introduced after the initial retrieval to dynamically determine whether supplementary retrieval is needed, and to fuse the original memory information and supplementary memory information when necessary to generate the final matching result; thus avoiding reasoning bias caused by missing key context in a single retrieval.

[0128] See Figure 5 Upon receiving a request, the system determines whether to trigger a query request and then converts it into a text vector. It then uses a vector model to perform multimodal retrieval and memory graph query to obtain a candidate memory set, i.e., memory information. The language big model then performs reasoning, conducts multiple rounds of reasoning and filtering, and determines whether supplementary memory information is needed. Finally, the system obtains the memory information and supplementary memory information, which are used as context in turn. Combined with the request, the language big model generates the output result and completes the response.

[0129] Based on any of the above embodiments, a visual memory generation and retrieval method based on a multimodal large model includes: 101: The system continuously receives multimodal input data, including video frames, audio streams, and text commands. 102: The video frames are subjected to object detection, face recognition, scene recognition, and behavior recognition to extract preliminary entities and their attributes (such as people, objects, locations, and actions). 103: The audio stream is subjected to speech recognition (ASR) and voice fingerprint recognition (VAD+SpeakerID) to extract the speaker's identity and speech content. 104: The face ID, voiceprint information, video frames, and input multimodal large model (joint encoding of image + audio + ID information) are used for deep semantic understanding to enhance the accuracy of entity recognition, action understanding, and scene relationship parsing. 105: The multimodal encoding model aligns and integrates image embedding, audio embedding, and text embedding to generate a unified semantic representation vector. 106: Based on time series and scene events, semantically structure the perceived information to generate a "visual event description," including the entities involved in the event (people / objects), the time and location of the event, and the relationships and behaviors between entities. 107: Store the events in an "entity-centric memory graph database," creating or updating nodes for each entity to achieve unified identification across scenes. 108: Generate meta-information for each memory, including timestamps, source fragment indexes, credibility scores, episodic memories, or semantic classification tags. 109: When a user submits a query or the system triggers a task, convert the text request into vector retrieval conditions. 110: Perform multimodal retrieval, searching for the most relevant entities, events, and context fragments from the memory graph, and returning a candidate memory set. 111: The large model performs multiple rounds of semantic reasoning and filtering on the retrieval results to construct the final recall context. 112: Output corresponding answers, decisions, or operational instructions to achieve reasoning and response driven by long-term visual memory.

[0130] By integrating multimodal large models with traditional visual algorithms, long-term structured memory of visual, audio, and text is achieved; memory graphs are constructed with entities at the center to support the storage and dynamic updating of plot and semantic memories; and by combining efficient retrieval and multi-round reasoning, information recall and intelligent decision-making across time and scenarios are realized, thereby significantly improving the continuous cognitive ability of AI and the level of intelligence in application scenarios.

[0131] The following describes a multimodal data processing apparatus provided in an embodiment of this application. The apparatus described below can be referred to in correspondence with the method described above. The apparatus in this embodiment is installed in an electronic device. Figure 6 , Figure 6 This is a structural block diagram of an apparatus according to one embodiment of this application, including: an acquisition module 210 for acquiring a user's request; a generation module 220 for generating a retrieval vector based on the request; a retrieval module 230 for determining memory information matching the retrieval vector from a memory graph; wherein the memory graph is constructed based on historical memory information obtained through semantic understanding of historical visual text information, audio text information, and text instructions; the visual text information is obtained by object detection of video frames in multimodal input data; the audio text information is obtained by speech recognition and voiceprint recognition of audio streams in multimodal input data; and a request processing module 240 for performing language reasoning based on a large language model, according to the memory information and the request, and outputting the request result.

[0132] In one possible implementation, the method further includes: a memory graph construction module, used to: acquire historical multimodal input data, including video frames, audio streams, and text commands; perform object detection on the video frames to obtain visual text information, including entity information and corresponding attribute information; perform speech recognition and voiceprint recognition on the audio stream to obtain audio text information, including speaker identity and speech content; and perform semantic understanding using a multimodal large model based on the visual text information, audio text information, and text commands to obtain historical memory information, and store the historical memory information in the memory graph.

[0133] In one feasible approach, the multimodal large model includes an encoder and a decoder, and a memory map construction module for: performing unified position encoding based on the encoder according to visual text information, audio text information, and text instructions to obtain encoded information; and decoding the encoded information using the decoder to obtain historical memory information.

[0134] In one possible implementation, the memory graph construction module is also used to: determine whether the historical memory information is critical information; if so, perform the step of storing the historical memory information into the memory graph; if not, discard the historical memory information.

[0135] In one possible implementation, it also includes: an update module for updating the memory graph based on a multimodal large model using update instructions, wherein the update instructions include at least one of the following: key information storage instructions and node update instructions.

[0136] In one feasible approach, a memory graph is a graph structure with entities as nodes, and temporal, spatial, and semantic relationships are established between the nodes.

[0137] In one possible implementation, the retrieval module 230 is configured to: determine whether a second retrieval is required; if so, determine supplementary memory information matching the retrieval vector from the memory graph; and, based on the language big model, determine the final memory information matching the retrieval vector according to the memory information and the supplementary memory information.

[0138] Figure 7 A structural diagram of an electronic device provided in an embodiment of the present invention, such as... Figure 7 As shown, the electronic device includes: a memory 60 for storing a computer program; and a processor 61 for executing the computer program to implement the steps of the method as described in the above embodiments.

[0139] The electronic devices provided in this embodiment may include, but are not limited to, smartphones, tablets, laptops, or desktop computers.

[0140] The processor 61 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 61 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 61 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 61 may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.

[0141] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 60 is used to store at least the following computer program 601, which, after being loaded and executed by the processor 61, is capable of implementing the relevant steps of the method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. The operating system 602 may include Windows, Unix, Linux, etc.

[0142] In some embodiments, the electronic device may further include a display screen 62, an input / output interface 63, a communication interface 64, a power supply 65, and a communication bus 66.

[0143] Those skilled in the art will understand that Figure 7 The structures shown do not constitute a limitation on electronic devices and may include more or fewer components than those shown.

[0144] It is understood that if the methods in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the current technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, magnetic disks, or optical disks, and other media capable of storing program code.

[0145] Based on this, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method described above.

[0146] Based on this, embodiments of the present invention also provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described method. It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0147] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application. < / voice1> < / face2> < / face1>

Claims

1. A method for processing multimodal data, characterized in that, include: Obtain the user's request and generate a retrieval vector based on the request; Determine the memory information that matches the retrieval vector from the memory graph; Based on the large language model, language reasoning is performed according to the memory information and the request, and the request result is output. The memory map is constructed based on historical memory information obtained through semantic understanding of historical visual text information, audio text information, and text instructions. The visual text information is obtained by target detection of video frames in multimodal input data, and the audio text information is obtained by speech recognition and voiceprint recognition of audio streams in multimodal input data.

2. The method according to claim 1, characterized in that, Also includes: Acquire historical multimodal input data, including video frames, audio streams, and text commands; Target detection is performed on the video frame to obtain visual text information, which includes entity information and attribute information corresponding to the entity information. Speech recognition and voiceprint recognition are performed on the audio stream to obtain audio text information, which includes the speaker's identity and the speech content; Based on the visual text information, the audio text information, and the text instructions, semantic understanding is performed using a multimodal large model to obtain historical memory information, and the historical memory information is stored in a memory graph.

3. The method according to claim 2, characterized in that, The multimodal large model includes an encoder and a decoder. Based on the visual text information, the audio text information, and the text instructions, semantic understanding is performed using a multimodal large model to obtain memory information, including: Based on the visual text information, the audio text information, and the text instructions, unified position encoding is performed using an encoder to obtain encoded information; The encoded information is decoded using a decoder to obtain historical memory information.

4. The method according to claim 2, characterized in that, Based on the visual text information, the audio text information, and the text instructions, after semantic understanding is performed using a multimodal large model to obtain historical memory information, the process further includes: Determine whether the historical memory information is critical information; If so, then perform the step of storing the historical memory information into the memory graph; If not, then discard the historical memory information.

5. The method according to claim 2, characterized in that, Also includes: Based on the multimodal large model, the memory graph is updated using update instructions, wherein the update instructions include at least one of the following: key information storage instructions and node update instructions.

6. The method according to claim 1, characterized in that, The memory graph is a graph structure with entities as nodes, and temporal, spatial, and semantic relationships are established between the nodes.

7. The method according to claim 1, characterized in that, Determining memory information that matches the retrieval vector from the memory graph includes: Based on the large language model, determine whether a second retrieval is needed; If so, then supplementary memory information matching the retrieval vector is determined from the memory graph; Based on the language big model, the final memory information that matches the retrieval vector is determined according to the memory information and supplementary memory information.

8. A multimodal data processing device, characterized in that, include: The acquisition module is used to acquire user requests; The generation module is used to generate a retrieval vector based on the request; The retrieval module is used to determine memory information that matches the retrieval vector from the memory graph; wherein the memory graph is constructed based on historical memory information obtained by semantic understanding of historical visual text information, audio text information and text instructions, the visual text information is obtained by target detection of video frames in multimodal input data, and the audio text information is obtained by speech recognition and voiceprint recognition of audio streams in multimodal input data; The request processing module is used to perform language reasoning based on the large language model, the memory information, and the request, and output the request result.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the method according to any one of claims 1 to 7 when running the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the method as described in any one of claims 1 to 7.