Intelligent storage method and device for multi-modal data and computer equipment

By transforming multimodal data into structured feature data with high semantic modality and constructing a dynamic labeling system, the problems of information silos and low retrieval efficiency in multimodal data storage are solved, and more efficient data retrieval and feedback are achieved.

CN121722754APending Publication Date: 2026-03-24BEIJING ZHUOYUE CREATIVE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In existing technologies, multimodal data suffers from problems such as information silos, lack of temporal continuity, and single media dimension during storage, resulting in high difficulty and low efficiency in data retrieval and an inability to effectively combine with databases for comprehensive analysis.

Method used

By employing a multimodal data acquisition strategy, each data modality is transformed into structured feature data with high semantic modality, generating summary data, constructing a dynamic tagging system, updating user profiles, and storing data.

Benefits of technology

It improves the semantic representation integrity and retrieval accuracy of data storage, optimizes the relevance of data content and the retrieval feedback effect, and enables the system to understand user intent more accurately and provide better feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121722754A_ABST
    Figure CN121722754A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent storage method and device for multi-modal data and computer equipment. The method comprises the steps of collecting data content of each data mode of each user through a multi-mode data collection strategy, and converting the data content of each data mode of each user into structured feature data of each high semantic mode for each user; generating abstract data of each high semantic mode through a context sensing strategy on the basis of the structured feature data of each high semantic mode, and constructing a dynamic label system of the user on the basis of the abstract data of each high semantic mode; and on the basis of the abstract data of each high semantic mode, through the dynamic label system of the user, performing updating processing on the historical user portrait of the user to obtain a new user portrait of the user, and performing data storage processing on the data to obtain a user database. By adopting the method, the calling feedback effect on the stored data content can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data intelligent storage technology, and in particular to an intelligent storage method, apparatus and computer equipment for multimodal data. Background Technology

[0002] In current mobile devices and PC (personal computer) systems, applications suffer from severe "context fragmentation": (1) Information silos across applications: users' historical operation data (such as document editing snippets, shopping cart records, chat keywords) are scattered across different apps; (2) Lack of temporal continuity: voice memos generated 24 hours ago cannot be retrieved in the email writing scenario at noon the next day; (3) Single media dimension: existing solutions (such as browser history) only support text log storage, while key user needs may exist in a screenshot of a video or a voice command snippet in a call. The above problems make it difficult and inefficient to retrieve stored data content. Therefore, how to improve the intelligent storage effect of data is the current research focus.

[0003] Existing technologies store data of different types in a comprehensive manner. However, the data stored in this way has poor usability, and the correlation information and calling relationships between the data are vague. It is impossible to combine the data with the database to effectively analyze the stored data, resulting in poor retrieval of the stored data. Summary of the Invention

[0004] Therefore, it is necessary to provide an intelligent storage method, device, computer equipment, computer-readable storage medium, and computer program product for multimodal data to address the aforementioned technical problems.

[0005] Firstly, this application provides an intelligent storage method for multimodal data, including:

[0006] Through a multimodal data acquisition strategy, the data content of each user's data modality is collected, and for each user, the data content of each user's data modality is transformed into structured feature data of each high semantic modality.

[0007] Based on the structured feature data of each of the aforementioned high semantic modalities, a context-aware strategy is used to generate summary data for each of the aforementioned high semantic modalities, and based on the summary data of each of the aforementioned high semantic modalities, a dynamic tagging system for the user is constructed.

[0008] Based on the summary data of each of the high semantic modalities, the historical user profile of the user is updated through the user's dynamic tagging system to obtain the user's new user profile. The structured feature data of each user's high semantic modalities, the summary data of each user's high semantic modalities, and the new user profile of each user are stored and processed to obtain the user database.

[0009] Optionally, the step of collecting data content for each user's various data modalities through a multimodal data acquisition strategy includes:

[0010] Acquire raw voice data, raw video data, raw image data, and raw text behavior data, and generate audio data modal data content from the raw voice data through a voice storage strategy;

[0011] The original video data is subjected to video frame extraction processing to obtain video frame data, and the video frame data is adjusted according to the video storage format to obtain the data content of the video data modality.

[0012] The original image data is processed to adjust the image quality according to a preset quality standard to obtain standard image data. Based on each standard image data, data processing is performed according to a preset storage logic to obtain the data content of the image data modality.

[0013] Identify the permission information corresponding to each sub-text behavior data in the original text behavior data, and based on the permission information corresponding to each sub-text behavior data, filter the data content of the text behavior modality through permission control strategies.

[0014] Optionally, for each user, converting the data content of each data modality of the user into structured feature data of each high semantic modality includes:

[0015] The audio data modality is converted into text to obtain the text content of the audio data modality, and the emotional features of the text content are extracted to obtain the structured feature data of the text data modality.

[0016] Extract the first text information corresponding to the data content of the video data modality and the second text information corresponding to the data content of the image data modality, and perform structural enhancement processing on the first text information and the second text information through a context enhancement parsing strategy to obtain the structured text data of the video data modality and the structured text data of the image data modality;

[0017] Based on the structured text data of the video data modality and the structured text data of the image data modality, structured feature data of the visual data modality is generated through a data structured storage strategy.

[0018] Obtain the user's historical user profile, and perform multi-source behavior parsing on the data content of the text behavior modality to obtain the structured behavior fields of the text behavior modality;

[0019] Based on the structured behavior fields and the historical user profiles, structured feature data of behavioral intent modalities are generated through an intent classification model.

[0020] Optionally, the step of generating summary data for each of the high-semantic modalities based on the structured feature data of each of the high-semantic modalities, using a context-aware strategy, includes:

[0021] The structured feature data of each of the high semantic modalities are sorted by timestamp to obtain the synchronous structured feature data of each of the high semantic modalities.

[0022] Based on the synchronous structured feature data of each of the aforementioned high semantic modalities, summary data of each of the aforementioned high semantic modalities is generated through a summary generation model.

[0023] Optionally, the step of constructing the user's dynamic tagging system based on the summary data of each of the high semantic modalities includes:

[0024] Based on the summary data of each of the aforementioned high semantic modalities, the semantic vectors of each of the aforementioned high semantic modalities are obtained through multimodal content vectorization processing;

[0025] The semantic vectors of each of the high semantic modalities are cross-modal aligned to obtain the synchronous semantic space vectors of each of the high semantic modalities.

[0026] Based on the synchronous semantic space vectors of each of the aforementioned high semantic modalities, a dynamic tag system for the user is constructed through dynamic update logic for each tag dimension.

[0027] Optionally, based on the summary data of each of the high semantic modalities, the user's historical user profile is updated using the user's dynamic tagging system to obtain the user's new user profile, including:

[0028] Based on the user's dynamic tag system, a weight value for each tag dimension is generated through a weight allocation strategy;

[0029] Based on the synchronous semantic space vectors of each of the aforementioned high semantic modalities, and through the user's dynamic tag system, the user's historical user profile is updated according to the weight values ​​of each tag dimension to obtain the user's new user profile.

[0030] Secondly, this application also provides an intelligent storage device for multimodal data, comprising:

[0031] The acquisition module is used to acquire the data content of each user's data modality through a multimodal data acquisition strategy, and to convert the data content of each user's data modality into structured feature data of each high semantic modality for each user.

[0032] The construction module is used to generate summary data of each of the high semantic modalities based on the structured feature data of each of the high semantic modalities through a context-aware strategy, and to construct the user's dynamic tag system based on the summary data of each of the high semantic modalities;

[0033] The update module is used to update the historical user profile of the user based on the summary data of each of the high semantic modalities and through the user's dynamic tag system to obtain the new user profile of the user. It also performs data storage processing on the structured feature data of each user's high semantic modalities, the summary data of each user's high semantic modalities, and the new user profile of each user to obtain the user database.

[0034] Optionally, the acquisition module is specifically used for:

[0035] Acquire raw voice data, raw video data, raw image data, and raw text behavior data, and generate audio data modal data content from the raw voice data through a voice storage strategy;

[0036] The original video data is subjected to video frame extraction processing to obtain video frame data, and the video frame data is adjusted according to the video storage format to obtain the data content of the video data modality.

[0037] The original image data is processed to adjust the image quality according to a preset quality standard to obtain standard image data. Based on each standard image data, data processing is performed according to a preset storage logic to obtain the data content of the image data modality.

[0038] Identify the permission information corresponding to each sub-text behavior data in the original text behavior data, and based on the permission information corresponding to each sub-text behavior data, filter the data content of the text behavior modality through permission control strategies.

[0039] Optionally, the acquisition module is specifically used for:

[0040] The audio data modality is converted into text to obtain the text content of the audio data modality, and the emotional features of the text content are extracted to obtain the structured feature data of the text data modality.

[0041] Extract the first text information corresponding to the data content of the video data modality and the second text information corresponding to the data content of the image data modality, and perform structural enhancement processing on the first text information and the second text information through a context enhancement parsing strategy to obtain the structured text data of the video data modality and the structured text data of the image data modality;

[0042] Based on the structured text data of the video data modality and the structured text data of the image data modality, structured feature data of the visual data modality is generated through a data structured storage strategy.

[0043] Obtain the user's historical user profile, and perform multi-source behavior parsing on the data content of the text behavior modality to obtain the structured behavior fields of the text behavior modality;

[0044] Based on the structured behavior fields and the historical user profiles, structured feature data of behavioral intent modalities are generated through an intent classification model.

[0045] Optionally, the building module is specifically used for:

[0046] The structured feature data of each of the high semantic modalities are sorted by timestamp to obtain the synchronous structured feature data of each of the high semantic modalities.

[0047] Based on the synchronous structured feature data of each of the aforementioned high semantic modalities, summary data of each of the aforementioned high semantic modalities is generated through a summary generation model.

[0048] Optionally, the building module is specifically used for:

[0049] Based on the summary data of each of the aforementioned high semantic modalities, the semantic vectors of each of the aforementioned high semantic modalities are obtained through multimodal content vectorization processing;

[0050] The semantic vectors of each of the high semantic modalities are cross-modal aligned to obtain the synchronous semantic space vectors of each of the high semantic modalities.

[0051] Based on the synchronous semantic space vectors of each of the aforementioned high semantic modalities, a dynamic tag system for the user is constructed through dynamic update logic for each tag dimension.

[0052] Optionally, the update module is specifically used for:

[0053] Based on the user's dynamic tag system, a weight value for each tag dimension is generated through a weight allocation strategy;

[0054] Based on the synchronous semantic space vectors of each of the aforementioned high semantic modalities, and through the user's dynamic tag system, the user's historical user profile is updated according to the weight values ​​of each tag dimension to obtain the user's new user profile.

[0055] Thirdly, this application provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described in any one of the first aspects.

[0056] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any one of the first aspects.

[0057] Fifthly, this application provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described in any one of the first aspects.

[0058] The aforementioned intelligent storage method, apparatus, and computer equipment for multimodal data, through a multimodal data acquisition strategy, collects data content from each user's various data modalities. For each user, the data content of each data modality is transformed into structured feature data for each high-semantic modality. Based on the structured feature data of each high-semantic modality, a context-aware strategy is used to generate summary data for each high-semantic modality. Based on the summary data of each high-semantic modality, the user's historical user profile is updated using the dynamic tagging system to obtain a new user profile. The structured feature data of each user's high-semantic modality, the summary data of each user's high-semantic modality, and the new user profile are then stored to obtain a user database. This solution, by first acquiring data content from each data modality through a multimodal data acquisition strategy, can unify the format and optimize the data content of each data modality. This reduces the amount of data stored while ensuring the semantic integrity of the data content, thereby improving the effectiveness of optimized data storage. Then, this solution transforms the data content of each data modality into structured feature data of each high-semantic modality and generates summary data for each high-semantic modality. This highlights the high semantic density of the structured features of each data modality, thereby improving the structured analysis and semantic parsing of the data to be stored. Furthermore, the terminal constructs a dynamic tagging system for users, dynamically updating user profiles, improving the correlation and optimization of data content and retrieval relationships. This effectively enhances the correlation between the stored data content and the user, allowing the system to more accurately understand the user's intent based on context and long-term profile when the user asks ambiguous questions via voice, text, images, or video, providing better feedback. Overall, this improves the accuracy of retrieval of stored data content and the effectiveness of retrieval feedback. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a flowchart illustrating an intelligent storage method for multimodal data in one embodiment;

[0061] Figure 2This is a flowchart illustrating an example of intelligent storage of multimodal data in one embodiment;

[0062] Figure 3 This is a structural block diagram of an intelligent storage device for multimodal data in one embodiment;

[0063] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0065] The intelligent storage method for multimodal data provided in this application can be applied to an intelligent storage system for multimodal data. This system can be applied to a terminal, which can be, but is not limited to, various personal computers, laptops, mid-range computers, etc. The terminal first collects data content from each data modality using a multimodal data acquisition strategy. It can then unify the format and optimize the data content, thereby reducing data storage volume while ensuring the semantic integrity of the data content, thus improving the effectiveness of optimized data storage. Then, this solution transforms the data content of each data modality into structured feature data of each high-semantic modality and generates summary data for each high-semantic modality. This highlights the high semantic density structured features of each data modality, thereby improving the structured analysis and semantic parsing effects of the data to be stored. Furthermore, the terminal constructs a dynamic tagging system for users, thereby dynamically updating user profiles. This improves the correlation between various data contents and optimizes the retrieval relationships, effectively enhancing the connection between the stored data content and the user. Consequently, when a user poses a vague question to the system via voice, text, image, or video, the system can more accurately understand the user's intent based on the context and long-term profile, providing better feedback. This comprehensively improves the accuracy of retrieving stored data content and the effectiveness of retrieval feedback.

[0066] In one exemplary embodiment, such as Figure 1 As shown, a method for intelligent storage of multimodal data is provided. Taking the application of this method to a terminal as an example, the method includes the following steps S101 to S103. Wherein:

[0067] Step S101: Through a multimodal data acquisition strategy, collect the data content of each user's data modality, and for each user, transform the data content of each user's data modality into structured feature data of each high semantic modality.

[0068] In this embodiment, the terminal collects the interaction content between the user and the system in real time, obtaining the data content of each user's data modality. Specifically, when the user interacts with data, the terminal triggers the recording of data content for each data modality through Session definition rules, and uses the Taskd function to acquire and record the interaction content, thereby obtaining the data content of each data modality. This data modality includes, but is not limited to, audio data modality, video data modality, image data modality, and text behavior data modality. The Session definition rules are as follows: Triggering conditions: Explicit trigger: User actively initiates a task (e.g., voice command "Start recording my needs"). Implicit trigger: Continuous inactivity timeout (e.g., session ends if there is no interaction within 5 minutes after the screen is off). Spatiotemporal boundaries: Time constraint: Event interval within the same Session ≤ 30 minutes (configurable). Spatial constraint: Based on the device's GPS positioning, if the user exceeds a preset geofence (e.g., within a 10-kilometer range), a new Session is forcibly created. The Taskd functionality includes: Task-level tracking: Each user request (such as voice input or click operation) and its corresponding response (system feedback) form an independent interaction unit (Task), assigned a unique TaskId. Enhanced contextual relevance: Supports cross-modal data (such as image upload triggered after a user mentions "I'll send you a screenshot" in a voice conversation) by binding semantic links through TaskId. Generation timing: When a user initiates an interaction: Whether it's voice input, image upload, or a behavioral operation, the system instantly generates a TaskId at the entry layer that receives the Request. Then, the terminal converts the data content of each data modality into high semantic density structured features, obtaining structured feature data for each high semantic modality. These high semantic modalities include, but are not limited to, text data modalities, visual data modalities, and behavioral intent modalities. The specific conversion process will be explained in detail later.

[0069] Step S102: Based on the structured feature data of each high semantic modality, a summary data of each high semantic modality is generated through a context-aware strategy, and a dynamic tagging system for users is constructed based on the summary data of each high semantic modality.

[0070] In this embodiment, the terminal generates summary data for each high-semantic modality based on the structured feature data of each high-semantic modality using a context-aware strategy. Based on this summary data, a dynamic tagging system for the user is constructed. The summary data for each high-semantic modality is used to represent key data content that interacts with and is associated with the user. Furthermore, this dynamic tagging system includes sub-tag systems for each tag dimension. These tag dimensions include, but are not limited to, tag systems for basic attributes (static tags), long-term preferences (low-frequency updates), short-term interests (dynamic decay), and goals and plans (task-driven). The specific construction process will be explained in detail later.

[0071] Step S103: Based on the summary data of each high semantic modality, the historical user profile of the user is updated through the user's dynamic tag system to obtain the new user profile of the user. The structured feature data of each high semantic modality of each user, the summary data of each high semantic modality of each user, and the new user profile of each user are stored and processed to obtain the user database.

[0072] In this embodiment, the terminal updates the user's historical user profile based on the summary data of each high semantic modality and through the user's dynamic tagging system to obtain the user's new user profile. It then performs data storage processing on the structured feature data of each user's high semantic modality, the summary data of each user's high semantic modality, and the new user profile of each user to obtain a user database. The specific update process will be described in detail later.

[0073] Based on the above scheme, firstly, by employing a multimodal data acquisition strategy, data content from each data modality is collected. This allows for format unification and data optimization, reducing data storage volume while ensuring the semantic integrity of the data content, thus improving the effectiveness of optimized data storage. Secondly, this scheme transforms the data content from each data modality into structured feature data for each high-semantic modality and generates summary data for each high-semantic modality. This highlights the high semantic density of the structured features of each data modality, thereby improving the structured analysis and semantic parsing of the data to be stored. Thirdly, the terminal constructs a dynamic user tagging system to dynamically update user profiles, improving the correlation and optimization of data content and retrieval relationships. This effectively enhances the correlation between the stored data content and the user, allowing the system to more accurately understand the user's intent based on context and long-term profile when the user poses ambiguous questions via voice, text, images, or video, providing better feedback. This comprehensively improves the accuracy of retrieval of stored data content and the effectiveness of retrieval feedback.

[0074] Optionally, a multimodal data acquisition strategy is used to collect data content for each user across various data modalities, including: acquiring raw voice data, raw video data, raw image data, and raw text behavior data; generating audio data modal content from the raw voice data using a voice storage strategy; performing frame extraction on the raw video data to obtain individual video frame data, and adjusting the format of each video frame data according to the video storage format to obtain video data modal content; adjusting the image quality of the raw image data according to a preset quality standard to obtain standard image data, and processing the standard image data according to a preset storage logic to obtain image data modal content; identifying the permission information corresponding to each sub-text behavior data in the raw text behavior data, and filtering the text behavior modal data content based on the permission information corresponding to each sub-text behavior data using a permission control strategy.

[0075] In this embodiment, the terminal acquires raw voice data, raw video data, raw image data, and raw text behavior data. It then uses the raw voice data to generate audio data modal content through a voice storage strategy. The technical parameters of the raw voice data are as follows: audio format: PCM (Pulse Code Modulation); sampling rate: 16kHz (balancing sound quality and storage efficiency); channel: mono (MONO) (reducing redundant data processing); bit depth: 16bit; data source: real-time input from the device microphone. The voice storage and recording process involves storing the raw voice data in blocks according to Session ID and TaksId as .wav files (filenames such as 093015250.wav), with the sampling rate and bit depth strictly adhering to the aforementioned parameters.

[0076] The terminal performs frame extraction on the raw video data to obtain individual video frame data. Each video frame is then formatted according to the video storage format to obtain the video data modality content. The frame extraction process uses the following strategy: Frequency: 1 frame / second continuous sampling (balancing dynamic information coverage and storage overhead). Resolution: Retains the original video resolution (720p / 1080p / 4K, dynamically adapted according to device capabilities). Keyframe priority: If the video encoding contains I-frames (keyframes), I-frames are extracted first to reduce decoding power consumption. Data stream: Video source: Real-time shooting from the user's camera, screen recording (authorization required). Preprocessing: The video stream is decomposed using FFmpeg and frame images are extracted (executed every second using `ffmpeg -iinput.mp4 -vf fps=1 frame_%04d.jpg`). The video storage format is: single frames are stored in JPEG format (80% compression ratio), with filename timestamps accurate to milliseconds (e.g., v_093015250.jpg).

[0077] The terminal processes the raw image data according to preset quality standards to obtain standard image data. Based on these standard image data, it performs data processing according to preset storage logic to obtain the data content of the image data modality. Specifically, the raw image data is collected by the user through active upload (e.g., selection from album, asking a question after taking a picture). Initially, a pop-up authorization window is required ("Allow the system to analyze this image to optimize the service"). Quality standards: Minimum resolution: 640×480 pixels (to avoid low-quality images affecting the OCR / visual model). Maximum file size: 5MB (if exceeded, gradient compression is triggered; the compression algorithm is WebP). The storage logic is as follows: sessionId and TaskId are stored in separate directories (path example: sessionId / TaskId / ), and metadata records the initial screening results of image semantic tags (e.g., image_text generated by the CLIP model: "close-up of a coffee cup").

[0078] The terminal identifies the permission information corresponding to each sub-text behavior data in the raw text behavior data, and based on the permission information corresponding to each sub-text behavior data, filters the data content of the text behavior modality through permission control policies. Specifically, the sources of raw text behavior data are: Text input: Sources include keyboard input (including soft keyboard and physical keyboard), OCR recognition results (such as scanned documents), and clipboard content (explicit authorization required). Specification: Encoding is uniformly UTF-8. Behavior logs: Data interface: Integrates third-party service APIs (such as Uber trip API, Meituan Waimai order API) through SDK. The permission control policy follows the principle of least privilege, collecting only user-authorized operations related to the current session.

[0079] Based on the above scheme, the data content of each data modality can be formatted and optimized, thereby reducing the amount of data storage while ensuring the semantic integrity of the data content, thus improving the effect of optimized data storage.

[0080] Optionally, for each user, the data content of each data modality is transformed into structured feature data of each high semantic modality, including: performing text conversion processing on the data content of the audio data modality to obtain the text content of the audio data modality, and performing sentiment feature extraction processing on the text content to obtain structured feature data of the text data modality; extracting the first text information corresponding to the data content of the video data modality and the second text information corresponding to the data content of the image data modality, and performing structure enhancement processing on the first text information and the second text information through a context enhancement parsing strategy to obtain structured text data of the video data modality and structured text data of the image data modality; generating structured feature data of the visual data modality and structured feature data of the image data modality based on the structured text data modality through a data structured storage strategy; obtaining the user's historical user profile, and performing multi-source behavior parsing processing on the data content of the text behavior modality to obtain structured behavior fields of the text behavior modality; and generating structured feature data of the behavior intent modality based on the structured behavior fields and the historical user profile through an intent classification model.

[0081] In this embodiment, the terminal performs text conversion processing on the audio data modality to obtain the text content of the audio data modality, and then performs sentiment feature extraction processing on the text content to obtain structured feature data of the text data modality. Specifically, the terminal first processes the audio data modality data content through audio processing strategies, thereby improving the comprehensiveness of the audio data modality data representation. The specific audio processing strategies are as follows: Frame segmentation and noise reduction: using WebRTC's 3A preprocessing of the speech signal. Voice activity detection (VAD): segmenting effective speech segments (minimum length 150ms) based on Silo-VAD. Speaker verification (Class Activation Mapping, CAM++): filtering out the voices of other speakers. Personalized hot word enhancement: extracting a preference word list (such as "oat milk" and "noise-canceling headphones") from the user's L2 profile, dynamically increasing the recognition weight of ASR (Automatic Speech Recognition). Then, the terminal uses the processed data content and ASR technology to recognize the text content of the audio data modality. Then, the terminal uses a text sentiment extraction model to extract positive / neutral / negative sentiment text data and the confidence level of the sentiment text data from the text content. The terminal then binds the positive / neutral / negative sentiment text data and the confidence level of the sentiment text data to TaskId and SessionId, and stores them in a session text feature library, obtaining structured feature data of the text data modality.

[0082] The terminal extracts first text information corresponding to the data content of the video data modality and second text information corresponding to the data content of the image data modality. Then, using a context-enhanced parsing strategy, it performs structural enhancement processing on the first and second text information to obtain structured text data for both the video and image data modalities. Specifically, the terminal first extracts image text information from each video frame in the video data modality or each image in the image data modality, such as "coffee cup," "laptop," "meeting room," "neutral background," and "daylight." Then, the terminal uses a context-enhanced parsing model to generate the first and second text information based on all associated text within the session (ASR results, historical messages), device sensor data (GPS, time), and the aforementioned image text information. This context-enhanced parsing model is a multimodal large model that adds contextual information.

[0083] The text information is, for example:

[0084] - Session time: 2023-10-15 09:30 (weekend morning);

[0085] - Recent user interests: Startup funding, coffee culture;

[0086] - Current dialogue excerpt: "Is this coffee shop suitable for a team meeting?"

[0087] Then, the terminal performs semantic parsing on the above structured text data to obtain parsed structured text data that is further analyzed and understood in conjunction with contextual information. For example, this structured text data includes:

[0088] Image shows:

[0089] - Three open MacBook Pros are placed on the central wooden table, displaying business plan PowerPoint presentations.

[0090] - On the right is a white ceramic coffee cup with the brand "Blue Bottle Coffee" printed on the coaster.

[0091] - The city skyline visible in the background, outside the glass window, is presumably in the city center business district.

[0092] Based on the context, it can be inferred that the user is holding a startup team meeting in a coffee shop.

[0093] Then, based on the structured text data of the video data modality and the structured text data of the image data modality, the terminal generates structured feature data of the visual data modality through a data structured storage strategy. Specifically, the data structured storage strategy involves storing the aforementioned structured text data at the Task granularity to obtain the structured feature data of the visual data modality.

[0094] The terminal acquires the user's historical user profile and performs multi-source behavior parsing on the text behavior modality data content to obtain structured behavior fields for the text behavior modality. This multi-source behavior parsing method includes standardized logic: API log parsing: providing a unified adapter for different service providers; cross-task behavior concatenation to obtain various structured behavior fields. Examples of these structured behavior fields include: for ride-hailing applications, "company → airport"; for coffee shops and restaurants, "oatmeal latte".

[0095] Then, based on structured behavioral fields and historical user profiles, the terminal generates structured feature data of behavioral intent modalities through an intent classification model. This intent classification model is a pre-trained BERT model (Bidirectional Encoder Representation from Transformers, a language representation model) fine-tuned based on the user's desired services. This structured feature data can be as follows: for example, if the structured behavioral field is a ride-hailing application, the structured feature data for "company → airport" is "business travel"; if the structured behavioral field is a coffee shop / restaurant, the structured feature data for "oatmeal latte" is "restaurant-coffee preference".

[0096] Based on the above scheme, structured feature data of each high semantic modality is obtained by structuring the data content of different data modalities, thereby improving the accuracy of structured analysis and semantic analysis of data content.

[0097] Optionally, based on the structured feature data of each high semantic modality, a context-aware strategy is used to generate summary data for each high semantic modality, including: sorting the structured feature data of each high semantic modality by timestamp to obtain synchronous structured feature data for each high semantic modality; and generating summary data for each high semantic modality based on the synchronous structured feature data of each high semantic modality through a summary generation model.

[0098] In this embodiment, the terminal timestamps and sorts the structured feature data of each high-semantic modality to obtain synchronized structured feature data for each high-semantic modality. Then, based on the synchronized structured feature data of each high-semantic modality, the terminal generates summary data for each high-semantic modality using a summary generation model. This summary generation model is based on the LangChain application framework and GPT-4-Turbo language model designed by our team. The generation method of the above summary generation model is as follows:

[0099] You are a professional user assistant and need to generate a summary of today's conversations based on the following timeline:

[0100] [Timestamp 09:30] User's voice: "I need to book a flight to Shanghai for tomorrow" (emotion: urgent).

[0101] [Timestamp 09:35] The screenshot shows information for China Eastern Airlines flight MU123.

[0102] [Related Behavior] Completed flight payment via Ctrip App at 10:00.

[0103] Please summarize the user's core needs and urgency, and infer potential intentions by combining long-term profile (occupation: entrepreneur).

[0104] Then, the terminal iteratively generates summary data through a dynamic update mechanism to obtain summary data for each high semantic modality. This dynamic update mechanism is as follows:

[0105] Incremental update: Whenever a new task completes the summary data extraction task, summary completion is triggered (similar to Git version merging).

[0106] Version control: Stores historical versions of the summary and supports backtracking differences (such as changes in user intent).

[0107] The abstract data is as follows:

[0108] Core requirement: Booking a flight to Shanghai on October 16 (related to business travel);

[0109] Urgency level: High (Sentiment analysis: Urgent + Multiple follow-up questions);

[0110] Potential Intent: May be related to the startup's fundraising roadshow plan (refer to "3-month fundraising goal" in the L2 profile).

[0111] Based on the above scheme, by extracting summary data from structured data, it is possible to further optimize the extraction and identification of the relationship between data content and users, while improving the optimization and simplification effect of each data content and establishing the relationship between users and data content.

[0112] Optionally, based on the summary data of each high semantic modality, a dynamic tagging system for users is constructed, including: obtaining semantic vectors for each high semantic modality through multimodal content vectorization processing based on the summary data of each high semantic modality; performing cross-modal alignment processing on the semantic vectors of each high semantic modality to obtain synchronous semantic space vectors for each high semantic modality; and constructing a dynamic tagging system for users based on the synchronous semantic space vectors of each high semantic modality through dynamic update logic for each tag dimension.

[0113] In this embodiment, the terminal obtains semantic vectors for each high-semantic modality based on the summary data of each high-semantic modality through multimodal content vectorization processing. Specifically, the semantically resonant vectors are generated through vectorization processing using a semantic embedding model.

[0114] Specifically:

[0115] Text data modalities, Sentence-BERT (all-mpnet-base-v2) generates 768-dimensional vectors;

[0116] Image data modality, CLIP-ViT image vector (512-dimensional).

[0117] Behavioral intent modality, one-hot encoding + TF-IDF weighting (e.g., "taxi: 0.7, business trip: 0.3").

[0118] The terminal performs cross-modal alignment processing on the semantic vectors of each high-semantic modality to obtain the synchronized semantic space vectors of each high-semantic modality. Specifically, the terminal generates synchronized semantic space vectors through a cross-modal alignment layer (such as contrastive learning in CLIP), which are the semantic vectors of each high-semantic modality after time synchronization.

[0119] Next, based on the synchronized semantic space vectors of each high-semantic modality, the terminal constructs a dynamic tag system for the user through dynamic update logic for each tag dimension. Specifically, the dynamic update logic for each tag dimension is as follows:

[0120] (1) Basic attributes (static tags):

[0121] Update conditions: triggered upon first registration or explicit user input; subsequent updates are only allowed through manual modification.

[0122] Inference method:

[0123] Gender: Speech analysis: Based on the Praat toolkit, the fundamental frequency (F0) distribution was extracted, with the mean for males ≈120Hz and for females ≈220Hz; Text analysis: Pronoun patterns ("my husband / girlfriend"), name segmentation (gender bias of Chinese surnames).

[0124] Occupation: Keyword matching: LinkedIn-style email (john.doe@ibm.com → "Computer industry"); Time pattern: Frequent conversations between 9-18 on weekdays → Inferring full-time worker.

[0125] 2) Long-term preferences (low-frequency updates):

[0126] Update frequency: Full data clustering analysis is triggered every 30 days;

[0127] Technical process:

[0128] 1. Data sampling: Extract high-confidence features from all historical L1 summaries (such as "coffee" mentions with confidence > 0.8).

[0129] 2. Hierarchical clustering: Algorithm: OPTICS (noise-resistant, suitable for unevenly dense data).

[0130] a. Features: Semantic vector + behavior frequency (e.g., number of coffee consumptions per month).

[0131] 3. Category generation: Core label: The central semantic meaning of each cluster (e.g., Cluster1 → "coffee culture").

[0132] a. Conflict resolution: If a user simultaneously expresses "buying coffee beans" (positive) and "insomnia complaints" (negative), manual review will be triggered.

[0133] (3) Short-term interest (dynamic decay):

[0134] Update rules: Real-time incremental update + time decay.

[0135] Mathematical model:

[0136] Decay function: exponential decay : Decay coefficient (short-term label λ=0.1 / hour, daily decay approximately 65%); Δt: Time difference from the last relevant event.

[0137] Reinforcement mechanism: Score accumulates each time a related event occurs. .

[0138] (4) Goals and Plans (Task-Driven):

[0139] Recognition techniques: Time-sensitive word extraction: Regular expression matching such as "within three months" and "before the end of next month"; Dependency parsing: Extracting the target predicate structure based on the Spacy model.

[0140] Update Logic: Conflict Detection: When a new goal conflicts with an existing planned time (e.g., both tasks need to be completed "next month"), a reminder is triggered: "Time conflict detected: Fundraising roadshow and family trip are both planned for November, should they be prioritized and adjusted?". Progress Tracking: Update the completion status of related behavior logs (e.g., document editing events).

[0141] Based on the above scheme, a dynamic tagging system with time sensitivity and context awareness is constructed based on the acquired multimodal data, which improves the comprehensiveness and accuracy of updating user profiles.

[0142] Optionally, based on the summary data of each high semantic modality, the user's historical user profile is updated through the user's dynamic tagging system to obtain the user's new user profile. This includes: generating weight values ​​for each tag dimension based on the user's dynamic tagging system through a weight allocation strategy; and updating the user's historical user profile according to the weight values ​​of each tag dimension based on the synchronous semantic space vectors of each high semantic modality through the user's dynamic tagging system to obtain the user's new user profile.

[0143] In this embodiment, the terminal generates weight values ​​for each tag dimension based on the user's dynamic tag system and through a weight allocation strategy. This weight allocation strategy dynamically adjusts the weight values ​​for each tag dimension using a weight allocation model, which is as follows:

[0144] Model architecture: DQN (Deep Q-Network) dynamically adjusts label influence;

[0145] State space: The user's current context (time, location, activity);

[0146] Action space: The adjustment range of each label weight (e.g., long-term preference +0.1);

[0147] Reward function: based on user feedback (clicking / ignoring recommendations);

[0148] Training data: 5 million <state, action, reward> triples from historical interaction logs.

[0149] Based on the synchronous semantic space vectors of various high-semantic modalities, and through the user's dynamic tagging system, the historical user profile is updated according to the weight values ​​of each tag dimension to obtain the user's new user profile. When there are priority conflicts in the order of data updates, the terminal updates sequentially according to preset priority rules, which are: time-sensitive priority: short-term interests > long-term preferences; transactional priority: goals and plans > ordinary behaviors.

[0150] This application also provides an example of intelligent storage for multimodal data, such as... Figure 2 As shown, the specific processing procedure includes the following steps:

[0151] Step S201: Obtain raw voice data, raw video data, raw image data, and raw text behavior data, and generate audio data modal data content from the raw voice data through a voice storage strategy.

[0152] Step S202: Perform video frame extraction on the original video data to obtain each video frame data, and adjust the format of each video frame data according to the video storage format to obtain the data content of the video data modality.

[0153] Step S203: The original image data is processed to adjust the image quality according to a preset quality standard to obtain standard image data. Based on the standard image data, data processing is performed according to a preset storage logic to obtain the data content of the image data modality.

[0154] Step S204: Identify the permission information corresponding to each sub-text behavior data in the original text behavior data, and based on the permission information corresponding to each sub-text behavior data, filter the data content of the text behavior modality through permission control strategy.

[0155] Step S205: The audio data modality data content is converted into text to obtain the text content of the audio data modality, and the emotional features of the text content are extracted to obtain the structured feature data of the text data modality.

[0156] Step S206: Extract the first text information corresponding to the data content of the video data modality and the second text information corresponding to the data content of the image data modality. Then, perform structural enhancement processing on the first text information and the second text information through a context enhancement parsing strategy to obtain the structured text data of the video data modality and the structured text data of the image data modality.

[0157] Step S207: Based on the structured text data of the video data modality and the structured text data of the image data modality, structured feature data of the visual data modality is generated through a data structured storage strategy.

[0158] Step S208: Obtain the user's historical user profile and perform multi-source behavior parsing on the text behavior modality data content to obtain the structured behavior fields of the text behavior modality.

[0159] Step S209: Based on structured behavioral fields and historical user profiles, structured feature data of behavioral intent modalities are generated through an intent classification model.

[0160] Step S210: The structured feature data of each high semantic modality is sorted by timestamp to obtain the synchronous structured feature data of each high semantic modality.

[0161] Step S211: Based on the synchronous structured feature data of each high semantic modality, the summary data of each high semantic modality is generated through the summary generation model.

[0162] Step S212: Based on the summary data of each high semantic modality, the semantic vectors of each high semantic modality are obtained through multimodal content vectorization processing.

[0163] Step S213: Perform cross-modal alignment processing on the semantic vectors of each high semantic modality to obtain the synchronous semantic space vectors of each high semantic modality.

[0164] Step S214: Based on the synchronous semantic space vectors of each high semantic modality, construct the user's dynamic tag system through the dynamic update logic of each tag dimension.

[0165] Step S215: Based on the user's dynamic tag system, generate weight values ​​for each tag dimension through a weight allocation strategy.

[0166] Step S216: Based on the synchronous semantic space vectors of each high semantic modality, the user's historical user profile is updated according to the weight values ​​of each tag dimension through the user's dynamic tag system to obtain the user's new user profile.

[0167] Step S217: Data storage processing is performed on the structured feature data of each user's high semantic modalities, the summary data of each user's high semantic modalities, and the new user profile of each user to obtain the user database.

[0168] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0169] Based on the same inventive concept, this application also provides an intelligent storage device for multimodal data to implement the intelligent storage method for multimodal data described above. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations in one or more embodiments of the intelligent storage device for multimodal data provided below can be found in the limitations of the intelligent storage method for multimodal data described above, and will not be repeated here.

[0170] In one exemplary embodiment, such as Figure 3 As shown, an intelligent storage device for multimodal data is provided, comprising: an acquisition module 310, a construction module 320, and an update module 330, wherein:

[0171] The acquisition module 310 is used to acquire the data content of each user's data modality through a multimodal data acquisition strategy, and for each user, convert the data content of each user's data modality into structured feature data of each high semantic modality.

[0172] The construction module 320 is used to generate summary data of each of the high semantic modalities based on the structured feature data of each of the high semantic modalities through a context-aware strategy, and to construct the user's dynamic tag system based on the summary data of each of the high semantic modalities.

[0173] The update module 330 is used to update the historical user profile of the user based on the summary data of each of the high semantic modalities and through the user's dynamic tag system to obtain the new user profile of the user. It also performs data storage processing on the structured feature data of each user's high semantic modalities, the summary data of each user's high semantic modalities, and the new user profile of each user to obtain the user database.

[0174] Optionally, the acquisition module 310 is specifically used for:

[0175] Acquire raw voice data, raw video data, raw image data, and raw text behavior data, and generate audio data modal data content from the raw voice data through a voice storage strategy;

[0176] The original video data is subjected to video frame extraction processing to obtain video frame data, and the video frame data is adjusted according to the video storage format to obtain the data content of the video data modality.

[0177] The original image data is processed to adjust the image quality according to a preset quality standard to obtain standard image data. Based on each standard image data, data processing is performed according to a preset storage logic to obtain the data content of the image data modality.

[0178] Identify the permission information corresponding to each sub-text behavior data in the original text behavior data, and based on the permission information corresponding to each sub-text behavior data, filter the data content of the text behavior modality through permission control strategies.

[0179] Optionally, the acquisition module 310 is specifically used for:

[0180] The audio data modality is converted into text to obtain the text content of the audio data modality, and the emotional features of the text content are extracted to obtain the structured feature data of the text data modality.

[0181] Extract the first text information corresponding to the data content of the video data modality and the second text information corresponding to the data content of the image data modality, and perform structural enhancement processing on the first text information and the second text information through a context enhancement parsing strategy to obtain the structured text data of the video data modality and the structured text data of the image data modality;

[0182] Based on the structured text data of the video data modality and the structured text data of the image data modality, structured feature data of the visual data modality is generated through a data structured storage strategy.

[0183] Obtain the user's historical user profile, and perform multi-source behavior parsing on the data content of the text behavior modality to obtain the structured behavior fields of the text behavior modality;

[0184] Based on the structured behavior fields and the historical user profiles, structured feature data of behavioral intent modalities are generated through an intent classification model.

[0185] Optionally, the building module 320 is specifically used for:

[0186] The structured feature data of each of the high semantic modalities are sorted by timestamp to obtain the synchronous structured feature data of each of the high semantic modalities.

[0187] Based on the synchronous structured feature data of each of the aforementioned high semantic modalities, summary data of each of the aforementioned high semantic modalities is generated through a summary generation model.

[0188] Optionally, the building module 320 is specifically used for:

[0189] Based on the summary data of each of the aforementioned high semantic modalities, the semantic vectors of each of the aforementioned high semantic modalities are obtained through multimodal content vectorization processing;

[0190] The semantic vectors of each of the high semantic modalities are cross-modal aligned to obtain the synchronous semantic space vectors of each of the high semantic modalities.

[0191] Based on the synchronous semantic space vectors of each of the aforementioned high semantic modalities, a dynamic tag system for the user is constructed through dynamic update logic for each tag dimension.

[0192] Optionally, the update module 330 is specifically used for:

[0193] Based on the user's dynamic tag system, a weight value for each tag dimension is generated through a weight allocation strategy;

[0194] Based on the synchronous semantic space vectors of each of the aforementioned high semantic modalities, and through the user's dynamic tag system, the user's historical user profile is updated according to the weight values ​​of each tag dimension to obtain the user's new user profile.

[0195] Each module in the aforementioned intelligent storage device for multimodal data can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0196] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a multimodal data intelligent storage method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0197] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0198] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of an intelligent storage method for multimodal data.

[0199] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the steps of which, when executed by a processor, implement a method for intelligent storage of multimodal data.

[0200] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of an intelligent storage method for multimodal data.

[0201] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0202] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0203] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0204] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for intelligent storage of multimodal data, characterized in that, The method includes: Through a multimodal data acquisition strategy, the data content of each user's data modal is collected, and for each user, the data content of each user's data modal is transformed into structured feature data of each high semantic modality. Based on the structured feature data of each of the aforementioned high semantic modalities, a context-aware strategy is used to generate summary data for each of the aforementioned high semantic modalities, and based on the summary data of each of the aforementioned high semantic modalities, a dynamic tagging system for the user is constructed. Based on the summary data of each of the high semantic modalities, the historical user profile of the user is updated through the user's dynamic tagging system to obtain the user's new user profile. The structured feature data of each user's high semantic modalities, the summary data of each user's high semantic modalities, and the new user profile of each user are stored and processed to obtain the user database.

2. The method according to claim 1, characterized in that, The multimodal data acquisition strategy collects data content for each user across various data modalities, including: Acquire raw voice data, raw video data, raw image data, and raw text behavior data, and generate audio data modal data content from the raw voice data through a voice storage strategy; The original video data is subjected to video frame extraction processing to obtain video frame data, and the video frame data is adjusted according to the video storage format to obtain the data content of the video data modality. The original image data is processed to adjust the image quality according to a preset quality standard to obtain standard image data. Based on each standard image data, data processing is performed according to a preset storage logic to obtain the data content of the image data modality. Identify the permission information corresponding to each sub-text behavior data in the original text behavior data, and based on the permission information corresponding to each sub-text behavior data, filter the data content of the text behavior modality through permission control strategies.

3. The method according to claim 2, characterized in that, For each user, the process of converting the data content of each data modality of the user into structured feature data of each high semantic modality includes: The audio data modality is converted into text to obtain the text content of the audio data modality, and the emotional features of the text content are extracted to obtain the structured feature data of the text data modality. Extract the first text information corresponding to the data content of the video data modality and the second text information corresponding to the data content of the image data modality, and perform structural enhancement processing on the first text information and the second text information through a context enhancement parsing strategy to obtain the structured text data of the video data modality and the structured text data of the image data modality; Based on the structured text data of the video data modality and the structured text data of the image data modality, structured feature data of the visual data modality is generated through a data structured storage strategy. Obtain the user's historical user profile, and perform multi-source behavior parsing on the data content of the text behavior modality to obtain the structured behavior fields of the text behavior modality; Based on the structured behavior fields and the historical user profiles, structured feature data of behavioral intent modalities are generated through an intent classification model.

4. The method according to claim 1, characterized in that, The structured feature data based on each of the high semantic modalities, through a context-aware strategy, generates summary data for each of the high semantic modalities, including: The structured feature data of each of the high semantic modalities are sorted by timestamp to obtain the synchronous structured feature data of each of the high semantic modalities. Based on the synchronous structured feature data of each of the aforementioned high semantic modalities, summary data of each of the aforementioned high semantic modalities is generated through a summary generation model.

5. The method according to claim 1, characterized in that, The construction of the user's dynamic tagging system based on the summary data of each of the aforementioned high semantic modalities includes: Based on the summary data of each of the aforementioned high semantic modalities, the semantic vectors of each of the aforementioned high semantic modalities are obtained through multimodal content vectorization processing; The semantic vectors of each of the high semantic modalities are cross-modal aligned to obtain the synchronous semantic space vectors of each of the high semantic modalities. Based on the synchronous semantic space vectors of each of the aforementioned high semantic modalities, a dynamic tag system for the user is constructed through dynamic update logic for each tag dimension.

6. The method according to claim 5, characterized in that, The summary data based on each of the aforementioned high semantic modalities, through the user's dynamic tagging system, updates the user's historical user profile to obtain the user's new user profile, including: Based on the user's dynamic tag system, a weight value for each tag dimension is generated through a weight allocation strategy; Based on the synchronous semantic space vectors of each of the aforementioned high semantic modalities, and through the user's dynamic tag system, the user's historical user profile is updated according to the weight values ​​of each tag dimension to obtain the user's new user profile.

7. An intelligent storage device for multimodal data, characterized in that, The device includes: The acquisition module is used to acquire the data content of each user's data modality through a multimodal data acquisition strategy, and to convert the data content of each user's data modality into structured feature data of each high semantic modality for each user. The construction module is used to generate summary data of each of the high semantic modalities based on the structured feature data of each of the high semantic modalities through a context-aware strategy, and to construct the user's dynamic tag system based on the summary data of each of the high semantic modalities; The update module is used to update the historical user profile of the user based on the summary data of each of the high semantic modalities and through the user's dynamic tag system to obtain the new user profile of the user. It also performs data storage processing on the structured feature data of each user's high semantic modalities, the summary data of each user's high semantic modalities, and the new user profile of each user to obtain the user database.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.