Blind person information memory enhancement system based on multi-modal vector database

Through the blind information memory enhancement system based on multimodal vector database, multiple technical problems of blind auxiliary equipment in the fields of contextual memory enhancement and life assistance are solved, efficient multimodal memory management and user privacy security are achieved, and the quality of life of blind people is improved.

CN120029464APending Publication Date: 2025-05-23LINKER
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510357187.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing blind auxiliary equipment has problems such as insufficient cross-modal memory association capabilities, separation of spatial information memory and physical information storage, difficulty in supporting unstructured multimodal data association retrieval in the fields of contextual memory enhancement and life assistance, and insufficient user privacy and security.

Method used

The blind information memory enhancement system based on a multimodal vector database is adopted. A variety of sensory information (voice, vision, space, text) is collected through the user terminal, and the information is encoded by the unified encoder, the vectorized database dynamically adjusts the weight of the modal data, and the adaptive interaction engine performs multimodal information fusion and feedback. The memory enhancement algorithm layer provides natural language query and context understanding.

Benefits of technology

It realizes efficient multimodal memory management, enhances the contextual memory ability of blind people, simplifies life assistance functions, ensures the privacy and security of user data, and improves the quality of life of blind people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029464A_ABST
    Figure CN120029464A_ABST
Patent Text Reader

Abstract

The invention discloses a blind person information memory enhancement system based on a multi-modal vector database, and the system comprises a user terminal which carries out the man-machine interaction with a user, recognizes a voice instruction, collects a visual scene, and broadcasts the memory content; the multi-mode acquisition layer is used for acquiring and processing voice, vision, space and text information; the unified encoder is used for encoding the acquired four types of information; the vectorization database is used for dynamically adjusting modal data weight coefficients based on situations and historical records and providing similarity search; the self-adaptive interaction engine is used for carrying out multi-modal information fusion and generating a result and feedback; and the memory enhancement algorithm layer provides natural language query. The beneficial effects of the invention lie in that through cooperative work of all the modules, all-directional information interaction and memory enhancement functions are provided for the blind, the blind is helped to obtain information more conveniently, and life convenience and independence are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of barrier-free assistance technology, and more particularly to an information memory enhancement system for the blind based on a multimodal vector database. Background Art

[0002] With the development of artificial intelligence technology, assistive devices for the blind have evolved from traditional guide sticks to smart navigation glasses, voice assistants and other products. Among these products, some will further optimize product performance by learning user preferences, thereby providing a personalized interactive experience. For example, patent number 202410656932.5, named a multimodal interactive intelligent control system, is a system that sets up an intelligent modal coordination module and uses the module to learn user preferences based on user interaction history and real-time user behavior data. However, there are still significant technical gaps in the application of the above technology in the field of situational memory enhancement and life assistance for the blind:

[0003] 1. Lack of cross-modal memory association ability, unable to trigger the memory of similar information of different senses through sensory features, such as triggering historical visual memory through voice and text.

[0004] 2. Spatial information memory is separated from physical information storage, which leads to difficulties in situational recall.

[0005] 3. Traditional database structures are difficult to support associative retrieval of unstructured multimodal data.

[0006] 4. Memory storage involves user privacy issues. Using vector storage can convert raw data such as voice and pictures into vector form, reducing the risk of sensitive data leakage. Summary of the invention

[0007] In view of the deficiencies in the prior art, the object of the present invention is to provide a blind memory enhancement system that combines multimodal data processing and vector database and has user privacy security.

[0008] To achieve the above-mentioned purpose, the present invention provides the following technical solution: a blind person information memory enhancement system based on a multimodal vector database, characterized in that it includes:

[0009] User terminal, which interacts with the user, is used to recognize the user's voice commands, collect the visual scene around the user, and convert the memory content into natural language and broadcast it to the user;

[0010] The multimodal acquisition layer is used to collect voice information, visual information, spatial information and text information, and process the collected four types of information;

[0011] A unified encoder, connected to the multimodal acquisition layer, is used to encode the four acquired information;

[0012] A vectorized database for dynamically encoding and adjusting the weight coefficients of each modality data in the final vector representation based on the user's current context and history, and providing similarity search;

[0013] Adaptive interaction engine, used to perform multimodal information fusion and generate results and feedback after fusion is completed;

[0014] A layer of memory-augmented algorithms for serving natural language queries.

[0015] As a further improvement of the present invention, the user terminal includes:

[0016] A microphone array is used to capture the user's voice commands and convert the voice into text format through automatic speech recognition technology;

[0017] The camera is used to capture the visual scene around the user in real time and use the deep learning model to pre-analyze the image and extract key feature points, which are the object category and location, to confirm whether to trigger visual information sampling;

[0018] The speaker array is used to convert the retrieved memory content into natural language through text-to-speech technology and broadcast it to the user.

[0019] As a further improvement of the present invention, the specific manner in which the multimodal acquisition layer processes the four collected information is as follows:

[0020] For speech information: extract Mel-spectrogram features or other acoustic feature extraction methods to capture important features such as pitch and rhythm;

[0021] For visual information: the camera of the terminal device extracts frames and the user inputs intentions to trigger the sampling of visual information image data, and a convolutional neural network is used to identify and extract key features;

[0022] For spatial information: fuse outdoor and indoor information to generate a spatiotemporal coding vector;

[0023] For text information: Use natural language processing technology to convert text into numerical vectors.

[0024] As a further improvement of the present invention, the unified encoder uses the following encoding formula to encode the four information:

[0025] V=α·V_audio+β·V_visual+γ·V_spatial+δ*V_Textual

[0026] Among them, V is the weighted sum of the corresponding dimensions of each modality for each dimension, α, β, γ, δ are weight coefficients, V_audio is the language information encoding, V_visual is the visual information encoding, V_spatial is the spatial information encoding, and V_Textual is the text information encoding.

[0027] As a further improvement of the present invention, the specific method of dynamically adjusting the encoding of the vectorized database is: based on scene classification, a classifier is used to identify the current scene, and then weights are assigned according to predefined rules.

[0028] As a further improvement of the present invention, the specific manner in which the adaptive interaction engine performs multimodal information fusion is as follows:

[0029] First, data synchronization and integration: synchronize and integrate data from different sensors into a unified timeline to ensure that all information can be accurately associated;

[0030] Then, intent recognition and classification are performed: natural language processing technology and context-aware algorithms are used to analyze user intentions, pre-trained language models are used to classify intents, and the user's past behavior patterns are combined to optimize prediction accuracy.

[0031] As a further improvement of the present invention, the specific manner in which the adaptive interaction engine generates results and provides feedback is as follows:

[0032] First, response generation is performed: appropriate response content is generated based on the retrieved information, including directly answering questions, providing navigation guidance, or reminding users of relevant matters.

[0033] Then multi-channel feedback is carried out: including voice feedback, tactile feedback, and audio prompts to provide information to users and enhance the interactive experience.

[0034] Beneficial effects of the present invention: The present invention proposes an innovative solution to solve the problems faced by the blind in situational memory enhancement. By integrating multiple sensory information and converting it into a unified vector representation, efficient memory management and convenient life-assistance functions are achieved. In addition, vectorized storage avoids the leakage of user personal information, image and voice data, and ensures the security and privacy of user data. This system not only improves the quality of life of the blind, but also provides a new development direction for future barrier-free assistive technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 The module block diagram of the blind person information memory enhancement system based on the multimodal vector database of the present invention. DETAILED DESCRIPTION

[0036] The present invention will be further described below in conjunction with the embodiments given in the accompanying drawings.

[0037] Reference Figure 1 As shown, a blind information memory enhancement system based on a multimodal vector database in this embodiment is mainly composed of the following six modules, namely, a user terminal, a multimodal acquisition layer, a unified encoder, a vectorized database, an adaptive interaction engine and a memory enhancement algorithm layer. When the system is working, the user terminal first performs human-computer interaction with the user to recognize the user's voice command, collect the visual scene around the user, and then transmit it to the multimodal acquisition layer. The multimodal acquisition layer collects voice information, visual information, spatial information and text information, and processes the collected four information. After the four information are processed, they are encoded by a unified encoder and then stored in the vectorized database. In the vectorized database, based on the user's current context and historical records, dynamic coding adjusts the weight coefficient of each modal data in the final vector representation and provides similarity search, so as to realize the information memory storage of the user. When the user needs to retrieve the memory, the natural language is sent to the adaptive interaction engine. The memory enhancement algorithm layer understands the natural language in context and then feeds it back to the adaptive interaction engine. The adaptive interaction engine recognizes the user's intention based on the understood natural language, and then retrieves the memory content corresponding to the intention through the adaptive interaction engine, and then feeds it back to the user terminal. The user terminal converts the memory content into natural language and broadcasts it to the user. The six modules are as follows:

[0038] User Terminal

[0039] Voice command recognition: Use a high-quality microphone array to capture the user's voice commands and convert the voice into text format through automatic speech recognition (ASR) technology.

[0040] Camera acquisition: Use high-resolution cameras on terminal devices (such as smart glasses) to capture the visual scene around the user in real time. On this basis, deep learning models (such as YOLO, SSD, etc.) can be used to pre-analyze the image and extract key feature points, such as object category and location, to confirm whether visual information sampling is triggered, reducing frequent transmission and communication.

[0041] Speaker array: The retrieved memory content is converted into natural language through text-to-speech (TTS) technology and broadcast to the user.

[0042] Multimodal acquisition layer

[0043] Speech information: Extract Mel-spectrogram features (MFCCs) or other acoustic feature extraction methods to capture important features such as pitch and rhythm.

[0044] Visual information: The visual information is sampled by the terminal device camera frame extraction and the user input intention triggering image data, and convolutional neural networks (CNNs) are used to identify and extract key features.

[0045] Spatial information: Fusion of GPS (outdoor) and UWB positioning (indoor) to generate space-time coding vectors

[0046] Text information: Use natural language processing techniques such as word embedding models (Word2Vec, GloVe) or Transformer architecture (BERT) to convert text into numerical vectors.

[0047] Unified Encoder

[0048] Coding formula:

[0049] V=α·V_audio+β·V_visual+γ·V_spatial+δ*V_Textual

[0050] (α, β, γ, δ are dynamically adjusted by user scenarios)

[0051] Voice information coding (V_audio)

[0052] Feature extraction: MFCC (Mel-frequency cepstral coefficient) is used to extract acoustic features.

[0053] Visual information encoding (V_visual)

[0054] Feature extraction: Use convolutional neural network (CNN) to extract image features.

[0055] Spatial information encoding (V_spatial)

[0056] Feature fusion: Fusion of GPS (outdoor) and UWB (indoor) positioning data.

[0057] Text information encoding (V_Textual)

[0058] Feature extraction: Generate text vectors using a pre-trained language model such as BERT.

[0059] Finally, each dimension of V is the weighted sum of the corresponding dimensions of each mode.

[0060] Vectorized Database

[0061] Dynamic encoding strategy: Dynamically adjust the weight coefficients (α, β, γ, δ) of each modal data in the final vector representation based on the user's current context and historical records. Dynamic adjustment strategy: Based on scene classification. Use classifiers (such as SVM, neural network) to identify the current scene (such as navigation, chat) and assign weights according to predefined rules.

[0062] Example: In the navigation scenario, the γ(spatial) weight is increased from 0.2 to 0.6.

[0063] Similarity search: Use efficient index structures (such as FAISS, Annoy) to accelerate the similarity search process in the vector database and quickly find the most relevant memory fragments.

[0064] Adaptive Interaction Engine:

[0065] Multimodal Information Fusion

[0066] Data synchronization and integration: Synchronize and integrate data from different sensors (such as voice, images, GPS positioning, etc.) into a unified timeline to ensure that all information can be accurately associated.

[0067] Intent recognition and classification: Use natural language processing (NLP) technology and context-aware algorithms to analyze user intent. Pre-trained language models (such as BERT, RoBERTa) can be used for intent classification, and the user's past behavior patterns can be combined to optimize prediction accuracy.

[0068] Result generation and feedback

[0069] Response generation: Generate appropriate response content based on the information retrieved. This may include directly answering the question, providing navigation instructions, or reminding the user of related matters.

[0070] Multi-channel feedback: In addition to voice feedback, information can also be provided to users through tactile feedback (vibration), audio prompts, and other methods to enhance the interactive experience.

[0071] Memory Enhancement Algorithm Layer

[0072] Natural language query

[0073] Complex query processing: Supports complex natural language queries, such as "Remember the name of the restaurant I ate at last Friday?" These are questions involving time, place, and personal experience.

[0074] Contextual understanding: The system should be able to understand context and maintain topic consistency in continuous conversations. For example, a user first asks "Where is the park I went to last week?" and then asks "What good food is nearby?" Multiple rounds of conversations should allow the system to understand that these two questions are related.

[0075] This embodiment provides the following examples:

[0076] Scenario description: Blind users hope to better manage and become familiar with their living environment, especially the layout and location of important items in unfamiliar environments, and be able to quickly find the items they need.

[0077] Operation process: Users use smart glasses to scan and record the layout of the current environment and the location of important items. The system uses the camera to collect image data and combines voice input during use (such as "This is my small courtyard") with GPS / UWB positioning technology to generate a spatiotemporal coding vector. When the user needs to find an item, he can query it through natural language (for example, "Where is my umbrella?"), and the system will provide precise location guidance based on the stored spatiotemporal memory and image data ("You have a black folding umbrella behind the door of the small courtyard").

[0078] In summary, the blind person information memory enhancement system based on the multimodal vector database of this embodiment effectively constructs a multimodal memory vector space for the blind person, and realizes efficient memory management and convenient life assistance functions.

[0079] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A blind information memory enhancement system based on a multimodal vector database, characterized in that: include: User terminal, which interacts with the user, is used to recognize the user's voice commands, collect the visual scene around the user, and convert the memory content into natural language and broadcast it to the user; The multimodal acquisition layer is used to collect voice information, visual information, spatial information and text information, and process the collected four types of information; A unified encoder, connected to the multimodal acquisition layer, is used to encode the four acquired information; A vectorized database for dynamically encoding and adjusting the weight coefficients of each modality data in the final vector representation based on the user's current context and history, and providing similarity search; Adaptive interaction engine, used to perform multimodal information fusion and generate results and feedback after fusion is completed; A memory-augmented algorithm layer for serving natural language queries.

2. The information memory enhancement system for the blind based on the multimodal vector database according to claim 1 is characterized in that: The user terminal comprises: A microphone array is used to capture the user's voice commands and convert the voice into text format through automatic speech recognition technology; The camera is used to capture the visual scene around the user in real time and use the deep learning model to pre-analyze the image and extract key feature points, which are the object category and location, to confirm whether to trigger visual information sampling; The speaker array is used to convert the retrieved memory content into natural language through text-to-speech technology and broadcast it to the user.

3. The blind person information memory enhancement system based on multimodal vector database according to claim 1 or 2, characterized in that: The specific way in which the multimodal acquisition layer processes the four collected information is as follows: For speech information: extract Mel-spectrogram features or other acoustic feature extraction methods to capture important features such as pitch and rhythm; For visual information: the camera of the terminal device extracts frames and the user inputs intentions to trigger the sampling of visual information image data, and a convolutional neural network is used to identify and extract key features; For spatial information: fuse outdoor and indoor information to generate a spatiotemporal coding vector; For text information: Use natural language processing technology to convert text into numerical vectors.

4. The blind person information memory enhancement system based on multimodal vector database according to claim 1 or 2, characterized in that: The unified encoder uses the following encoding formula to encode the four information: V=α·V_audio+β·V_visual+γ·V_spatial+δ*V_Textual Among them, V is the weighted sum of the corresponding dimensions of each modality for each dimension, α, β, γ, δ are weight coefficients, V_audio is the language information encoding, V_visual is the visual information encoding, V_spatial is the spatial information encoding, and V_Textual is the text information encoding.

5. The blind person information memory enhancement system based on multimodal vector database according to claim 1 or 2, characterized in that: The specific method of dynamically adjusting the encoding of the vectorized database is: based on scene classification, a classifier is used to identify the current scene, and then weights are assigned according to predefined rules.

6. The information memory enhancement system for the blind based on a multimodal vector database according to claim 1 or 2, characterized in that: The specific method of the adaptive interaction engine for multimodal information fusion is as follows: First, data synchronization and integration: synchronize and integrate data from different sensors into a unified timeline to ensure that all information can be accurately associated; Then, intent recognition and classification are performed: natural language processing technology and context-aware algorithms are used to analyze user intentions, pre-trained language models are used to classify intents, and the user's past behavior patterns are combined to optimize prediction accuracy.

7. The information memory enhancement system for the blind based on the multimodal vector database according to claim 6 is characterized in that: The specific method of the adaptive interaction engine to generate results and provide feedback is as follows: First, response generation is performed: appropriate response content is generated based on the retrieved information, including directly answering questions, providing navigation guidance, or reminding users of relevant matters. Then multi-channel feedback is carried out: including voice feedback, tactile feedback, and audio prompts to provide information to users and enhance the interactive experience.

Citation Information

Patent Citations

  • Multi-mode interactive intelligent control system

    CN118226967A