Voice, video and text training method suitable for blind glasses

Through the multi-stage multi-modal alignment training method, the visual and voice interaction capabilities of blind glasses are improved, and the existing intelligent blind guide system cannot provide specific obstacle information and insufficient voice interaction is solved, achieving an efficient and smooth multi-modal interaction experience of blind glasses.

CN120299343APending Publication Date: 2025-07-11LINKER
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510379574.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing intelligent blinding system cannot effectively provide specific information about obstacles, and the voice interaction performance is insufficient, resulting in poor daily life experience for blind users.

Method used

The multi-stage multi-modal alignment training method is adopted to gradually train the large language model (LLM) to understand visual and speech information, so that it establishes a close connection between visual and speech modes. Through visual-language alignment training and audio alignment training, the multi-modal understanding and response capabilities of the model are improved.

Benefits of technology

It realizes efficient and smooth visual and voice interactions of blind glasses, enhances the daily life experience of blind users, improves response speed and user autonomy, and reduces dependence on external help.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005334145530000031
    Figure BDA0005334145530000031
  • Figure BDA0005334145530000055
    Figure BDA0005334145530000055
  • Figure FDA0005334145520000021
    Figure FDA0005334145520000021
Patent Text Reader

Abstract

The invention discloses a voice, video and text training method suitable for glasses for blind people, which comprises the following steps of: 1, carrying out vision-language alignment training on a large language model to establish a preliminary relation between vision and language modalities, so that the large language model can understand image semantics and generate corresponding description through languages; step 2, performing audio alignment training on the large language model, and introducing audio information into a multi-modal model, so that the multi-modal model has a voice input processing capability; the method has the advantages that through multi-mode alignment training, the large language model has the capacity of understanding image semantics, generating description and processing voice input, an effective training method is provided for blind glasses to achieve voice, video and text processing, and the blind is helped to better perceive external information in a voice mode and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a pair of glasses for the blind, and more specifically to a training method for speech, video, and text applicable to glasses for the blind. Background Art

[0002] Currently, with the progress and development of technology, some intelligent blind guiding systems have gradually emerged on the market. However, these intelligent blind guiding systems generally utilize the already mature ultrasonic technology to achieve. Although using ultrasonic technology can determine whether there are obstacles in front of the blind, it cannot let the blind know what the obstacle is, how far the obstacle is, and other information.

[0003] For this reason, in the prior art, a patent with the patent number 202111172571.X and the name of an invention patent for a blind obstacle avoidance method, system, device, and readable storage medium discloses a method of collecting a first image and a second image through a binocular vision module, and then performing recognition and analysis on the collected first image and second image for obstacles, so that the blind can know what the obstacle is, how far the obstacle is, and other information. However, the above method realizes the interaction with the blind by using a voice device to broadcast fixed voices during the interaction with the blind. Therefore, there is a problem that the performance of the glasses for the blind in visual and voice interaction is insufficient. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a multi-stage multi-modal alignment training method to gradually train the LLM model to understand visual and voice information, so that each modality does not interfere with the performance of other modalities during the enhancement process, thereby providing efficient and smooth voice and visual interaction capabilities for glasses for the blind. The purpose of the present invention is to optimize the performance of glasses for the blind in visual and voice interaction, significantly improve the response speed, and achieve near-real-time visual and voice interaction to enhance the daily life experience of blind users.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A training method for speech, video, and text applicable to glasses for the blind, including the following steps:

[0006] Step 1, perform visual-language alignment training on the large language model to establish a preliminary connection between the visual and language modalities, so that the large language model can understand the semantics of the image and generate corresponding descriptions through language.

[0007] Step 2, perform audio alignment training on the large language model, introduce audio information into the multi-modal model, and enable it to have the ability to process voice input.

[0008] As a further improvement of the present invention patent, the specific steps of performing visual-language alignment training in Step 1 are as follows:

[0009] Step 1-1: First, align visual features with language features;

[0010] Step 1-2: Use a large amount of descriptive Caption data to train the visual encoder, visual MLP, and LLM, so that the LLM can not only receive visual inputs but also understand image content by generating natural language descriptions, thereby enhancing the visual understanding ability of the large language model;

[0011] Step 1-3: Introduce Q&A data to further improve the performance of the large language model in handling visual and language instruction tasks, so as to enhance the visual instruction following ability of the large language model.

[0012] As a further improvement of this invention patent, the specific method of aligning visual features with language features in Step 1-1 is: Extract visual features by using a pre-trained visual encoder and train them in combination with Caption descriptive text data;

[0013] Among them, only the visual MLP module is trained in this stage, and other modules remain frozen. The purpose is to enable the model to initially master the feature representation of visual content through learning image descriptions and align these features with language descriptions.

[0014] As a further improvement of this invention patent, the specific steps of performing audio alignment training in Step 2 are as follows:

[0015] Step 2-1: Use a speech-transcription dataset and train the speech encoder using the connectionist temporal classification loss function;

[0016] Step 2-2: Introduce a speech frequency modulator, and after training the speech encoder in Step 1, use the speech frequency modulator to downsample the speech features;

[0017] Step 2-3: Add the Q&A function of speech questions and text answers.

[0018] As a further improvement of this invention patent, the connectionist temporal classification loss function in Step 2-1 is specifically as follows:

[0019]

[0020] Among them, is the conditional probability of the given input sequence x and the target label sequence and is to map the target label sequence to the set of all possible label alignment methods.

[0021] Advantages of this invention:

[0022] Enhanced Multimodal Understanding: By gradually introducing visual, language, and speech modalities, the method enables the model to better understand and process multimodal information. For blind glasses users, the model can generate accurate image descriptions through speech, enhancing the usability of the product.

[0023] Tight Integration of Vision and Language: In the vision-language alignment phase, through the learning of visual descriptions, the model can accurately understand the content of images and generate corresponding natural language descriptions, helping users understand the surrounding environment through speech.

[0024] Reinforced Instruction Following Ability: By introducing question-and-answer data, the model not only understands visual information but also can perform tasks according to speech instructions, enhancing the speech control ability of blind glasses and enabling users to obtain more accurate responses through speech.

[0025] Efficient Speech Input and Output: Through audio alignment and frequency modulation techniques, the model can quickly and accurately process speech input and generate responses, improving the response speed and fluency of blind glasses in speech interaction.

[0026] Excellent Multimodal Adaptability: The fusion of multimodals during the training process enables the model to flexibly handle different input forms, improving the performance of the product in various usage scenarios and enhancing the user experience.

[0027] Enhanced User Autonomy: This method makes blind glasses have a higher level of intelligence. Users can interact with the device efficiently through speech, reducing dependence on external help and greatly improving the quality of life. Detailed Implementation Manner

[0028] The following will further elaborate on the present invention with the given embodiments.

[0029] A training method for speech, video, and text applicable to blind glasses in this embodiment is carried on a system with a visual encoder, a large language model LLM (such as LLAMA), a visual MLP, a speech encoder, and a speech frequency modulator. Then, based on the key modules of the above system, the following training method is adopted to realize the training of the large language model LLM (such as LLAMA), improve the performance of blind glasses in visual and speech interaction, and optimize the interaction experience, especially in terms of real-time response and accuracy. The specific training method is as follows:

[0030] Vision-Language Alignment Training

[0031] The goal of this stage is to establish a preliminary connection between the visual and language modalities, enabling the large language model to understand the semantics of images and generate corresponding descriptions through language.

[0032] Alignment of Visual and Linguistic Features: Visual features are extracted using a pre-trained visual encoder (e.g., various variants based on the Vision Transformer (ViT) model) and trained in combination with Caption descriptive text data. Only the visual MLP module is trained in this stage, and other modules remain frozen. The purpose is to enable the model to initially master the feature representation of visual content through learning image descriptions and align these features with linguistic descriptions.

[0033] Enhancement of Visual Understanding: In this stage, a large amount of descriptive Caption data is used to train the visual encoder, visual MLP, and LLM, enabling the LLM to not only receive visual inputs but also understand image content by generating natural language descriptions. Through this training, the model can generate accurate language descriptions based on visual information, further strengthening the connection between the visual and linguistic modalities.

[0034] Enhancement of Visual Instruction-Following Ability: In this stage, question-and-answer data is introduced to further improve the model's performance in handling visual and linguistic instruction tasks. By continuing to train the visual module, visual MLP, and LLM and combining a small amount of descriptive Caption data, the model's ability to follow instructions based on visual understanding is enhanced. At this time, the model can not only understand visual content but also generate appropriate responses according to given linguistic instructions, improving its performance in visual question-and-answer tasks.

[0035] Audio Alignment Training

[0036] After the visual-linguistic alignment training is completed, the next goal is to introduce audio information into the multi-modal model to enable it to handle speech inputs.

[0037] Speech Encoder Training: Use a speech-transcription dataset and train the speech encoder using the Connectionist Temporal Classification (CTC) loss function. This loss function can effectively model the correspondence between speech inputs and transcription texts, enabling the speech encoder to extract speech features and map these features to the text representation space. In this stage, the training of the speech encoder ensures that the model can effectively understand speech signals and convert them into texts.

[0038] where is the conditional probability of the given input sequence x and the target label sequence and is the set that maps the target label sequence to all possible label alignment ways.

[0039]

[0040] Introduce a voice frequency modulator: After the voice encoder, use the voice frequency modulator to downsample the voice features to reduce the frequency of the voice features, which can accelerate the processing speed of the large language model (LLM) and optimize the efficiency of the model. The role of the voice frequency modulator is to help the audio features enter the input layer of the model in a way suitable for LLM processing, ensuring that the LLM can effectively understand and process voice input. The training objective at this stage is to enable the LLM to not only understand the transcribed text of the voice data but also effectively process and generate corresponding language outputs.

[0041] Question-answering function for voice questions and text answers: At this stage, add the question-answering function for voice questions and text answers. To do this, convert some text-based question data into corresponding voice versions and use this data for training the model. At this time, the visual encoder, visual MLP, voice encoder, voice frequency modulator, and LLM will all participate in the training, thereby improving the model's adaptability to multimodal inputs. Through the training at this stage, the model can not only process images and text but also effectively handle voice inputs, enhancing its performance and interaction capabilities in a multimodal environment.

[0042] In summary, the multi-stage training method of this embodiment gradually improves the processing ability of each modality by introducing visual, language, and voice modalities in stages, while ensuring that the capabilities of existing modalities are not weakened during the introduction of new modalities. Through this method, the model can perform excellently in complex multimodal tasks. Especially in application scenarios such as blind glasses that require visual and voice interaction, it can effectively understand and generate image content, interact through voice, and execute instructions, ultimately achieving a more efficient user experience.

[0043] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. Any technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A training method for speech, video and text applicable to blind glasses, characterized in that: It includes the following steps: Step 1: Conduct visual-language alignment training on the large language model to establish a preliminary connection between the visual and language modalities, enabling the large language model to understand the semantics of images and generate corresponding descriptions through language; Step 2: Conduct audio alignment training on the large language model to introduce audio information into the multi-modal model, enabling it to have the ability to process speech input.

2. The training method for speech, video, and text applicable to blind glasses according to claim 1, characterized in that: The specific steps for visual-language alignment training in Step 1 are as follows: Step 1-1: First, align visual features with language features; Step 1-2: Use a large amount of descriptive Caption data to train the visual encoder, visual MLP, and LLM, enabling the LLM to not only receive visual input but also understand image content by generating natural language descriptions, thereby enhancing the visual understanding ability of the large language model; Step 1-3: Introduce question-and-answer data to further improve the performance of the large language model in processing visual and language instruction tasks, thereby enhancing the visual instruction following ability of the large language model.

3. The training method for speech, video and text applicable to blind glasses according to claim 2, characterized in that: The specific method for aligning visual features with language features in Step 1-1 is as follows: Extract visual features using a pre-trained visual encoder and train in combination with Caption descriptive text data; among them, only the visual MLP module is trained in this stage, and other modules remain frozen. The purpose is to enable the model to initially master the feature representation of visual content through learning image descriptions and align these features with language descriptions.

4. The training method for speech, video, and text applicable to blind glasses according to claims 1 to 3, characterized in that: The specific steps for audio alignment training in Step 2 are as follows: Step 2-1: Use a speech-transcription dataset and train the speech encoder using the connectionist temporal classification loss function; Step 2-2: Introduce a speech frequency modulator and, after training the speech encoder in Step 1, use the speech frequency modulator to downsample the speech features; Step 2-3: Add the question-and-answer function of speech questions and text answers.

5. The training method for speech, video and text applicable to blind glasses according to claim 4, characterized in that: The connectionist temporal classification loss function in Step 2-1 is specifically as follows: where, is the conditional probability of the given input sequence x and the target label sequence , is to map the target label sequence to the set of all possible label alignments.

Citation Information

Patent Citations

  • Obstacle avoidance method, system, device and readable storage medium for the blind

    CN113893142B