Multi-mode integrated auxiliary system and method for special demand group
Through the embedded wearable device working in concert with the cloud server, the problem of existing devices being unable to collect facial expressions and lack of active navigation is solved, and efficient and low-cost solutions for two-way sign language translation and visually impaired navigation are realized, improving the comfort and flexibility of the equipment.
Patent Information
- Application Number
- CN202510321545.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing equipment cannot effectively collect facial expression information, resulting in sign language translation errors, high hardware dependence, single functions, inability to achieve two-way interaction, and the blinding device lacks active navigation and poor adaptability to complex environments.
It adopts embedded wearable devices, combined with cameras, microphones, speakers and IoT modules, and collaboratively works with cloud servers through the embedded processing unit to realize two-way sign language translation, visually impaired navigation and text speech conversion, and uses contrast learning and large language models for sign language translation, and combines semantic segmentation and path planning for navigation.
It realizes two-way seamless communication between deaf and mute people and healthy listeners, precise navigation for visually impaired people, improves equipment comfort and flexibility, and reduces hardware costs and power consumption.
Smart Images

Figure CN120356391A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of auxiliary devices, and particularly to a multimodal integrated auxiliary system and method for special needs groups. Background Art
[0002] Current mainstream devices (such as the sign language translation gloves described in publication number CN221554755A and the wristband translator described in publication number CN119358697A) mainly collect hand movement data through inertial measurement units (IMUs) or electromyography sensors (EMGs), and combine traditional algorithms or deep learning models to achieve sign language translation. The main limitations of such devices are as follows:
[0003] 1. Semantic ambiguity: Sign language expressions often need to be combined with facial expressions and body movements. However, such devices only capture hand postures and cannot collect facial expression information, resulting in incorrect translations due to differences in expressions for the same gesture.
[0004] 2. Strong hardware dependence: Users need to wear high-precision sensor devices (such as gloves and wristbands), which are costly and inconvenient for daily use (such as being stuffy to wear in summer and affecting hand flexibility).
[0005] 3. Functional singularity: Only one-way translation (sign language → speech / text) is supported, and two-way interaction between hearing people and deaf-mute people cannot be achieved.
[0006] 4. Limited expandability: The sign language system is relatively complex. Collecting sign language data only through wearing high-precision sensor devices is costly and has low collection efficiency.
[0007] Existing blind guiding devices (such as the intelligent blind guiding cane described in publication number CN119345004A) mostly rely on pressure sensors, infrared sensors or millimeter wave radars to detect obstacles and prompt users through vibration or voice. The main limitations of such devices are as follows:
[0008] 1. Passive detection: It can only provide immediate feedback on obstacles and lacks active path planning and autonomous navigation functions (such as being unable to generate a detour route when the blind path is occupied).
[0009] 2. Poor scene adaptability: It has insufficient semantic understanding ability for complex environments (such as dynamic obstacles and missing blind paths), and mostly relies on local device computing power and cannot utilize cloud big data analysis. Summary of the Invention
[0010] The purpose of the present invention is to provide a multimodal integrated auxiliary system and method for special needs groups, aiming to solve the above problems in the prior art.
[0011] An embodiment of the present invention provides a multi-modal integrated assistance system for special needs groups, including:
[0012] A basic function subsystem, connected to an embedded processing unit, for collecting basic information and transmitting it to the embedded processing unit, and providing assistance services for special needs groups according to relevant content from the embedded processing unit; wherein, the basic information includes image data and audio data; the relevant content includes sign language translation text, sign language translation voice data, optimal path guidance text, optimal path guidance voice data, speech conversion text, and 3D digital sign language action data;
[0013] An embedded processing unit, connected to the basic function subsystem and a cloud server, for driving each module in the basic function subsystem, transmitting the received basic information to the cloud server, and downloading the relevant content from the cloud server to the basic function subsystem;
[0014] A cloud server, connected to the embedded processing unit and a user terminal, is deployed with a two-way sign language translation subsystem, a visually impaired navigation subsystem, and a text-speech two-way conversion subsystem that communicate with each other, for performing two-way sign language translation, visually impaired navigation, and text-speech two-way conversion based on the basic information through the two-way sign language translation subsystem, the visually impaired navigation subsystem, and the text-speech two-way conversion subsystem, and sending the generated relevant content to the embedded processing unit or the user terminal;
[0015] A user terminal, connected to the cloud server, for visually displaying the relevant content.
[0016] An embodiment of the present invention provides a multi-modal integrated assistance method for special needs groups, including:
[0017] Collect basic information through the basic function subsystem and transmit it to the embedded processing unit, and provide assistance services for special needs groups according to relevant content from the embedded processing unit; wherein, the basic information includes image data and audio data; the relevant content includes sign language translation text, sign language translation voice data, optimal path guidance text, optimal path guidance voice data, speech conversion text, and 3D digital sign language action data;
[0018] Drive each module in the basic function subsystem through the embedded processing unit, transmit the received basic information to the cloud server, and download the relevant content from the cloud server to the basic function subsystem;
[0019] Through the two-way sign language translation subsystem, visually impaired navigation subsystem, and text-to-speech and speech-to-text conversion subsystem in the cloud server, perform two-way sign language translation, visually impaired navigation, and text-to-speech and speech-to-text conversion based on the basic information, and send the generated relevant content to the embedded processing unit or the user terminal;
[0020] Visualize the relevant content through the user terminal.
[0021] Adopting the embodiments of the present invention may include the following beneficial effects: The multi-modal assistance system for deaf-mute and visually impaired persons based on the embedded wearable device proposed in the embodiments of the present invention adopts the form of a wearable device hardware architecture, which has better comfort and flexibility. Among them, the two-way sign language translation subsystem included in the system can help the deaf-mute group communicate seamlessly with the outside world; the visually impaired navigation subsystem can effectively solve the problem of inconvenient travel for visually impaired persons. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in one or more embodiments or the prior art of this specification, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic diagram of the multi-modal integrated assistance system for special needs groups in the embodiments of the present invention;
[0024] Figure 2 It is a schematic diagram of the overall architecture in the embodiments of the present invention;
[0025] Figure 3 It is a schematic diagram of the hand bone graph perception prior pre-training model based on contrast learning and the morpheme conversion model based on the large language model in the embodiments of the present invention;
[0026] Figure 4 It is a schematic diagram of the 3D digital human engine in the embodiments of the present invention;
[0027] Figure 5 It is a schematic diagram of the visually impaired navigation subsystem in the embodiments of the present invention;
[0028] Figure 6 It is a schematic diagram of the text-to-speech and speech-to-text conversion subsystem in the embodiments of the present invention;
[0029] Figure 7 It is a schematic diagram of the natural connection relationship of the hand key point sequences in the embodiments of the present invention;
[0030] Figure 8Schematic diagram of the natural connection relationship of the facial key point sequence in the embodiment of the present invention;
[0031] Figure 9 Flowchart of the multi-modal integration assistance method for special needs groups in the embodiment of the present invention. Detailed implementation manners
[0032] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
[0033] System embodiment
[0034] According to an embodiment of the present invention, there is provided a multi-modal integration assistance system for special needs groups. Figure 1 Schematic diagram of the multi-modal integration assistance system for special needs groups in the embodiment of the present invention, as Figure 1 shown, the multi-modal integration assistance system for special needs groups according to an embodiment of the present invention specifically includes:
[0035] The basic function subsystem 10, connected to the embedded processing unit, is used to collect basic information and then transmit it to the embedded processing unit, and provide assistance services for special needs groups according to the relevant content from the embedded processing unit; wherein, the basic information includes image data and audio data; the relevant content includes sign language translation text, sign language translation voice data, optimal path guidance text, optimal path guidance voice data, speech conversion text, and 3D digital human sign language action data;
[0036] Among them, the basic function subsystem includes a camera module, a microphone module, a speaker module, an Internet of Things communication module, and a power management module that communicate with each other;
[0037] The camera module is used to collect image data and then transmit it to the embedded processing unit; wherein, the image data includes the sign language hand movements, facial expressions of special needs groups, and image data of the blind path environment;
[0038] The microphone module is used to collect audio data and then transmit it to the embedded processing unit; wherein, the audio data includes the conversation content of the conversants and the external environment sounds;
[0039] The speaker module is used to perform voice broadcast on the sign language translation voice data and the optimal path guidance voice data from the embedded processing unit;
[0040] The Internet of Things communication module is used to provide traffic networking communication services for the system;
[0041] The power management module is used to provide power management services for the system.
[0042] The embedded processing unit 12 is connected to the basic function subsystem and the cloud server, and is used to drive each module in the basic function subsystem, transmit the received basic information to the cloud server, and download the relevant content from the cloud server to the basic function subsystem;
[0043] The cloud server 14 is connected to the embedded processing unit and the client, and is deployed with a two-way sign language translation subsystem, a visual impairment navigation subsystem, and a text-to-speech two-way conversion subsystem that communicate with each other. It is used to perform two-way sign language translation, visual impairment navigation, and text-to-speech two-way conversion based on the basic information through the two-way sign language translation subsystem, the visual impairment navigation subsystem, and the text-to-speech two-way conversion subsystem, and send the generated relevant content to the embedded processing unit or the client;
[0044] Among them, the two-way sign language translation subsystem specifically includes:
[0045] The hand bone graph perception module is connected to the morpheme conversion module, and is used to process the image data of sign language hand movements and facial expressions through a hand bone graph perception prior pre-training model based on contrast learning, generate a corresponding sign language morpheme sequence, and transmit the sign language morpheme sequence to the morpheme conversion module; among them, the hand bone graph perception prior pre-training model includes a visual encoder, a text encoder, and a metric network, and is specifically used for:
[0046] Input the image data of sign language hand movements and facial expressions into the human pose estimation model in the visual encoder to extract key point coordinates, obtain a key point coordinate sequence, construct a corresponding dynamic graph according to the natural connection relationship of the key point coordinate sequence, and perform multiple aggregations on the dynamic graph through a graph neural network to generate a visual vector encoding representing hand bones and facial expressions;
[0047] Use the text encoder to perform multiple calculations on the text description of the sign language action to generate a text vector encoding representing the sign language action;
[0048] Align and calculate the similarity between the visual vector encoding and the text vector encoding through the metric network, and output the sign language morpheme sequence corresponding to the sign language action text description with the highest similarity to the visual vector encoding;
[0049] A morpheme conversion module, connected to the hand bone map perception module and the text-to-speech and speech-to-text conversion subsystem, is used to convert the sign language morpheme sequence through a morpheme conversion model based on a large language model to generate a sign language translation text, and transmit the sign language translation text to the text-to-speech and speech-to-text conversion subsystem;
[0050] A 3D digital human engine module, connected to the text-to-speech and speech-to-text conversion subsystem and the user terminal, is used to generate corresponding 3D digital human sign language action data according to the speech conversion text from the text-to-speech and speech-to-text conversion subsystem, and send the 3D digital human sign language action data to the user terminal.
[0051] The visually impaired navigation subsystem specifically includes:
[0052] A blind path positioning module, connected to the path planning module, is used to segment the image data of the blind path environment into a binary image of the blind path by using a blind path positioning model based on a semantic segmentation algorithm, and output a sequence of blind path center point coordinates, and transmit the sequence of blind path center point coordinates to the path planning module;
[0053] An obstacle positioning module, connected to the path planning module, is used to detect the image data of the blind path environment through an obstacle positioning model based on an object detection algorithm, output a sequence of obstacle center point coordinates, and transmit the sequence of obstacle center point coordinates to the path planning module;
[0054] A monocular depth estimation module, connected to the path planning module, is used to convert the image data of the blind path environment into a depth image through a monocular depth estimation model, and calculate the straight-line distance between each pixel in the depth image and the current viewing angle after calculation;
[0055] A path planning module, connected to the blind path positioning module, the obstacle positioning module, the monocular depth estimation module and the text-to-speech and speech-to-text conversion subsystem, is used to obtain a spatial distance sequence corresponding to the sequence of blind path center point coordinates and the sequence of obstacle center point coordinates based on the straight-line distance between each pixel in the image and the current viewing angle, calculate the current optimal obstacle avoidance path by using a path planning model designed based on the A* algorithm according to the spatial distance sequence, and generate a corresponding optimal path guidance text according to the current optimal obstacle avoidance path, and send the optimal path guidance text to the text-to-speech and speech-to-text conversion subsystem.
[0056] The text-to-speech and speech-to-text conversion subsystem specifically includes:
[0057] The text-to-speech module is connected to the morpheme conversion module, the path planning module, and the embedded processing unit, and is used to convert the sign language translation text and the optimal path guidance text into sign language translation voice data and optimal path guidance voice data through a text-to-speech model, and download the sign language translation voice data and the optimal path guidance voice data to the embedded processing unit;
[0058] The voice-to-text module is connected to the embedded processing unit and the 3D digital human engine module, and is used to convert the audio data from the embedded processing unit into voice conversion text through a voice-to-text model, and send the voice conversion text to the 3D digital human engine module.
[0059] The client 16 is connected to the cloud server and is used to visually display the relevant content. Specifically, it is used for:
[0060] Visually display the sign language translation text, the optimal path guidance text, and the voice conversion text, and dynamically display the 3D digital human sign language action data.
[0061] The following combines the specific situation of the multi-modal integrated assistance system for special needs groups in the embodiments of the present invention, as Figure 2 shown, to elaborate on the above technical solutions of the embodiments of the present invention in detail.
[0062] The embodiments of the present invention aim to solve the two-way real-time and accurate sign language translation needs between deaf-mute people and hearing people, the blind path navigation and accurate obstacle detection needs of visually impaired people, and the balance problem of wearable devices among low power consumption, lightweight, and high performance. Therefore, the embodiments of the present invention propose a multi-modal assistance system for deaf-mute and visually impaired people based on an embedded wearable device, which includes a device hardware architecture, a two-way sign language translation subsystem, a visually impaired navigation subsystem, and a text-voice two-way conversion subsystem.
[0063] 1. The device hardware architecture includes an embedded processing unit, a camera module, a speaker module, a microphone module, an Internet of Things communication module, and a cloud server.
[0064] Among them, the camera module, the speaker module, the microphone module, and the Internet of Things communication module are all driven by the embedded processing unit and are jointly integrated on a printed circuit board (PCB).
[0065] The information collected by the camera module includes but is not limited to sign language hand movements, facial expressions, and video data including the blind path environment.
[0066] The information collected by the microphone module includes but is not limited to the content spoken by the converser and the environmental sound.
[0067] The embedded processing unit is built with a WiFi module and provides network access through the Internet of Things communication module, so as to communicate with the cloud server.
[0068] The data collected by the camera module and the microphone module will be transmitted to the cloud server through the embedded processing unit.
[0069] The speaker module will receive the voice data sent back by the cloud server and perform voice broadcast.
[0070] 2. As Figure 3 、 Figure 4 shown, the two-way sign language translation subsystem includes a hand skeleton map perception prior pre-training model based on contrast learning, a morpheme conversion model based on a large language model, and a 3D digital human engine.
[0071] (1) The hand skeleton map perception prior pre-training model based on contrast learning mainly includes a visual encoder, a text encoder, and a metric network.
[0072] A. The visual encoder includes a human pose estimation model obtained by knowledge distillation of MediaPipe and a sequence-to-sequence (Seq2Seq) model constructed based on a graph neural network.
[0073] The specific process is as follows: The color image (RGB) collected by the camera is input into the human pose estimation model to extract the key points of the hand and face, generating a sequence of key point coordinates. A graph is constructed according to the natural connection relationship of the key points, and the graph is aggregated multiple times through the graph neural network to generate a vector encoding representing the hand skeleton and facial expressions.
[0074] B. The text encoder adopts an efficient state space model (SSM) architecture based on Mamba-2. This encoder receives the action text descriptions of all sign languages and generates vector encodings of the text through multiple calculations. Among them, these action text descriptions are obtained through tools such as web crawlers on relevant sign language material platforms (such as the national sign language general dictionary), and then proofread by professional sign language teachers for the collected sign language action data.
[0075] C. The metric network is responsible for aligning and calculating the similarity of the vectors output by the above visual encoder and text encoder. Through contrast learning, the metric network will optimize the representation spaces of the visual and text modalities, and finally output the sign language morpheme sequence corresponding to the action text description with the highest similarity to the visual vector.
[0076] (2) The morpheme conversion model based on the large language model is mainly based on large language model technology and is fine-tuned on the sign language morpheme to Chinese sentence task in combination with DeepSeek-R1-Distill-Qwen-7B. The core function of the model is to further convert the input sign language morpheme sequence into a fluent Chinese sequence.
[0077] The cloud server will receive the video data containing sign language hand movements and facial expressions sent by the embedded processing unit. After being inferred by the two-way sign language translation subsystem, the corresponding Chinese text will be generated, and then the speech data will be generated through the text-to-speech model and sent back to the embedded processing unit. Finally, the speech will be broadcast through the speaker module, thus realizing the translation from sign language to speech.
[0078] 3. As Figure 5 shown, the visually impaired navigation subsystem includes a blind path positioning model based on the semantic segmentation algorithm, an obstacle positioning model based on the object detection algorithm, a monocular depth estimation model, and a path planning model.
[0079] (1) The blind path positioning model based on the semantic segmentation algorithm is an improved model based on MVANet, which can segment the above RGB image into a binary image of the blind path and output the central point coordinate sequence of the blind path.
[0080] (2) The obstacle positioning model based on the object detection algorithm is a lightweight model based on Hyper-YOLO, which can detect the obstacles in the above RGB image and output the central point coordinate sequence of the obstacles.
[0081] (3) The monocular depth estimation model is the open-source model DepthFM, which can convert the above RGB image into a depth image, calculate and output the straight-line distance between each pixel in the image and the current viewing angle.
[0082] (4) The path planning model is a path planning model designed based on the A* algorithm, which can calculate the optimal obstacle avoidance path according to the spatial distance sequence corresponding to the above central point coordinate sequence of the blind path and the central point coordinate sequence of the obstacles, and output the guiding text generated according to this path.
[0083] 4. As Figure 6 shown, the text-speech two-way conversion subsystem includes a text-to-speech model and a speech-to-text model. Among them, the text-to-speech model is the open-source model Fast-TTS, and the speech-to-text model is the open-source model Whisper.
[0084] The above-mentioned two-way sign language translation subsystem, visually impaired navigation subsystem, and text-to-speech two-way conversion subsystem are all deployed in the cloud server. The cloud server will receive the audio data containing the content spoken by the converser and the environmental sounds sent by the embedded processing unit. After being processed by the speech-to-text model, the corresponding Chinese text is generated. After being inferred by the two-way sign language translation subsystem, the corresponding 3D digital human actions are generated. Then, the Chinese text and the 3D digital human action data are transmitted to the mobile client for display, thus realizing the translation from speech to sign language.
[0085] The cloud server will receive the video data containing the blind path environment sent by the embedded processing unit. After being inferred by the visually impaired navigation subsystem, the specific positions of the blind path and obstacles are obtained. It analyzes whether avoidance is needed and by what means, and converts the result through the text-to-speech model and transmits it back to the embedded drive unit. Finally, the voice is broadcast by the speaker module, thus realizing visually impaired navigation. That is, the visually impaired navigation subsystem will detect the positions of the blind path and related obstacles. If the obstacle is not on the blind path, there is no need to voice prompt the blind person to avoid it. Otherwise, it will calculate the avoidance route through the path planning algorithm and then voice prompt the blind person to avoid it, such as voice prompting "There is a shared bicycle 5m ahead, please turn left to avoid it".
[0086] The specific implementation process of the embodiments of the present invention is as follows:
[0087] 1. Hardware Design and Implementation
[0088] 1.1 Selection of the Hardware Architecture of the Wearable Device
[0089] The embedded processing unit is ESP32-S3-WROOM-1-N16R8, which integrates a WiFi / BLE module inside and is responsible for data processing and communication with other modules.
[0090] The camera module is OmniVision OV5640, a 5-megapixel CMOS camera module, which is used to collect RGB video data of sign language actions, facial expressions, and the blind path environment.
[0091] The microphone module is INMP441, an I2S digital microphone, which is used to collect the voices of the converser and environmental sounds.
[0092] The speaker module is a TPA2012D1 audio power amplifier and an 8Ω / 0.5W micro speaker, which supports mono audio output and has a low-power design. It is used for voice broadcast to output translation results or navigation prompts.
[0093] The Internet of Things communication module is a C417MeSIM chip, which supports remote configuration and management and provides network connection traffic for the embedded processing unit.
[0094] The power management module is an MP2315 buck DC-DC converter, which supports 3.3V / 5V output and low quiescent current, provides stable power management for the entire system, and supports multiple input voltages (such as battery or USB charging).
[0095] Housing design: Made of carbon fiber and silicone materials, with a weight ≤ 50g, dimensions of 3cm × 2.5cm × 1cm, and an IP67 waterproof rating.
[0096] 1.2 Hardware circuit design and assembly
[0097] Schematic design: Use EDA tools (such as Altium Designer) to draw the schematic diagram, with the embedded processing unit as the core, connect the camera module, microphone module, speaker module, Internet of Things communication module, and power management module, and ensure that the interfaces of all modules are correctly connected, such as the DVP interface of the camera, the I2S interface of the microphone, the UART / SPI interface of the communication module, etc.
[0098] PCB layout design: Design a compact PCB layout according to the size requirements of the wearable device. Place the embedded processing unit in the center of the PCB, and then layout the camera interface, microphone interface, speaker interface, and communication module around it in sequence. The power management module is placed near the power input terminal to ensure the stability of power distribution. Adopt a multi-layer PCB design to ensure signal integrity (SI) and power integrity (PI).
[0099] Component soldering and assembly: Select a suitable soldering process for component soldering. Use surface mount devices (SMD) to reduce the PCB size and improve reliability. Conduct electrical tests after soldering to ensure that all components are correctly connected and there are no short circuits or open circuits.
[0100] Housing design and integration: Design the housing of the wearable device according to the PCB size, using lightweight materials (such as carbon fiber and silicone materials).
[0101] Ensure that the opening designs of modules such as the camera, microphone, and speaker are reasonable and do not affect their functions. Design replaceable battery or USB charging openings for the long-term use of the device.
[0102] 2. Software implementation and system operation process
[0103] 2.1 Two-way sign language translation subsystem
[0104] A. Sign language to speech translation:
[0105] a. Data acquisition: The camera continuously captures sign language action videos and transmits them to the cloud server.
[0106] b. Key point extraction: Use the open-source MediaPipe model to extract hand and face key points and generate a sequence of key point coordinates.
[0107] c. Graph construction and encoding: Construct a dynamic graph based on the natural connection relationship of the key point sequence, and generate vector encoding through hypergraph convolution.
[0108] Among them, the natural connection relationship of the hand key point sequence is as Figure 7 shown, and the natural connection relationship of the face key point sequence is as Figure 8 shown.
[0109] Text encoding: Encode the action text description based on the Mamba-2 State Space Model (SSM). For example, the action text description "restore" corresponds to "stretch both hands flat, palms down, then flip, palms up"; "benefit" corresponds to "benefit, (1) pat the left palm with the right hand, then extend the thumb upward; (2) extend one finger horizontally and keep it still".
[0110] d. Alignment and similarity measurement: Align and calculate the similarity between the above visual encoding and text encoding through a measurement network, and finally output the text vector with the highest similarity to the visual encoding.
[0111] e. Optimization output of the large language model: Use the fine-tuned large language model DeepSeek-R1-Distill-Qwen-7B to map the output text vector to the corresponding sign language morpheme, and further convert the sign language morpheme into fluent Chinese text.
[0112] For example, the sign language morpheme sequence "what / person / shout / hero / can / ?", corresponds to the Chinese text "What kind of people can be called heroes?"; "he / child / time / start / do / lawyer / hope / .", corresponds to the Chinese text "He has wanted to be a lawyer since childhood."
[0113] f. Speech synthesis: Convert the text into speech through the open-source Fast-TTS model and send it back to the device speaker for playback.
[0114] B. Speech to sign language translation:
[0115] a. Speech collection: Record speech data with a microphone and transmit it to the cloud server.
[0116] b. Speech to text: Use the open-source Whisper model to convert speech into Chinese text.
[0117] c. Sign language generation: Generate a sign language digital human animation from the Chinese text through a 3D digital human engine and transmit it to the mobile client for display.
[0118] 2.2 Visual impairment navigation subsystem
[0119] A. Data collection: The camera collects the video of the blind path environment in real time and transmits it to the cloud server.
[0120] B. Blind path positioning: Based on the open-source MVANet model, a binary image of the blind path is generated, and the center line of the blind path is extracted.
[0121] C. Obstacle detection: The open-source Hyper-YOLO model is used to detect obstacles, and the distance is estimated by combining monocular ranging.
[0122] D. Path planning: Based on the A* algorithm, an obstacle avoidance path and navigation prompts are generated.
[0123] E. Speech synthesis: The navigation prompts are converted into speech through the Fast-TTS model and sent back to the device speaker for broadcasting.
[0124] 2.3 Text-to-speech and speech-to-text bidirectional conversion subsystem
[0125] Text-to-speech (TTS): Input Chinese text, generate speech through the Fast-TTS model, and broadcast it through the speaker.
[0126] Speech-to-text (STT): Input speech data, generate Chinese text through the Whisper model for subsequent processing.
[0127] 3. User interaction process
[0128] 3.1 Usage process for deaf-mute users
[0129] Sign language to speech translation: The user makes sign language gestures, the device collects the video and transmits it to the cloud server, and the device broadcasts the translated speech prompts.
[0130] Speech to sign language translation: The user inputs speech through the microphone, the device collects the audio and transmits it to the cloud server, and the mobile client displays the 3D digital human sign language animation.
[0131] 3.2 Usage process for visually impaired users
[0132] Navigation startup: The user starts the system through a voice command (such as "start navigation").
[0133] Environmental perception: The device collects the video of the blind path environment and transmits it to the cloud server.
[0134] Voice prompt: The device broadcasts navigation prompts (such as "There is a shared bicycle 2 meters ahead, please turn right").
[0135] 4. Testing and verification
[0136] 4.1 Functional testing
[0137] Sign language translation test: Input 100 groups of sign language actions, and test that the translation accuracy rate is ≥92%.
[0138] Navigation test: Test the navigation accuracy in indoor and outdoor environments, and the obstacle avoidance success rate is ≥95%.
[0139] 4.2 Performance test
[0140] Latency test: The latency from sign language to speech is ≤300 ms, and the latency from speech to sign language is ≤20 ms.
[0141] Battery life test: The continuous running time of the device is ≥8 hours.
[0142] Method embodiments
[0143] According to an embodiment of the present invention, a multi-modal integrated assistance method for special needs groups is provided. Figure 9 It is a flowchart of the multi-modal integrated assistance method for special needs groups according to the embodiment of the present invention. As Figure 9 shown, the multi-modal integrated assistance method for special needs groups according to the embodiment of the present invention specifically includes:
[0144] Step S901, after collecting basic information through the basic function subsystem, transmit it to the embedded processing unit, and provide assistance services for special needs groups according to the relevant content from the embedded processing unit; wherein, the basic information includes image data and audio data; the relevant content includes sign language translation text, sign language translation voice data, optimal path guidance text, optimal path guidance voice data, speech conversion text, and 3D digital sign language action data; specifically including:
[0145] Collect image data through the camera module in the basic function subsystem and transmit it to the embedded processing unit; wherein, the image data includes the sign language hand movements, facial expressions of special needs groups, and image data of the blind path environment.
[0146] Collect audio data through the microphone module in the basic function subsystem and transmit it to the embedded processing unit; wherein, the audio data includes the conversation content of the conversants and the external environment sounds.
[0147] Perform voice broadcast of the sign language translation voice data and the optimal path guidance voice data from the embedded processing unit through the speaker module in the basic function subsystem.
[0148] Provide traffic networking communication services for the system through the Internet of Things communication module in the basic function subsystem.
[0149] Provide power management services for the system through the power management module in the basic function subsystem.
[0150] Step S902: Drive each module in the basic function subsystem through the embedded processing unit, transmit the received basic information to the cloud server, and download the relevant content from the cloud server to the basic function subsystem;
[0151] Step S903: Perform two-way sign language translation, visually impaired navigation, and two-way text-to-speech conversion through the two-way sign language translation subsystem, visually impaired navigation subsystem, and text-to-speech two-way conversion subsystem in the cloud server and based on the basic information, and send the generated relevant content to the embedded processing unit or the client; specifically including:
[0152] The hand skeleton map perception module in the two-way sign language translation subsystem processes the image data of sign language hand movements and facial expressions using a hand skeleton map perception prior pre-training model based on contrastive learning to generate corresponding sign language morpheme sequences, and transmits the sign language morpheme sequences to the morpheme conversion module; wherein, the hand skeleton map perception prior pre-training model includes a visual encoder, a text encoder, and a metric network;
[0153] The morpheme conversion module in the two-way sign language translation subsystem converts the sign language morpheme sequences using a morpheme conversion model based on a large language model to generate sign language translation texts, and transmits the sign language translation texts to the text-to-speech two-way conversion subsystem;
[0154] The 3D digital human engine module in the two-way sign language translation subsystem generates corresponding 3D digital human sign language action data according to the speech conversion text from the text-to-speech two-way conversion subsystem, and sends the 3D digital human sign language action data to the client;
[0155] Step S904: Visualize the relevant content through the client.
[0156] The embodiment of the present invention is a method embodiment corresponding to the above system embodiment. The specific operations of each step can be understood with reference to the description of the system embodiment, and will not be elaborated here.
[0157] In summary, the embodiment of the present invention adopts the form of a wearable device hardware architecture, which has better comfort and flexibility compared with existing sign language translation tools. Among them, the two-way sign language translation subsystem proposed in the embodiment of the present invention can help the deaf and mute population communicate seamlessly with the outside world; the visually impaired navigation subsystem can effectively solve the problem of inconvenient travel for visually impaired people.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A multimodal integrated assistance system for special needs groups, characterized in that, Adopt an embedded wearable hardware architecture, and the system includes: A basic function subsystem, connected to the embedded processing unit, for collecting basic information and then transmitting it to the embedded processing unit, and providing auxiliary services for special needs groups according to relevant content from the embedded processing unit; wherein, the basic information includes image data and audio data; the relevant content includes sign language translation text, sign language translation voice data, optimal path guidance text, optimal path guidance voice data, speech conversion text, and 3D digital human sign language action data; An embedded processing unit, connected to the basic function subsystem and the cloud server, for driving each module in the basic function subsystem, and transmitting the received basic information to the cloud server, and downloading the relevant content from the cloud server to the basic function subsystem; A cloud server, connected to the embedded processing unit and the client, and deployed with a two-way sign language translation subsystem, a visually impaired navigation subsystem, and a text-to-speech and speech-to-text conversion subsystem that communicate with each other, for performing two-way sign language translation, visually impaired navigation, and text-to-speech and speech-to-text conversion based on the basic information through the two-way sign language translation subsystem, the visually impaired navigation subsystem, and the text-to-speech and speech-to-text conversion subsystem, and sending the generated relevant content to the embedded processing unit or the client; A client, connected to the cloud server, for visually displaying the relevant content.
2. The system according to claim 1, wherein The basic function subsystem includes a camera module, a microphone module, a speaker module, an Internet of Things communication module, and a power management module that communicate with each other; The camera module is used for collecting image data and then transmitting it to the embedded processing unit; wherein, the image data includes image data of the sign language hand movements, facial expressions of special needs groups, and the blind path environment; The microphone module is used for collecting audio data and then transmitting it to the embedded processing unit; wherein, the audio data includes the conversation content of the conversants and the external environment sounds; The speaker module is used for voice broadcasting the sign language translation voice data and the optimal path guidance voice data from the embedded processing unit; The Internet of Things communication module is used for providing traffic networking communication services for the system; The power management module is used for providing power management services for the system.
3. The system according to claim 2, wherein The two-way sign language translation subsystem specifically includes: A hand skeleton map perception module, connected to the morpheme conversion module, for processing the image data of sign language hand movements and facial expressions through a hand skeleton map perception prior pre-training model based on contrastive learning to generate corresponding sign language morpheme sequences, and transmitting the sign language morpheme sequences to the morpheme conversion module; wherein, the hand skeleton map perception prior pre-training model includes a visual encoder, a text encoder, and a metric network; A morpheme conversion module, connected to the hand skeleton map perception module and the text-to-speech and speech-to-text conversion subsystem, for converting the sign language morpheme sequences through a morpheme conversion model based on a large language model to generate sign language translation text, and transmitting the sign language translation text to the text-to-speech and speech-to-text conversion subsystem; The 3D digital human engine module is connected to the text-to-speech and speech-to-text conversion subsystem and the user terminal, and is used to generate corresponding 3D digital human sign language action data according to the speech conversion text from the text-to-speech and speech-to-text conversion subsystem, and send the 3D digital human sign language action data to the user terminal.
4. The system according to claim 3, wherein The hand bone map perception module is specifically used for: Inputting the image data of sign language hand actions and facial expressions into the human pose estimation model in the visual encoder to extract key points, obtaining a sequence of key point coordinates, constructing a corresponding dynamic graph according to the natural connection relationship of the sequence of key point coordinates, and performing multiple aggregations on the dynamic graph through a graph neural network to generate a visual vector encoding representing hand bones and facial expressions; Using the text encoder to perform multiple calculations on the text description of the sign language action to generate a text vector encoding representing the sign language action; Aligning and calculating the similarity between the visual vector encoding and the text vector encoding through a metric network, and outputting a sign language morpheme sequence corresponding to the sign language action text description with the highest similarity to the visual vector encoding.
5. The system according to claim 3, wherein The visually impaired navigation subsystem specifically includes: The blind path positioning module is connected to the path planning module, and is used to segment the image data of the blind path environment into a binary image of the blind path by using a blind path positioning model based on a semantic segmentation algorithm, and output a sequence of blind path center point coordinates, and transmit the sequence of blind path center point coordinates to the path planning module; The obstacle positioning module is connected to the path planning module, and is used to detect the image data of the blind path environment through an obstacle positioning model based on an object detection algorithm, output a sequence of obstacle center point coordinates, and transmit the sequence of obstacle center point coordinates to the path planning module; The monocular depth estimation module is connected to the path planning module, and is used to convert the image data of the blind path environment into a depth image through a monocular depth estimation model, and calculate the straight-line distance between each pixel in the depth image and the current viewing angle; The path planning module is connected to the blind path positioning module, the obstacle positioning module, the monocular depth estimation module and the text-to-speech and speech-to-text conversion subsystem, and is used to obtain a sequence of spatial distances corresponding to the sequence of blind path center point coordinates and the sequence of obstacle center point coordinates based on the straight-line distance between each pixel in the depth image and the current viewing angle, calculate the current optimal obstacle avoidance path by using a path planning model designed based on the A* algorithm according to the sequence of spatial distances, and generate a corresponding optimal path guidance text according to the current optimal obstacle avoidance path, and send the optimal path guidance text to the text-to-speech and speech-to-text conversion subsystem.
6. The system according to claim 5, wherein The text-to-speech and speech-to-text conversion subsystem specifically includes: The text-to-speech module is connected to the morpheme conversion module, the path planning module and the embedded processing unit, and is used to convert the sign language translation text and the optimal path guidance text into sign language translation speech data and optimal path guidance speech data through a text-to-speech model, and download the sign language translation speech data and the optimal path guidance speech data to the embedded processing unit; A speech-to-text module, connected to the embedded processing unit and the 3D digital human engine module, is used to convert audio data from the embedded processing unit into speech-converted text through a speech-to-text model, and send the speech-converted text to the 3D digital human engine module.
7. The system according to claim 1, characterized in that, The client is specifically used for: Visually display the sign language translation text, the optimal path guidance text, and the speech-converted text, and dynamically display the 3D digital human sign language action data.
8. A multimodal integrated assistance method for special needs groups, characterized in that, It includes: After collecting basic information through the basic function subsystem, transmit it to the embedded processing unit, and provide auxiliary services for special needs groups according to the relevant content from the embedded processing unit; among them, the basic information includes image data and audio data; the relevant content includes sign language translation text, sign language translation voice data, optimal path guidance text, optimal path guidance voice data, speech-converted text, and 3D digital human sign language action data; Drive each module in the basic function subsystem through the embedded processing unit, transmit the received basic information to the cloud server, and download the relevant content from the cloud server to the basic function subsystem; Through the two-way sign language translation subsystem, visually impaired navigation subsystem, and text-speech two-way conversion subsystem in the cloud server, and based on the basic information, perform two-way sign language translation, visually impaired navigation, and text-speech two-way conversion, and send the generated relevant content to the embedded processing unit or the client; Visually display the relevant content through the client.
9. The method according to claim 8, characterized in that After collecting basic information through the basic function subsystem and transmitting it to the embedded processing unit, providing auxiliary services for special needs groups according to the relevant content from the embedded processing unit specifically includes: Collect image data through the camera module in the basic function subsystem and transmit it to the embedded processing unit; among them, the image data includes image data of the sign language hand movements, facial expressions of special needs groups, and the blind path environment; Collect audio data through the microphone module in the basic function subsystem and transmit it to the embedded processing unit; among them, the audio data includes the conversation content of the converser and the external environment sound; Play the sign language translation voice data and the optimal path guidance voice data from the embedded processing unit through the speaker module in the basic function subsystem; Provide traffic networking communication services for the system through the Internet of Things communication module in the basic function subsystem; Provide power management services for the system through the power management module in the basic function subsystem.
10. The method according to claim 9, wherein Performing two-way sign language translation based on the basic information through the two-way sign language translation subsystem in the cloud server specifically includes: The hand skeleton map perception module in the two-way sign language translation subsystem uses a hand skeleton map perception prior pre-training model based on contrast learning to process the image data of sign language hand movements and facial expressions, generate the corresponding sign language morpheme sequence, and transmit the sign language morpheme sequence to the morpheme conversion module; among them, the hand skeleton map perception prior pre-training model includes a visual encoder, a text encoder, and a metric network; The morpheme conversion module in the two-way sign language translation subsystem uses a morpheme conversion model based on a large language model to convert the sign language morpheme sequence, generate a sign language translation text, and transmit the sign language translation text to the text-to-speech and speech-to-text conversion subsystem; The 3D digital human engine module in the two-way sign language translation subsystem generates corresponding 3D digital human sign language action data according to the speech conversion text from the text-to-speech and speech-to-text conversion subsystem, and sends the 3D digital human sign language action data to the user terminal.
Citation Information
Patent Citations
Intelligent interaction system and intelligent blind guiding walking stick
CN119345004A
Sign language translation method based on motion sensor of intelligent wrist-worn device
CN119358697A
Multifunctional sign language translation glove
CN221554755U
Body language translation system and method
CN108766433A
Wearable visual auxiliary system and method based on deep learning and fuzzy control
CN115317329A
Cited By
Dynamic tracking detection method and device for disabled people
CN121122708A
Array speech enhancement method and system based on Hemma optimization and Mama
CN121506167A