Barrier-free reply system
The barrier-free response system recognizes external voices and scenes, converts them into sound output to output the deaf-mute person's response content, solves the communication barriers of the deaf-mute and enables effective communication.
Patent Information
- Application Number
- CN202510618542.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-09-26
AI Technical Summary
Deaf-mute people are unable to play output content in the form of sound when communicating, resulting in communication barriers.
An accessible reply system is adopted, including a recognizer, a voice processor, a reply content storage, a reply content selector and a reply content outputter, which recognizes external voices and scenes and converts them into sound data output.
It enables deaf-mute people to communicate effectively with others, outputs the deaf-mute people's responses through voice, and reduces communication barriers.
Smart Images

Figure CN120708607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent devices, and in particular to a barrier-free reply system. Background Art
[0002] Because deaf-mute people cannot hear or speak, they face many communication barriers in their life, study and work.
[0003] Deaf-mute people face many challenges in finding work due to language barriers. AR glasses, smartphones, and artificial intelligence technologies can significantly reduce these communication barriers. However, while existing electronic devices like smartphones can help deaf-mute people communicate, they cannot provide the necessary voice input.
[0004] Therefore, in the prior art, there is a technical problem that the content output by the deaf-mute cannot be played in the form of sound when the deaf-mute person is communicating. Summary of the Invention
[0005] The barrier-free answering system provided by the present invention solves the technical problem in the prior art that the content output by the deaf-mute person cannot be played in the form of sound during a conversation between the deaf-mute person and the system.
[0006] Some implementation plans adopted to solve the above technical problems include: An accessible reply system includes a recognizer; a speech processor in communication with the recognizer; a reply content memory, wherein the reply content memory stores reply content data; a reply content selector, the reply content selector being used to select the reply content data stored in the reply content memory; and a reply content outputter, wherein the reply content outputter outputs the reply content in a voice output manner; wherein the identifier identifies the current scene; When there is voice output from the outside world, the recognizer recognizes the voice from the outside world; when there is no voice output from the outside world, the recognizer continuously monitors whether there is voice output from the outside world; The speech processor receives the external speech recognized by the recognizer and determines the reply content data according to the current scene and the external speech; The voice processor converts the reply content data into sound data after determining the reply content data; The reply content outputter plays sound data according to the reply content selected by the reply content selector.
[0007] Preferably, the identifier comprises a microphone.
[0008] Preferably, the identifier further includes a camera.
[0009] Preferably, the voice processor communicates with a cloud server, and the camera communicates with the cloud server through the voice processor.
[0010] Preferably, the cloud server determines the current scene based on the image captured by the camera.
[0011] Preferably, the barrier-free answering system also includes a display.
[0012] Preferably, the display displays the reply content data selected by the reply content selector in text form.
[0013] Preferably, the reply content selector is a keyboard, or the reply content selector is a touch screen.
[0014] Preferably, the reply content outputter is a speaker.
[0015] Preferably, the speaker is provided with a fixer for fixing the speaker to the human body.
[0016] Compared with the prior art, the present invention has the following advantages: According to the application scenario and context, several replies are preset for the user's questions for the deaf-mute person to choose from. After the deaf-mute person selects a reply, the reply is converted into voice and played to the other party in the form of sound through the reply content output device carried by the deaf-mute person, so that the deaf-mute person can communicate and interact with others effectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] For the purpose of explanation, several embodiments of the present invention are described in the following figures. The following figures are incorporated into this document and constitute a part of the detailed description. In some cases, well-known structures and components are shown in block diagram form to avoid obscuring the concepts of the present invention.
[0018] Figure 1 Schematic diagram of the present invention.
[0019] Figure 2 Flowchart of the present invention. DETAILED DESCRIPTION
[0020] The specific embodiment shown below is intended to be a description of the various configurations of the subject technology of the present invention, and is not intended to represent that the subject technology of the present invention can be put into practice. The specific embodiment includes that specific details are intended to provide a thorough understanding of the subject technology of the present invention. However, it will be clear and apparent to those skilled in the art that the subject technology of the present invention is not limited to the specific details shown herein, and can be put into practice without these specific details.
[0021] It will be understood that, herein, relational terms such as “first” and “second” are intended to distinguish one entity or operation from another entity or operation, and are not intended to express or imply any actual relationship or order between these entities or operations.
[0022] The terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0023] Reference Figures 1 to 2 As shown, a barrier-free reply system includes a recognizer; a speech processor in communication with the recognizer; a reply content memory, wherein the reply content memory stores reply content data; a reply content selector, the reply content selector being used to select the reply content data stored in the reply content memory; and a reply content outputter, wherein the reply content outputter outputs the reply content in a voice output manner; wherein the identifier identifies the current scene; When there is voice output from the outside world, the recognizer recognizes the voice from the outside world; when there is no voice output from the outside world, the recognizer continuously monitors whether there is voice output from the outside world; The speech processor receives the external speech recognized by the recognizer and determines the reply content data according to the current scene and the external speech; The voice processor converts the reply content data into sound data after determining the reply content data; The reply content outputter plays sound data according to the reply content selected by the reply content selector.
[0024] In some embodiments, the identifier comprises a microphone.
[0025] In some embodiments, the identifier further includes a camera.
[0026] In some embodiments, the voice processor communicates with a cloud server, and the camera communicates with the cloud server through the voice processor.
[0027] In some embodiments, the cloud server determines the current scene based on the image captured by the camera.
[0028] In some embodiments, the accessible response system further includes a display.
[0029] In some embodiments, the display displays the reply content data selected by the reply content selector in text form.
[0030] In some embodiments, the reply content selector is a keyboard, or the reply content selector is a touch screen.
[0031] In some embodiments, the reply content outputter is a speaker.
[0032] In some embodiments, the speaker is provided with a fixer for fixing the speaker to a human body.
[0033] The recognizer serves as the system's "ears" and "eyes," responsible for capturing and interpreting information from the external environment. It primarily consists of the following components: Microphone: This captures external voice input. When there's voice output, the microphone quickly responds and captures these sound signals. Furthermore, when there's no voice output, the microphone continuously monitors to ensure that no immediate voice commands or inquiries are missed. Camera: This captures the surrounding environment, providing the system with richer context. This is crucial for identifying specific scenes, human actions, or objects, helping the system make more precise responses.
[0034] The speech processor is the "brain" of the system, responsible for processing the information transmitted by the recognizer and making corresponding decisions. It performs the following tasks: Receive and parse speech: Receive the speech signal captured by the microphone and parse it to extract meaningful content.
[0035] Scenario-to-Speech Matching: Based on the current scenario identified by the recognizer or cloud server (e.g., camera-assisted recognition), combined with the parsed speech content, the most appropriate response data is intelligently selected and determined. The selected response data is converted into audio data for subsequent playback via the response output device.
[0036] The response storage is a vast knowledge base storing data on a wide range of possible responses. This data covers a wide range of fields and topics, ensuring the system can respond to a variety of scenarios and questions. The response selector allows users to manually adjust or select responses, adding flexibility and personalization to the system. It can be: Keyboard: Traditional keyboard input, suitable for users familiar with keyboard operations. Touchscreen: An intuitive and easy-to-use touchscreen interface that supports tapping or swiping to select responses, particularly suitable for visually impaired users. The response output device is responsible for playing the processed audio data to the user. It primarily consists of the following components: Speaker: A high-fidelity speaker ensures that responses are clearly and accurately conveyed to the user. To enhance ease of use, the speaker is equipped with a mount that allows it to be conveniently attached to a user's body part, such as the shoulder, waist, or backpack. The voice processor's ability to communicate with a cloud server adds additional intelligence and flexibility to the system. Camera images can be uploaded to the cloud server through the voice processor for in-depth analysis and scene recognition. This helps the system more accurately understand the current environment and provide more appropriate responses. To further enhance the system's accessibility, a display can be added as an auxiliary output device. The display shows the reply content data selected by the reply content selector in text form, providing convenience for users with visual impairments or scenarios requiring double confirmation.
[0037] One application scenario could be a deaf-mute cafe clerk. When a customer approaches, the clerk wears a loudspeaker and actively speaks, such as "Good morning" or "Good afternoon." When the customer needs to place an order, the clerk uses a small numeric keypad to select the answer, such as 1: Do you want sugar? 2. Do you want milk? 3. Large or small cup, etc., and then automatically announces it through the loudspeaker.
[0038] The above describes the subject technical solution and corresponding details of the present invention. It can be understood that the above description is only some implementation plans of the subject technical solution of the present invention, and some details may be omitted during its specific implementation.
[0039] In addition, in some embodiments of the above invention, multiple embodiments may be implemented in combination. Due to space limitations, various combination schemes are not listed one by one. Those skilled in the art can freely combine and implement the above embodiments as needed in specific implementation to obtain a better application experience.
[0040] When implementing the subject technical solution of the present invention, those skilled in the art can obtain other detailed configurations or drawings based on the subject technical solution of the present invention and the drawings. Obviously, these details still fall within the scope covered by the subject technical solution of the present invention without departing from the subject technical solution of the present invention.
Claims
1. A barrier-free answering system, characterized by: It includes a recognizer; a speech processor, the speech processor communicates with the recognizer; a reply content memory, the reply content memory storing reply content data; a reply content selector, the reply content selector being used to select the reply content data stored in the reply content memory; and a reply content outputter, the reply content outputter outputs the reply content in the form of sound output; wherein, the recognizer recognizes the current scene; when there is voice output from the outside world, the recognizer recognizes the voice from the outside world, and when there is no voice output from the outside world, the recognizer continuously monitors whether there is voice output from the outside world; the voice processor receives the external voice recognized by the recognizer, and determines the reply content data according to the current scene and the external voice; after determining the reply content data, the voice processor converts the reply content data into sound data; the reply content outputter plays the sound data according to the reply content selected by the reply content selector.
2. The barrier-free answering system according to claim 1, characterized in that: The identifier includes a microphone.
3. The barrier-free answering system according to claim 2, characterized in that: The identifier further includes a camera.
4. The barrier-free answering system according to claim 3, wherein: The voice processor communicates with the cloud server, and the camera communicates with the cloud server through the voice processor.
5. The barrier-free answering system according to claim 4, characterized in that: The cloud server determines the current scene based on the image captured by the camera.
6. The barrier-free answering system according to claim 1, wherein: The barrier-free answering system also includes a display.
7. The barrier-free answering system according to claim 6, characterized in that: The display displays the reply content data selected by the reply content selector in text form.
8. The barrier-free answering system according to claim 1, wherein: The reply content selector is a keyboard, or the reply content selector is a touch screen.
9. The barrier-free answering system according to claim 1, characterized in that: The reply content output device is a speaker.
10. The barrier-free answering system according to claim 9, characterized in that: The speaker is provided with a fixer for fixing the speaker to a human body.