Smart glasses system including smartphone for the conversation of the disabled persons
Patent Information
- Application Number
- KR1020250053887
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2026-08-11
- Estimated Expiration
- 2045-04-24
Smart Images

Figure 112025046821091-PAT00005_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a smart glasses system for conversation with disabled people, and to a technology that enables ordinary people to converse with disabled people without an intermediary assistant for conversation or communication, and enables disabled people with different disabilities to converse directly without an intermediary assistant. Background Technology
[0003] In most conversations between people with speech, hearing, or visual impairments and the general public, or among people with disabilities themselves, an intermediary is absolutely necessary. This intermediary is proficient in sign language and plays a crucial role as a communicator. Conversation is impossible without an intermediary, and their role is already indispensable in television and broadcast media. However, the inconveniences faced by people with disabilities in personal activities and daily life are indescribable, and the general public often fails to understand their intentions.
[0004] Hearing-impaired and speech-impaired individuals who use sign language cannot converse or communicate with visually impaired people, the general public, or hearing-impaired individuals who use lip-reading or hearing aids without an intermediary. Although assistive devices such as sign language interpretation gloves have been developed, they cause inconvenience during conversation due to a lack of convenience and dysfunction.
[0005] Existing Google Glass and Meta's smart glasses (Meta Ray-Ban Smart Glasses) are not designed for conversation with people with disabilities and are not smart glasses intended for people with disabilities, so they cannot be used for conversation between able-bodied people and people with disabilities, or even for conversation between people with other disabilities. Prior art literature
[0006] Republic of Korea Registered Patent 10-1885473 The problem to be solved
[0007] The present invention was devised to solve the above-mentioned problems, and aims to provide a smart glasses system that enables ordinary people to converse with disabled people without an intermediary assistant for conversation or communication, and enables disabled people with different disabilities to converse directly without an intermediary assistant.
[0008] The objectives of the present invention are not limited to those mentioned above, and other unmentioned objectives will be clearly understood by a person skilled in the art from the description below. means of solving the problem
[0010] The present invention for achieving such objectives relates to a smart glass system, comprising smart glasses that are worn on a user's face, collect data related to a conversation between the user and another party, transmit the collected data to a smartphone, and output result data received from the smartphone, and a smartphone that analyzes and processes data received from the smart glasses based on an AI agent function and transmits the processed result data to the smart glasses.
[0011] The smart glasses may comprise an input unit for collecting conversation-related data of a user or a counterpart, a user adjustment selection unit for processing input signals input by the user, an output unit for outputting result data received from the smartphone, a first Bluetooth communication unit for communicating with the smartphone via Bluetooth, a first transmission / reception signal adjustment processing unit for multiplexing, demultiplexing, and packet processing of signals transmitted and received with the smartphone, and a rechargeable battery for supplying power to the smart glasses.
[0012] The smart glasses may further include an input tab located in a part of the smart glasses, which has a built-in touch sensor, and for receiving operation control commands for the smart glasses from a user.
[0013] The above input unit may include a camera sensor for collecting video data related to the conversation between the user and the other party or video data of the surrounding environment, and a microphone sensor for collecting audio data related to the conversation between the user and the other party or audio data of the surrounding environment.
[0014] The above output unit may be capable of displaying in an AR (Augmented Reality) manner and may comprise an AR display for outputting text, sign language images, or motion images processed by the smartphone, and a bone conduction speaker for outputting voice and vibration processed by the smartphone.
[0015] The smartphone may comprise a second Bluetooth communication unit for communicating with the smart glasses via Bluetooth, a second transmission / reception signal adjustment processing unit for multiplexing, demultiplexing, and packet processing of signals transmitted and received with the smart glasses, and a disabled person conversation AI processing unit that performs conversion and processing necessary for conversation between disabled people or conversation between disabled people and non-disabled people based on an AI agent function for data collected from the input unit according to an input signal input from the smart glasses.
[0016] The above disabled person conversation AI processing unit may comprise an input conversion AI processing unit for converting data collected from the input unit into text or generating and processing user surrounding guidance information data using the collected data, and an output conversion AI processing unit for converting the text or surrounding guidance information data processed by the input conversion AI processing unit into an output form according to the setting of the input signal input from the user adjustment selection unit and transmitting it to the output unit.
[0017] The above input conversion AI processing unit performs image analysis and object detection on image data collected through the camera sensor, generates object information guidance through object recognition to recognize objects including obstacles and stairs around the user for surrounding guidance, recognizes sign language and performs Natural Language Processing (NLP) to convert it into text, performs audio analysis on audio data collected through the microphone sensor, performs noise canceling and speech recognition, and performs Natural Language Processing to convert it into text.
[0018] The above input conversion AI processing unit, in recognizing sign language, detects the hands of the conversation partner through image analysis and object detection to guide the conversation content in a situation where a speech-using user is conversing with a sign-using person, recognizes gestures and movements according to the movements of the detected hands, recognizes the sign language corresponding to the recognized gestures and movements, translates the recognized sign language, and converts it into text through natural language processing.
[0019] The above input conversion AI processing unit, in object recognition, can generate object information guidance data by detecting objects including roads and buildings on a path through image analysis and object detection to guide the user with information about the surrounding environment and objects in daily life when the user is visually impaired, and by determining obstacles including stairs around the user through object recognition.
[0020] The above input conversion AI processing unit can detect the voice signal of the other party along the path through audio analysis to guide the conversation content in a situation where a user using sign language is conversing with a person using spoken language in the context of speech recognition, recognize the conversation content through noise canceling and speech recognition on the detected voice signal, and convert it into text through natural language processing.
[0021] The above input conversion AI processing unit can recognize ambient sounds and ambient voices by classifying them according to the nature of sounds through audio analysis to provide guidance on situations and dangers occurring in the surroundings during daily life when the user is hearing impaired, determine whether an accident has occurred including the presence of a siren sound using the recognized ambient sounds and ambient voices, and generate guidance data regarding situations and dangers based on the determined accident occurrence.
[0022] The output conversion AI processing unit can set the magnitude and location of vibration according to the direction, size, and type of the object, generate an audio signal including voice and sound for voice guidance, generate a vibration signal for vibration guidance, and transmit the generated audio signal and vibration signal to the bone conduction speaker through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit.
[0023] The output conversion AI processing unit can convert the text processed by the input conversion AI processing unit into a display signal for display when outputting in text form according to the setting of the input signal input from the user adjustment selection unit, and can transmit the converted display signal to the AR display through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit so as to provide an Augmented Reality (AR) screen by applying an AR overlay.
[0024] The output conversion AI processing unit can convert the text processed by the input conversion AI processing unit into a display signal for displaying the text as a sign language image and motion according to the setting of the input signal input from the user adjustment selection unit, and can transmit the converted display signal to the AR display through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit so as to provide an Augmented Reality (AR) screen by applying an AR overlay.
[0025] When the output conversion AI processing unit outputs in the form of voice audio according to the setting of the input signal input from the user adjustment selection unit, it converts the text processed by the input conversion AI processing unit into an audio signal including voice and sound, and can transmit the converted audio signal to the bone conduction speaker through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit.
[0026] The output conversion AI processing unit can convert surrounding guidance information data, including surrounding conditions and danger situations processed by the input conversion AI processing unit, to generate guidance signals in the form of images and text, and can transmit the generated guidance signals to the AR display through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit so as to provide them as an Augmented Reality (AR) screen by applying an AR overlay to the generated guidance signals.
[0027] The above user adjustment selection unit can process the user's input signal input from the above input tab or the above microphone sensor and transmit it to the above disabled person conversation AI processing unit.
[0028] The above-mentioned smartphone may further include a wireless communication unit for connecting to an internet network or communicating with other devices.
[0029] The above wireless communication unit may comprise a Wi-Fi communication unit for accessing an internet network via a public Wi-Fi AP or a private Wi-Fi AP, and a cellular communication unit for communicating with other devices via a cellular network.
[0030] The above disabled person conversation AI processing unit can provide a basic setting function that sets default values for the input unit and the output unit according to an input signal input from the above user adjustment selection unit.
[0031] In providing the above-mentioned disability conversation AI processing unit, when an input signal indicating that the user is a hearing-impaired or speech-impaired person who uses sign language is received from the above-mentioned user adjustment selection unit, the above-mentioned disability conversation AI processing unit may provide the above-mentioned basic setting functions as voice recognition and ambient sound analysis functions through the microphone sensor of the input unit, text and sign language output functions through the AR display of the output unit, and guidance functions through vibration via the bone conduction speaker of the output unit.
[0032] In providing the above-mentioned disability conversation AI processing unit, when an input signal indicating that the user is a hearing impaired person using lip reading or a hearing aid is input from the above-mentioned user adjustment selection unit, the above-mentioned disability conversation AI processing unit may provide, as basic setting functions, a sign language recognition function including gestures and movements through the camera sensor of the input unit, an ambient sound analysis function through the microphone sensor of the input unit, a text output function through the AR display of the output unit, and a guidance function through vibration via the bone conduction speaker of the output unit.
[0033] In providing the above-mentioned disability conversation AI processing unit, when an input signal indicating that the user is visually impaired is input from the above-mentioned user adjustment selection unit, the above-mentioned disability conversation AI processing unit may provide sign language recognition and object recognition functions including gestures and movements through the camera sensor of the input unit, and audio output through the bone conduction speaker of the output unit and guidance functions through vibration as basic setting functions.
[0034] In providing the above-mentioned disabled conversation AI processing unit, when an input signal indicating "ordinary person" is input from the above-mentioned user adjustment selection unit, the above-mentioned disabled conversation AI processing unit may provide a sign language recognition function including gestures and movements through the camera sensor of the input unit, a text output function through the AR display of the output unit, and an audio output function through the bone conduction speaker of the output unit as basic setting functions.
[0035] The above disabled person conversation AI processing unit can access a specific website through the above wireless communication unit to download and install software related to disabled person conversation or AI agent functions.
[0036] The above disabled person conversation AI processing unit can access a specific website through the above wireless communication unit to download and install update software that includes additional functions and improves functions related to disabled person conversation or AI agent functions. Effects of the invention
[0038] According to the present invention, ordinary people can converse with disabled people without an intermediary assistant for conversation or communication, and disabled people with different disabilities can converse directly without an intermediary assistant. Brief explanation of the drawing
[0040] Figure 1 shows an example of the use and application of smart glasses for conversation with disabled people according to the present invention. Figure 2 shows the conversion function and operation example of the smart glasses for disabled persons' conversation of the present invention. FIG. 3 shows an example of the structure and appearance of smart glasses for disabled persons' conversation according to the present invention. Figure 4 shows examples of text and sign language that can be output through the display of the smart glasses system for conversation with disabled people according to the present invention. Figure 5 shows a configuration diagram of user adjustment and selection operations of the smart glasses system for conversation with disabled people according to the present invention, in a cooperative operation mode with a smartphone. FIG. 6 shows an example of user default settings for setting input / output devices and modes to be used according to the user in the smart glasses system for conversation with disabled people according to the present invention. Figure 7 shows a system flowchart of smart glasses for disabled persons' conversation of the present invention when the user is hearing impaired. Figure 8 shows a system flowchart of smart glasses for disabled persons' conversation of the present invention when the user is visually impaired. FIG. 9 shows an example of a flowchart for converting sign language (gestures, movements) into text in a smart glasses system for conversation with disabled people according to the present invention. FIG. 10 shows an example of a flowchart for converting voice recognition results into text in a smart glasses system for conversation with disabled people according to the present invention. FIG. 11 shows an example of a flowchart for outputting text as voice through a bone conduction speaker in a smart glasses system for disabled persons' conversation of the present invention. FIG. 12 shows an example of a flowchart for outputting text in sign language via AR in a smart glasses system for disabled persons' conversation of the present invention. FIG. 13 shows the key transformation algorithm and configuration of an AI agent system for conversation with disabled people when the smart glasses of the present invention are based on a standalone operation mode. Figure 14 shows a comparison of system features and advantages and disadvantages according to the implementation form of the smart glasses of the present invention. Specific details for implementing the invention
[0041] The present invention is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present invention to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention.
[0042] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0043] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.
[0044] Furthermore, in the description referring to the attached drawings, identical components are assigned the same reference numeral regardless of drawing symbols, and redundant descriptions thereof are omitted. In describing the present invention, if it is determined that a detailed description of related prior art could unnecessarily obscure the essence of the present invention, such detailed description is omitted.
[0045] The present invention relates to a smart glasses system that enables conversation between people with disabilities and the general public, or between people with disabilities. That is, to enable people with hearing and speech disabilities to converse or communicate with people with visual impairments, the general public, and people with hearing impairments who use lip reading or hearing aids without an intermediary, the smart glasses enable object recognition, sign language recognition, and voice recognition through a camera and microphone mounted on the smart glasses, and also enable conversation by outputting and providing information in a form desired by the user through an AR (Augmented Reality) display and a bone conduction speaker (154).
[0046] The present invention consists of artificial intelligence processing and system configuration for five basic conversions and four complex application conversions required for conversation with disabled people. That is, the five basic conversions are 1) text-to-speech conversion, 2) speech-to-text conversion, 3) sign language (gesture, movement)-to-text conversion, 4) text-to-display conversion, and 5) text-to-sign language display (image, motion) conversion, and the four complex application conversions are 1) sign language (gesture, movement)-to-speech conversion, 2) sign language (gesture, movement)-to-display conversion, 3) speech-to-display conversion, and 4) speech-to-sign language display (image, motion) conversion.
[0047] The present invention presents a smart glasses system capable of guiding the content of a conversation in a form that the user can intuitively understand by recognizing camera and audio collected by the smart glasses during a conversation with a counterpart using a different communication method (sign language, oral communication). In addition, through the analysis results of camera data and audio data, it is possible to recognize surrounding objects, situations, and the occurrence of dangers that are difficult for the hearing impaired or visually impaired to perceive, and to provide guidance to the user.
[0048] The present invention relates to a smart glasses system, comprising smart glasses that are worn on a user's face, collect data related to a conversation between the user and another party, transmit the collected data to a smartphone, and output result data received from the smartphone, and a smartphone that analyzes and processes data received from the smart glasses based on an AI agent function and transmits the processed result data to the smart glasses.
[0049] The smart glasses may comprise an input unit for collecting conversation-related data of a user or a counterpart, a user adjustment selection unit for processing input signals input by the user, an output unit for outputting result data received from the smartphone, a first Bluetooth communication unit for communicating with the smartphone via Bluetooth, a first transmission / reception signal adjustment processing unit for multiplexing, demultiplexing, and packet processing of signals transmitted and received with the smartphone, and a rechargeable battery for supplying power to the smart glasses.
[0050] The smart glasses may further include an input tab located in a part of the smart glasses, which has a built-in touch sensor, and for receiving operation control commands for the smart glasses from a user.
[0051] The above input unit may include a camera sensor for collecting video data related to the conversation between the user and the other party or video data of the surrounding environment, and a microphone sensor for collecting audio data related to the conversation between the user and the other party or audio data of the surrounding environment.
[0052] The above output unit may be capable of displaying in an AR (Augmented Reality) manner and may comprise an AR display for outputting text, sign language images, or motion images processed by the smartphone, and a bone conduction speaker for outputting voice and vibration processed by the smartphone.
[0053] The smartphone may comprise a second Bluetooth communication unit for communicating with the smart glasses via Bluetooth, a second transmission / reception signal adjustment processing unit for multiplexing, demultiplexing, and packet processing of signals transmitted and received with the smart glasses, and a disabled person conversation AI processing unit that performs conversion and processing necessary for conversation between disabled people or conversation between disabled people and non-disabled people based on an AI agent function for data collected from the input unit according to an input signal input from the smart glasses.
[0054] The above disabled person conversation AI processing unit may comprise an input conversion AI processing unit for converting data collected from the input unit into text or generating and processing user surrounding guidance information data using the collected data, and an output conversion AI processing unit for converting the text or surrounding guidance information data processed by the input conversion AI processing unit into an output form according to the setting of the input signal input from the user adjustment selection unit and transmitting it to the output unit.
[0055] The above input conversion AI processing unit performs image analysis and object detection on image data collected through the camera sensor, generates object information guidance through object recognition to recognize objects including obstacles and stairs around the user for surrounding guidance, recognizes sign language and performs Natural Language Processing (NLP) to convert it into text, performs audio analysis on audio data collected through the microphone sensor, performs noise canceling and speech recognition, and performs Natural Language Processing to convert it into text.
[0056] The above input conversion AI processing unit, in recognizing sign language, detects the hands of the conversation partner through image analysis and object detection to guide the conversation content in a situation where a speech-using user is conversing with a sign-using person, recognizes gestures and movements according to the movements of the detected hands, recognizes the sign language corresponding to the recognized gestures and movements, translates the recognized sign language, and converts it into text through natural language processing.
[0057] The above input conversion AI processing unit, in object recognition, can generate object information guidance data by detecting objects including roads and buildings on a path through image analysis and object detection to guide the user with information about the surrounding environment and objects in daily life when the user is visually impaired, and by determining obstacles including stairs around the user through object recognition.
[0058] The above input conversion AI processing unit can detect the voice signal of the other party along the path through audio analysis to guide the conversation content in a situation where a user using sign language is conversing with a person using spoken language in the context of speech recognition, recognize the conversation content through noise canceling and speech recognition on the detected voice signal, and convert it into text through natural language processing.
[0059] The above input conversion AI processing unit can recognize ambient sounds and ambient voices by classifying them according to the nature of sounds through audio analysis to provide guidance on situations and dangers occurring in the surroundings during daily life when the user is hearing impaired, determine whether an accident has occurred including the presence of a siren sound using the recognized ambient sounds and ambient voices, and generate guidance data regarding situations and dangers based on the determined accident occurrence.
[0060] The output conversion AI processing unit can set the magnitude and location of vibration according to the direction, size, and type of the object, generate an audio signal including voice and sound for voice guidance, generate a vibration signal for vibration guidance, and transmit the generated audio signal and vibration signal to the bone conduction speaker through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit.
[0061] The output conversion AI processing unit can convert the text processed by the input conversion AI processing unit into a display signal for display when outputting in text form according to the setting of the input signal input from the user adjustment selection unit, and can transmit the converted display signal to the AR display through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit so as to provide an Augmented Reality (AR) screen by applying an AR overlay.
[0062] The output conversion AI processing unit can convert the text processed by the input conversion AI processing unit into a display signal for displaying the text as a sign language image and motion according to the setting of the input signal input from the user adjustment selection unit, and can transmit the converted display signal to the AR display through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit so as to provide an Augmented Reality (AR) screen by applying an AR overlay.
[0063] When the output conversion AI processing unit outputs in the form of voice audio according to the setting of the input signal input from the user adjustment selection unit, it converts the text processed by the input conversion AI processing unit into an audio signal including voice and sound, and can transmit the converted audio signal to the bone conduction speaker through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit.
[0064] The output conversion AI processing unit can convert surrounding guidance information data, including surrounding conditions and danger situations processed by the input conversion AI processing unit, to generate guidance signals in the form of images and text, and can transmit the generated guidance signals to the AR display through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit so as to provide them as an Augmented Reality (AR) screen by applying an AR overlay to the generated guidance signals.
[0065] The above user adjustment selection unit can process the user's input signal input from the above input tab or the above microphone sensor and transmit it to the above disabled person conversation AI processing unit.
[0066] The above-mentioned smartphone may further include a wireless communication unit for connecting to an internet network or communicating with other devices.
[0067] The above wireless communication unit may comprise a Wi-Fi communication unit for accessing an internet network via a public Wi-Fi AP or a private Wi-Fi AP, and a cellular communication unit for communicating with other devices via a cellular network.
[0068] The above disabled person conversation AI processing unit can provide a basic setting function that sets default values for the input unit and the output unit according to an input signal input from the above user adjustment selection unit.
[0069] In providing the above-mentioned disability conversation AI processing unit, when an input signal indicating that the user is a hearing-impaired or speech-impaired person who uses sign language is received from the above-mentioned user adjustment selection unit, the above-mentioned disability conversation AI processing unit may provide the above-mentioned basic setting functions as voice recognition and ambient sound analysis functions through the microphone sensor of the input unit, text and sign language output functions through the AR display of the output unit, and guidance functions through vibration via the bone conduction speaker of the output unit.
[0070] In providing the above-mentioned disability conversation AI processing unit, when an input signal indicating that the user is a hearing impaired person using lip reading or a hearing aid is input from the above-mentioned user adjustment selection unit, the above-mentioned disability conversation AI processing unit may provide, as basic setting functions, a sign language recognition function including gestures and movements through the camera sensor of the input unit, an ambient sound analysis function through the microphone sensor of the input unit, a text output function through the AR display of the output unit, and a guidance function through vibration via the bone conduction speaker of the output unit.
[0071] In providing the above-mentioned disability conversation AI processing unit, when an input signal indicating that the user is visually impaired is input from the above-mentioned user adjustment selection unit, the above-mentioned disability conversation AI processing unit may provide sign language recognition and object recognition functions including gestures and movements through the camera sensor of the input unit, and audio output through the bone conduction speaker of the output unit and guidance functions through vibration as basic setting functions.
[0072] In providing the above-mentioned disabled conversation AI processing unit, when an input signal indicating "ordinary person" is input from the above-mentioned user adjustment selection unit, the above-mentioned disabled conversation AI processing unit may provide a sign language recognition function including gestures and movements through the camera sensor of the input unit, a text output function through the AR display of the output unit, and an audio output function through the bone conduction speaker of the output unit as basic setting functions.
[0073] The above disabled person conversation AI processing unit can access a specific website through the above wireless communication unit to download and install software related to disabled person conversation or AI agent functions.
[0074] The above disabled person conversation AI processing unit can access a specific website through the above wireless communication unit to download and install update software that includes additional functions and improves functions related to disabled person conversation or AI agent functions.
[0076] Figure 1 shows an example of the use and application of smart glasses for conversation with disabled people according to the present invention.
[0077] FIG. 1 illustrates a 1:1 conversation between a user and a counterpart, but the present invention is not limited thereto. That is, the smart glasses system for conversation with disabled people according to the present invention can also support a multi-party conversation method in which multiple people can converse simultaneously.
[0078] Referring to FIG. 1, in the present invention, the user or the other party in the conversation includes not only ordinary people but also people with hearing impairments who use sign language, people with speech impairments, people with visual impairments, and people with hearing impairments who use lip reading or hearing aids. The conversationalists converse without an intermediary using smart glasses that perform the role of an Artificial Intelligence (AI) agent for conversation with people with disabilities, which collects and analyzes conversation information through the other party's language, sign language, lip reading, and gestures to recognize and inform the content of the conversation.
[0079] In FIG. 1, the smart glasses system for disabled persons' conversation according to the present invention can be used for conversation between a hearing-impaired or speech-impaired user who uses sign language, a visually impaired person, a non-visually impaired person, and a hearing-impaired user who uses lip reading or a hearing aid.
[0080] Smart glasses can be worn by a user to analyze and recognize the voice or sign language (gestures, movements) of a conversation partner according to the user's settings, and can provide the analyzed conversation content to the user in the form of text or sign language through an AR (Augmented Reality) display, or in the form of vibrations of audio such as voice and sound through bone conduction speakers.
[0081] In Fig. 1, if the user is a hearing-impaired or speech-impaired person who uses sign language, the smart glasses can recognize the voice of the conversation partner through a microphone mounted on the front and provide text or sign language images to the user through an AR display. In the case where the user is a visually impaired person, a non-visually impaired person, or a hearing-impaired person who uses lip reading or a hearing aid, the smart glasses can recognize the sign language (gestures, movements) of the conversation partner through a camera mounted on the front and provide text through an AR display or voice audio through a bone conduction speaker mounted on the temple of the glasses.
[0083] Figure 2 shows the conversion function and operation example of the smart glasses for disabled persons' conversation of the present invention.
[0084] Referring to FIG. 2, the sign language (gesture, movement) recognition function and the sign language-to-text conversion function are functions for recognizing the sign language of a conversation partner who is a hearing-impaired or speech-impaired person using sign language in situations where the user is a visually impaired person, a non-visually impaired person, or a hearing-impaired person using oral communication or a hearing aid. The gestures and movements of the person using sign language are captured through a camera mounted on smart glasses, the video data is recognized and analyzed, and the analysis results can be converted into text. The converted text is transmitted to an appropriate system and output unit to be output in a format set by the user.
[0085] The object recognition function is designed to provide guidance on the surrounding environment and object information in situations where the user is visually impaired. By capturing the user's surroundings through a camera mounted on the smart glasses and analyzing the resulting video data, it can perform object recognition tasks such as distinguishing between sidewalks and roadways, identifying structures that may act as obstacles, and recognizing people. The recognition results can be provided to the user through audio generation and output functions.
[0086] The voice recognition and speech-to-text conversion functions are designed to recognize the spoken language of a conversation partner when the user is a hearing-impaired person using sign language or a person with a speech disability, or when the conversation partner is a visually impaired person, a sighted person, or a hearing-impaired person using lip reading or hearing aids. The voice of the conversation partner collected using a microphone sensor mounted on the smart glasses is distinguished, recognized, and analyzed through noise canceling, and the analysis results can be converted into text. The converted text is transmitted to an appropriate system and output device to be displayed in a format set by the user.
[0087] The function of receiving, collecting, and recognizing ambient audio (voice, sound, etc.) is designed to provide guidance on the surrounding environment and the occurrence of dangers in situations where the user is hearing impaired. By receiving and collecting audio from the surrounding environment through a microphone mounted on the smart glasses, it distinguishes and analyzes voices and sounds; based on the analysis results, it recognizes the surrounding situation and human voices, and can determine the occurrence of danger or accidents. Depending on the situation and danger, the analysis results can be provided to the user via an AR display or audio vibrations.
[0088] The AR display function, text-to-display, and text-to-sign language display (image, motion) conversion functions are designed to provide users with conversations and surrounding guidance in the form of visual information via an Augmented Reality (AR) display. They receive text generated via sign language-to-text or speech-to-text, as well as situational and danger information generated through ambient audio recognition, and can output this information according to user settings. If the user selects text output, the text can be overlaid onto the AR display using the text-to-display conversion function; if the user selects sign language output, sign language images or motion can be overlaid onto the AR display using the text-to-sign language display (image, motion) conversion function.
[0089] In the case of user adjustment and selection functions in the form of an AI agent, users can set and adjust the input / output devices to be used according to their disability and situation through user input, and select modes for system operation and network connection. If the user is hearing impaired, the microphone sensor can be set as the default input device and the AR display as the default output device; if the user is visually impaired, the camera sensor can be set as the default input device and the bone conduction speaker as the default output device. Input / output devices can be configured depending on the situation during daily life and the situation during conversation. Depending on the implementation method of the smart glasses-based AI agent for conversation with the disabled and the type of communication connection, users can select a smart glasses-only operation mode, a smartphone-via mode, or a smartphone-cooperative operation mode where the smartphone handles data processing.
[0090] The text-to-speech conversion and audio generation functions are designed to generate audio to provide users with conversations and surrounding guidance in the form of auditory information. By receiving text generated via sign language-to-text or speech-to-text, as well as object information generated through object recognition functions, it can generate audio, such as voice or sound effects, to be output for guidance according to user settings.
[0091] In the case of the vibration conversion and output function of generated audio, audio generated from the text-to-speech conversion and audio generation function can be output to the user through the bone conduction speaker of the smart glasses. In the case of audio generated about object information, the volume and position of the audio can be adjusted to provide information about the orientation, size, and type of the object.
[0092] Smart glasses may include an artificial intelligence computing device that performs various transformations and AI functions. The artificial intelligence computing device may be composed of an AI CPU, memory, an AI GPU, etc., and the smart glasses equipped with AI including various transformation functions in the present invention are configured to have the characteristic of guiding the conversation and surroundings of a person with a disability. The AI functions may include an AI assistant function to smoothly assist in the setup and execution of the user's smart glasses, and an AI agent function to perform transformations and processing related to the conversation of a person with a disability.
[0094] FIG. 3 shows an example of the structure and appearance of smart glasses for disabled persons' conversation according to the present invention.
[0095] Referring to FIG. 3, the smart glasses (10) of the present invention may include a camera sensor (112), a microphone sensor (114), an AR display (152), a bone conduction speaker (154), an input tab (111), a wireless communication unit (140), a rechargeable battery (160), and a disabled person conversation AI processing unit (130), and have the characteristics of an AI agent that provides conversation with disabled people and guidance on the surroundings. These characteristics are implemented through the disabled person conversation AI processing unit (130). The disabled person conversation AI processing unit (130) may include an artificial intelligence CPU, memory, and an artificial intelligence GPU. The artificial intelligence CPU is the center of all artificial intelligence computing processing and is responsible for overall system control and conversion. The artificial intelligence GPU is an assistant to the artificial intelligence CPU and is responsible for processing AI computations. The memory stores data collected through the corresponding AI and data processing and input unit.
[0096] The camera sensor (112) can be composed of a depth camera, etc., and can generate depth information based on the angle and size difference of the same object and analyze the distance and shape of the space. Accordingly, in the sign language (gesture, movement) recognition function, the position and joints of the hand can be detected to recognize the hand gestures and movements, and in the object recognition function, the accurate direction and distance of the object can be analyzed.
[0097] In the embodiment of FIG. 3, the camera sensor (112) may be implemented as multiple cameras to generate depth information based on angle and size differences for the same object and to analyze the distance and shape of the space.
[0098] The microphone sensor (114) can be configured as a microphone using MEMS (Micro-Electro-Mechanical Systems) or a micro-piezoelectric element, and can clearly extract and recognize the voice of a conversation partner through noise canceling in the voice recognition function, and can identify the direction from which a specific sound originated by analyzing sounds coming from various directions through beamforming in the ambient audio (voice, sound, etc.) reception, collection, and recognition function.
[0099] In the embodiment of FIG. 3, the microphone sensor (114) can be implemented as a micro array microphone for performing noise canceling and sound direction analysis.
[0100] The AR display (152) is a lens of the smart glasses (10), and in the AR display (152) function, it can project digital information guided by the smart glasses (10) into the user's field of vision and provide it in the form of AR (Augmented Reality). The AR display (152) projects digital information into the user's field of vision and provides it in the form of augmented reality.
[0101] The bone conduction speaker (154) is embedded in the upper ear portion of the temple end of the frame of the smart glasses (10) and can output audio in the form of vibration and voice guidance as the primary function of converting and outputting generated audio into vibration. Additionally, since it transmits sound without blocking the ears, it is possible to receive conversation and surrounding guidance while simultaneously recognizing the surrounding situation, and it can provide guidance as if sound is coming from a specific location through directional audio. The bone conduction speaker (154) outputs audio in the form of vibration and voice guidance.
[0102] The input tab (111) may be located on the frame (rim) of the smart glasses (10) and is a device for transmitting user input to the smart glasses (10) in various transformations in the form of an AI assistant (agent) and user adjustment and selection of AI functions. Through the tab, the user can select a mode for the operation method and switch the user input method to voice recognition. The input tab (111) transmits user input in the form of a touch tab to the smart glasses (10).
[0103] The wireless communication unit (140) is a device for enabling the smart glasses (10) to connect to the internet to perform web searches, cloud services, software updates, or to connect with a smartphone via Bluetooth. The wireless communication unit (140) enables the smart glasses (10) to connect to the internet and communicate with a smartphone and a server.
[0104] The rechargeable battery (160) is a component for supplying stable power to electronic components and devices within the smart glasses (10). Since the smart glasses (10) require a small and lightweight design, a rechargeable battery (160) capable of high efficiency, low power consumption, and long usage can be used. The rechargeable battery (160) serves to supply power to the smart glasses (10).
[0105] The disabled person conversation AI processing unit (130) controls the overall operation of the smart glasses (10) and can perform various conversions and AI functions. In the embodiment of FIG. 3, the disabled person conversation AI processing unit (130) is an artificial intelligence computing device that performs various conversions and AI functions, and can be implemented by including an artificial intelligence CPU, memory, artificial intelligence GPU, artificial intelligence TPU, artificial intelligence NPU, artificial intelligence AI ASIC, etc.
[0106] The present invention relates to smart glasses having the characteristics of an AI agent that performs various conversions and AI functions, such as a conversion function that recognizes sign language for conversation and guidance for the surroundings of the disabled, and can be utilized to enable free conversation between disabled people or between disabled people and the general public.
[0108] FIG. 4 shows examples of text and sign language that can be output through the display of the smart glasses system for communication with disabled people according to the present invention. The smart glasses (10) can implement AR (Augmented Reality) by overlaying text or sign language on the display.
[0109] Referring to FIG. 4, the smart glasses (10) can provide text generated through the sign language-to-text conversion function and the speech-to-text conversion function to the user in the form of text, sign language images, or gestures according to user settings. The text output example on the left is an example in which the generated text is received, undergoes text-to-display conversion, and is output through an AR overlay, and the sign language image output example on the right is an example in which the generated text is received, undergoes text-to-sign language display (image, motion) conversion, and is output through an AR overlay.
[0111] Figure 5 shows a configuration diagram of user adjustment and selection operations of the smart glasses system for conversation with disabled people according to the present invention, in a cooperative operation mode with a smartphone.
[0112] Although a smartphone (20) is illustrated in the embodiment of FIG. 5, the present invention is not limited thereto and can be implemented in various ways, such as a mobile communication terminal capable of communication including a smartphone, PDA, tablet PC, etc., a portable wireless smart device, and related devices.
[0113] Referring to FIG. 5, the smartphone and cooperative operation mode is a form in which an AI agent function that performs conversion and processing related to conversations with disabled people is processed and supported in a personal portable wireless smart device such as a smartphone or a portable wireless smart device dedicated to disabled people, and is connected to smart glasses (10) via Bluetooth communication.
[0114] In a cooperative operation mode with a smartphone, the smart glasses (10) may include an input unit (110), an output unit (150), a user adjustment selection unit (120), a first transmission / reception signal adjustment processing unit (170), a first Bluetooth communication unit (146), and a rechargeable battery (160).
[0115] In a cooperative operation mode with a smartphone, the smartphone (20) may include a second Bluetooth communication unit (174), a second transmission and reception signal adjustment processing unit (180), a disabled person conversation AI processing unit (130), and a wireless communication unit (140).
[0116] In cooperative operation mode with a smartphone, the configuration of the smart glasses is simplified and the device becomes lighter. Additionally, the battery consumption of the smart glasses is reduced, extending the usage time. In this mode, a personal wireless smart device such as a smartphone is essential, and Bluetooth communication connection and maintenance affect overall performance.
[0117] In cooperative operation mode with a smartphone, the smart glasses rely on the smartphone to perform data conversion and processing, and network access. The smart glasses (10) transmit data collected through the input unit (110) to the smartphone (20) via the first transmission / reception signal adjustment processing unit (170) and the first Bluetooth communication unit (146) of the wireless communication unit (140). The smartphone (20) transmits the received data to the input conversion AI processing unit (132) of the disabled conversation AI processing unit (130) via the second Bluetooth communication unit (174) and the second transmission / reception signal adjustment processing unit (180), converts the conversation content into text, detects surrounding objects, situations, and dangers, and generates guidance information. The generated text and guidance information are transmitted to the output conversion AI processing unit (134), processed into a form that the user can understand according to the user's settings, and transmitted to the smart glasses (10) via the second transmission / reception signal adjustment processing unit (180) and the second Bluetooth communication unit (174). The smart glasses (10) provide the received processing results to the user through the first Bluetooth communication unit (146), the first transmission / reception signal adjustment processing unit (170), and the output unit (150). The network connection of the smart glasses (10) is made by connecting to the internet network via a smartphone (20).
[0118] In a cooperative operation mode with smartphones, a smart glasses-supported AI agent system can be installed not only on smartphones but also on wireless mobile terminals such as PDAs and tablets, intelligent mobile terminals, and AI personal devices to perform various transformations and AI functions.
[0119] The input unit (110) is the entity that collects information about the user's surrounding environment. The input unit (110) may include a camera sensor (112) for collecting video data and a microphone sensor (114) for collecting audio data. The input unit (110) may transmit the collected video data and audio data to the first transmission / reception signal adjustment processing unit (170).
[0120] The output unit (150) is the entity that provides conversation content and surrounding guidance to the user. The output unit (150) may include an AR (Augmented Reality) display (152) for outputting conversation content, surrounding conditions, and danger guidance signals received from the first transmission / reception signal adjustment processing unit (170) in the form of text or sign language, and a bone conduction speaker (154) for outputting audio generated through text-to-speech and vibration signals generated for object information guidance.
[0121] The user adjustment selection unit (120) is the subject of a type of AI assistant function to support and assist the user's input and output. The user adjustment selection unit (120) can select a network connection method and an operating mode of the smart glasses according to the user's input via touch tap or user voice recognition, and can control signals for the execution of the input / output device and necessary conversion algorithms set for each mode through the first transmission / reception signal adjustment processing unit (170) in response to the user's mode selection according to the situation. In the case of FIG. 5, as a cooperative operating mode with a smartphone, the smart glasses perform network connection via Bluetooth through the smartphone, and the first transmission / reception signal adjustment processing unit (170) uses the Bluetooth mode of the wireless communication unit (140) to transmit packets so that the disabled conversation AI processing unit (130) of the smart glasses (10) performs the role of the disabled conversation AI processing unit (130) of the smartphone (20).
[0122] The first transmission / reception signal adjustment processing unit (170) is responsible for multiplexing, demultiplexing, and packet processing of signals transmitted and received between the smart glasses (10) and the smartphone (20) through the first Bluetooth communication unit (146). In the case of the first transmission / reception signal adjustment processing unit (170) of the smart glasses (10), it can receive video data and audio data from the input unit (110) and a signal regarding the user's mode selection from the user adjustment selection unit (120), create a packet, and transmit it to the first Bluetooth communication unit (146). It can also analyze the packet from the smartphone (20) received through the first Bluetooth communication unit (146) and transmit the conversation content, surrounding situation, and danger guidance signal to the output unit (150). In the case of the second transmission and reception signal adjustment processing unit (180) of the smartphone (20), it can analyze the packet of the smart glasses (10) received through the second Bluetooth communication unit (174) of the smartphone (20) and transmit the video data, audio data, and signals regarding the user's mode selection to the disabled conversation AI processing unit (130), and can analyze the packet received from the disabled conversation AI processing unit (130) to create a packet of conversation content, surrounding situation, and danger guidance signals and transmit it to the second Bluetooth communication unit (174).
[0123] The disabled person conversation AI processing unit (130) is the subject of the AI agent function that performs conversion and processing related to disabled person conversation. The disabled person conversation AI processing unit (130) performs the role of the disabled person conversation AI processing unit (130) of the smart glasses (10). The disabled person conversation AI processing unit (130) may include an input conversion AI processing unit (132) for converting video data and audio data of the smart glasses (10) received from the second transmission / reception signal adjustment processing unit (180) into text or generating surrounding guidance information, and may include an output conversion AI processing unit (134) for receiving the converted text and generated surrounding guidance information, converting them into an output form according to user settings, and transmitting them to the second transmission / reception signal adjustment processing unit (180). The system and algorithm selected by the user can be controlled and executed through the user adjustment selection unit (120).
[0124] The Bluetooth communication unit (146, 174) is the entity that performs Bluetooth wireless communication between the smart glasses (10) and the smartphone (20). The Bluetooth communication unit (146, 174) can transmit packets received from the transmission / reception signal adjustment processing unit (170, 180) to the smartphone (20) or the smart glasses (10), and can transmit packets received from the smart glasses (10) or the smartphone (20) to the transmission / reception signal adjustment processing unit (170, 180). Through Bluetooth wireless communication, it is possible to connect to the internet network via the smartphone (20) in a manner similar to a Bluetooth bridge.
[0125] The advantage of the cooperative operation mode with the smartphone is that the smartphone can bear a significant portion of the device and processor required to implement the above-mentioned functions, thereby reducing the weight of the smart glass device, making the implementation of the device and program simpler, and extending the battery life (operation time) due to reduced power consumption.
[0126] The disadvantage of the smartphone and cooperative operation mode is that the operation of the system relies entirely on the smartphone, and the complexity of the first transmission and reception signal adjustment processing unit (170) of the smart glasses (10) increases. Also, some time delay may occur during the processing.
[0128] FIG. 6 shows an example of user default settings for setting input / output devices and modes to be used according to the user in the smart glasses system for conversation with disabled people according to the present invention.
[0129] Referring to FIG. 6, when the smart glasses are started, the smart glasses perform user settings and adjustments according to the input of the user. The user can set the input / output devices according to their disability status and can set the operating mode according to the components, performance, and network of the smart glasses.
[0130] In the case where the user is a person with a disability, the user can set basic input / output devices for daily life and conversation situations with a partner using sign language or spoken language, depending on their type of disability and whether they use spoken or sign language.
[0131] If the user is a hearing-impaired or speech-impaired person who uses sign language, voice recognition and ambient sound analysis through the microphone sensor (114) and guidance through text and sign language output through the AR display (152) or vibration through the bone conduction speaker (154) can be set as the default.
[0132] If the user is a hearing impaired person who uses speech or a hearing aid, the default settings may include sign language recognition such as gestures and movements via a camera sensor (112), ambient sound analysis via a microphone sensor (114), text output via an AR display (152), or guidance via vibration via a bone conduction speaker (154).
[0133] If the user is visually impaired, sign language recognition and object recognition such as gestures and movements through the camera sensor (112), audio output through the bone conduction speaker (154), and guidance through vibration can be set as default.
[0134] If the user is a general person, sign language recognition such as gestures and movements through the camera sensor (112), sign language-to-text conversion, text output through the AR display (152), or audio generation through sign language-to-speech conversion and audio output through the bone conduction speaker (154) can be set as the default.
[0135] In setting the operation mode of the smart glasses, the user can set the smart glasses standalone operation mode, smartphone-via mode, smartphone-via operation mode, smart glasses direct server connection mode, and smartphone-via server connection mode depending on the components, performance, and network connection method of the smart glasses.
[0136] In the case of the smart glasses standalone operation mode, the smart glasses are configured to perform data analysis of the input unit (110) and transmit the analysis results, generated audio signals, and display signals to the output unit (150). Regarding network connection, the smart glasses can be configured to directly connect to and communicate with the internet through the wireless communication unit (140). When using the Wi-Fi communication unit (142) of the wireless communication unit (140), the internet network can be accessed through a public / private Wi-Fi AP, and when using the cellular communication unit (144), the internet network can be accessed through a cellular base station.
[0137] In the case of the smartphone-based mode, the smart glasses (10) are configured to perform data analysis of the input unit (110) and transmit the analysis results, generated audio signals, and display signals to the output unit (150). Regarding network connection, the smart glasses can be configured to perform internet connection and communication via the smartphone. By connecting to the smartphone (20) through the Bluetooth communication unit (146) of the wireless communication unit (140), the smartphone can connect to the internet network and transmit desired information to the smart glasses.
[0138] In the case of a smartphone and cooperative operation mode, data from the input unit (110) is transmitted to the smartphone through the first transmission / reception signal adjustment processing unit (170) and the first Bluetooth communication unit (146), and the smart glasses-supporting AI agent system inside the smartphone is set and adjusted to operate, so that the generated audio signal and display signal are received through the first Bluetooth communication unit (146), transmitted to the first transmission / reception signal adjustment processing unit (170), and transmitted to the output unit (150).
[0139] In the case of the direct server connection mode of the smart glasses, data from the input unit (110) is transmitted to the server through the first transmission / reception signal adjustment processing unit (170) and the Wi-Fi communication unit (142) or the cellular communication unit (144), and the input conversion AI processing unit (132) and output conversion AI processing unit (134) of the server are set and adjusted to operate so that the generated audio signal and display signal are received through the Wi-Fi communication unit (142) or the cellular communication unit (144) and transmitted to the first transmission / reception signal adjustment processing unit (170) and transmitted to the output unit (150).
[0140] In the case of the server connection mode via smartphone, data from the input unit (110) is transmitted to the smartphone (20) via the first transmission / reception signal adjustment processing unit (170) and the Bluetooth communication unit (146), and the smartphone (20) transmits it to the server (30). The input conversion AI processing unit (132) and output conversion AI processing unit (134) of the server (30) can be configured and adjusted to operate so that the generated audio signal and display signal are received via the smartphone (20) through the Bluetooth communication unit (146), transmitted to the first transmission / reception signal adjustment processing unit (170), and transmitted to the output unit (150).
[0142] Figure 7 shows a system flowchart of smart glasses for disabled persons' conversation of the present invention when the user is hearing impaired.
[0143] Referring to FIG. 7, a hearing-impaired user launches the smart glasses (S101) and, depending on the situation, selects and launches a mode set on the smart glasses via touch tap or voice recognition (S103).
[0144] In the case of a mode for conversation with a hearing-impaired or visually impaired person or a general person who uses speech or a hearing aid (S105), audio analysis is performed on voice data collected through a microphone sensor (114) (S107), and speech-to-text (STT) conversion is performed to convert audio into text through noise canceling and real-time speech recognition (S109, S111, S113).
[0145] When a text output command is input from the user, a text-to-display conversion is performed on the generated text (S115, S117), and the text is output through an AR display (152) via a text overlay through AR (S119, S121).
[0146] When a sign language output command, such as an image or motion, is input from the user, a text-to-sign language display (image, motion) conversion is performed (S115, S123), an image or motion corresponding to the sign language is generated (S125), and output through an AR display (152) via an image overlay through AR (S127, S129).
[0147] In the case of a mode for daily life (S131), an analysis is performed on audio signals such as voice and sound around the user collected through a microphone sensor (114) (S133), and through ambient sound and ambient voice recognition, it is checked whether an audio signal indicating danger, such as a siren or horn, is detected in the user's surrounding situation (S135, S137).
[0148] When an audio signal indicating danger is detected (S137), or when a command to use the vibration function is input from the user (S141), a signal is generated according to the volume and type of the audio signal (S143), and the volume and location where vibration will occur are set according to the direction and volume of the sound (S145), and vibration is output through the bone conduction speaker (154) and surrounding guidance is performed (S147).
[0149] Alternatively, when a command to use the AR display (152) is input from the user, an image and text guidance signal is generated according to the volume, type, and direction of the audio signal (S151), a text-to-display conversion is performed on the text guidance signal (S153), and output through the AR display (152) via an overlay through AR (S155, S157).
[0151] Figure 8 shows a system flowchart of smart glasses for disabled persons' conversation of the present invention when the user is visually impaired.
[0152] Referring to FIG. 8, the smart glasses system is started (S201), and the visually impaired user selects and executes a mode set on the smart glasses via touch tap or voice recognition depending on the situation (S203).
[0153] In the case of a mode for conversation with a hearing-impaired or speech-impaired person who uses sign language (S205), an analysis of video data collected through a camera sensor (112) is performed (S207), sign language such as gestures and movements of the target is recognized through real-time video recognition (S209, S211), and sign language-to-text conversion is performed to convert the sign language into text (S213).
[0154] When a command to use the bone conduction speaker (154) is input from the user (S215), text-to-speech conversion (TTS) is performed on the generated text (S217), and an audio signal is generated according to user settings such as voice and sound and output through the bone conduction speaker (154) (S219, S221).
[0155] In the case of a mode for daily life (S223), an analysis of image data of the user's surrounding environment collected through a camera sensor (112) is performed (S225), and through object detection and object recognition, it is checked whether dangerous elements such as obstacles and stairs are detected around the user (S227, S229). If danger is detected, the danger is determined according to the degree of the size, speed, etc. of the dangerous element (S231), and if there is a command to use the bone conduction speaker (154) from the user (S233), audio and vibration signals for guidance are generated according to the direction and size of the danger (S235), and vibration is output through the bone conduction speaker (154) and surrounding guidance is performed (S237).
[0157] FIG. 9 shows an example of a flowchart for converting sign language (gestures, movements) into text in a smart glasses system for conversation with disabled people according to the present invention.
[0158] In Figure 9, CNN (Convolutional Neural Network) is a convolutional neural network, RNN (Recurrent Neural Network) is a recurrent neural network, LSTM (Long Short-Term Memory) is a long-term short-term memory neural network, GRU (Gated Recurrent Unit) is a technique for analyzing actions that change over time, TCN (Temporal Convolutional Network) is a technique for learning dependencies over a longer range by processing time-series data in parallel, Transformer-based models (BERT, GPT, ViT, etc.) are techniques for contextually interpreting entire sign language sentences, NLP (Natural Language Processing) is a natural language processing technique, and Seq2Seq (Sequence-to-Sequence) is a sequence-to-sequence technique.
[0159] Referring to FIG. 9, the sign language (gesture, motion)-to-text conversion can be broadly divided into a data collection and preprocessing step through a camera (S301, S303), a sign language analysis and text conversion step (S305~S311), and a natural language processing (NLP) and text output step (S313, S315).
[0160] In the data collection and preprocessing steps (S301, S303) through the camera, the smart glasses collect video data of the other party's sign language (gesture, movement) in real time through the camera sensor (112) (S301), and remove backgrounds and unnecessary objects through video preprocessing (S303).
[0161] In the sign language analysis and text conversion steps (S305~S311), the hand position and finger joints are detected using R-CNN (Regions with Convolutional Neural Networks features), and changes in movement between frames are distinguished using segmentation (S305). The hand shape, finger position, and movement pattern are analyzed through the analysis of finger joint positions (S307), and static hand shapes are analyzed using CNN (Convolutional Neural Network) to match them with corresponding sign language such as alphabets and numbers (S309). Continuous hand movements are analyzed using RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), or TCN (Temporal Convolutional Network) to match them with corresponding sign language such as words and sentences (S311).
[0162] In the natural language processing (NLP) and text output steps (S313, S315), the sequence of the sign language is analyzed using a Seq2Seq model and converted into natural language such as Korean or English (S313). The converted natural language is then used with a Transformer model to correct the structure of the natural language sentences, which were originally the sentence structure of the sign language, to conform to natural language grammar, and the corrected text is output (S315).
[0164] FIG. 10 shows an example of a flowchart for converting voice recognition results into text in a smart glasses system for conversation with disabled people according to the present invention.
[0165] In Figure 10, VAD (Voice Activity Detection) is a voice activity detection technology, MFCC (Mel-Frequency Cepstral Coefficients) is a Mel frequency cepstral coefficient, CNN (Convolutional Neural Network) is a convolutional neural network, RNN (Recurrent Neural Network) is a recurrent neural network, LSTM (Long Short-Term Memory) is a long short-term memory neural network, CTC (Connectionist Temporal Classification) is connectionist time series classification, RNN-T (Recurrent Neural Network Transducer) is a sequence model suitable for real-time STT, and NLP (Natural Language Processing) is a natural language processing technology.
[0166] Referring to FIG. 10, speech-to-text conversion can be broadly divided into voice input and preprocessing steps through a microphone (S401~S405), voice recognition and text conversion steps (S407~S411), and natural language processing (NLP) and text output steps (S413, S415).
[0167] In the voice input and preprocessing steps (S401~S405) through the microphone, the smart glasses (10) collect voice data through the microphone sensor (114) (S401), and distinguish between background noise in the user's surrounding environment and the voice of the conversation partner according to voice activity level through VAD (Voice Activity Detection) and process only the necessary parts (S403). Then, noise in the audio is removed, and the quality of the voice data is improved through voice normalization and acoustic modeling (S405).
[0168] In the speech recognition and text conversion steps (S407~S411), frequency characteristics of the speech are extracted using Mel-Frequency Cepstral Coefficients (MFCC) to convert the speech data into a vector form so that it can be input into an artificial intelligence model (S407). Then, the speech signal is converted into phonemes that fit the frame using Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and Long Short-Term Memory (LSTM) (S409). Then, the speech data is converted into text in real time based on the speech data using Connectionist Temporal Classification (CTC) and Recurrent Neural Network Transducer (RNN-T) (S411).
[0169] In the natural language processing (NLP) and text output steps (S413, S415), words are predicted and transformed by considering the context through a Transformer (S413). Then, the text is divided into sentence units using natural language processing (NLP), typos are corrected, the context is corrected, and the corrected text is output (S415).
[0171] FIG. 11 shows an example of a flowchart for outputting text as voice through a bone conduction speaker (154) in a smart glasses system for disabled person conversation of the present invention.
[0172] In Figure 11, Neural Text Normalization is neural text normalization, NLP (Natural Language Processing) is a natural language processing technique, Dependency Parsing is dependency parsing, G2P (Grapheme-to-Phoneme) is a technique for converting characters into phonemes, Transformer-based models are techniques capable of parallel processing for fast conversion, RNN (Recurrent Neural Network) is a recurrent neural network, and Mel spectrogram is a Mel spectrogram representing acoustic features.
[0173] Referring to FIG. 11, text-to-speech conversion can be broadly divided into a text preprocessing step (S501, S503), a language and pronunciation conversion step (S505~S509), and a speech signal conversion and output step (S511, S513).
[0174] In the text preprocessing step (S501, S503), the text is normalized through neural text normalization (S501). Then, grammatical structures such as sentence morphemes are analyzed through an NLP model and Dependency Parsing (S503).
[0175] In the language and pronunciation conversion steps (S505~S509), the characters (Grapheme) of the text are converted into phonemes, which are the smallest units of speech, through G2P (Grapheme-to-Phoneme) (S505). Then, intonation, stress, pauses, etc., are determined according to the context through a Transformer (S507). Then, the phonemes are converted into speech signals through acoustic modeling of the RNN (Recurrent Neural Network) and Transformer model (S509).
[0176] In the voice signal conversion and output step (S511, S513), the spectrogram generated from the acoustic model is converted into an actual voice signal and synthesized through a vocoder (S511). Then, noise generated during the voice synthesis process is removed, and vibration is generated and output through a bone conduction speaker (154) (S513).
[0178] FIG. 12 shows an example of a flowchart for outputting text in sign language via AR in a smart glasses system for disabled persons' conversation of the present invention.
[0179] In Figure 12, NLP (Natural Language Processing) is a natural language processing technology, Seq2Seq (Sequence-to-Sequence) is a sequence-to-sequence technology, ASL Dataset (American Sign Language Dataset) is an American sign language dataset, KSL Dataset (Korean Sign Language Dataset) is a Korean sign language dataset, Unity is the Unity engine, a technology used for 2D and 3D rendering, and Unreal Engine is the Unreal Engine, a technology used for 3D rendering.
[0180] Referring to FIG. 12, the text-to-sign language display (image, motion) conversion can be broadly divided into text preprocessing and sign language translation steps (S601~S605), sign language image and motion conversion steps (S607), and generated image and motion AR output steps (S609, S611).
[0181] In the text preprocessing and sign language translation steps (S601~S605), the grammatical structure of the text, such as sentence morphemes, is analyzed using an NLP model, and words unnecessary for sign language translation are removed (S601). Then, the analyzed sentence is converted to conform to sign language grammar using Seq2Seq (Sequence-to-Sequence) (S603). Then, the converted sentence is linked to sign language gestures (sign language symbols) using ASL (American Sign Language) and KSL (Korean Sign Language) datasets (S605).
[0182] Then, in the sign language image and motion conversion step (S607), the converted sign language movement, the sign language image and motion corresponding to the symbol sentence are produced as a sequence.
[0183] In the generated image and motion AR output steps (S609, S611), hand motions of sign language produced through Unity and Unreal are placed in an AR (augmented reality) environment (S609). Then, hand motions are converted into graphics through AR rendering and output to an AR display (152) after undergoing an optimization process (S611).
[0185] FIG. 13 illustrates the key transformation algorithms and configuration of an AI agent system for conversation with disabled people when the smart glasses of the present invention are based on a standalone operation mode. FIG. 13 shows examples of algorithms required to implement each transformation algorithm, detailed technical specifications, basic device configurations, and minimum specifications required to run the system smoothly.
[0186] Text-to-speech conversion is a conversion for guiding text of conversation content into speech. Text preprocessing techniques may include rule-based conversion, Transformer-based Neural Text Normalization, etc. Phoneme conversion techniques may include G2P conversion, etc. Acoustic modeling techniques may include MFCC, LPC, etc. Vocoder techniques may include WaveNet, WaveGlow, etc. Basic device configurations for executing and outputting conversion may include an artificial intelligence computing device and a bone conduction speaker (154), etc.
[0187] Speech-to-text conversion is a conversion for generating text by recognizing the content of a conversation that takes place in speech, that is, spoken language. Audio preprocessing techniques may include noise canceling, etc. Phoneme conversion techniques may include CNN, RNN, Transformer-based models, etc. Acoustic modeling techniques may include MFCC, LPC, etc. Natural language processing techniques may include Transformer, BERT-based models, etc. Basic device configurations for input and conversion execution may include a microphone sensor (114) and an artificial intelligence computing device, etc.
[0188] Sign language (gestures, movements)-to-text conversion is a transformation designed to generate text by recognizing conversational content in sign language, such as gestures and movements. Image input preprocessing technologies may include R-CNN, DeepLab, etc. Finger joint detection and position analysis technologies may include Finger Keypoints, HOG, etc. Hand shape and movement analysis technologies may include CNN, RNN, LSTM, TCN, etc. Natural language processing technologies may include Seq2Seq, Transformer, etc. Basic device configurations for input and transformation execution may include vision sensors, IMU sensors, and artificial intelligence computing devices.
[0189] Text-to-display conversion is a conversion for guiding text about conversation content in text form through an AR display (152). Text preprocessing techniques may include rule-based conversion, Transformer-based Neural Text Normalization, etc. Text rendering techniques may include Billboarding, SDF, etc. Augmented reality (AR) overlay techniques may include Depth Estimation, ATW, etc. Basic device configurations for executing and outputting conversion may include an artificial intelligence computing device and an AR display (152), etc.
[0190] Text-to-sign language display (image, motion) conversion is a conversion intended to guide text regarding conversation content into sign language forms such as images and motion via an AR display (152). Text preprocessing techniques may include rule-based conversion and Transformer-based Neural Text Normalization. Sign language datasets may include ASL and KSL datasets. Image and motion rendering techniques may include Unity and Adobe Animate. Augmented reality (AR) overlay techniques may include Depth Estimation and ATW. Basic device configurations for executing and outputting conversion may include an artificial intelligence computing device and an AR display (152).
[0191] Sign language (gestures, movements)-to-speech conversion is a conversion designed to recognize conversational content in sign language, such as gestures and movements, and provide voice guidance. It can be implemented through the connection of two conversion algorithms and processing systems: sign language (gestures, movements)-to-text and text-to-speech.
[0192] Sign language (gesture, motion)-to-display conversion is a conversion for recognizing conversation content in sign language, such as gestures and motions, and guiding it in the form of text through an AR display (152). It can be implemented through the connection of two conversion algorithms and processing systems: sign language (gesture, motion)-to-text and text-to-display.
[0193] Speech-to-display conversion is a conversion for recognizing conversational content that is spoken, i.e., spoken language, and guiding it in text form through an AR display (152). It can be implemented through the connection of two conversion algorithms and processing systems: speech-to-text and text-to-display.
[0194] Speech-to-sign language display (image, motion) conversion is a conversion to recognize conversation content consisting of speech, i.e., spoken language, and to provide guidance in the form of sign language, such as images and motion, through an AR display (152). It can be implemented through the connection of two conversion algorithms and a processing system: speech-to-text and text-to-sign language display (image, motion).
[0195] In the basic device configuration, the artificial intelligence computing device is hardware for implementing AI agent characteristics that perform various transformations and AI functions of smart glasses, and may include an artificial intelligence CPU, memory, an artificial intelligence GPU, an artificial intelligence TPU, an artificial intelligence NPU, an artificial intelligence AI ASIC, etc.
[0196] The AI CPU is the center of all AI computing processing and is responsible for the control and transformation of the entire system.
[0197] The memory is responsible for storing data collected through the AI and data processing and input unit (110).
[0198] Artificial intelligence GPUs serve as assistants to AI CPUs and are responsible for the computational processing of general-purpose AI. The main advantages of AI GPUs are that they reduce latency, minimize bandwidth requirements, and enhance security by rapidly processing data, such as sign language, close to the source where it is generated. While general GPUs are used for high-performance computing, large-scale AI training, and high-quality graphics rendering, AI GPUs are specialized for real-time data processing, IoT applications, and mobile AI, playing an important role in IoT devices such as smart glasses and high-speed mobile communication networks.
[0199] Artificial intelligence TPUs are responsible for the computational processing of specific AI, are primarily optimized for inference tasks, and play a role in accelerating the execution speed of deep learning models. The main advantage of artificial intelligence TPUs is that they can process AI models quickly while operating at lower power than artificial intelligence GPUs.
[0200] AI ASICs are optimized for specific AI tasks and are designed to be more precise than AI TPUs for specific functions such as speech recognition and image analysis, and can provide the highest level of power efficiency and performance.
[0202] FIG. 14 illustrates a comparison of system features and advantages and disadvantages according to the implementation form of the smart glasses of the present invention. FIG. 14 includes various implementation methods and network connection features according to the implementation form of the smart glasses system, various advantages and disadvantages occurring during the system operation process, and remarks for reference.
[0203] Referring to Fig. 14, the smartphone cooperative operation realization form (Fig. 5) is a form of cooperative operation mode with a smartphone, in which the AI agent system functions for many conversions and processing of conversations for the disabled are executed on a personal portable wireless smart device such as a smartphone. That is, the AI agent functions that perform conversions and processing related to conversations for the disabled are processed and supported by a personal portable wireless smart device such as a smartphone, and are connected to the smart glasses via short-range communication such as Bluetooth communication. In the smartphone cooperative operation realization form, the configuration of the smart glasses becomes simpler and the device becomes lighter. On the other hand, the battery consumption of the smart glasses is reduced, extending the usage time, and a personal wireless smart device such as a smartphone is essential, and Bluetooth communication connection and maintenance affect the overall performance.
[0205] The present invention relates to a smart glasses system for communication with persons with disabilities. Specifically, it concerns a technological invention that enables ordinary people to converse with persons with disabilities without an intermediary for conversation or communication, and allows persons with disabilities of various types to converse directly without an intermediary. It is currently impossible for hearing-impaired and speech-impaired individuals who use sign language to converse or communicate with visually impaired people, ordinary people, or hearing-impaired individuals who use lip reading or hearing aids without an intermediary. Although assistive devices such as sign language interpretation gloves have been developed to address this, they are inconvenient due to a lack of convenience and functionality, making convenient forms of conversation impossible.
[0206] Existing Google Glass and Meta's smart glasses (Meta Ray-Ban Smart Glasses) do not have the function and configuration for conversation with people with disabilities, and since they are not smart glasses primarily intended for people with disabilities, conversation between able-bodied people and people with disabilities, and even between people with different disabilities, is impossible.
[0207] The present invention relates to a smart glasses system that enables conversation between people with disabilities and the general public, or between people with disabilities. That is, to enable people with hearing and speech disabilities to converse or communicate with people with visual impairments, the general public, and people with hearing impairments who use lip reading or hearing aids without an intermediary, the smart glasses enable object recognition, sign language recognition, and voice recognition through a camera and microphone mounted on the smart glasses, and also enable conversation by outputting and providing information in a form desired by the user through an AR (Augmented Reality) display and a bone conduction speaker (154).
[0208] The smart glasses of the present invention utilize the characteristics of an AI agent that performs various conversions and AI functions to collect and convert audio data of spoken language and video data of sign language through a camera and microphone embedded in the smart glasses, thereby assisting in conversation between a hearing-impaired person who uses sign language and a visually impaired person who uses sign language, and between a hearing-impaired person who uses spoken language or a hearing aid and a hearing-impaired person who uses sign language.
[0209] To enable hearing and visually impaired users to cope with various situations and dangers in their surroundings, the smart glasses utilize cameras and microphones embedded in them to perform object recognition (such as obstacles and stairs), sound recognition (such as sirens and horns), and voice recognition (such as calling). The perceived surroundings are then converted into guidance that the user can intuitively understand and provided to them.
[0210] In the present invention, a smart glasses-supported AI agent system can be installed on wireless AI personal devices such as PDAs (Personal Digital Assistants) and tablets, in addition to smartphones, to perform various transformations and AI functions. That is, depending on the implementation method and communication connection type of the smart glasses-based conversational AI agent for the disabled, wireless AI personal devices and servers can be made to perform functions such as data analysis, recognition, and transformation.
[0211] In the present invention, regarding the six conversion systems of an AI agent system for conversation with disabled people, six conversion functions are implemented, including text-to-speech conversion for guiding text about conversation content into voice, speech-to-text conversion for recognizing conversation content made of spoken language and generating text, sign language (gesture, movement)-to-text conversion for recognizing conversation content made of sign language such as gestures and movements and generating text, text-to-display conversion for guiding text about conversation content in text form through an AR display, text-to-sign language display (image, motion) conversion for guiding text about conversation content in sign language form such as images and motions through an AR display, and sign language (gesture, movement)-to-speech conversion for recognizing conversation content made of sign language such as gestures and movements and guiding it into voice.
[0212] The present invention can be implemented in smart glasses by configuring individual applications for the six conversion functions and a combination of the same to support an AI agent for conversation with disabled people in six conversion functions, namely text-to-speech conversion, speech-to-text conversion, sign language (gesture, movement)-to-text conversion, text-to-display conversion, text-to-sign language display (image, motion) conversion, and sign language (gesture, movement)-to-speech conversion.
[0213] In the present invention, the device configuration and processing method may be included, which can perform or apply various conversions and AI functions by using not only smartphones but also personal portable smart devices such as PDAs (Personal Digital Assistants) and tablets, and portable smart wireless devices exclusively for the disabled, and by installing a smart glasses-type conversational AI agent system for the disabled.
[0214] In the present invention, the smart glasses operate in a cooperative operation mode with a smartphone. Before using the smart glasses, an AI agent system and related software that support conversation with disabled people are installed in advance to perform various transformations and AI functions. This allows the user to use the smart glasses through simple initial setup, or the user can access a specific website to access or download the AI agent system and related software related to conversation with disabled people to use the smart glasses. Additionally, the smart glasses system can be implemented in a form that allows the user to use the smart glasses by upgrading the AI agent system and related software related to conversation with disabled people, thereby improving characteristics or enabling additional functions through updates.
[0215] Although the present invention has been described above using several preferred embodiments, these embodiments are illustrative and not limiting. Those skilled in the art will understand that various changes and modifications can be made without departing from the spirit of the invention and the scope of rights set forth in the appended claims. Explanation of the symbols
[0217] 110 Input Unit 112 Camera Sensor 114 Microphone Sensor 150 Output Unit 152 AR Display 154 Bone Conduction Speaker 111 Input Tap 140 Wireless Communication Unit 160 Rechargeable Battery 130 AI Processing Unit for Dialogue with the Disabled 132 Input Conversion AI Processing Unit 134 Output Conversion AI Processing Unit 10 Smart Glasses 20 Smartphone
Claims
Claim 1 Smart glasses that are worn on a user's face, collect data related to a conversation between the user and a counterpart, transmit the collected data to a smartphone, and output result data received from the smartphone; and a smartphone that analyzes and processes data received from the smart glasses based on an AI agent function and transmits the processed result data to the smart glasses, wherein the smart glasses include: an input unit for collecting data related to a conversation between a user or a counterpart; a user adjustment selection unit for processing an input signal input by a user; an output unit for outputting result data received from the smartphone; a first Bluetooth communication unit for communicating with the smartphone via Bluetooth; and a first transmission / reception signal adjustment processing unit for multiplexing, demultiplexing, and packet processing of signals transmitted and received with the smartphone. The smart glasses are composed of a rechargeable battery, comprising a rechargeable battery for supplying power to the smart glasses, wherein the smart glasses have a touch sensor built into a part of the smart glasses and further comprise an input tab for receiving operation control commands for the smart glasses from a user, wherein the input unit comprises a camera sensor for collecting video data related to a conversation between the user and the other party or video data of the surrounding environment; and a microphone sensor for collecting audio data related to a conversation between the user and the other party or audio data of the surrounding environment, wherein the output unit comprises an AR display capable of displaying in an AR (Augmented Reality) manner and outputting text, sign language images, or motion images processed by the smartphone; and a bone conduction speaker for outputting voice and vibration processed by the smartphone, wherein the smartphone comprises a second Bluetooth communication unit for communicating with the smart glasses via Bluetooth; and multiplexing of signals transmitted and received with the smart glasses.It comprises: a second transmission / reception signal adjustment processing unit for demultiplexing and packet processing; and a disabled person conversation AI processing unit that performs conversion and processing necessary for conversation between disabled people or conversation between a disabled person and a non-disabled person based on an AI agent function for data collected from the input unit according to an input signal input from the smart glasses, wherein the disabled person conversation AI processing unit includes an input conversion AI processing unit for converting data collected from the input unit into text or generating and processing user surrounding guidance information data using the collected data; The system comprises an output conversion AI processing unit for converting text or surrounding guidance information data processed by the input conversion AI processing unit into an output form according to the setting of an input signal input from the user adjustment selection unit and transmitting it to the output unit. The input conversion AI processing unit performs image analysis and object detection on image data collected through the camera sensor, generates object information guidance through object recognition to recognize objects including obstacles and stairs around the user for surrounding guidance, recognizes sign language and performs Natural Language Processing (NLP) to convert it into text, performs audio analysis on audio data collected through the microphone sensor, performs noise canceling and speech recognition, and performs natural language processing to convert it into text. In recognizing sign language, the input conversion AI processing unit detects the hands of the conversation partner through image analysis and object detection to provide guidance on the conversation content in a situation where a speech-reading user converses with a sign-reading person, recognizes gestures and movements according to the movement of the detected hands, recognizes sign language corresponding to the recognized gestures and movements, translates the recognized sign language, and converts it into text through natural language processing. The input conversion AI processing unit converts, and in object recognition, when the user is visually impaired, in order to provide guidance on the situation of the surrounding environment and object information in daily life,Through image analysis and object detection, objects including roads and buildings on the path are detected, and obstacles including stairs around the user are determined through object recognition to generate object information guidance data. The input conversion AI processing unit, in terms of speech recognition, detects the voice signal of the counterpart on the path through audio analysis to guide the conversation content in a situation where a sign language user is conversing with a speech-speaking counterpart, recognizes the conversation content of the detected voice signal through noise canceling and speech recognition, and converts it into text through natural language processing. The input conversion AI processing unit, in the case where the user is hearing impaired, classifies sounds according to their properties through audio analysis to recognize ambient sounds and ambient voices to guide situations and dangers occurring in the surroundings during daily life, determines whether an accident has occurred including the presence of a siren sound using the recognized ambient sounds and voices, and generates guidance data regarding situations and dangers based on the determined accident occurrence. The output conversion AI processing unit sets the magnitude and location of vibration according to the direction, size, and type of the object, generates an audio signal including voice and sound for voice guidance, generates a vibration signal for vibration guidance, and the generated audio signal and vibration signal are the The output conversion AI processing unit transmits to the bone conduction speaker through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit, and when the output conversion AI processing unit outputs in text form according to the setting of the input signal input from the user adjustment selection unit, the text processed by the input conversion AI processing unit is converted into a display signal for display, and the converted display signal is transmitted to the AR display through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit to provide an Augmented Reality (AR) screen by applying an AR overlay.When the output conversion AI processing unit outputs in a sign language form including images and motion according to the settings of the input signal input from the user adjustment selection unit, it converts the text processed by the input conversion AI processing unit into a display signal for displaying sign language images and motion, and transmits the converted display signal to the AR display through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit to provide an Augmented Reality (AR) screen by applying an AR overlay; when the output conversion AI processing unit outputs in a voice audio form according to the settings of the input signal input from the user adjustment selection unit, it converts the text processed by the input conversion AI processing unit into an audio signal including voice and sound, and transmits the converted audio signal to the bone conduction speaker through the second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit; and when the output conversion AI processing unit converts surrounding guidance information data including surrounding conditions and danger situations processed by the input conversion AI processing unit to generate a guidance signal in the form of images and text, and applies an AR overlay to the generated guidance signal to provide an Augmented Reality (AR) screen. The second transmission / reception signal adjustment processing unit and the second Bluetooth communication unit transmit to the AR display to provide a Reality screen, and the user adjustment selection unit processes the user's input signal input from the input tab or the microphone sensor and transmits it to the disabled conversation AI processing unit. The smartphone further includes a wireless communication unit for connecting to an internet network or communicating with other devices, and the wireless communication unit comprises a Wi-Fi communication unit for connecting to an internet network via a public Wi-Fi AP or a private Wi-Fi AP, and a cellular communication unit for communicating with other devices via a cellular network.The above-mentioned disability conversation AI processing unit provides a basic setting function that sets default values for the input unit and the output unit according to an input signal input from the above-mentioned user adjustment selection unit, and in providing the above-mentioned disability conversation AI processing unit, when an input signal is input from the above-mentioned user adjustment selection unit indicating that the user is a hearing impaired person or a speech impaired person who uses sign language, the above-mentioned disability conversation AI processing unit provides, as basic setting functions, a voice recognition and ambient sound analysis function through the microphone sensor of the input unit, a text and sign language output function through the AR display of the output unit, and a guidance function through vibration via the bone conduction speaker of the output unit, and in providing the above-mentioned disability conversation AI processing unit, when an input signal is input from the above-mentioned user adjustment selection unit indicating that the user is a hearing impaired person who uses lip reading or a hearing aid, the above-mentioned disability conversation AI processing unit provides as basic setting functions a sign language recognition function including gestures and movements through the camera sensor of the input unit, a ambient sound analysis function through the microphone sensor of the input unit, a text output function through the AR display of the output unit, and a guidance function through vibration via the bone conduction speaker of the output unit, and in providing the above-mentioned disability conversation AI processing unit, when an input signal is input from the above-mentioned user adjustment selection unit indicating that the user is a visually impaired person When input is received, the above-mentioned disabled conversation AI processing unit provides sign language recognition and object recognition functions including gestures and movements through the camera sensor of the input unit, audio output through the bone conduction speaker of the output unit, and guidance functions through vibration as basic setting functions, wherein when an input signal indicating "ordinary person" is received from the user adjustment selection unit, the above-mentioned disabled conversation AI processing unit provides sign language recognition functions including gestures and movements through the camera sensor of the input unit, text output functions through the AR display of the output unit, and audio output functions through the bone conduction speaker of the output unit as basic setting functions, and the above-mentioned user adjustment selection unit,It performs an AI assistant function to support and assist user input and output, selects a network connection method and an operating mode of the smart glasses according to user input through the input tab or the microphone sensor, controls signals for the execution of input / output devices and necessary conversion algorithms set for each selected operating mode through the first transmission / reception signal adjustment processing unit, wherein the operating mode includes a cooperative operating mode with a smartphone, wherein in the cooperative operating mode with the smartphone, the smart glasses transmit data from the input unit to the smartphone through the first transmission / reception signal adjustment processing unit and the first Bluetooth communication unit, and the disabled conversation AI processing unit of the smartphone is configured to perform conversion and processing corresponding to the settings of the selected user type, operating mode, and output form for the data of the input unit, and is configured to receive audio signals and display signals generated from the smartphone through the first Bluetooth communication unit and transmit them to the output unit via the first transmission / reception signal adjustment processing unit, and wherein the disabled conversation AI processing unit performs text-to-speech conversion, speech-to-text conversion, sign language-to-text conversion, text-to-display conversion, text-to-sign language display conversion, and A smart glasses system characterized by executing, individually or in combination, a conversion function among a plurality of conversion functions including sign language-to-speech conversion, corresponding to a user type and output form set by the user adjustment selection unit and whether the conversation partner uses spoken language or sign language determined by video analysis or audio analysis by the input conversion AI processing unit. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 delete Claim 7 delete Claim 8 delete Claim 9 delete Claim 10 delete Claim 11 delete Claim 12 delete Claim 13 delete Claim 14 delete Claim 15 delete Claim 16 delete Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete Claim 21 delete Claim 22 delete Claim 23 delete Claim 24 delete Claim 25 delete Claim 26 A smart glasses system according to claim 1, wherein the disabled person conversation AI processing unit accesses a specific website through the wireless communication unit to download and install software related to disabled person conversation or AI agent functions. Claim 27 A smart glasses system according to claim 26, wherein the disabled person conversation AI processing unit accesses a specific website through the wireless communication unit to download and install update software that improves functions and includes additional functions related to disabled person conversation or AI agent functions.
Citation Information
Patent Citations
Extramissive spatial imaging digital eyeglass apparatus for virtual or augmediated vision
JP2019149202A
Smart glass system for hearing-impaired communication
KR101830908B1
Glass type terminal
KR1020160027813A
Augmented reality smart glass and method for controlling the output of smart glasses
KR1020230079846A
Wearable glass-type device and method of controlling the device
KR102195773B1