Multimodal AI-Based Senior Care Robot System

KR103003846B1Active Publication Date: 2026-08-14MIRACLE AG I CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020250159866
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-08-14
Estimated Expiration
2045-10-30

Smart Images

  • Figure 112025120977681-PAT00002_ABST
    Figure 112025120977681-PAT00002_ABST
Patent Text Reader

Abstract

A multimodal AI-based senior care robot system according to the present invention comprises an input unit for collecting voice and video data to indicate the user's condition, an AI processing unit for generating control signals by analyzing and inferring input data collected from the input unit, and an output unit for providing multimodal feedback according to the control of the AI ​​processing unit, wherein the AI ​​processing unit is implemented in an on-device structure and configured to process the input data in a local environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to the field of intelligent care robot technology that simultaneously supports the safety of the elderly and emotional connection. Specifically, a multimodal AI processing unit integrates and analyzes voice and video data and processes them in real-time through an on-device optimization structure, thereby enabling early detection of emergency situations such as falls and unresponsiveness. Furthermore, the invention relates to an AI-based intelligent senior care robot system capable of providing customized care services in various environments, such as homes or nursing facilities, by offering healthcare functions like medication management and health monitoring, while simultaneously performing entertainment functions such as experiential content, singing along, and storytelling. Background Technology

[0003] As the aging society intensifies, the need for care technology targeting vulnerable groups such as the elderly living alone, residents of nursing facilities, and long-term hospitalized patients is rapidly increasing. Existing care robots and wearable devices have been limited to detecting falls based primarily on single sensor data or providing simple notification functions, which has limited their ability to ensure reliability in complex living environments.

[0004] To address this, safety and healthcare platforms utilizing RGB-D sensors, thermal imaging sensors, and IoT sensors, as well as helper robot systems linked to user terminals, have recently been proposed. These technologies have contributed to a certain level of safety improvement by providing various functions such as fall prediction, abnormal body temperature detection, and environmental information monitoring. However, most of these systems are limited to simple data collection and notification delivery, and intelligent care robot technology that integrates real-time interaction and emotional connection by optimizing multimodal AI on-device has yet to be presented.

[0005] In this regard, Korean Registered Patent No. 10-2425941 (Prior Art 1) discloses a helper system using a helper robot. This document presents a configuration in which a helper robot collects voice, video, and contact information, and a user terminal analyzes this information to provide customized helper services. For example, it includes methods such as outputting exercise content and evaluating the user's performance results, or providing response services through tendency analysis and abnormal state determination. However, this document does not disclose intelligent interaction control that reflects characteristics of the elderly, such as speech delay, cognitive load, and reduced concentration, by fusionally analyzing multimodal signals such as VAD, gaze, and reaction time.

[0006] In addition, Korean Registered Patent No. 10-2546460 (Prior Art 2) discloses an AI-based physical activity recognition and service integration platform. This document presents a configuration that detects the physical activity and risk factors of a managed subject through various IoT sensors and provides services such as fall prediction, detection of abnormal body temperature, and detection of environmental abnormalities based on a learning model. However, this document does not disclose the integrated adaptive service structure of the present invention, which implements a multimodal AI processing unit in an on-device optimized structure to fuse and analyze raw data such as VAD, gaze, and reaction time, and based on this, provides interaction control, medication management, and experiential entertainment that corrects speech delay, cognitive load, and reduced concentration in real time.

[0007] The aforementioned prior art technologies are significant in that they collect status information of users, such as the elderly or patients, and provide customized services based on this information. However, these technologies are mainly limited to providing notifications or basic services by simply collecting and analyzing status information, and there was a limitation in that intelligent care robot technology was not disclosed that implements a multimodal AI processing unit in an on-device optimized structure to integrate and analyze detailed raw data such as VAD, gaze, and reaction time, and performs interaction control to correct characteristics of the elderly, such as speech delay, cognitive load, and reduced concentration, in real time. Prior art literature

[0009] Prior Art 1: Korean Registered Patent No. 10-2425941 (Published June 16, 2021) Prior Art 2: Korean Registered Patent No. 10-2546460 (Published June 14, 2023) The problem to be solved

[0010] The objective of the present invention is to provide a multimodal AI-based care robot that integrates voice and video data to enable the elderly to form emotional bonds through natural interaction with the robot. In particular, the main objective of the present invention is to provide a senior care robot system capable of simultaneously ensuring latency-free responsiveness and high security by processing data in real-time through an on-device optimized structure without relying on the cloud. means of solving the problem

[0012] A multimodal AI-based senior care robot system according to the present invention comprises an input unit for collecting voice and video data to indicate the user's condition, an AI processing unit for generating control signals by analyzing and inferring input data collected from the input unit, and an output unit for providing multimodal feedback according to the control of the AI ​​processing unit, wherein the AI ​​processing unit is implemented in an on-device structure and configured to process the input data in a local environment.

[0013] In addition, the AI ​​processing unit is mounted on a main robot terminal operating in an on-device structure, and the main robot terminal is connected to a plurality of client robot terminals via a network to expand into a hub-type structure.

[0014] In addition, the main robot terminal is equipped with an artificial intelligence model configured to perform voice recognition, visual analysis, and language processing in real time in an on-device environment, and the client robot terminal is characterized by being responsible for input data transmission and output interface functions in conjunction with the main robot terminal.

[0015] In addition, the AI ​​processing unit is characterized by including a central processing unit for recognizing a user state by analyzing the voice and video data, respectively, and a multimodal control unit for generating a control signal corresponding to the user state based on the analysis result of the central processing unit.

[0016] In addition, the central processing unit is characterized by including a voice processing module for preprocessing voice data collected from a user to calculate parameters indicating whether speech has occurred and voice characteristics, a landmark processing module for recognizing facial position information from a video signal to extract landmark data necessary for interaction, and a language processing module for interpreting conversational meaning and managing contextual information based on textualized voice data.

[0017] In addition, the multimodal control unit is characterized by including an interaction control engine that dynamically controls the conversation flow and level of participation with the user based on the analysis results of the central processing unit, a healthcare control engine that detects changes in the user's biological state and behavior to determine and respond to health-related abnormalities, and an entertainment control engine that adjusts the output of immersive content by considering the user's cognitive state and emotional response.

[0018] In addition, the interaction control engine is characterized by including an attention guarantee controller that determines the concentration state based on an analysis indicator of the central processing unit, a turn-taking adaptive controller that dynamically determines the start and end times of speech, and a cognitive load adaptive controller that adjusts the output difficulty and conversation speed according to the user's cognitive state.

[0019] In addition, the healthcare control engine is characterized by including a medication support mode that determines whether the user is delaying medication by referring to the VAD signal of the voice processing module and the response time data of the language processing module, and verifies the actual act of taking medication by cross-verifying the gaze direction and mouth movement amount of the landmark processing module, and an emergency detection mode that determines the user's emergency situation by integrating and analyzing the VAD signal of the voice processing module, the signal-to-noise ratio (SNR) of the environment, the head posture and gaze direction of the landmark processing module, and the response time of the language processing module.

[0020] In addition, the entertainment control engine is characterized by including: an experiential content mode that determines whether a user performs an action using body motion information and gaze direction extracted from the landmark processing module and provides corresponding feedback; a sing-along mode that determines whether the user speaks and the intensity of speech using the VAD signal and mouth motion (MAR) indicator of the voice processing module, and controls music playback and subtitle output according to the determination result; and a storytelling mode that evaluates the user's cognitive load state using the reaction time (RT) and language difficulty indicator of the language processing module, and controls the length of the story and speech speed according to the evaluation result.

[0021] In addition, the input unit includes a microphone for collecting voice and a webcam for recognizing the user's state, and the output unit includes a display for displaying visual information to support the user's understanding and response, and a speaker for outputting synthesized voice.

[0022] In addition, the multimodal AI-based senior care robot system according to the present invention is characterized by further comprising a wheel installed at the bottom of the robot body to guide the movement of the robot. Effects of the invention

[0024] According to the proposed invention, since the user's speech intent, gaze, facial expressions, etc., can be comprehensively identified through a multimodal AI processing unit that simultaneously analyzes voice and video data, there is an advantage of enabling much more natural and intuitive emotional interaction compared to existing single-modal-based robots.

[0025] In addition, according to the present invention, since data can be processed in real time through an on-device optimization structure based on GPU and NPU without relying on a cloud server, it is possible to provide safe and highly reliable care services by minimizing response delays and reducing the risk of personal information leakage.

[0026] In addition, according to the present invention, since it can integrally provide not only healthcare functions such as medication management and health monitoring but also various entertainment functions such as experiential content, singing along, and storytelling, it has the effect of simultaneously providing safety and enjoyment to the daily lives of the elderly.

[0027] In addition, according to the present invention, since a structure for wirelessly linking an on-device hub robot and a plurality of lightweight terminal robots can be adopted, there is an effect of implementing a network-type care service that can be economically scalable even in multi-story buildings or large-scale nursing facilities. Brief explanation of the drawing

[0029] FIG. 1 is a front and rear view illustrating the exterior of a multimodal AI-based senior care robot according to the present invention. FIG. 2 is an exploded perspective view of a multimodal AI-based senior care robot according to the present invention, illustrating the configuration of the main modules. FIG. 3 is a block diagram showing the overall control configuration of a multimodal AI-based senior care robot according to the present invention. FIG. 4 is a block diagram illustrating the relationship between the input unit, central processing unit, multimodal control unit, and output unit of a multimodal AI-based senior care robot according to the present invention. FIG. 5 is a block diagram illustrating the detailed configuration of the central processing unit and the multimodal control unit according to the present invention. FIG. 6 is a block diagram illustrating the mode configuration of the healthcare control engine and the entertainment control engine of the multimodal control unit according to the present invention. FIG. 7 is a conceptual diagram illustrating an example of a network connection of a multimodal AI-based senior care robot according to the present invention. Specific details for implementing the invention

[0030] The present invention is susceptible to various modifications and may take various forms, and embodiments are to be described in detail in the text. However, this is not intended to limit the invention to the specific disclosed forms, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention. Similar reference numerals have been used for similar components in the description of each figure. Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms.

[0031] The above terms are used solely for the purpose of distinguishing one component from another. The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise.

[0032] In this application, terms such as "comprising" or "consisting of" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0033] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0034] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the drawings.

[0035] Referring collectively to FIGS. 1 to 3, the multimodal AI-based senior care robot (10) according to the present invention is an integrated robot device that simultaneously supports the safety of the elderly and emotional interaction, and is implemented with an external hardware configuration, an internal processing module, and an on-device AI processing unit (150) that operates based thereon.

[0036] First, looking at the exterior shown in FIG. 1, a head unit (20) is positioned on the upper part of the robot, and the head unit (20) is integrally equipped with a display (21), a webcam (22), feedback lighting (23), and a microphone (24). The display (21) goes beyond simply conveying visual information and serves as a multi-layered interface, such as providing medication time guidance, health status notifications, video calls with a guardian, and entertainment content. Additionally, the webcam (22) can check the user's facial expressions, gaze direction, and head movements, and the microphone (24) collects voice commands and conversation inputs. Along with this, the feedback lighting (23) is configured to intuitively inform the user of the robot's status; for example, it flashes red when an emergency is detected and lights up blue when a medication reminder is given, thereby helping the user understand. This configuration is designed so that the elderly can immediately perceive the robot's intentions without any separate complex operations.

[0037] Additionally, a speaker (32) and an NFC reader (31) are installed on the robot body (30). The speaker (32) provides a clear auditory response by outputting not only simple voice guidance but also music, notification sounds, or alarm sounds in emergency situations. The NFC reader (31) is utilized in a way that tags a medicine container or medicine card in medication support mode, allowing the user to verify whether the correct medicine has been taken at the correct time. Furthermore, a support (40) and a wheel (15) are provided on the bottom of the robot, ensuring not only stable installation but also ease of movement. In particular, the wheel (15) is designed so that the user can easily pull or push the robot to change its position, allowing it to be conveniently used even in situations where organization or movement is required.

[0038] Meanwhile, FIG. 2 is an exploded perspective view illustrating the internal structure of a robot, in which a main processing module (35), a power and battery module (36), and a sub-board module (37) are installed inside the robot. The main processing module (35) is core hardware dedicated to AI computation, including high-performance computing devices such as a CPU, GPU, and NPU, and implements the function corresponding to the on-device AI processing unit (150) of FIG. 3. In other words, the on-device AI processing unit (150) is specifically realized by the main processing module (35). Through this, the robot can directly process complex AI computations, such as voice recognition, image analysis, and natural language understanding, in a local environment without relying on a cloud server.

[0039] This on-device approach fundamentally reduces cloud latency issues, ensuring immediate responsiveness and offering the advantage of stable operation even in environments with unstable or limited internet connections. In particular, since personal and biometric data is not transmitted to external servers, it is excellent in terms of security and privacy protection.

[0040] However, high-performance on-device computing inevitably entails high heat generation. To address this, the present invention may install one or more pairs of blower fans inside the robot body. Specifically, a fan housing equipped with blower fans is provided at the rear of the main body (30), so that the airflow can be designed to rapidly dissipate heat generated by the main processing module (35). This allows the internal temperature to be maintained stably even while performing AI computing for a long time, and enables safe use even in indoor spaces with poor ventilation, such as the living environment of the elderly. This plays an important role not only in maintaining performance but also in ensuring fire safety.

[0041] The above sub-board module (37) includes an I / O port section, a wireless communication module, and various auxiliary control circuits, enabling the robot to interact with external devices and networks. Through this, data can be exchanged in real time with a terminal of a guardian or medical institution, and can be expanded into a group-type care service in which multiple robots interact with each other through a network.

[0042] FIG. 3 is a block diagram of the above configuration, which is divided into an input unit (100), an on-device AI processing unit (150), a central processing unit (200), a multimodal control unit (300), and an output unit (400). The input unit (100) includes a microphone (24) and a webcam (22) to simultaneously collect voice and video signals. The output unit (400) provides results audibly and visually through a speaker (32) and a display (21). Additionally, the on-device AI processing unit (150) is composed of a central processing unit (200) and a multimodal control unit (300) to enable real-time analysis and inference of input data. Specifically, the central processing unit (200) derives a suitable service scenario based on these results. Subsequently, the multimodal control unit (300) controls various engines such as interaction, healthcare, and entertainment to direct the final output.

[0043] As such, the robot (10) of the present invention according to FIGS. 1 to 3 is an integrated system in which external hardware (Fig. 1), internal module configuration (Fig. 2), and control block diagram (Fig. 3) are organically combined. In particular, the main processing module (35) implements an on-device AI processing unit (150) to ensure fast response speed and high security, and a blower fan structure to resolve heat generation issues can ensure stability. Furthermore, through the input unit (100) and output unit (400), the user can interact with the robot in an intuitive manner, and various care functions are provided in real time through the central processing unit (200) and multimodal control unit (300). In this way, the present invention can be implemented as a next-generation intelligent care system that goes beyond a simple information guidance device and simultaneously satisfies the safety, health, and emotional connection of the elderly.

[0044] The control relationships between each component and the detailed control configuration are illustrated in more detail in FIGS. 4 and FIGS. 5. Referring to the drawings, the AI-based senior care robot system (10) of the present invention can be broadly divided into an input unit (100), a central processing unit (200), a multimodal control unit (300), and an output unit (400). Each component operates independently yet is closely interconnected to form an integrated structure that simultaneously analyzes voice and video signals, estimates the user's condition in real time based on this, and outputs an adaptive response. In particular, the present invention has a technical distinction in that it enables a response tailored to the user's condition rather than a simple response, taking into account special situations such as speech delay, reduced concentration, and cognitive load in the elderly.

[0045] First, the input unit (100) includes a microphone (24) and a webcam (22). The microphone (24) performs a Voice Activity Detection (VAD) function that detects the possibility of human speech amidst ambient sounds. For example, it can distinguish only the speaker's voice even in an environment mixed with TV noise or everyday noise, and if the signal-to-noise ratio (SNR) is low, it increases the noise suppression weight to amplify only the voice signal. The voice processed in this way is immediately transmitted to the Speech-to-Text (STT) module of the central processing unit (200) and converted into text data. Meanwhile, the webcam (22) captures the user's face and part of the upper body in real time and tracks the location and movement of major landmarks such as eyes, nose, and mouth. Through this, the user's gaze direction, head tilt, and degree of mouth opening and closing are quantitatively calculated, and, for example, if the user turns their head to look at the screen or closes their eyes, it can be determined that the level of concentration is low.

[0046] Data collected from the input unit (100) is analyzed step-by-step in the central processing unit (200). The central processing unit (200) is composed of a voice processing module (210), a landmark processing module (220), and a language processing module (230). The voice processing module (210) combines the VAD signal and the STT conversion result to analyze not only simple text conversion but also speech intensity, speech speed, and speech continuity. For example, if there is a long pause before an elderly person starts speaking or if there is a long pause in the middle, the corresponding pattern is recorded as a speech delay characteristic. Then, the landmark processing module (220) tracks the tilt of the face, the frequency of eye blinking, and mouth movements (MAR index) from the video to determine whether the current user is in a state ready to speak or is simply responding with facial expressions. At this time, the voice signal and the landmark signal can be analyzed in synchronization.

[0047] Additionally, the language processing module (230) interprets the text obtained through STT using an LLM-based language model and generates a response appropriate to the context. During the generation process, the user's cognitive state is reflected; for example, if signals of failure to understand ("Huh?", "Say that again"), reaction time delay, or gaze deviation are detected, the difficulty of the response is automatically lowered to prioritize outputting a short and simple sentence.

[0048] The results of such analysis are transmitted to a multimodal control unit (300), and a final interaction strategy is determined. The multimodal control unit (300) includes an interaction control engine (310), a healthcare control engine (320), and an entertainment control engine (330). In particular, the interaction control engine (310) performs functions such as ensuring attention, adapting to turn-taking, and alleviating cognitive load in an integrated manner. For example, if the user's attention level is calculated to be low, the control engine slows down the TTS speed and simultaneously increases the size of the subtitles displayed on the display (21) to enhance visual emphasis. Conversely, if the user's attention level is high, it maintains a short and fast conversation flow. In addition, the turn-taking adaptation function corrects the speech patience window according to the individual profile in response to speech initiation delays or long pauses in speech by the elderly. Thus, it more accurately recognizes whether the user has finished speaking or not, thereby preventing unnecessary interruptions in conversation or misrecognition. In addition, the cognitive load controller immediately detects when the user has difficulty with a long explanation and summarizes the explanation or switches it to a question-and-answer format.

[0049] Additionally, the healthcare control engine (320) provides medication support and emergency detection functions. For example, when the time for the user to take the medication arrives, the engine simultaneously outputs a voice notification and a visual caption to guide the elderly person so that they do not miss it. At this time, whether the user has actually taken the medication can be verified using an NFC reader or camera (webcam) sensing, and if the unconfirmed state persists for a certain period of time or longer, an automatic notification is sent to the guardian. In addition, in the emergency detection mode, if the webcam (22) captures a sudden change in posture or a state of non-response for a long time, and the microphone (24) simultaneously detects a shock sound or a groan, an abnormality score is calculated using multimodal gating. If the score exceeds a threshold, an emergency response procedure is carried out step by step. First, the user is called out directly to check for a response, and if there is no response, contact with the guardian is made, and finally, an emergency rescue request can be made.

[0050] The entertainment control engine (330) executes customized leisure content, such as experiential content, singing along to songs, and providing old stories (storytelling). For example, if a user says, "I want to sing the song OO," the language processing module (230) interprets this as a command. Then, the entertainment control engine (330) simultaneously performs music playback and displays lyrics. If the user is unable to follow along with the lyrics subtitles, the system automatically slows down the output speed or provides repeat playback to maintain engagement. Additionally, the "storytelling" content enhances immersion by adjusting the length of the story based on the user's facial expressions or whether their gaze is fixed.

[0051] Finally, the output unit (400) consists of a speaker (32) and a display (21). The speaker (32) converts text generated by the LLM into TTS and outputs naturally synthesized voice, while the display (21) simultaneously provides subtitles, graphics, and avatar expressions. These two forms of output are synchronized with lip-sync technology, so that voice and video are provided in a synchronized state, allowing the elderly to understand more intuitively. In particular, since voice and visual information are provided simultaneously, even users with impaired hearing can sufficiently recognize information through visual assistance.

[0052] As such, FIG. 4 illustrates the process in which the senior care robot system (10) of the present invention tracks the user's status in real time and calculates and outputs a customized adaptive response within the overall flow leading from the input unit, the central processing unit, the multimodal control unit, and the output unit. By realizing intelligent interaction that considers the characteristics of the elderly, such as attention, speech delay, and cognitive load, rather than merely limiting to voice recognition or video output, it is possible to provide a much more natural and safe user experience than existing single-modal-based systems. A detailed example of this is additionally illustrated in FIG. 5.

[0053] FIG. 5 is a block diagram showing the linkage structure between the central processing unit (200) and the multimodal control unit (300) of a multimodal AI-based senior care robot according to the present invention, which is particularly detailed with the internal controller operation of the interaction control engine (310). The central processing unit (200) is composed of a voice processing module (210), a landmark processing module (220), and a language processing module (230), and each module calculates complex indicators such as whether speech is made, gaze, head posture, and conversation context in real time.

[0054] First, the voice processing module (210) calculates a VAD (Voice Activity Detection) signal, a STT (Speech-to-Text) converted text result, and an environment signal-to-noise ratio (SNR). The VAD signal is expressed as a speech probability score and distinguishes whether the user shows an actual intention to speak or if it is simply ambient noise. For example, in a quiet living room environment, the VAD alone can clearly determine whether speech is being made, but if TV noise is mixed in, the SNR is calculated to be low and is cross-analyzed together with the mouth momentum indicator of the landmark processing module (220). Therefore, the output of the voice processing module (210) is primarily utilized by the turn-taking adaptive controller (312) to precisely determine the start and end of speech.

[0055] Next, the landmark processing module (220) extracts gaze direction, head posture, and mouth aspect ratio (MAR) in real time from the camera image. Gaze direction and head posture are used as key indicators when the attention guarantee controller (311) calculates the level of concentration. For example, if the user does not look at the screen, looks away, or lowers their head, the concentration score is lowered. On the other hand, the mouth aspect ratio is input into the turn-taking adaptive controller (312) to serve as a basis for distinguishing whether speech has ended or is continuing. For example, even if the voice signal is cut off, if the user's lips are moving slightly, the turn end judgment is delayed, and the robot maintains a waiting state for a natural conversational flow.

[0056] The above language processing module (230) interprets the conversational context based on the STT converted text, records the response time (RT), and calculates the LLM output difficulty. RT is measured as the time taken from the question to the answer, and as the time increases, it is interpreted as a signal of decreased concentration or increased cognitive load. Additionally, the LLM output difficulty is calculated as an indicator subdivided into sentence length (L), vocabulary level (W), and content complexity (C), which serves as a standard for the cognitive load adaptive controller (313) to adjust the response difficulty in real time.

[0057] These indicators are transmitted to the interaction control engine (310) of the multimodal control unit (300) and are fused and analyzed step-by-step in the attention guarantee controller (311), turn-taking adaptive controller (312), and cognitive load adaptive controller (313).

[0058] The above attention guarantee controller (311) can calculate a real-time attention score S(t) by integrating the VAD score, gaze direction, and RT indicator through the following formula.

[0059] S(t) = w1·VAD(t) + w2·Gaze(t) + w3·RT(t)

[0060] Here, VAD(t) is the speech probability index, Gaze(t) is the gaze fixation rate, and RT(t) is the reaction speed score, while weights w1, w2, and w3 are adjusted according to user-specific characteristics. For example, when VAD = 70 points, Gaze = 50 points, and RT = 40 points, applying w1 = 0.4, w2 = 0.35, and w3 = 0.25 yields S(t) = 55.5 points. In such cases, it is determined that the user is in a state of reduced concentration, and subtitle enlargement or avatar gesture emphasis is executed in stages (e.g., a score of 70 or lower is determined to be in a state of reduced concentration). This structure, in which output is gradually adjusted according to concentration, can be considered an important mechanism for maintaining the participation of the elderly.

[0061] Additionally, the turn-taking adaptive controller (312) determines the start and end times of speech by combining the VAD signal, mouth momentum, head fine movement, and environment SNR. The present invention is particularly unique in that it accommodates delayed speech by the elderly by adding a speech patience window (PW). The speech termination reliability T_end can be expressed by the following formula.

[0062] T_end = VAD_off + α·MAR + β·PW

[0063] Here, VAD_off represents the silence duration, MAR represents the mouth momentum, and PW represents the speech endurance time index, and α and It is adjusted according to the environment and individual characteristics. For example, if VAD_off = 40 points, MAR = 70 points, PW = 90 points, α = 0.4, and β = 0.6, T_end = 122 points is calculated, and it is determined that the speech has not ended (e.g., if the score is 100 points or more, it is determined that the speech is continuing). Accordingly, the robot maintains the speech turn and waits for the user's subsequent speech, which is a processing method optimized for the slow speech characteristics of the elderly.

[0064] In addition, the cognitive load adaptive controller (313) can calculate the cognitive load score D by combining the STT text, conversation context, comprehension failure signal, RT, and LLM output difficulty using the following formula.

[0065] D = f (L, W, C)

[0066] Here, L represents sentence length, W represents vocabulary difficulty, and C represents content complexity. If D is measured as high, the system simplifies the LLM output sentences and slows down the TTS speed to 0.7 to 0.8 times. Conversely, if D is low, it increases the sentence length, raises the vocabulary level, and adjusts the TTS speed to 1.2 times to provide lively conversation.

[0067] In this way, the attention guarantee controller (311), turn-taking adaptive controller (312), and cognitive load adaptive controller (313) each operate independently but are integrated in the interaction control engine (310). For example, when attention is low and the end of speech is unclear, the system extends the speech waiting time and enlarges the subtitle size to induce participation in the conversation. As another example, in a situation with high cognitive load, the turn-taking controller increases the waiting time to guarantee thinking time, and the attention controller simplifies avatar gestures to reduce visual burden.

[0068] Ultimately, the present invention provides an interaction pipeline that simultaneously optimizes three axes: participation, flow, and understanding. Unlike existing conversation systems that rely on a single modal, this invention is technically differentiated in that it enables customized conversation optimization that reflects the characteristics of the elderly by implementing real-time scoring and stepwise control algorithms based on multimodal indicators.

[0069] As described in FIG. 5 with the controller of the interaction control engine, FIG. 6 shows in detail how the healthcare and entertainment engines receive and interpret data streams from the central processing unit (200) to perform specific mode-specific functions. In particular, since healthcare and entertainment functions must simultaneously satisfy the safety and emotional satisfaction of the elderly, multimodal feature fusion and adaptive feedback can serve as core design principles beyond simple event trigger levels.

[0070] The emergency detection mode (321) of the healthcare control engine (320) simultaneously receives the VAD signal from the voice processing module (210), the environment SNR, the head posture and gaze direction from the landmark processing module (220), and the response time indicator from the language processing module (230) to identify critical situations such as falls and non-response at an early stage. For example, assuming a situation where a user falls in the living room, impact sounds and groans are detected by the microphone, causing the VAD to rise sharply; as the impact sound is loud, the noise-to-voice ratio decreases, causing the SNR to drop. At the same time, in the webcam video, the head posture tilts sharply toward the floor, the gaze direction remains fixed for a long time, and the response time of the language processing module is abnormally extended. When these multiple indicators satisfy the gating conditions, the system performs the escalation protocol step by step.

[0071] Initially, the robot directly calls out to the user to check for a response; if there is no response, it sends a warning message to the guardian's terminal, and finally, sends an emergency rescue request. At this stage, instead of making a judgment based on a single indicator, the urgency score is calculated using a weighted sum of VAD, line of sight, and RT indicators, thereby effectively suppressing the false positive rate.

[0072] Additionally, the medication support mode (322) is implemented as an intelligent management mode that cross-verifies the actual act of taking medication beyond simple notifications. When it is time to take the medication, the robot outputs medication guidance voice through the speaker and simultaneously displays notification subtitles on the display. At this time, if the user's response time is excessively long or if negative utterances such as "later" or "I won't take it" are detected in the STT recognition results, the possibility of non-compliance with medication is raised.

[0073] At this time, the landmark processing module cross-checks whether the gaze is directed toward the location of the medication and whether the mouth activity level (MAR) reflects the actual act of intake by analyzing the results. If there is no fixed gaze at all and the mouth activity level converges to zero, it is determined that the medication has failed, and the notification intensity is gradually increased. Initially, a gentle voice guidance is provided, but after a certain period, an alarm sound is output, and finally, an automatic message is sent to the guardian. This can be considered an advanced management system that goes beyond conventional technology based on simple notifications, as it can verify medication compliance in real time through a multimodal feedback loop.

[0074] Additionally, the entertainment control engine (330) is composed of an experiential content mode (331), a sing-along mode (332), and a storytelling mode (333), and plays a role in enhancing user immersion and emotional connection. First, when, for example, a gymnastics video is played on the screen during the experiential content mode (331), the webcam tracks the user's arm and leg movements and analyzes the direction of gaze and head posture to evaluate whether the user is actually looking at the screen and following the movements. The degree of movement matching is quantified by a landmark-based movement matching algorithm, and feedback messages such as "Raise your arms a little higher" are displayed on the display in real time. Furthermore, the speaker outputs coaching voice in parallel to provide an experience similar to an exercise instructor, which can be expanded into an experiential interaction that induces active participation rather than simple content consumption.

[0075] In the sing-along mode, the VAD signal and mouth momentum indicator are combined to determine whether the user is actually speaking. If the speech intensity is low or the STT error rate is high, the system slows down the music playback speed to 0.8 times or repeats a specific section. Additionally, if the response time indicator provided by the language processing module becomes long, the subtitle display time is increased or the lyrics are simplified to maintain both the learning effect and the fun.

[0076] Storytelling mode dynamically combines LLM-based generated responses with the user's cognitive state indicators. When response times are short and signals of understanding are captured smoothly, the story gradually becomes longer and richer in description. Conversely, if responses are delayed and signals of misunderstanding occur frequently, the system abbreviates sentences, lowers vocabulary difficulty, and slows down the TTS speed to aid comprehension. This adaptive difficulty adjustment is realized by an attention-based multimodal interpretation structure and is applied differentially according to each user's profile.

[0077] In this way, the healthcare control engine (320) and entertainment control engine (330) of FIG. 6 go beyond simply receiving voice and video signals from the central processing unit and materialize them into a final service through multimodal gating and an adaptive feedback loop. Specifically, the emergency detection mode (321) precisely determines a fall or non-response situation through the simultaneous analysis of multiple indicators, and the medication support mode (322) checks whether medication is taken in real time by cross-verifying gaze, speech, and actions. In addition, in the gymnastics, singing, and storytelling modes, the output difficulty and feedback are dynamically adjusted by reflecting the intensity of participation and cognitive state. This structure is distinctly different from conventional care systems that rely on a single modal and can function as an integrated senior care solution that simultaneously satisfies the three axes of safety, health, and emotion.

[0078] In addition, the present invention can be implemented in a network expansion structure beyond single robot unit operation, and an example thereof is illustrated in FIG. 7. Specifically, FIG. 7 illustrates a network expansion structure of a multimodal AI-based senior care robot system (10) according to the present invention, and in particular, shows a hub client type architecture in which a main robot terminal with an embedded on-device AI processing unit (150) and a plurality of lightweight client robot terminals with on-device modules omitted are connected via a wireless communication network.

[0079] The central main robot terminal (10) is equipped with an on-device AI processing unit (150) composed of a GPU, NPU, and ARM-based computational core, and a Transformer-based lightweight model is run to perform all high-load computations locally, such as real-time voice recognition, visual landmark analysis, natural language understanding, and emergency detection. Since such on-device computation can fundamentally reduce cloud latency issues, it demonstrates excellent performance in situations requiring immediacy, such as fall detection, emergency voice calls, and real-time conversation responses. In addition, since personal information and biometric data are not transmitted to an external server, it offers significant advantages in terms of security and privacy protection.

[0080] Meanwhile, the sub-robot terminals (client robot terminals) deployed in the vicinity are lightweight client devices equipped with the same input / output devices such as displays, microphones, cameras, and speakers, but with on-device computing units omitted. They transmit voice and video signals collected from the user to the main robot terminal via a wireless communication network, receive the results analyzed and processed by the main robot terminal, and output them through the display and speakers. Therefore, all computationally intensive functions are performed by the main robot terminal, and the client robot terminals focus on serving as an interface for input and output.

[0081] For example, assuming a senior care center consisting of a multi-story building, a main robot terminal is installed on the 5th floor to serve as an on-device AI hub, and client robot terminals deployed on each floor (1st to 4th) are wirelessly connected to the main robot terminal. If the terminal on the 2nd floor detects a user's fall, the microphone detects impact sounds and groans (VAD rises), the camera recognizes sudden changes in posture, and simultaneously, a signal indicating an abnormal delay in the user's reaction time is recorded by the language processing module. This multi-input data is immediately transmitted to the main robot terminal on the 5th floor, which calculates an emergency detection score in real time using a Transformer-based lightweight model. If the result exceeds a threshold, a step-by-step response procedure is executed immediately, including calling a guardian, sounding an alarm within the facility, and, if necessary, linking with 119. Since this process is completed within the local network without passing through a cloud server, an immediate response is possible without delays of a few seconds.

[0082] Furthermore, the present invention maximizes the utilization of computational resources through an on-device optimization strategy. The Transformer model is lightweighted to fit ARM-based processors and operates efficiently even in small robot hardware environments through GPU and NPU computations and distributed memory optimization. Through this, the main robot terminal can perform large-scale computations independently without relying on the cloud, and can also switch to a hybrid structure that operates in parallel with a cloud server as needed. For example, normal conversations and monitoring can be sufficiently handled solely through on-device processing, while long-term data accumulation or statistical analysis can be processed in parallel by linking with a cloud server. By doing so, it is possible to flexibly respond to various device environments and operating conditions.

[0083] Furthermore, the distributed structure between the main robot terminal and client robot terminals facilitates facility expansion. The service coverage can be extended not only to the same building but also to adjacent buildings sharing the same network, while significantly reducing hardware costs. In other words, since a single main robot can support multiple client robot terminals without the need to redundantly equip every robot with high-performance computing units, it offers distinct advantages in terms of cost-effectiveness and maintainability.

[0084] Ultimately, the structure illustrated in Fig. 7 goes beyond simple single-robot operation to realize an on-device hub-type network architecture. This can provide technical benefits that simultaneously ensure real-time performance and scalability in multi-story buildings and large-scale facility environments, while complementing the latency and security limitations of existing cloud-based models.

[0085] Although the invention has been described above with reference to embodiments, a person skilled in the art will understand that various modifications and changes can be made to the invention without departing from the spirit and scope of the invention as described in the following claims. Explanation of the symbols

[0087] 10: AI-based senior care robot 15: Wheel 20: Head section 21: Display 22: Webcam 24: Mike 30: Main body 32: Speaker 100: Input section 150: On-device AI processing unit 200: Central Processing Unit 210: Speech processing module 220: Landmark Processing Module 230: Language processing module 300: Multimodal control unit 310: Interaction Control Engine 320: Healthcare Control Engine 330: Entertainment Control Engine 400: Output section

Claims

Claim 1 The system includes: an input unit for collecting voice and video data to indicate the user's state; an AI processing unit for generating control signals by analyzing and inferring input data collected from the input unit; and an output unit for providing multimodal feedback according to the control of the AI ​​processing unit, wherein the AI ​​processing unit is mounted on a main robot terminal operating in an on-device structure and is configured to process the input data in a local environment, wherein the main robot terminal is networked with a plurality of client robot terminals not equipped with the AI ​​processing unit and is expanded into a hub-type structure, wherein the client robot terminal is configured to transmit the input data to the main robot terminal and receive and output results analyzed and processed by the main robot terminal, and wherein the AI ​​processing unit comprises a central processing unit for recognizing the user's state by analyzing the voice and video data respectively; and includes a multimodal control unit for generating a control signal corresponding to a user state based on the analysis results of the central processing unit, wherein the multimodal control unit includes an interaction control engine that dynamically controls the conversation flow and degree of participation with the user based on the analysis results of the central processing unit, and the interaction control engine includes an attention guarantee controller that calculates a user's real-time concentration score by weighted summing a VAD (Voice Activity Detection) score, an eye fixation rate, and a reaction speed score, and controls to enlarge the size of subtitles output through the output unit or emphasize avatar gestures when the calculated concentration score is below a preset threshold; and a turn-taking adaptive controller that determines whether to maintain the turn of the conversation by determining whether to end the utterance by combining the silence duration of the voice signal, the mouth aspect ratio (MAR), and a Patience Window (PW) index set according to the characteristics of the elderly.A multimodal AI-based senior care robot system characterized by including a cognitive load adaptive controller that calculates a cognitive load score based on cognitive load indicators including the length of conversation sentences, vocabulary difficulty, and content complexity, and dynamically varies the output speed of synthesized speech according to the calculated cognitive load score. Claim 2 delete Claim 3 A multimodal AI-based senior care robot system according to claim 1, wherein the main robot terminal is equipped with an artificial intelligence model configured to perform voice recognition, visual analysis, and language processing in real time in an on-device environment, and the client robot terminal is linked with the main robot terminal to perform input data transmission and output interface functions. Claim 4 delete Claim 5 A multimodal AI-based senior care robot system according to claim 1, wherein the central processing unit comprises: a voice processing module for preprocessing voice data collected from a user to calculate parameters indicating whether speech has occurred and voice characteristics; a landmark processing module for recognizing facial position information from a video signal to extract landmark data necessary for interaction; and a language processing module for interpreting conversational meaning based on textualized voice data and managing contextual information. Claim 6 A multimodal AI-based senior care robot system according to claim 5, wherein the multimodal control unit comprises: a healthcare control engine that detects changes in the user's biological state and behavior to determine and respond to health-related abnormalities; and an entertainment control engine that adjusts the output of immersive content by considering the user's cognitive state and emotional response. Claim 7 delete Claim 8 A multimodal AI-based senior care robot system according to claim 6, wherein the healthcare control engine includes: a medication support mode that determines whether the user is delaying medication by referring to the VAD signal of the voice processing module and the response time data of the language processing module, and verifies the actual act of taking medication by cross-verifying the gaze direction and mouth movement amount of the landmark processing module; and an emergency detection mode that determines the user's emergency situation by integrating and analyzing the VAD signal of the voice processing module, the environment signal-to-noise ratio (SNR), the head posture and gaze direction of the landmark processing module, and the response time of the language processing module. Claim 9 A multimodal AI-based senior care robot system according to claim 6, wherein the entertainment control engine comprises: an experiential content mode that determines whether a user performs an action using body motion information and gaze direction extracted from the landmark processing module and provides corresponding feedback; a sing-along mode that determines whether the user speaks and the intensity of speech using the VAD signal and mouth motion (MAR) index of the voice processing module, and controls music playback and subtitle output according to the determination result; and a storytelling mode that evaluates the user's cognitive load state using the reaction time (RT) and language difficulty index of the language processing module, and controls the length of the story and speech speed according to the evaluation result. Claim 10 A multimodal AI-based senior care robot system according to claim 1, wherein the input unit includes a microphone for collecting voice and a webcam for recognizing the user's state, and the output unit includes a display for displaying visual information to support the user's understanding and response, and a speaker for outputting synthesized voice. Claim 11 A multimodal AI-based senior care robot system according to claim 1, further comprising a wheel installed at the lower part of the robot body to guide the movement of the robot.

Citation Information

Patent Citations

  • Robot Apparatus and Method for Providing of Care Service Using Multi-Modal Artificial Intelligence

    KR1020240105601A

  • System and method for controlling robot using on-device ai

    KR1020250153641A

  • Care robot for providing care service

    KR102752005B1