Care robot for providing care service

The care robot collects and processes diverse user information to provide customized services by generating prompts for language models, improving service accuracy and reducing costs.

WO2025230042A1PCT designated stage Publication Date: 2025-11-06ROBOCARE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/010807
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-03
Filing Date
2024-07-25
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Conventional care robots struggle to collect and process diverse user information effectively, leading to challenges in providing customized services and increasing costs due to frequent use of advanced language models.

Method used

A care robot equipped with sensors and AI capabilities to collect various types of information, generate prompts for language models, and determine appropriate care services, while controlling the timing of input to reduce model usage.

Benefits of technology

Enhances the accuracy of service provision by accurately determining user needs and reduces costs through optimized language model usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024010807_06112025_PF_FP_ABST
    Figure KR2024010807_06112025_PF_FP_ABST
Patent Text Reader

Abstract

A care robot according to an embodiment of the present invention comprises: a robot collection information generation unit for generating robot collection information related to a user by means of at least one sensing device installed in the care robot; a language model input unit for generating a prompt on the basis of the robot collection information and inputting the generated prompt into a language model; and a service determination unit for determining a care service to be provided to the user on the basis of an output of the language model.
Need to check novelty before this filing date? Find Prior Art

Description

Care robots that provide care services

[0001] The present invention relates to a care robot that collects information from a user and provides care services to the user based on the collected information.

[0002] Care robots are robotic technologies designed for the elderly, people with disabilities, and individuals requiring special care in an aging society. They provide a variety of services, including assistance with daily living, health monitoring, and emotional support, contributing to improving the quality of life for users. While care robots play a crucial role in making users' lives more convenient and safer, conventional care robots have several limitations in providing customized services that adequately reflect the individual needs and circumstances of users.

[0003] Artificial intelligence technologies, particularly conversational AI and natural language processing, play a key role in enhancing user-robot interactions. These technologies empower robots to understand the user's verbal input and generate appropriate responses, facilitating natural user communication and enabling deeper interactions. For example, recognizing when a user expresses specific emotions and responding appropriately is crucial for robots to support the user's emotional needs.

[0004] However, integrating these advanced interactive capabilities into robots presents various technical challenges. First, powerful data processing capabilities are required to effectively process and interpret the complexity of data generated from diverse users and environments. Furthermore, dynamic decision-making algorithms are needed to adapt and respond appropriately to changing user conditions and environments in real time. This requires the integration of advanced algorithms and machine learning models into the robot system, which increases the design and production costs of the robot.

[0005] Recent advancements in artificial intelligence (AI) have been remarkable, and efforts to integrate these technologies into robotic systems are increasing. LLMs can perform complex language understanding and generation tasks based on natural language data, and their application to robotics can make communication between robots and users more natural and effective. In particular, interactions utilizing LLMs play a crucial role in enabling robots to more accurately understand user needs and provide appropriate responses.

[0006] For the effective use of LLM, the quality of the prompts provided as input is crucial. Prompts are questions or commands presented by the robot to the LLM, and they determine the context and accuracy of the output generated by the LLM. Therefore, to generate appropriate and accurate prompts, the robot must be able to comprehensively collect and analyze diverse information, including not only the user's words but also their behavior, environmental context, and emotional state.

[0007] However, conventional robots have difficulty collecting various types of information related to users, and even if they do collect information, there are problems in processing this information to input it into language models such as LLM.

[0008] The present invention aims to solve the above-mentioned problems by providing a care robot that can collect information related to a user and determine the care service required by the user based on a language model.

[0009] In addition, the present invention seeks to provide a care robot that collects various types of information from a user so that a language model can more accurately determine the user's situation.

[0010] In addition, the present invention seeks to provide a care robot that can mitigate increased costs due to frequent use of a language model by controlling the timing at which information collected by the robot is input into the language model.

[0011] However, the technical tasks that this embodiment seeks to accomplish are not limited to the technical tasks described above, and other technical tasks may exist.

[0012] As a means for achieving the above-described technical task, one embodiment of the present invention may include a robot collection information generation unit that generates robot collection information related to a user through at least one sensing device installed in a care robot; a language model input unit that generates a prompt based on the robot collection information and inputs the generated prompt into a language model; and a service determination unit that determines a care service to be provided to the user based on an output of the language model.

[0013] The above-described problem-solving methods are merely exemplary and should not be construed as limiting the present invention. In addition to the exemplary embodiments described above, additional embodiments may exist, as described in the drawings and detailed description of the invention.

[0014] According to any one of the above-described problem solving means of the present invention, the present invention can provide a care robot that can collect information related to a user and determine a care service required by the user based on a language model to solve the above-described problems.

[0015] Additionally, the present invention collects various types of information from users, enabling the language model to more accurately determine the user's situation.

[0016] In addition, the present invention has the effect of alleviating the increase in cost due to frequent use of the language model by controlling the point in time when information collected by the robot is input into the language model.

[0017] Figure 1 is a configuration diagram of a care robot and a care system according to one embodiment of the present invention.

[0018] Figure 2 is a configuration diagram of a care robot according to one embodiment of the present invention.

[0019] FIGS. 3 to 5 are drawings illustrating a process for generating robot collection information according to one embodiment of the present invention.

[0020] FIG. 6 and FIG. 7 are drawings illustrating a process of moving a robot to generate robot collection information according to one embodiment of the present invention.

[0021] FIGS. 8 to 11 are diagrams illustrating a process of generating a prompt for input to a language model based on collected information according to one embodiment of the present invention.

[0022] Below, with reference to the attached drawings, embodiments of the present invention are described in detail so that those skilled in the art can easily implement them. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, irrelevant parts have been omitted for clarity of description, and similar reference numerals have been used throughout the specification to indicate similar elements.

[0023] Throughout the specification, when a part is said to be "connected" to another part, this includes not only the case where it is "directly connected" but also the case where it is "electrically connected" with another element in between. Furthermore, when a part is said to "include" a component, this should be understood to mean that, unless specifically stated to the contrary, it may include other components rather than excluding them, and does not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0024] In this specification, the term "unit" includes a unit realized by hardware, a unit realized by software, and a unit realized using both. Furthermore, a single unit may be realized using two or more pieces of hardware, and two or more units may be realized by a single piece of hardware.

[0025] Some of the operations or functions described herein as being performed by a terminal or device may instead be performed by a server connected to the terminal or device. Similarly, some of the operations or functions described as being performed by a server may also be performed by a terminal or device connected to the server.

[0026] The functions realized by the components described in this specification may be implemented in processing circuitry including general purpose processors, special purpose processors, integrated circuits, Application Specific Integrated Circuits (ASICs), Central Processing Units (CPUs), circuits and / or combinations thereof programmed to realize the described functions. A processor includes transistors or other circuits and is considered a circuit or processing circuit. The processor may be a programmed processor that executes a program stored in a memory.

[0027] In this specification, a circuit, part, unit, or means is hardware programmed to realize the described function or hardware that executes the function. The hardware may be any hardware disclosed in this specification or any hardware known to be programmed or executed to realize the described function.

[0028] If the hardware is a processor considered to be a circuit type, the circuit, the part, means or unit is a combination of hardware and software used to configure the hardware and / or the processor.

[0029] Hereinafter, an embodiment of the present invention will be described in detail with reference to the attached drawings.

[0030] FIG. 1 is a drawing for explaining a care robot (100) and a care system according to one embodiment of the present invention.

[0031] Referring to FIG. 1, the care system may include a care robot (100) and a server (40) for providing care services to a user (20).

[0032] The care robot (100) of FIG. 1 can be connected to a server (40) through a network (30). As illustrated in FIG. 1, the care robot (100) can be connected to the network (30) simultaneously or at intervals. The network (30) refers to a connection structure that enables information exchange between each node, such as terminals and servers, and includes a local area network (LAN), a wide area network (WAN), the Internet (WWW), wired and wireless data communication networks, telephone networks, wired and wireless television communication networks, and the like. Examples of wireless data communication networks include, but are not limited to, 3G, 4G, 5G, 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), WIMAX (World Interoperability for Microwave Access), Wi-Fi, Bluetooth communication, infrared communication, ultrasonic communication, visible light communication (VLC), LiFi, and the like.

[0033] The care robot (100) is a robot that includes its own means of transportation, and may include a camera for photographing a user (20), a microphone for receiving sound input from the user, a display for displaying information to the user, a speaker for outputting sound to the user, etc. In addition to the camera and microphone mentioned above, the care robot (100) may include a sensor for detecting the surrounding temperature, humidity, illuminance, pressure, etc. The care robot (100) includes an integrated control unit for controlling the above mentioned devices, and the integrated control unit may mean a device having a computational processing capability by having a memory and a processor.

[0034] In an embodiment of the present invention, the camera of the care robot (100) can capture an object and generate information in the form of an image. A representative example is an RGB camera that generates a pixel image with RGB (Red Green Blue) properties. The camera can also be a depth camera that provides distance information to the object, or an infrared camera that can capture images in a dark environment. The camera can take a photo to create a single image of the user (20) described below, or perform video capture to create an image composed of multiple frames.

[0035] Memory is a device that stores information and may include various types of memory, such as high-speed random access memory, magnetic disk storage devices, flash memory devices, and non-volatile memory such as other non-volatile solid-state memory devices. In the care robot (100) described below, memory may be described in the form of a database.

[0036] The care robot (100) can generate a prompt to be input into a language model based on information collected from a user (20), determine a care service to be provided to the user based on the output of the language model, and provide the determined care service to the user. A more detailed description of this will be provided later.

[0037] A user (20) is a person who receives care services provided by a care robot (100), and may primarily refer to patients with difficulty moving, the elderly, or the elderly. However, the user (20) is not limited to this, and may include various people who need to use a care robot (100).

[0038] The server (40) may include a memory in which a plurality of modules are stored, a processor connected to the memory and responding to the plurality of modules, and processing service information provided to the care robot (100) or action information controlling the service information, a communication means, and a UI (user interface) display means.

[0039] On the other hand, in the present invention, the server (40) may correspond to an external server capable of operating a language model. Information processed by the server (40) may include robot collection information, IoT collection information, cloud collection information, and service information available for provision of the care robot (100), which will be described later. This information may be converted into service information or action information necessary to direct and coordinate the actions of the care robot (100).

[0040] The care robot (100) can be linked to an external server including a cloud-based language model via a network (30). This external server performs complex natural language processing tasks and generates appropriate linguistic responses based on data received from the care robot (100). The care robot (100) can utilize these responses to interact with the user more naturally and effectively. The language model of the server (40) can provide the care robot (100) with information about the user's (20) situation, the types of care services that the care robot (100) can provide to the user, etc. in text format in response to data input from the care robot (100) in the form of a prompt. To this end, the server (40) can analyze and process data generated from multiple modules. In the above-described process, the communication means of the server (40) enables continuous data exchange with the care robot (100), and can support the user or administrator to monitor the status of the server and perform necessary operations through a user interface (UI) display means.

[0041] This structure ensures smooth information flow between the server (40) and an external server, and ultimately enables the provision of customized care services to users through complex data processing and response generation processes.

[0042] Figure 2 is a schematic diagram for explaining a care robot (100) according to one embodiment of the present invention.

[0043] Referring to FIG. 2, the care robot (100) may include a robot collection information generation unit (210), an IoT collection information generation unit (220), a cloud collection information generation unit (230), a service information generation unit (240), a language model input unit (250), a situation information generation unit (260), and a service decision unit (270).

[0044] And the robot collection information generation unit (210) may include a sensor unit (211), a collection information analysis unit (212), a camera control unit (213), and a robot movement unit (214). In addition, the language model input unit (250) may include an input time determination unit (251).

[0045] More specifically, the care robot (100) may include a robot collection information generation unit (210) that generates robot collection information related to a user through a sensor unit (211) including at least one sensing device installed in the care robot (100), a language model input unit (250) that generates a prompt based on the robot collection information and inputs the generated prompt into a language model, and a service determination unit (270) that determines a care service to be provided to the user based on the output of the language model.

[0046] The robot collection information generation unit (210) may include a sensor unit (211) including a camera that photographs a user, and a collection information analysis unit (212) that analyzes sensing information generated from the sensor unit (211) to generate the robot collection information.

[0047] And the robot collection information generation unit (210) may include a camera control unit (213) that controls at least one operation among vertical rotation of the camera, horizontal rotation of the camera, and height change of the camera, and a robot movement unit (214) that moves the care robot (100) to change the shooting position of the care robot (100).

[0048] In addition, the care robot (100) further includes an IoT (Internet on Thing) collection information generation unit (220) that generates IoT collection information based on information collected from an IoT device connected to the care robot (100), and the language model input unit (250) can generate a prompt based on the robot collection information and the IoT collection information.

[0049] And the care robot (100) further includes a cloud collection information generation unit (230) that generates cloud collection information from a cloud in which at least one of profile information about the user and pattern information about the user's daily pattern is stored, and the language model input unit (250) can generate the prompt based on the robot collection information and the cloud collection information.

[0050] And the care robot (100) further includes a service information generation unit (240) that generates service information that includes a list of care services that can be provided to a user by the care robot (100), and the language model input unit (250) can generate the prompt based on the robot collection information and the service information that can be provided.

[0051] And the language model input unit (250) includes an input point determination unit (251) that generates a prompt and determines an input point in time for inputting the prompt into the language model, and the language model input unit (250) can input the prompt into the language model in response to the determined input point in time.

[0052] And the care robot (100) may include a situation information generation unit (260) that generates situation information about the user's surrounding environment based on robot collected information.

[0053] In addition to the above-described configuration, the care robot (100) may include various configurations for checking the user's status and providing appropriate care services.

[0054] For example, the care robot (100) may include a display that displays the service provided to the user (20) and a speaker that outputs sounds generated in the process of providing the service and voice messages to the user (20) in the form of sounds.

[0055] The display is a type of liquid crystal display device, and can output one or more of text, images, and videos of predetermined information. Here, the predetermined information may include status information of the care robot (100), such as communication status strength information, remaining battery level information, and wireless Internet ON / OFF information. The display can display content related to the service described below, and can also display information converted into text from voice output through the speaker. For example, when the care robot (100) provides a voice corresponding to service-related content to the user and gives an action instruction, text such as "Raise your right arm higher" may be output on the display. In addition, as described below, when the care robot (100) performs identity recognition of the user (20), text such as "Come closer" or "Look straight at me" may be output through the display to instruct the user to take a necessary action.

[0056] The text output on this display may be one of the previously described pieces of information repeatedly output, or multiple pieces of information may be output alternately, or specific information may be set as default and output. For example, status information of the care robot (100), such as communication status strength information, remaining battery level information, and wireless Internet ON / OFF information, may be set as small text at the top or bottom of the display and continuously output by default, and other information may be output alternately.

[0057] Meanwhile, since the display can output one or more of images and videos, in this case, it is preferable to implement the display as a large-sized liquid crystal display with high resolution to improve visibility, rather than a display that only outputs text. As described below, the display may be configured on the exterior or interior of the care robot, or may be configured as a separate device outside the care robot (100).

[0058] Meanwhile, the display is positioned at the front of the care robot (100), so that a user (20) looking at the front of the care robot (100) can view the contents of the screen displayed on the display unit together with the care robot (100).

[0059] The speaker can output various sounds, including voice. Here, voice corresponds to auditory information output by the care robot (100) for interaction with the user, and the type of voice can be set in various ways through a media-only application installed on the user's (20) terminal (not shown) or through direct control of the care robot (100).

[0060] For example, you can choose the type of voice output through the speaker from among various voices, such as male voice, female voice, adult voice, and child voice, and you can even choose the type of language, such as Korean, English, Japanese, and French.

[0061] Meanwhile, although the speaker outputs sound, it can also perform the functions of a typical speaker, such as general sound output. For example, if a user (20) wishes to listen to music through the care robot (100), the music can be output through the speaker. In addition, if a video is output through the display, sound synchronized with the video can be output through the speaker.

[0062] FIGS. 3 to 5 are drawings illustrating a process for generating robot collection information according to one embodiment of the present invention.

[0063] The care robot (100) can capture at least a portion of the user's (20) body (e.g., face) through a camera included in the sensor unit (211) of the robot collection information generation unit (210).

[0064] The care robot (100) can detect a face in an image being captured by a camera through a face detection model. The face detection model can be used in a similar manner to the object detection model described above. Unlike the case where a specific object is detected in an image through the object detection model, the face detection model can identify the location of a face in the image and generate a bounding box corresponding to the location. The face detection model can be a detection model based on a convolutional neural network (CNN), and for example, object detection algorithms such as RCNN, Fast RCNN, YOLO, Single Shot Detector (SSD), Retina-Net, and Pyramid Net can be utilized.

[0065] To improve the accuracy of the results, the face detection model can input information about the user's height, either pre-entered or estimated using a skeletal extraction model. This allows the care robot (100) to rotate the camera vertically to face the user's face, enabling it to accurately recognize the user's face even at close range while standing. A more detailed description of this will be provided later.

[0066] FIG. 3 is a drawing for explaining an identity recognition model (304) for applying a face detection model (302) to an image (301) and recognizing a user's identity from a face image (303) generated based on the location of the face determined through the face detection model (302).

[0067] A feature value (305) can be extracted from a face image (303) processed as a bounding box by an identity recognition model (304) and stored in a database (306). Then, a face detection model and an identity recognition model are applied to a new image (307), and the extracted feature (309) is compared with the feature points already stored in the database, thereby performing identity recognition based on the user's face image. A model based on a convolutional neural network can be applied as the identity recognition model (304), and examples thereof include algorithms such as VGG-Face, FaceNet, OpenFace, DeepFace, and ArcFace.

[0068] Referring to FIG. 3, it can be seen that a face detection model (302) is applied to an image (301) of a person to create a bounding box for an area where a face is located, thereby extracting a face image (303), and inputting this face image (303) into an identity recognition model (304) to extract its feature values ​​(305) and store them in a database (306). Thereafter, when a new image (307) is input, the extracted features (309) are applied to the above-described face detection model and identity recognition model (308) to compare and match them with the feature values ​​stored in the existing database (306), thereby recognizing the identity.

[0069] To improve the accuracy of the results, the identity recognition model (304) can provide voice guidance to the user (20) to direct the face of the user toward the camera of the care robot (100). This allows the care robot (100) to generate an image of the front of the user's (20) face, thereby enabling more accurate recognition of the user's identity. Whether the user's face is facing the camera can be confirmed through the result data generated by the face direction recognition model described below.

[0070] Figures 4 and 5 are drawings illustrating a face direction recognition model for recognizing a face direction from a face image generated based on the position of the determined face.

[0071] First, a face detection model (402) can extract a face image (403) of a user (20) based on an image (401) of the user. Thereafter, a landmark extraction model (404) can extract a facial landmark (405) from the face image (403). In this process, a convolutional neural network can be used, and a face direction recognition model can output a result value regarding the direction in which the face is facing based on the position of the eyes, the position of the nose, and the facial outline.

[0072] Referring to FIG. 4, it can be seen that a bounding box for a face image (403) is generated from an input image (401) through a face detection model (402), and the face image (403) is input to a landmark extraction model (404) to extract a face landmark (405).

[0073] Referring to Figure 5, the process of outputting a result value regarding the direction in which the face is facing from the extracted facial landmark can be seen.

[0074] The robot collection information generation unit (210) of the care robot (100) may include a sensor unit (211) and a collection information analysis unit (212) that analyzes the information collected from the sensor unit (211). The sensor unit (211) may include various sensing devices including a camera. As described with reference to FIGS. 3 to 5, the user's face may be captured through the camera of the sensor unit (211), and the user's identity may be confirmed based on the captured facial image of the user.

[0075] The above camera can capture at least one of the user's (20) face, the user's (20) walking scene, the user's (20) body shape, and the clothing worn by the user (20).

[0076] And the collection information analysis unit (212) can analyze images captured by the camera to generate robot collection information.

[0077] To this end, images of the user's (20) face, walking scene, body type, and clothing worn, or feature values ​​derived from the images, may be stored in a database in which the care robot (100) is placed or in another database accessible by the care robot (100) through a wired or wireless connection, and matched with the user (20).

[0078] The care robot (100) can move around the user (20) and change the camera shooting direction, as illustrated in the example of FIG. 6. Furthermore, the care robot can provide messages to the user (20) via voice or text output, such as "Walk in front of me" or "Bring your face closer to the camera."

[0079] More specifically, the care robot (100) may include a sensor unit (211) and a collection information analysis unit (212) capable of recognizing and analyzing various physical characteristics and behaviors of the user (20). The camera included in the sensor unit (211) may be a primary means of collecting information related to the user by capturing images of the user's face, walking scenes, body shape, and clothing worn. Images captured by this camera play a key role in analyzing the user's identity.

[0080] The care robot (100) can capture a walking scene of a user (20). The care robot (100) can analyze the walking speed, gait pattern, and body movements of the user (20) from the captured walking scene. The care robot (100) can extract the characteristic values ​​of the walking pattern from the captured walking image using a motion analysis algorithm. This is based on the unique walking style of each user. In addition, the care robot (100) can analyze the user's information by comparing the extracted characteristic values ​​with the walking pattern stored in the database.

[0081] The care robot (100) can capture the overall body shape of the user (20) and analyze the silhouette, body shape ratio, size, etc. The care robot (100) can extract body shape feature values ​​through a body shape recognition algorithm and use these to identify the user or analyze the type of clothing worn by the user.

[0082] The care robot (100) can capture the color, pattern, style, etc. of the clothing worn by the user (20). Thereafter, the care robot (100) can use a clothing recognition algorithm to derive feature values ​​from the captured clothing image. Using the feature values ​​of the clothing, the care robot (100) can identify that the user is wearing a specific piece of clothing at a specific time.

[0083] The collected information analysis unit (212) can analyze the captured images to generate robot collected information. The facial recognition model, motion analysis algorithm, body type recognition algorithm, and clothing recognition algorithm described above can be any known image analysis algorithm. Furthermore, these algorithms or models can be performed by the collected information analysis unit (212) of the care robot (100). During the analysis of the collected information, comparative analysis can be performed with user information stored in a database within the care robot (100) or an external database to which the robot is connected.

[0084] The information analyzed by the collected information analysis unit (212) can be used to generate a prompt in the language model input unit (250), as described later.

[0085] The sensor unit (211) may include a microphone for recognizing the voice of the user (20).

[0086] And the collection information analysis unit (212) can generate robot collection information based on the sensing information generated by the sensor unit (211).

[0087] The microphone can receive input from the user or the user's surroundings to capture the user's voice and analyze the user's voice characteristics through voice recognition technology. The collected information analysis unit (212) can analyze the user's voice's unique pitch, tone, and stress based on the sound input through the microphone, and convert the user's voice or voice input into text. A STT (Speech-To-Text) algorithm can be used in this process.

[0088] The configuration of these various sensor units (211) can be installed in a care robot (100) to collect information about the user in various environments and situations and contribute to generating specific robot-collected information.

[0089] FIG. 6 illustrates a care robot photographing a user (20) at a first location (601), then moving to a second location (602) to photograph the user (20) again to more accurately analyze the situation through the photographed image. The care robot (100) can provide the user (20) with messages such as "Please take off your glasses" or "You must not smile" via voice or text output. This can induce certain behaviors from the user (20) and derive more accurate robot-collected information.

[0090] In order for the care robot (100) to more accurately recognize and photograph the user's condition or surrounding situation, a camera control unit (213) or a robot moving unit (214) may be used to change the position or shooting angle of the camera.

[0091] As described above, the robot collection information generation unit (210) may include at least one of a camera control unit (213) that controls at least one operation of a vertical rotation of the camera, a horizontal rotation of the camera, and a height change of the camera, and a robot movement unit (214) that moves the care robot (100) to change the shooting position of the care robot (100).

[0092] The camera control unit (213) may include a tilting unit (not shown) that rotates the camera vertically, a panning unit (not shown) that rotates the camera horizontally, and an elevating unit (not shown) that changes the shooting height of the camera. The camera control unit (213) may change the shooting direction of the camera. The tilting unit, panning unit, and elevating unit may be configured to be directly connected to the camera, but depending on the shape or size of the care robot (100), the tilting unit, panning unit, and elevating unit may be placed in a state physically separated from the camera. In addition, at least one function of the tilting unit, panning unit, and elevating unit may be replaced by the rotational movement of the care robot (100) itself or another device that elevates the height of the care robot (100).

[0093] The robot collection information generation unit (210) improves the pose recognition rate of the care robot (100) by controlling the camera control unit (213) if the pose recognition rate is below a predetermined standard, and the camera control unit (213) can change the shooting direction of the camera by controlling at least one of the tilting unit, the panning unit, and the lifting unit.

[0094] The robot moving unit (214) includes a configuration for moving the care robot (100), and can change the position of the care robot (100) or rotate it through the movement. Accordingly, the camera can capture the user (20) at the changed position of the care robot (100), thereby improving the pose recognition rate in the process of generating hand gesture information or full body pose information. In addition, the robot moving unit (214) can move the care robot (100) so that it can follow the user (20) and move as needed.

[0095] The robot moving unit (214) provides a means for the care robot (100) to move within a specific space in relation to driving according to a movement command of the control device. More specifically, the robot moving unit (214) includes a motor and a plurality of wheels, which, when combined, can perform the functions of driving, changing direction, and rotating the care robot (100).

[0096] Fig. 7 illustrates the relative positions between the user (20) and the care robot illustrated in Fig. 6 expressed on a virtual map. The care robot can calculate coordinate information from the starting point A to the destination B on the virtual map and an angle (θ) to look at the user (20) after moving to the destination. Thereafter, the care robot can move to the coordinate information (702) of the destination, and upon arriving at the destination, control the care robot to rotate by the angle (θ) to take a picture of the user (20).

[0097] In the example of Fig. 7, a care robot (701) located at A can have a coordinate value (0,0) on a virtual map corresponding to location A, and can move to location B by receiving a coordinate value (-300,300) of location B. Thereafter, it can rotate by the input rotation angle (θ) to photograph the user (20), and at this time, the coordinate value (-300, 0) of the user (20) and the coordinate value of the care robot can be continuously updated by the location tracking unit.

[0098] In another embodiment of the present invention, unlike the aforementioned case, the care robot can control the movement of the exercise assistance service robot by generating a virtual map solely based on the photographic information generated by the vision sensor, rather than relying on the photographic information generated by the vision sensor and the lidar. For this purpose, Vision SLAM (Visual Simultaneous Localization and Mapping) technology can be used.

[0099] The robot-collected information may be information in the form of an image or text, as each piece of information directly collected by the care robot (100) through the language model input unit (250). The language model input unit (250) may include an image captured by the camera of the sensor unit (211) in the robot-collected information. In addition, the language model input unit (250) may convert user-related information generated by the collected information analysis unit (212) into text to generate a prompt.

[0100] The robot-collected information may include the user's analyzed biometric information (the user's heartbeat, pulse, blood pressure, breathing, stress, etc.), vision information generated by analyzing an image captured by a camera (the user's identity, facial expression, posture, behavior, surrounding objects), spatial information (location on a 3D map) of the care robot's (100) location, content information (game score, number of games, etc.) provided by the care robot (100), conversation information generated by analyzing the conversation history between the user and the care robot (100) (health information, emotional information, hobby information, etc., information that can be obtained through conversation), the user's schedule information (medication, meals, going out, other schedules), the user's health information (personal illness, sleep information, gait analysis information, pain area, etc.), and other information (weather, news, notices, etc.).

[0101] The language model input unit (250) of the care robot (100) can convert the robot collection information generated by the above method into text to generate a prompt.

[0102] For example, the language model input unit (250) can generate text in the form of "Location: living room, facial expression: pain, posture: lying down, surrounding object: water bottle, health information: complained of dizziness 1 hour ago" based on robot-collected information, and input the generated text into a prompt.

[0103] In another embodiment of the present invention, the language model input unit (250) can generate a prompt based on at least one of robot collection information, IoT collection information, cloud collection information, and service information that can be provided.

[0104] IoT collection information can be generated by the IoT collection information generation unit (220). The IoT information generation unit (220) can generate IoT collection information based on information collected from an IoT device connected to the care robot (100).

[0105] More specifically, the care robot (100) can receive various types of information from IoT devices connected wirelessly or wired. The care robot (100) can be located near the IoT device and connected wirelessly via Bluetooth or infrared communication, or directly via wired connection. Furthermore, the care robot (100) can receive information from other external IoT devices that transmit and receive information to a server connected to the care robot (100) via a network, via the server.

[0106] These IoT devices may be exemplified by various smart electronic products installed indoors (e.g., TVs, air conditioners, lighting devices, artificial intelligence speakers, smart watches, etc.), and each IoT device may provide various sensing information related to the user's behavior or the environment around the user (e.g., temperature, illuminance, weather, content being viewed by the user, the user's IoT device usage history, etc.) to the care robot (100). The care robot (100) may process information received from the connected IoT device into a text format to generate IoT collection information. At this time, the generated IoT collection information may include the creator of the corresponding IoT creation information, the label of the created information, numerical values, etc. For example, IoT collected information converted into text can be generated as "Smart TV: Currently watching channel - 8, TV volume size: 15, Viewing time: 20 minutes" or "Smart air conditioner: Current room temperature: 20 degrees, Target room temperature: 15 degrees, Air conditioner operation time: 30 minutes, Wind speed: Maximum" or "Living room temperature: 27 degrees, Living room humidity: 50%, Fine dust: Normal, Air conditioner 1: Off, Air purifier: On, Pulse: 120, Maximum blood pressure: 140".

[0107] Cloud collection information can be generated by a cloud collection information generation unit (230). More specifically, the cloud collection information generation unit (230) can generate cloud collection information from a cloud in which at least one of profile information about a user and pattern information about the user's daily patterns is stored.

[0108] And the cloud collection information generation unit (230) may include a pattern analysis model that outputs expected pattern information on activities that the user is expected to perform at the current time when activity information including the user's activity history by time zone and information on the pattern generation period are input.

[0109] Cloud-collected information may include user basic information (user name, gender, age, date of birth, health status, family information, height, weight, etc.), user daily pattern information (events that occur frequently at specific times, events that occur periodically, events that occur continuously, events that occur depending on the environment or emotions, etc.).

[0110] Here, user daily pattern information can be generated by extracting events according to the following rules, taking into account time, frequency, environment, emotion, and event relationships based on information acquired over a specific period (e.g., one month).

[0111] The first daily pattern can be derived by dataifying events that occur at specific times (e.g., waking up, going to bed, eating, exercising, watching TV, bathing, etc. - time, behavior) and extracting daily patterns based on frequency (e.g., if they occur more than n times during a specific period).

[0112] Secondary daily patterns can be derived by extracting daily patterns based on frequency (e.g., when an event occurs more than n times during a specific period) by dataifying time and behavior, such as events that occur at a specific cycle (e.g., a user visits the bathroom once every two hours, exercises once every three days, and visits the hospital once a month).

[0113] The third daily pattern can be derived by extracting the daily pattern by frequency (if it occurs more than n times during a specific period) by dataifying continuous behaviors such as events that occur continuously (e.g., going to the bathroom after exercising, going to the bathroom after waking up).

[0114] The fourth daily pattern can be derived by extracting daily patterns by frequency (e.g., if they occur more than n times during a specific period) by dataifying time, environment, and behavior, such as events that occur depending on the environment (e.g., doing laundry when the weather is good, eating buchimgae when it rains, canceling going out when it snows).

[0115] The fifth daily pattern can be derived by extracting daily patterns by frequency (for example, if it occurs more than n times during a specific period) by dataifying time, emotion, and behavior, such as events that occur according to the user's emotions (for example, if you are in a good mood, go for a walk at 3 PM, if you are in a bad mood, go to bed an hour earlier than usual).

[0116] Here, the specific period can vary depending on the pattern to be extracted. For example, for the first daily pattern, the specific period can be set to one month, as similar trends appear every month. For the fourth daily pattern, the specific period can be set to one year, as different trends appear every month.

[0117] The priority of each daily pattern can also be customized to suit the user. For example, even if a user exercises every day at 3 PM (the first daily pattern), if it rains at that time, they might go out to eat buchimgae (the fourth daily pattern). This information can be digitized, allowing the user to determine the daily pattern based on the higher frequency of the event when the probability of occurrence of certain events overlaps.

[0118] As users' behavior or environment changes, daily patterns can be continuously updated, and based on this, the analysis model for daily patterns can be continuously optimized.

[0119] These anomaly patterns can be extracted rule-based or by leveraging artificial intelligence, such as machine learning. When utilizing AI, the user's schedule, current time, and recently occurring events are input into the AI, and the output can be the event with the highest probability of occurrence. A pattern analysis model that outputs predicted patterns of activities expected to occur at the current time can refer to a model based on artificial intelligence, such as machine learning, as described above.

[0120] The expected pattern information on the daily pattern generated by the cloud collection information generation unit (230) is included in the cloud collection information and can be used to generate subsequent prompts.

[0121] The service information available for provision can be generated by the service information generation unit (240).

[0122] The service information generation unit (240) can generate service information that includes a list of care services that can be provided to a user by a care robot (100).

[0123] The information on available services may vary depending on the type, shape, size, function, or malfunction of the care robot (100). Furthermore, the information on available services may vary depending on the user's age, health, situation, or the current location or location of the care robot (100). In other words, the information on available services may include, in addition to all care services that the care robot (100) can provide, a list of services that the care robot (100) can provide to the user in the current situation.

[0124] For example, the list of services included in the controllable service information may include services for providing content for user cognitive training or exercise assistance, games, emotional conversation, music and video playback, emergency response services, diary writing services, schedule alarm services, meal reminders, medication reminders, security services, video chat services, weather information provision services, educational content provision services, biometric information acquisition, health management services, cleaning services, IoT control, SNS search services, etc.

[0125] The information on the available services further includes information on the priority among the care services that can be provided to the user, and the service decision unit (270) can determine the care service to be provided to the user based on the output of the language model and the priority.

[0126] In addition to simply listing the services available to users in parallel, the information on available services may further include information on priorities for determining which services should be provided to users first.

[0127] For example, if the information on the available services includes a medication reminder service and a weather provision service, and the medication reminder service has a higher priority than the weather provision service, the service decision unit (270) may decide to provide the medication reminder service to the user as a priority rather than the weather provision service.

[0128] On the other hand, priority information is not specified in advance for each care service, and can be set to be provided differently depending on the user's situation information.

[0129] For example, the service decision unit (270) may decide to give priority to providing care services related to the first-priority situation when a risky situation such as falling is the first priority and exercising in the living room at 3 PM is the second priority.

[0130] In such cases, rule-based and AI-based methods can be used to obtain priority information.

[0131] Rule-based methods can involve generating prioritized information based on human-defined rules. For example, meal reminders might be prioritized for users with diabetes. Similarly, exercise might be prioritized for users with obesity or high blood pressure. For example, elderly users might prioritize emergency situation awareness and methods. In environments with many stairs, crime prevention and emergency situation awareness might be prioritized.

[0132] An AI-based method may involve AI determining priorities based on all information acquired by the care robot (100) (robot-collected information, IoT-collected information, cloud-collected information, and available service information) and generating optimal priority information tailored to each space and user. Priority information may be updated periodically or whenever a situational assessment is made. The AI ​​may be machine learning-based or include a generative language model.

[0133] Conversely, rule-based and AI-based methods can be utilized together. This means that human-defined rules can be given top priority, while AI-defined priorities can be leveraged for situations where no rules are defined.

[0134] FIGS. 8 to 11 are diagrams illustrating screens that display a process of generating a prompt for inputting into a language model based on collected information according to one embodiment of the present invention.

[0135] In the present invention, a prompt may refer to text information in natural language format input to a language model. Furthermore, with the recent development of generative language models, there is a growing trend toward developing generative models capable of multimodal input, such as images, document files, and sound files, in addition to text information. Accordingly, prompts are not limited to text and can include various forms of information, such as images, sound, documents, and web links.

[0136] The language model used in the present invention may refer to a highly developed generative model or a large language model. The language model of the present invention may refer to a model based on the underlying technology of ChatGPT, such as GPT-3 and GPT-4, a generational transformer network developed by OpenAI, BERT (Bidirectional Encoder Representations from Transformers), a model developed by Google, RoBERTa (Robustly Optimized BERT approach), a variant of BERT developed by Facebook, T5 (Text-to-Text Transfer Transformer), etc.

[0137] Although these models employ different architectures and learning methods, they all possess the ability to perform advanced natural language processing, excelling at understanding complex linguistic contexts and generating appropriate output.

[0138] The language model of the present invention possesses the ability to process multimodal data. The language model can accept input data in various forms, such as text, images, and voice, and generate appropriate text or voice output based on the input. This multimodal capability not only facilitates natural and effective interaction with the user, but also provides the care robot (100) with the ability to more accurately understand and respond to the user's needs.

[0139] Language models are trained on large datasets, embedding extensive knowledge that can respond to a variety of linguistic situations. With natural language understanding (NLU) and natural language generation (NLG) capabilities, language models can generate human-like conversations, playing a crucial role in identifying the nuanced needs of individuals in need of care and providing tailored services.

[0140] The language model of the present invention can be implemented as a cloud-based system installed and operated externally to the care robot (100). This structure utilizes high-performance computing resources to quickly and efficiently perform complex calculations and data processing tasks, and can continuously improve the model's performance through continuous updates and learning. Since the language model operates on a cloud basis, the care robot (100) can access the language model via a network, exchange data in real time, and receive necessary support.

[0141] Language models offer a wide range of capabilities beyond simple text processing. In particular, their multimodal data processing capabilities significantly help robots understand not only the user's speech but also nonverbal cues and environmental context. This significantly improves the accuracy and efficiency of services required to care for users.

[0142] Meanwhile, the above-mentioned prompt is generated by the language model input unit (250) and may be generated based on at least one of robot collection information, IoT collection information, cloud collection information, and service information that can be provided.

[0143] Prompts can consist of multimodal data, including images as well as text, enabling language models to generate richer and more contextually relevant responses. For example, prompts can include image data of the user's facial expressions or surroundings, allowing the language model to respond more sensitively to the user's emotional state or needs.

[0144] Prompts may also include guidance information. Guidance information specifies the desired response format for the care robot (100) and includes specific instructions to assist in determining which service options the robot can execute. This information serves to tailor the responses generated by the language model to the context in which the care robot operates, and provides a crucial basis for recommending the most appropriate service for the user's current needs and circumstances among the services the robot can provide.

[0145] Figure 8 illustrates a screen where a prompt (800) generated by a language model input unit (250) is displayed.

[0146] Referring to FIG. 8, the prompt (800) may include guide information (810), basic information about the user (820), text corresponding to robot collection information (830), text corresponding to IoT collection information (840), text corresponding to service information available for provision (850), and text corresponding to priority information (860). Each text may be generated by the language model input unit (250) based on the robot collection information, IoT collection information, service information available for provision, and priority information.

[0147] The prompt (800) may include instructional text in the preamble, for example, "Based on the information below, determine the current situation and select the most appropriate service to provide to the user. The format of the answer is as follows."

[0148] Guide information (810) may include text such as, for example, "1. Situation assessment: Answer 2. Appropriate service: Answer."

[0149] Basic information about the user (820) may include text such as, for example, "[Basic information] Name: Kim Young-hee, Gender: Female, Age: 72, Health status: Diabetes."

[0150] The text (830) corresponding to the robot-collected information may include, for example, text such as "living room temperature 27 degrees, living room humidity: 50%, fine dust: normal, air conditioner 1: off, air purifier: on, shock sensor: shock detected, pulse: 120, maximum blood pressure: 140, minimum blood pressure: 90."

[0151] The text (830) corresponding to the service that can be provided may include, for example, text such as "cognitive training, exercise content, game, emotional conversation, music and video playback, emergency response, 119 report, diary writing, schedule notification, meal notification, medication notification, crime prevention, video chat, weather information, education, detailed information acquisition, health management, cleaning, IoT control, SNS."

[0152] Text (860) corresponding to priority information may include, for example, text such as "Priority 1: Emergency response, 119 report, Priority 2: Meal reminder, Meal reminder, Priority 3: Other."

[0153] Fig. 9 illustrates an answer (900) generated based on a prompt input from a care robot (100) into which a language model is input. The answer of Fig. 9 may be an example of an answer generated by the language model in response to the prompt of Fig. 8. As shown in Fig. 9, the language model analyzes the content of the prompt to analyze the user's situation in response to the instructions of the guide information (810) and generates an answer regarding the type of care service that should be provided.

[0154] The answer (900) generated in Figure 9 is, for example, "Situational Assessment: Kim Young-hee (female, 72 years old) is lying in the living room with a painful expression, and complained of dizziness an hour ago. Her pulse was recorded as 120, her systolic blood pressure as 140, her diastolic blood pressure as 90, and the shock sensor detected a shock. This information suggests that Kim Young-hee may be suffering from a serious health problem. In addition, she has diabetes and needs to be more careful about her health.

[0155] Appropriate services: Emergency response and 119 reporting: The current situation is assessed as an emergency. Kim Young-hee's health, recent symptoms of dizziness, high pulse and blood pressure, and shock indicate the need for immediate medical attention. Therefore, it is of utmost importance to immediately call 119 and request emergency medical assistance.

[0156] Depending on your situation, you may need to utilize health management services for ongoing monitoring and management. Meal reminders and medication reminders for diabetes management will also be important.

[0157] Figure 10 illustrates a care robot (100) inputting an image into a language model and an answer output from the language model.

[0158] Referring to Fig. 10, the query content (1000) between the care robot (100) and the language model is exemplified.

[0159] In the example of Figure 10, the query content (1000) is, for example, "Language model: The picture shows two elderly people having a meal. On the left is a white-haired grandmother with a scarf on her head, wearing a beige sweater, eating. On the right is an elderly person wearing a gray sweater. The elderly person is wearing glasses and is eating. There is food on a plate on the table and she is eating. The elderly person seems to be talking to another elderly person, and this is inside the community center of a nursing home, which is a home. The overall atmosphere is warm and quiet.

[0160] Care Robot: What are you eating?

[0161] Language model: The food in the photo appears to be a typical Western meal. On the plate is a single dish that appears to be fried potato slices or rice, with several carrot slices and shredded salad next to it. A beverage, likely a soft drink or water, is in front of the elderly person. The elderly people in the photo are seen lifting their food to the table one bite at a time, and the two are seen having a pleasant conversation during the meal. This could include text like this:

[0162] When an image captured through a camera of a care robot (100) and guide information requesting analysis of the image are input together into a language model, the language model can analyze the image and generate an answer about the situation, clothing, positioning of the person, and atmosphere of the person in the image.

[0163] Fig. 11 illustrates a prompt (1100) generated according to one embodiment of the present invention. Referring to Fig. 11, the prompt (1100) may include image information (1110), first guide information (1120), basic information about the user (1130), text corresponding to IoT collection information (1140), text corresponding to service information available for provision (1150), text corresponding to priority information (1160), and second guide information (1170).

[0164] It can be seen that the example of Fig. 11, unlike Fig. 8, includes image information (1110) and second guide information (1160) in the prompt (1100). The image information (1110) may include an image captured by the camera of the care robot (100). In addition, the second guide information (1160) may be composed of text preset by the designer or user of the care robot (100).

[0165] In FIG. 11, the first guide information (1120) may include text such as, for example, "Please choose the most suitable service that can be provided to the user by providing an example with the information below. Also refer to the guidelines. Please answer as follows: 1. Situation assessment: Answer, 2. Available services: Answer, 3. Include details: Answer."

[0166] And basic information about the user (1130) may include text such as, for example, “Situation judgment: Answer, Name: Kim Young-hee, Gender: Female, Age: 72, Health status: Diabetes.”

[0167] And the text (1140) corresponding to the IoT collected information may include text such as, for example, "Smart device progress information, living room temperature: 27 degrees, living room humidity: 50%, message screen: unknown, earphones: off, shock detection status: off, pulse status: not detected, pulse: 70, maximum pulse: 100, minimum pulse: 70."

[0168] And the text (1150) corresponding to the service information that can be provided may include text such as, for example, “cognitive training, exercise content, game, emotional conversation, music and video playback, emergency response, 119 report, diary writing, schedule notification, meal notification, medication notification, crime prevention, video chat, weather information, education, detailed information acquisition, health management, cleaning, IoT control, SNS.”

[0169] And the text (1160) corresponding to the priority information may include text such as, for example, "Priority 1: Emergency response, 119 report, Priority 2: Meal reminder, Meal reminder, Priority 3: Other."

[0170] And the second guide information (1170) may include text such as, for example, "Guideline: If it is determined that the user in the photo is eating, the meal menu is analyzed and a combination of meal menus is displayed. The goal is to provide a meal menu that matches the user's health condition. For example, a low-sugar diet is recommended for a diabetic patient. In case of an emergency, the user is recommended to receive immediate first aid and call 119. It is important to keep the user stable and safe during the time until the emergency rescue team arrives after detecting an emergency. In addition, for the user's safety, it is necessary to set daily meal reminders and medication reminders and help the user take their medication at the set time. It is also important to record and analyze conversations with the user through a smart device to provide regular health checkups and diagnoses. If necessary, additional information about the meal menu is added, and points are provided to help the user make better decisions about his or her health."

[0171] The language model into which the prompt of Fig. 11 is input can be a language model that can receive and analyze data in the form of images as well as text, and can support data in a multi-modal format.

[0172] The care robot (100) can determine which image among the multiple images to input into the prompt without including all of the multiple captured images in the prompt to be input into the language model.

[0173] The care robot (100) can select photos highly relevant to the situation or event of a plurality of input candidate images. In this process, the artificial intelligence (AI) installed in the care robot (100) can be utilized. The care robot (100) inputs the input candidate images into an AI model, and the AI ​​model can output a numerical value indicating the degree of event occurrence (e.g., in the range of 0.0 to 1.0) as a result value. Thereafter, the image with the highest result value can be included in a prompt for input to the language model.

[0174] Unlike the first guide information (1120), the second guide information (1160) may include more specific guidance information for the language model's answers. In another embodiment of the present invention, the guidance information may be used for fine-tuning the initial language model without being entered into the prompt. In this case, the language model can be utilized without having to enter the guidance information into the prompt each time.

[0175] The language model input unit (250) may include an input point determination unit (251) that generates a prompt and determines an input point in time for inputting the generated prompt into the language model. In addition, the language model input unit (250) may input the prompt into the language model in response to the determined input point in time.

[0176] The input point determination unit (251) can determine the input point based on pattern information about the user's daily pattern.

[0177] In addition, the care robot (100) may further include a situation information generation unit (260) that generates situation information about the user's surroundings based on robot-collected information. In addition, the input time determination unit (251) may determine the input time based on the situation information.

[0178] The process of the care robot (100) inputting prompts into the language model and receiving responses from the language model may incur fees or consume computing resources related to the use of the language model. Therefore, if the care robot (100) inputs prompts into the language model frequently, excessive fees or consumption of computing resources may occur. Conversely, setting a long interval between inputs may prevent the provision of appropriate care services in critical situations. Therefore, there is a need to control the input timing of prompts as needed.

[0179] In one embodiment of the present invention, the input time determination unit (251) may determine the input time according to a preset cycle. For example, the input time determination unit (251) may determine the input time so that a prompt is input to the language model every hour.

[0180] In another embodiment of the present invention, the input time determination unit (251) may determine the input time according to the user's schedule and daily routine. The input time determination unit (251) may determine the input time at a preset time according to the user's daily routine. For example, if the user has a pattern of waking up at 7 o'clock, the input time determination unit (251) may determine the input time so that the care robot (100) moves to the user's bedroom at that time, assesses the situation, generates a prompt, and inputs the prompt into the language model. As another example, if the user has a pattern of taking medicine at 1 o'clock in the afternoon, the care robot (100) may move near the user at that time, assesses the situation, generates a prompt, and inputs the prompt into the language model.

[0181] In another embodiment of the present invention, the input timing determination unit (251) can determine the situation and input timing based on information collected in real time. The care robot (100) can generate a prompt based on robot-collected information and IoT-collected information acquired in real time and make a primary determination on whether to input the prompt into the language model. This primary situation determination can be performed based on situation information derived from the care robot (100) or situation information generated by the server based on information provided from the care robot (100).

[0182] The situation information generation unit (260) included in the care robot can generate situation information used to determine the input time of the input time determination unit (251).

[0183] FIG. 12 is a drawing illustrating a process in which a care robot according to one embodiment of the present invention recognizes a user's situation through a situation information generation unit.

[0184] Referring to FIG. 12, it can be seen that the care robot (1201) takes a picture of the user and generates the picture information (1202). Thereafter, the picture information (1202) can be input into a space classification model (1203), an object detection model (1204), a skeleton extraction model (1205), and an action recognition model (1206) to generate element information including the result values ​​for each analysis model.

[0185] Next, it can be seen that integrated information (1209) is generated by integrating element information, user location information (1207), and time information (1208). Thereafter, the integrated information (1209) is input into the situation information generation unit (1210) to generate situation information (1211) regarding the current situation.

[0186] The spatial classification model (1203) for generating the above-mentioned integrated information (1209) is an object detection model (1204), a skeleton extraction model (1205), and an action recognition model (1206) that receive shooting information (1202) generated by the camera of the care robot (1201) as input and provide a preset output for each model, and may correspond to a machine learning or deep learning model trained for graphic processing.

[0187] More specifically, the spatial classification model (1203) is a model that outputs the type of space in which the user is located based on the captured information, and can output the class of the captured space. The spatial classification model can utilize a known classification model based on a convolutional neural network. For this purpose, a convolutional neural network architecture such as AlexNet, VGG-16, Inception, ResNet, or MobileNet can be used.

[0188] The object detection model (1204) can output the class of the object being photographed based on the photographing information. The object detection model can use a classification model based on a convolutional neural network (CNN), and can output a bounding box indicating the type of the detected object and the area in which the object is located. For this purpose, a convolutional neural network architecture such as AlexNet, VGG-16, Inception, ResNet, and MobileNet can be used. Representative object detection models include RCNN, Fast RCNN, YOLO, Single Shot Detector (SSD), Retina-Net, and Pyramid Net.

[0189] The skeleton extraction model (1205) may be a model for extracting the skeleton of a user within an image included in the shooting information. For skeleton extraction, a skeleton extraction algorithm based on a convolutional neural network may be used on a 2D or 3D image to extract 2D / 3D skeletal information. A feature detection model based on a convolutional neural network may be applied as the skeletal information extraction model, and for example, algorithms such as MoveNet, PoseNet, OpenPose, MediaPipe, AlphaPose, Nuitrack, and Kinecct SDK may be utilized. Skeletal information and values ​​calculated based on the skeletal information do not change in a short period of time, and thus can be effectively utilized for identifying a tracking target.

[0190] The action recognition model (1206) may be a model for analyzing what kind of action a user takes from joint data generated through the skeleton extraction model (1205). The action recognition model may recognize actions using 2D / 3D skeletal information extracted from 2D / 3D images by the skeleton extraction model. In this case, a rule-based method that classifies actions using a threshold-based method may be exemplified. As another method, a machine learning-based method may be used, which utilizes a recurrent neural network (RNN) that is suitable for action recognition based on time-series skeletal information as a powerful neural network for time-series data processing, or a GCN (Graph Convolutional Network) that utilizes the graph structure of skeletal data.

[0191] That is, the situation information generation unit (1210) can generate situation information (1211) about the user's situation by integrating time information (1208) about the space class, object class, user action, and current time derived from various models based on the shooting information (1202) and user location information (1207) about the user's estimated location.

[0192] Here, the user location information (1207) may be information about the user's location on an indoor map drawn by the care robot. Here, the location information may mean a two-dimensional coordinate value based on a specific point on the indoor map.

[0193] As described above, the context information generation unit (260) can input the collected information into a rule-based or machine learning-based model to generate context information. For example, if a sound exceeding a certain decibel level is detected in the middle of the night, the care robot can move to the corresponding location and generate context information. For another example, if the care robot discovers a person collapsed during security, the context information generation unit (260) can generate related context information. For another example, if the temperature in a room, as acquired through IoT collected information, is abnormally high, the context information generation unit (260) can move to the corresponding location and generate context information.

[0194] Returning to Figure 2, the service decision unit (270) can determine the care service to be provided to the user based on the output of the language model. The service decision unit (270) can extract text related to the care service available from the output of the language model and determine the care service corresponding to the extracted text. Thereafter, the care robot (100) can provide the determined care service to the user.

[0195] The care robot (100) may provide the user with the care service determined by the service decision unit (270) and then upload the resulting information to a server. The resulting information may include contextual information about the user recognized by the care robot (100), information about the care service determined in response to the context, and user feedback information about the care service. The resulting information may be uploaded to a server or cloud connected to the care robot (100).

[0196] The above feedback information can be received from the user via chat, voice conversation, or screen touch after the care robot (100) provides care services, such as service suitability or satisfaction. This feedback information can then be used as learning material for the language model to output more appropriate responses.

[0197] Additionally, the care robot (100) can separately designate information to be uploaded to the server during the process of determining the care service to be provided to the user. Important information from the responses output by the language model based on the image can be separately extracted, for example, to determine information to be uploaded to the server. As shown in Fig. 11, the second guide information (1160) designates the information to be separately uploaded to the server as "server upload information." Thereafter, the care robot (100) can extract the information marked as "server upload information" from the responses output by the language model, store it separately from the care service decision, and upload it to the server.

[0198] The method for providing care services performed by the above-described care robot may also be implemented in the form of a computer program stored on a computer-readable recording medium executed by a computer or a recording medium containing computer-executable commands. Furthermore, the method may also be implemented in the form of a computer program stored on a computer-readable recording medium executed by a computer.

[0199] A computer-readable recording medium may be any available medium that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media. Furthermore, a computer-readable recording medium may include computer storage media. Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.

[0200] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will readily appreciate that the present invention can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single entity may be implemented in a distributed manner, and similarly, components described as distributed may be implemented in a combined manner.

[0201] The scope of the present invention is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present invention.

Claims

1. A robot collection information generation unit that generates robot collection information related to a user through at least one sensing device installed in a care robot; A language model input unit that generates a prompt based on the robot collection information and inputs the generated prompt into a language model; and A care robot comprising a service decision unit that determines a care service to be provided to the user based on the output of the language model.

2. In paragraph 1, The above robot collection information generation unit a sensor unit including a camera for photographing the user; and A care robot, comprising a collection information analysis unit that analyzes sensing information generated from the sensor unit to generate robot collection information.

3. In paragraph 2, The above robot collection information generation unit A care robot further comprising a camera control unit that controls at least one operation among vertical rotation of the camera, horizontal rotation of the camera, and height change of the camera.

4. In paragraph 2, The above robot collection information generation unit A care robot including a robot moving unit that moves the care robot to change the shooting position of the care robot.

5. In paragraph 1, Further comprising an IoT (Internet on Thing) collection information generation unit that generates IoT collection information based on information collected from an IoT device connected to the care robot; A care robot, wherein the language model input unit generates the prompt based on the robot collection information and the IoT collection information.

6. In paragraph 1, Further comprising a cloud collection information generation unit that generates cloud collection information from a cloud in which at least one of profile information about the user and pattern information about the user's daily pattern is stored; A care robot, wherein the language model input unit generates the prompt based on the robot collection information and the cloud collection information.

7. In paragraph 6, The above cloud collection information generation unit includes a pattern analysis model that outputs expected pattern information on activities that the user is expected to perform at the current time when activity information including the user's activity history by time zone and information on the pattern generation period are input.

8. In paragraph 1, Further comprising: a provisional service information generation unit that generates provisional service information including a list of care services that can be provided to the user by the care robot; The above language model input section A care robot that generates the prompt based on the robot collection information and the service information that can be provided.

9. In paragraph 8, The above information on available services is Further including information about the priority among care services that may be provided to the user; A care robot wherein the service decision unit determines the care service to be provided to the user based on the output of the language model and the priority.

10. In paragraph 1, The language model input unit includes an input point determination unit that generates the prompt and determines the input point in time for inputting the generated prompt into the language model; A care robot, wherein the language model input unit inputs the prompt to the language model in response to the determined input point in time.

11. In paragraph 10, The above input point determination part A care robot that determines the input point based on pattern information about the user's daily pattern.

12. In paragraph 10, Further comprising a context information generation unit that generates context information about the user's surroundings based on the robot-collected information; A care robot, wherein the input point determination unit determines the input point based on the situation information.

Citation Information

Patent Citations

  • Dialogue system using knowledge base and language model for automotive systems and application

    JP2024043564A

  • Module for moral decision making, robot comprising the same, and method for moral decision making

    KR1020180058563A

  • W / O emulsion cosmetic composition with improved long-lasting and feeling of moisture

    KR1020190098733A

  • Tactile Sensor

    KR1020230152941A

  • KR20200087337A