system

A system using generative AI to generate personalized video and audio interactions and determine urgency addresses staff shortages in medical facilities, improving response efficiency and patient satisfaction.

JP2026070250APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

In medical facilities, chronic staff shortages due to population decline and aging lead to difficulties in quickly and individually responding to patient calls, resulting in increased burden on medical and care workers and decreased patient satisfaction.

Method used

A system that utilizes a receiving means to detect patient calls, generates personalized video and audio interactions using generative AI, determines urgency based on patient information, and dispatches staff or equipment as needed to ensure timely and appropriate responses.

Benefits of technology

Reduces the burden on healthcare professionals and improves patient satisfaction by enabling rapid, individualized responses to patient needs and emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070250000001_ABST
    Figure 2026070250000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A receiving means for receiving call signals transmitted by a user, A generation means that provides generated video and audio to a user terminal based on the received signal, A determination means for determining the priority of urgency based on the generated video and audio and user information, If the aforementioned determination means determines that it is an emergency, the instruction means dispatches personnel or equipment promptly, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, in the medical field, chronic staff shortages due to population decline and aging are an issue. Especially in elderly care facilities, it is difficult to respond quickly and individually, increasing the burden on medical and care workers. Also, the prioritization of responses to patient calls may not be appropriately carried out, leading to a decrease in satisfaction. As a result, there is a risk that patients cannot receive sufficient medical care. Therefore, an object of the present invention is to respond quickly and appropriately to calls from patients, reduce the burden on medical staff, and improve patient satisfaction.

Means for Solving the Problems

[0005] This invention provides a receiving means for receiving call signals transmitted by users, thereby enabling rapid recognition of patient calls. It also includes a generating means for providing video and audio generated based on the received signals to the user terminal, facilitating real-time interactive communication with patients. Furthermore, a determination means for prioritizing urgency based on the generated video and audio and patient information allows for more efficient resource allocation. If an emergency is determined, the system includes an instruction means for quickly dispatching staff or equipment, ensuring immediate and appropriate action. This system, incorporating these functions, aims to reduce the burden on healthcare professionals and improve patient satisfaction.

[0006] "Receiving means" refers to a device or method that has the function of detecting call signals transmitted by a user and incorporating their content into the system.

[0007] "Generation means" refers to a device or method that has the function of generating video and audio similar to that of an actual nurse on a computer based on a received call signal, and displaying and playing it on a user terminal.

[0008] "Determination means" refers to a device or method that has the function of calculating and evaluating the urgency and priority of response to a patient's request based on the generated video and audio, as well as patient information.

[0009] "Instruction means" refers to a device or method that has the function of forming and executing orders to dispatch personnel or medical equipment as needed, based on the priority determined by the determination means. [Brief explanation of the drawing]

[0010] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0011] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0012] First, let's explain the terminology used in the following explanation.

[0013] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0014] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0015] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0016] In the following embodiments, the labeled communication I / F (Interface) is an interface that includes a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0017] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0018] [First Embodiment]

[0019] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0020] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0021] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0022] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0023] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0024] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0025] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0026] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0027] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0028] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0029] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0030] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0031] This invention relates to a system using generative AI for receiving and automatically responding to nurse calls initiated by users in medical facilities. This system can reduce the burden on medical staff and improve user satisfaction.

[0032] First, when a user presses a nurse call button installed in the patient's room, the terminal receives the signal. The terminal sends this information to the server and instructs it to begin responding. The server processes the received data and uses generative AI to generate video and audio of a simulated nurse. This generation is based on past history and user preferences, providing personalized interaction.

[0033] The generated video and audio are sent back to the terminal and displayed on the user's screen. The user can converse with the generated nurse through an interactive interface. This interaction collects information about the user's requests and physical condition.

[0034] The server analyzes the user's conversation and data from connected sensor devices to determine the urgency and priority of the response. For example, if a user complains of chest pain, it is immediately determined to be an "emergency," and instructions are generated to request immediate assistance from medical personnel.

[0035] Instructions are fed back to the user via the terminal. In some cases, physical personnel may be dispatched or medical equipment may be provided. This ensures that the user receives appropriate support without any omissions or delays.

[0036] Furthermore, data collected during everyday conversations can be used to improve future services. For example, the results of users' regular health checks and the frequency with which specific caregiving activities are needed can be analyzed, enabling the provision of more personalized services. The introduction of this system is expected to improve efficiency in medical settings and enhance the quality of medical care provided.

[0037] The following describes the processing flow.

[0038] Step 1:

[0039] The user presses the nurse call button. The terminal detects the button press and receives the signal.

[0040] Step 2:

[0041] The terminal sends the received call signal to the server and generates a request to initiate a response.

[0042] Step 3:

[0043] The server receives the request and retrieves the user's ID, location information, and past history data from the database.

[0044] Step 4:

[0045] Based on the acquired data, the server uses a generation AI to create video and audio of a nurse character.

[0046] Step 5:

[0047] The server transmits the generated video and audio to the terminal, presenting the user with a virtual nurse.

[0048] Step 6:

[0049] The terminal displays the generated video and audio of the nurse to the user and initiates an interactive conversation.

[0050] Step 7:

[0051] The user converses with a nurse on the screen. The terminal continuously transmits the user's input and voice data to the server.

[0052] Step 8:

[0053] The server analyzes the conversation content and sensor data to determine the urgency and priority of the response.

[0054] Step 9:

[0055] The server generates necessary responses and instructions based on its judgment, and issues instructions regarding the dispatch of personnel and equipment.

[0056] Step 10:

[0057] The terminal provides feedback to the user based on instructions from the server, displaying them on the screen and via sound.

[0058] Step 11:

[0059] If necessary, staff will be dispatched to the user's location. The terminal reports the completion of the process to the server.

[0060] Step 12:

[0061] The server records the content of conversations and updates the database to help improve subsequent responses.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] Responding quickly and accurately to inquiries and requests from users in medical facilities is crucial for significantly reducing the burden on healthcare professionals and improving user satisfaction. However, current systems struggle to efficiently manage the enormous number of calls, and the time and resources available for providing appropriate individual responses are limited. This can lead to delays in responding to urgent cases, causing risks and dissatisfaction among users.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes a device for receiving communication signals transmitted by a user, a generation device for providing visual and auditory information generated using a generation AI model to a terminal device, and an analysis device for determining the urgency of the situation. This makes it possible to provide quickly generated individual responses to user inquiries based on multiple information resources, reducing the workload of medical professionals and improving the accuracy and speed of services to users.

[0067] A "communication signal" is electrical data that a user transmits from their terminal, indicating a call or request.

[0068] "Device" refers to an instrument or machine used to perform a specific function, and in this invention, it refers to a mechanism that receives communication signals and manipulates generated AI models.

[0069] A "generative AI model" is a form of artificial intelligence that constructs visual and auditory information based on past data and personal preferences.

[0070] "Visual and auditory information" refers to images displayed on a screen and sounds emitted from speakers, created by a generative AI model.

[0071] A "terminal device" is a device that serves as an interface with the user, displaying information and playing audio.

[0072] An "analysis device" is a device that has the function of analyzing and determining the urgency level using user information and data collected from sensors.

[0073] A "commanding device" refers to a mechanism that issues commands to workers or equipment to carry out specific actions based on analysis results.

[0074] An "information resource access device" is a device that has the function of accessing a database to generate an adapted response by referring to the user's past history and individual preferences.

[0075] This invention relates to a system for responding quickly and individually to inquiries from users in medical facilities. The following specifically describes embodiments for carrying out this invention.

[0076] This system uses terminals installed in patient rooms to receive communication signals transmitted by users and utilizes a server to process that data. The terminal functions as an interface for user operation, and when a nurse call button is pressed, it sends a signal to the server.

[0077] The server analyzes the received signal to determine the user's room number and the content of the call. Next, it uses a generative AI model to generate personalized visual and auditory information for each user. This information incorporates data reflecting past medical history and the user's specific preferences. Specifically, the server utilizes a cloud-based AI platform for data analysis and video / audio generation.

[0078] For example, if a user presses the nurse call button and requests water, the device sends the voice message to the server. The server, using a generative AI model, generates video and audio responses saying, "We'll bring you water right away," and sends them back to the device. The user can then see this interaction through the device's screen and speaker.

[0079] A concrete example of a prompt message would be, "The patient has pressed the nurse call button. Please generate a response." This allows the server to quickly take an appropriate action via the AI ​​model.

[0080] Furthermore, the server collects data from connected sensor devices to monitor the user's situation and health status. This data helps determine the urgency of a situation and the appropriate response, enabling efficient medical support.

[0081] This system is expected to significantly improve efficiency in healthcare settings and user satisfaction by enabling users to receive prompt and appropriate support and reducing the burden on healthcare professionals.

[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0083] Step 1:

[0084] The user presses the nurse call button on the terminal. This action causes the terminal to generate a communication signal, which is then sent to the server as input data. Specifically, pressing the button causes the terminal to construct a data packet containing information such as the room number and the reason for the call. The output is the transmission of a signal to the server.

[0085] Step 2:

[0086] The server analyzes the signals received from the terminal. The input is a data packet containing the user's room number and requests. The server transforms this data into a prompt for the AI ​​model and inputs it into the generative AI model. The prompt might be, for example, "The patient says they want water. Generate an appropriate response." The output of this process is a template of the visual and auditory information that the generative AI model generates.

[0087] Step 3:

[0088] The server generates visual and auditory information using a generative AI model. The input is a generation template based on the prompts generated in step 2. The AI ​​considers past historical data and user preferences when generating video and audio based on this template. Specifically, the AI ​​model adjusts the tone of the voice and the facial expressions in the video to suit the user. The output is visual and auditory media files sent to the terminal.

[0089] Step 4:

[0090] The terminal receives visual and auditory information transmitted from the server and presents it to the user. The input is a generated media file, and the output is the provision of visual and auditory information to the user. Specifically, the terminal displays images on the screen and plays audio through the speaker, providing information to the user in an intuitive and easy-to-understand manner.

[0091] Step 5:

[0092] The user interacts with a nurse generated through the device. The input for this step is the user's voice and further requests. The output is the user's speech being sent to the server as text data. Specifically, the device's microphone captures the user's voice and converts it to text using speech recognition technology.

[0093] Step 6:

[0094] The server analyzes interaction data from the user. Inputs include text data sent from the terminal and additional information from sensor devices. The server analyzes this data to determine urgency. The results are expressed as work instructions and prioritized response protocols. Outputs are the next instructions and alerts sent to the terminal.

[0095] Step 7:

[0096] The terminal notifies the user of instructions from the server and adjusts physical responses as needed. Inputs are work instructions and alert information from the server. Outputs are notifications of appropriate actions to the user and medical staff. Specifically, the terminal displays emergency messages on the screen and issues alarms to medical staff if necessary.

[0097] (Application Example 1)

[0098] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0099] In industrial settings, there is a need to respond quickly and accurately to abnormal signals and worker requests while maintaining production efficiency and safety. However, conventional systems have struggled to make real-time decisions and respond to these requests. This has resulted in challenges such as disruptions to on-site work, production delays, and even safety risks.

[0100] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0101] In this invention, the server includes a device for receiving signals transmitted by a user, a generation device for providing generated visual media and sound to a terminal based on the received signals, a determination device for determining priority based on the generated visual media and sound and user information, an instruction device for quickly dispatching personnel or equipment when the determination device determines the situation to be important, and a device for receiving abnormal signals from equipment and work stations at an industrial site and responding as a virtual maintenance person using a generated AI model. This enables immediate judgment and response to abnormal situations, thereby improving the efficiency and safety of production at industrial sites.

[0102] A "signal" is a means of transmitting data or requests originating from a user or device.

[0103] A "device" is a machine or electronic device that has a specific function and is used to receive, generate, judge, or direct signals.

[0104] A "generative AI model" is a mathematical or algorithmic model that uses artificial intelligence technology to analyze signals and generate appropriate responses.

[0105] "Visual media" refers to digital content presented in the form of images or videos that contain visual information.

[0106] "Acoustics" refers to information that includes audio data and is used as sound for communication.

[0107] A "terminal" is an electronic device or part of a computer that a user interacts with directly.

[0108] A "decision-making device" is a device or software used to analyze the content of received data and determine its importance and the priority of actions to be taken.

[0109] A "commanding device" is a device that generates or transmits commands to initiate necessary actions or processes based on a judgment.

[0110] An "abnormal signal" is a signal that includes warnings or error messages indicating a condition that is different from the normal state.

[0111] A "virtual maintenance staff" is a maintenance staff system artificially created by a generative AI model, capable of handling problems without human intervention.

[0112] The system that realizes this invention aims to efficiently handle abnormalities in industrial settings. First, a terminal receives a signal transmitted by the user. The terminal sends this signal to a server and starts processing the response.

[0113] The server is equipped with a generative AI model that analyzes abnormal signals and user requests. The generated video and audio are sent back to the terminal as visual and auditory media. This allows users to interact with a virtual maintenance representative through the screen and audio, and quickly receive the necessary information and support.

[0114] The server analyzes the received data, and a determination device is activated to assess the urgency. If it is deemed critical, an instruction device is activated, issuing orders to dispatch personnel and equipment as needed. As a result, an appropriate response is quickly taken at the scene.

[0115] For example, if a worker reports that "the machine is making an unusual noise," the AI ​​model will predict the possibility of an anomaly and suggest appropriate countermeasures. This process is initiated by entering a prompt message into the user's terminal. An example of a prompt message is, "The machine is making an unusual noise. Please tell me how to deal with it."

[0116] The hardware used includes smartphones and head-mounted displays, while the software utilizes TENSORFLOW® and OpenCV for signal processing and data analysis. This allows for the efficient execution of the process from signal reception to processing.

[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0118] Step 1:

[0119] When a user sends a signal, the terminal receives that signal. The input is the call signal from the user, and the output is the digital data of the signal. This allows the terminal to recognize an anomaly or a request.

[0120] Step 2:

[0121] The terminal forwards the received signal to the server. The input is digital data of the signal, and the output is prompt data directed to the server. This data forms the basis for the next processing step.

[0122] Step 3:

[0123] The server inputs the received prompt data into the generating AI model. The input is prompt data, and the output is the analysis results. The generating AI model determines the nature and priority of the anomalies.

[0124] Step 4:

[0125] The server generates visual and auditory media for a virtual maintenance worker based on the analysis results of the generated AI model. The input is the analysis results from the AI ​​model, and the output is the generated visual and auditory media. This is then used as feedback to the user.

[0126] Step 5:

[0127] The generated visual and auditory media are sent to the terminal and presented to the user. The input is the generated visual and auditory media, and the output is video and audio for the user. This allows the user to interact with the system.

[0128] Step 6:

[0129] If the user provides further interaction, the server analyzes the data based on that interaction and generates the necessary instructions. The input is the additional interaction with the user, and the output is further instruction data. If necessary, instructions are given to quickly dispatch personnel and equipment to the site.

[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0131] This invention relates to a generative AI system that combines an emotion engine to automatically and individually respond to nurse calls from users in medical facilities. This system can further improve user satisfaction by recognizing the emotional state of users and providing personalized responses accordingly.

[0132] When a user presses the nurse call button, the terminal receives the signal and sends the information to the server. Based on this information, the server uses generative AI to generate video and audio of a virtual nurse. Because this generation is based on the user's past history and preferences, personalized interaction is possible.

[0133] The server then analyzes the user's emotional state through an emotion engine. This engine uses voice analysis and facial recognition technology to determine whether the user is stressed, at ease, or confused. The analyzed emotional data is then sent to the server and used to adjust the conversation.

[0134] The generated video and audio are sent to the device, initiating interaction with the user. For example, if the user expresses feelings such as "I'm a little anxious, but I don't feel any major problems," the generating AI will respond calmly with something like, "I'll help you feel more at ease. Please let me know if there's anything specific I can do to help."

[0135] The terminal transmits the conversation with the user to the server in real time, and the server determines the urgency level based on the conversation content. Along with the emotional state, appropriate instructions are generated. For example, if the emotional analysis indicates a low level of urgency but the user is feeling anxious, additional counseling via voice call can be provided.

[0136] Another important function of this system is to improve the quality of care and medical treatment by tracking changes in users' emotions and conducting long-term emotional analysis. This allows for the provision of care tailored to each user and enhances their sense of security. In this way, the system is a powerful tool for reducing the burden on healthcare professionals and implementing comprehensive medical care.

[0137] The following describes the processing flow.

[0138] Step 1:

[0139] The user presses the nurse call button. The terminal receives this signal and prepares a request to the server.

[0140] Step 2:

[0141] The terminal sends call information to the server, initiating a processing request along with the user's ID and location information.

[0142] Step 3:

[0143] Based on the received data, the server generates video and audio of a virtual nurse tailored to the user's history and preferences.

[0144] Step 4:

[0145] The server activates the emotion engine, analyzes emotional data from the user's voice and video, and determines the user's emotional state.

[0146] Step 5:

[0147] The server sends the generated screen and emotion data to the terminal, initiating interaction with the user.

[0148] Step 6:

[0149] The terminal uses the received video and audio to display the conversation between the user and the virtual nurse, providing interactive support.

[0150] Step 7:

[0151] As the user converses with the virtual nurse, they express a variety of reactions and emotions. The device transmits the conversation data to the server in real time.

[0152] Step 8:

[0153] The server analyzes the conversation content and emotional data, determines the urgency level, and decides on a flexible response based on the emotions, if necessary.

[0154] Step 9:

[0155] Based on the assessment results, the server generates instruction data to dispatch personnel or provide additional remote support if necessary.

[0156] Step 10:

[0157] Based on instructions from the server, the terminal continuously provides feedback to the user and offers additional support to provide a sense of security.

[0158] Step 11:

[0159] Upon receiving a user response, the terminal reports the response result and sentiment data back to the server, completing the entire process.

[0160] Step 12:

[0161] The server records all conversation and sentiment data, and updates the database to help with future improvements and long-term analysis.

[0162] (Example 2)

[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0164] In healthcare facilities, there is a need for a system that can appropriately respond to calls from users, understand their emotional state, and provide individualized interactions to increase user satisfaction, while also enabling rapid response to emergencies. Conventional systems have the problem of not being able to provide satisfactory service to users because they are not capable of analyzing emotional states or generating individualized responses.

[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0166] In this invention, the server includes a receiving means, a generating means, an analysis means, an adjustment means, and an instruction means. This enables rapid analysis of the user's emotional state and the provision of personalized services. Specifically, it becomes possible to generate a virtual supporter that meets the needs of each user by utilizing a generation AI model and to provide interaction based on emotional analysis.

[0167] A "receiving means" is a function that detects signals transmitted by the user and incorporates them into the system.

[0168] The "generation means" is a function that generates information about a virtual supporter based on the signal information received by the receiving means and supplies this information to the information processing device.

[0169] The "analysis means" is a function that determines the user's emotional state by analyzing data such as the user's voice and video.

[0170] The "adjustment means" is a function that plays the role of generating individually adjusted information based on the results of the analysis means.

[0171] The "determination means" is a function that determines the urgency of the current situation based on generated information and user history information.

[0172] An "instruction mechanism" is a function that issues instructions to promptly take necessary measures based on the results of a determination mechanism.

[0173] A "recording means" is a function that sequentially collects user conversation data and generates information to track changes in emotional state based on this data.

[0174] The "response generation means" is a function that enables the generation of personalized responses based on the user's history data and emotional data.

[0175] This invention relates to a system for responding quickly and appropriately to user calls in medical facilities. The system involves multiple components, including a server, terminals, and users. The server utilizes a generative AI model to generate video and audio of a virtual supporter tailored to each user's needs. The generative AI model used is vendor-independent but implements state-of-the-art machine learning algorithms.

[0176] The terminal is responsible for receiving call signals from the user and transmitting them to the server. The server uses the received signals and past user data to perform voice analysis using an emotion engine and facial recognition technology. This analysis determines the user's emotional state, and personalized interactions are set accordingly. The server can also use general natural language processing techniques for voice analysis and existing image analysis techniques for facial recognition.

[0177] For example, Microsoft® Azure® facial recognition APIs and Google® Cloud natural language processing APIs can be used. The generated video and audio are sent to the device, and a conversation with the user begins. During this interaction, the device continuously sends the conversation content to the server. The server determines the urgency based on the conversation content and emotional state, and generates appropriate instructions.

[0178] For example, if a user feels anxious, the system uses AI to provide a response such as, "We'll help you feel more at ease. Please let us know if there's anything specific we can do to help." Furthermore, this system tracks the user's emotional changes over the long term, providing data that can be used to improve medical or nursing care services.

[0179] An example of a prompt message would be: "Explain the process for generating an appropriate virtual caregiver based on past history information and sentiment analysis when a patient presses the nurse call button, and for optimizing the interaction with the patient." This allows the system to provide user-optimized medical care and reduce the burden on healthcare professionals.

[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0181] Step 1:

[0182] When a user presses the nurse call button, the terminal receives the signal as input data. This signal is in digital format and contains the user's ID and location information. The terminal forwards this input to the server, which then sends a trigger to prepare for response.

[0183] Step 2:

[0184] The server generates video and audio of a virtual supporter based on signals received from the terminal, utilizing a generative AI model. Specifically, it searches a database for past history information and preferences identified by the user's ID and supplies this as input to the generative AI model. The generative AI model customizes individual interactions based on prompt messages and creates video and audio output.

[0185] Step 3:

[0186] The server sends the generated virtual supporter's video and audio data to the terminal. The terminal receives this data and displays the video and plays the audio for the user. The user experiences this interaction in real time.

[0187] Step 4:

[0188] During user interaction, the device sequentially collects the user's voice and facial expressions and sends them to the server as input data. The server analyzes this data using an emotion engine to determine the user's emotional state. At this time, stress levels and feelings of security are quantified using voice recognition and facial recognition technologies.

[0189] Step 5:

[0190] The server adjusts conversations and responses based on the analyzed emotional data. Specifically, if the user is feeling anxious, it uses a generative AI model to generate additional support messages and delivers them to the device.

[0191] Step 6:

[0192] The terminal sequentially sends the content of the interaction with the user to the server, which then processes the conversation to determine its urgency. The server integrates sentiment data with the conversation content, applies urgency determination logic, and determines the priority.

[0193] Step 7:

[0194] If the situation is deemed highly urgent, the server will instruct the terminal to take immediate action based on rules defined within the system. This includes instructing the dispatch of medical staff as needed and preparing relevant equipment.

[0195] (Application Example 2)

[0196] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0197] Conventional security systems have the challenge of being unable to adequately consider the emotional state of users and respond appropriately to diverse situations. Furthermore, they struggle to provide users with customized, real-time feedback for ensuring their safety. As a result, they are unable to implement optimal security measures and adequately alleviate user anxiety.

[0198] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0199] In this invention, the server includes means for receiving signals transmitted by a user, generation means for providing generated visual and auditory information to an information terminal based on the received signals, and analysis means for analyzing the user's emotional state. This makes it possible to analyze the user's emotions in real time and provide optimal visual and auditory feedback tailored to that situation.

[0200] "Receiving means" refers to devices or systems for receiving signals transmitted by users.

[0201] A "generation means" is a mechanism for providing information terminals with generated visual and auditory content based on received signals.

[0202] A "determination means" is a mechanism that determines priority based on generated visual and auditory information and user information.

[0203] "Instruction means" refers to a means of quickly dispatching personnel or equipment when the determination means determines that the matter has a high priority.

[0204] "Analysis methods" refer to processes and techniques for analyzing the emotional state of users.

[0205] "Adjustment means" refers to means that have the function of adjusting the response content based on the analysis results obtained by the analysis means.

[0206] A "recording system" is a mechanism that sequentially collects dialogue data with users and generates information for use in new responses.

[0207] A "database access means" is a means of obtaining the data necessary to provide a customized response through an information terminal.

[0208] This invention relates to a personal security system that utilizes emotion analysis and was developed with the aim of enhancing user safety. In this system, the user's device, specifically smart glasses, is used as the main interface.

[0209] The server first receives the user's visual and audio data transmitted from the smart glasses. This data is processed to analyze facial expressions and voice tone. Specifically, facial recognition technology is used for facial expressions, and voice analysis technology is used for voice. The hardware used includes the camera and microphone within the smart glasses. This makes it possible to evaluate the user's emotional state in real time.

[0210] Subsequently, the server uses a generative AI model based on the sentiment analysis results to generate a response tailored to the user. At this stage, specific advice and warnings are generated to alleviate the user's anxiety and fear. For example, if the user expresses the sentiment, "I've been feeling uneasy on the street lately," the generative AI will provide advice such as, "Pay attention to your surroundings and be wary of suspicious activity."

[0211] The generated advice and warnings are delivered to the user as audio and visual feedback through smart glasses, and the user can receive them in real time. A key feature of this system is that the responses are personalized to the user and optimized according to the user's emotional state.

[0212] An example of a prompt would be: "Sentiment analysis result: User's sentiment data Voice analysis result: User's voice text Provide the best advice for this user." By inputting this prompt into the generating AI model, a customized response tailored to the user can be obtained.

[0213] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0214] Step 1:

[0215] The server receives visual data (images) and audio data transmitted from the terminal (smart glasses). This input data serves as material for analyzing the user's emotional state. Upon receiving this data, the server sends the visual and audio data to separate analysis processes.

[0216] Step 2:

[0217] The server uses visual data to perform facial recognition technology for expression analysis. Specifically, it extracts facial features from the image and determines the user's emotional state. The output obtained here is an emotion label such as "reassured" or "anxious."

[0218] Step 3:

[0219] Simultaneously, the server applies speech analysis technology to the audio data, evaluating the emotional state by analyzing tone and pitch. This results in the output of speech-based emotion labels such as "calm" or "tense."

[0220] Step 4:

[0221] The server integrates the emotional labels obtained from the visual and auditory data to assess the user's overall emotional state. This integrated emotional data is then used in the next step.

[0222] Step 5:

[0223] The server uses a generative AI model based on integrated sentiment data to create prompt messages. Specifically, it embeds the results of sentiment analysis into the prompt messages and provides them as input to generate optimal feedback for the user.

[0224] Step 6:

[0225] The generative AI model receives prompt messages generated by the server as input and generates advice and warnings to provide to the user. The output of this step is a personalized message tailored to the specific situation.

[0226] Step 7:

[0227] The server sends generated advice and warnings to the device, presenting them to the user as visual and audio feedback. The device plays this back and provides it to the user in real time, thereby improving the user's safety awareness.

[0228] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0229] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0230] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0231] [Second Embodiment]

[0232] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0233] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0234] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0235] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0236] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0237] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0238] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0239] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0240] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0241] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0242] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0243] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0244] This invention relates to a system using generative AI for receiving and automatically responding to nurse calls initiated by users in medical facilities. This system can reduce the burden on medical staff and improve user satisfaction.

[0245] First, when a user presses a nurse call button installed in the patient's room, the terminal receives the signal. The terminal sends this information to the server and instructs it to begin responding. The server processes the received data and uses generative AI to generate video and audio of a simulated nurse. This generation is based on past history and user preferences, providing personalized interaction.

[0246] The generated video and audio are sent back to the terminal and displayed on the user's screen. The user can converse with the generated nurse through an interactive interface. This interaction collects information about the user's requests and physical condition.

[0247] The server analyzes the user's conversation and data from connected sensor devices to determine the urgency and priority of the response. For example, if a user complains of chest pain, it is immediately determined to be an "emergency," and instructions are generated to request immediate assistance from medical personnel.

[0248] Instructions are fed back to the user via the terminal. In some cases, physical personnel may be dispatched or medical equipment may be provided. This ensures that the user receives appropriate support without any omissions or delays.

[0249] Furthermore, data collected during everyday conversations can be used to improve future services. For example, the results of users' regular health checks and the frequency with which specific caregiving activities are needed can be analyzed, enabling the provision of more personalized services. The introduction of this system is expected to improve efficiency in medical settings and enhance the quality of medical care provided.

[0250] The following describes the processing flow.

[0251] Step 1:

[0252] The user presses the nurse call button. The terminal detects the button press and receives the signal.

[0253] Step 2:

[0254] The terminal sends the received call signal to the server and generates a request to initiate a response.

[0255] Step 3:

[0256] The server receives the request and retrieves the user's ID, location information, and past history data from the database.

[0257] Step 4:

[0258] Based on the acquired data, the server uses a generation AI to create video and audio of a nurse character.

[0259] Step 5:

[0260] The server transmits the generated video and audio to the terminal, presenting the user with a virtual nurse.

[0261] Step 6:

[0262] The terminal displays the generated video and audio of the nurse to the user and initiates an interactive conversation.

[0263] Step 7:

[0264] The user converses with a nurse on the screen. The terminal continuously transmits the user's input and voice data to the server.

[0265] Step 8:

[0266] The server analyzes the conversation content and sensor data to determine the urgency and priority of the response.

[0267] Step 9:

[0268] The server generates necessary responses and instructions based on its judgment, and issues instructions regarding the dispatch of personnel and equipment.

[0269] Step 10:

[0270] The terminal provides feedback to the user based on instructions from the server, displaying them on the screen and via sound.

[0271] Step 11:

[0272] If necessary, staff will be dispatched to the user's location. The terminal reports the completion of the process to the server.

[0273] Step 12:

[0274] The server records the content of conversations and updates the database to help improve subsequent responses.

[0275] (Example 1)

[0276] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0277] Responding quickly and accurately to inquiries and requests from users in medical facilities is crucial for significantly reducing the burden on healthcare professionals and improving user satisfaction. However, current systems struggle to efficiently manage the enormous number of calls, and the time and resources available for providing appropriate individual responses are limited. This can lead to delays in responding to urgent cases, causing risks and dissatisfaction among users.

[0278] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0279] In this invention, the server includes a device for receiving communication signals transmitted by a user, a generation device for providing visual and auditory information generated using a generation AI model to a terminal device, and an analysis device for determining the urgency of the situation. This makes it possible to provide quickly generated individual responses to user inquiries based on multiple information resources, reducing the workload of medical professionals and improving the accuracy and speed of services to users.

[0280] A "communication signal" is electrical data that a user transmits from their terminal, indicating a call or request.

[0281] "Device" refers to an instrument or machine for performing a specific function, and in this invention, it refers to a mechanism for receiving communication signals and operating a generative AI model.

[0282] "Generative AI model" is a form of artificial intelligence that constructs visual and auditory information based on past data and personal preferences.

[0283] "Visual and auditory information" refers to the video displayed on the screen and the sound emitted from the speaker, which are created by the generative AI model.

[0284] "Terminal device" is a device that serves as an interface with the user and is capable of displaying information and playing sound.

[0285] "Analysis device" is a device that has the function of analyzing and judging the urgency using the information of the user and the data collected from sensors.

[0286] "Instruction device" refers to a mechanism that issues commands to an operator or a device to execute specific actions based on the analysis results.

[0287] "Information resource access device" is a device that has the function of accessing a database for generating an adapted response by referring to the user's past history and individual preferences.

[0288] The present invention relates to a system for quickly and individually responding to inquiries from users in a medical facility. Hereinafter, the embodiments for implementing this invention will be specifically shown.

[0289] This system receives the communication signals transmitted by the user using a terminal installed in the hospital room, and utilizes a server to process the data. The terminal functions as an interface for the user to operate, and when the nurse call button is pressed, the signal is transmitted to the server.

[0290] The server analyzes the received signal to determine the user's room number and the content of the call. Next, it uses a generative AI model to generate personalized visual and auditory information for each user. This information incorporates data reflecting past medical history and the user's specific preferences. Specifically, the server utilizes a cloud-based AI platform for data analysis and video / audio generation.

[0291] For example, if a user presses the nurse call button and requests water, the device sends the voice message to the server. The server, using a generative AI model, generates video and audio responses saying, "We'll bring you water right away," and sends them back to the device. The user can then see this interaction through the device's screen and speaker.

[0292] A concrete example of a prompt message would be, "The patient has pressed the nurse call button. Please generate a response." This allows the server to quickly take an appropriate action via the AI ​​model.

[0293] Furthermore, the server collects data from connected sensor devices to monitor the user's situation and health status. This data helps determine the urgency of a situation and the appropriate response, enabling efficient medical support.

[0294] This system is expected to significantly improve efficiency in healthcare settings and user satisfaction by enabling users to receive prompt and appropriate support and reducing the burden on healthcare professionals.

[0295] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0296] Step 1:

[0297] The user presses the nurse call button on the terminal. This action causes the terminal to generate a communication signal, which is then sent to the server as input data. Specifically, pressing the button causes the terminal to construct a data packet containing information such as the room number and the reason for the call. The output is the transmission of a signal to the server.

[0298] Step 2:

[0299] The server analyzes the signals received from the terminal. The input is a data packet containing the user's room number and requests. The server transforms this data into a prompt for the AI ​​model and inputs it into the generative AI model. The prompt might be, for example, "The patient says they want water. Generate an appropriate response." The output of this process is a template of the visual and auditory information that the generative AI model generates.

[0300] Step 3:

[0301] The server generates visual and auditory information using a generative AI model. The input is a generation template based on the prompts generated in step 2. The AI ​​considers past historical data and user preferences when generating video and audio based on this template. Specifically, the AI ​​model adjusts the tone of the voice and the facial expressions in the video to suit the user. The output is visual and auditory media files sent to the terminal.

[0302] Step 4:

[0303] The terminal receives visual and auditory information transmitted from the server and presents it to the user. The input is a generated media file, and the output is the provision of visual and auditory information to the user. Specifically, the terminal displays images on the screen and plays audio through the speaker, providing information to the user in an intuitive and easy-to-understand manner.

[0304] Step 5:

[0305] The user interacts with a nurse generated through the terminal. The input for this step is the user's voice and further requests. The output is that the content of the user's speech is sent to the server as text data. As a specific operation, the microphone of the terminal captures the user's voice and converts it into text using speech recognition technology.

[0306] Step 6:

[0307] The server analyzes the interaction data from the user. The input is the text data sent from the terminal and additional information from the sensor device. The server analyzes these data and determines the urgency. The result appears as a work instruction or a prioritized response protocol. The output is the next instruction or alert information sent to the terminal.

[0308] Step 7:

[0309] The terminal notifies the user of the instruction from the server and adjusts the physical response as needed. The input is the work instruction or alert information from the server. The output is the notification of appropriate actions to the user and medical staff. As a specific operation, the terminal displays an emergency message on the screen or issues an alarm to the medical staff if necessary.

[0310] (Application Example 1)

[0311] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0312] In an industrial site, it is required to respond quickly and accurately to abnormal signals and requests from workers while maintaining production efficiency and safety. However, in conventional systems, it has been difficult to make real-time judgments and responses to such requests. For this reason, there have been problems such as obstacles to on-site work, production delays, and even safety risks.

[0313] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0314] In this invention, the server includes a device for receiving signals transmitted by a user, a generation device for providing generated visual media and sound to a terminal based on the received signals, a determination device for determining priority based on the generated visual media and sound and user information, an instruction device for quickly dispatching personnel or equipment when the determination device determines the situation to be important, and a device for receiving abnormal signals from equipment and work stations at an industrial site and responding as a virtual maintenance person using a generated AI model. This enables immediate judgment and response to abnormal situations, thereby improving the efficiency and safety of production at industrial sites.

[0315] A "signal" is a means of transmitting data or requests originating from a user or device.

[0316] A "device" is a machine or electronic device that has a specific function and is used to receive, generate, judge, or direct signals.

[0317] A "generative AI model" is a mathematical or algorithmic model that uses artificial intelligence technology to analyze signals and generate appropriate responses.

[0318] "Visual media" refers to digital content presented in the form of images or videos that contain visual information.

[0319] "Acoustics" refers to information that includes audio data and is used as sound for communication.

[0320] A "terminal" is an electronic device or part of a computer that a user interacts with directly.

[0321] A "decision-making device" is a device or software used to analyze the content of received data and determine its importance and the priority of actions to be taken.

[0322] A "commanding device" is a device that generates or transmits commands to initiate necessary actions or processes based on a judgment.

[0323] An "abnormal signal" is a signal that includes warnings or error messages indicating a condition that is different from the normal state.

[0324] A "virtual maintenance staff" is a maintenance staff system artificially created by a generative AI model, capable of handling problems without human intervention.

[0325] The system that realizes this invention aims to efficiently handle abnormalities in industrial settings. First, a terminal receives a signal transmitted by the user. The terminal sends this signal to a server and starts processing the response.

[0326] The server is equipped with a generative AI model that analyzes abnormal signals and user requests. The generated video and audio are sent back to the terminal as visual and auditory media. This allows users to interact with a virtual maintenance representative through the screen and audio, and quickly receive the necessary information and support.

[0327] The server analyzes the received data, and a determination device is activated to assess the urgency. If it is deemed critical, an instruction device is activated, issuing orders to dispatch personnel and equipment as needed. As a result, an appropriate response is quickly taken at the scene.

[0328] For example, if a worker reports that "the machine is making an unusual noise," the AI ​​model will predict the possibility of an anomaly and suggest appropriate countermeasures. This process is initiated by entering a prompt message into the user's terminal. An example of a prompt message is, "The machine is making an unusual noise. Please tell me how to deal with it."

[0329] The hardware used includes smartphones and head-mounted displays, while the software utilizes TensorFlow and OpenCV for signal processing and data analysis. This allows for the efficient execution of the process from signal reception to processing.

[0330] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0331] Step 1:

[0332] When a user sends a signal, the terminal receives that signal. The input is the call signal from the user, and the output is the digital data of the signal. This allows the terminal to recognize an anomaly or a request.

[0333] Step 2:

[0334] The terminal forwards the received signal to the server. The input is digital data of the signal, and the output is prompt data directed to the server. This data forms the basis for the next processing step.

[0335] Step 3:

[0336] The server inputs the received prompt data into the generating AI model. The input is prompt data, and the output is the analysis results. The generating AI model determines the nature and priority of the anomalies.

[0337] Step 4:

[0338] The server generates visual and auditory media for a virtual maintenance worker based on the analysis results of the generated AI model. The input is the analysis results from the AI ​​model, and the output is the generated visual and auditory media. This is then used as feedback to the user.

[0339] Step 5:

[0340] The generated visual and auditory media are sent to the terminal and presented to the user. The input is the generated visual and auditory media, and the output is video and audio for the user. This allows the user to interact with the system.

[0341] Step 6:

[0342] If the user provides further interaction, the server analyzes the data based on that interaction and generates the necessary instructions. The input is the additional interaction with the user, and the output is further instruction data. If necessary, instructions are given to quickly dispatch personnel and equipment to the site.

[0343] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0344] This invention relates to a generative AI system that combines an emotion engine to automatically and individually respond to nurse calls from users in medical facilities. This system can further improve user satisfaction by recognizing the emotional state of users and providing personalized responses accordingly.

[0345] When a user presses the nurse call button, the terminal receives the signal and sends the information to the server. Based on this information, the server uses generative AI to generate video and audio of a virtual nurse. Because this generation is based on the user's past history and preferences, personalized interaction is possible.

[0346] The server then analyzes the user's emotional state through an emotion engine. This engine uses voice analysis and facial recognition technology to determine whether the user is stressed, at ease, or confused. The analyzed emotional data is then sent to the server and used to adjust the conversation.

[0347] The generated video and audio are sent to the device, initiating interaction with the user. For example, if the user expresses feelings such as "I'm a little anxious, but I don't feel any major problems," the generating AI will respond calmly with something like, "I'll help you feel more at ease. Please let me know if there's anything specific I can do to help."

[0348] The terminal transmits the conversation with the user to the server in real time, and the server determines the urgency level based on the conversation content. Along with the emotional state, appropriate instructions are generated. For example, if the emotional analysis indicates a low level of urgency but the user is feeling anxious, additional counseling via voice call can be provided.

[0349] Another important function of this system is to improve the quality of care and medical treatment by tracking changes in users' emotions and conducting long-term emotional analysis. This allows for the provision of care tailored to each user and enhances their sense of security. In this way, the system is a powerful tool for reducing the burden on healthcare professionals and implementing comprehensive medical care.

[0350] The following describes the processing flow.

[0351] Step 1:

[0352] The user presses the nurse call button. The terminal receives this signal and prepares a request to the server.

[0353] Step 2:

[0354] The terminal sends call information to the server, initiating a processing request along with the user's ID and location information.

[0355] Step 3:

[0356] Based on the received data, the server generates video and audio of a virtual nurse tailored to the user's history and preferences.

[0357] Step 4:

[0358] The server activates the emotion engine, analyzes emotional data from the user's voice and video, and determines the user's emotional state.

[0359] Step 5:

[0360] The server sends the generated screen and emotion data to the terminal, initiating interaction with the user.

[0361] Step 6:

[0362] The terminal uses the received video and audio to display the conversation between the user and the virtual nurse, providing interactive support.

[0363] Step 7:

[0364] As the user converses with the virtual nurse, they express a variety of reactions and emotions. The device transmits the conversation data to the server in real time.

[0365] Step 8:

[0366] The server analyzes the conversation content and emotional data, determines the urgency level, and decides on a flexible response based on the emotions, if necessary.

[0367] Step 9:

[0368] Based on the assessment results, the server generates instruction data to dispatch personnel or provide additional remote support if necessary.

[0369] Step 10:

[0370] Based on instructions from the server, the terminal continuously provides feedback to the user and offers additional support to provide a sense of security.

[0371] Step 11:

[0372] Upon receiving a user response, the terminal reports the response result and sentiment data back to the server, completing the entire process.

[0373] Step 12:

[0374] The server records all conversation and sentiment data, and updates the database to help with future improvements and long-term analysis.

[0375] (Example 2)

[0376] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0377] In healthcare facilities, there is a need for a system that can appropriately respond to calls from users, understand their emotional state, and provide individualized interactions to increase user satisfaction, while also enabling rapid response to emergencies. Conventional systems have the problem of not being able to provide satisfactory service to users because they are not capable of analyzing emotional states or generating individualized responses.

[0378] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0379] In this invention, the server includes a receiving means, a generating means, an analysis means, an adjustment means, and an instruction means. This enables rapid analysis of the user's emotional state and the provision of personalized services. Specifically, it becomes possible to generate a virtual supporter that meets the needs of each user by utilizing a generation AI model and to provide interaction based on emotional analysis.

[0380] A "receiving means" is a function that detects signals transmitted by the user and incorporates them into the system.

[0381] The "generation means" is a function that generates information about a virtual supporter based on the signal information received by the receiving means and supplies this information to the information processing device.

[0382] The "analysis means" is a function that determines the user's emotional state by analyzing data such as the user's voice and video.

[0383] The "adjustment means" is a function that plays the role of generating individually adjusted information based on the results of the analysis means.

[0384] The "determination means" is a function that determines the urgency of the current situation based on generated information and user history information.

[0385] An "instruction mechanism" is a function that issues instructions to promptly take necessary measures based on the results of a determination mechanism.

[0386] A "recording means" is a function that sequentially collects user conversation data and generates information to track changes in emotional state based on this data.

[0387] The "response generation means" is a function that enables the generation of personalized responses based on the user's history data and emotional data.

[0388] This invention relates to a system for responding quickly and appropriately to user calls in medical facilities. The system involves multiple components, including a server, terminals, and users. The server utilizes a generative AI model to generate video and audio of a virtual supporter tailored to each user's needs. The generative AI model used is vendor-independent but implements state-of-the-art machine learning algorithms.

[0389] The terminal is responsible for receiving call signals from the user and transmitting them to the server. The server uses the received signals and past user data to perform voice analysis using an emotion engine and facial recognition technology. This analysis determines the user's emotional state, and personalized interactions are set accordingly. The server can also use general natural language processing techniques for voice analysis and existing image analysis techniques for facial recognition.

[0390] For example, Microsoft Azure's facial recognition API and Google Cloud's natural language processing API can be used. The generated video and audio are sent to the device, and a conversation with the user begins. During this interaction, the device continuously sends the conversation content to the server. The server determines the urgency based on the conversation content and emotional state, and generates appropriate instructions.

[0391] For example, if a user feels anxious, the system uses AI to provide a response such as, "We'll help you feel more at ease. Please let us know if there's anything specific we can do to help." Furthermore, this system tracks the user's emotional changes over the long term, providing data that can be used to improve medical or nursing care services.

[0392] An example of a prompt message would be: "Explain the process for generating an appropriate virtual caregiver based on past history information and sentiment analysis when a patient presses the nurse call button, and for optimizing the interaction with the patient." This allows the system to provide user-optimized medical care and reduce the burden on healthcare professionals.

[0393] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0394] Step 1:

[0395] When a user presses the nurse call button, the terminal receives the signal as input data. This signal is in digital format and contains the user's ID and location information. The terminal forwards this input to the server, which then sends a trigger to prepare for response.

[0396] Step 2:

[0397] The server generates video and audio of a virtual supporter based on signals received from the terminal, utilizing a generative AI model. Specifically, it searches a database for past history information and preferences identified by the user's ID and supplies this as input to the generative AI model. The generative AI model customizes individual interactions based on prompt messages and creates video and audio output.

[0398] Step 3:

[0399] The server sends the generated virtual supporter's video and audio data to the terminal. The terminal receives this data and displays the video and plays the audio for the user. The user experiences this interaction in real time.

[0400] Step 4:

[0401] During user interaction, the device sequentially collects the user's voice and facial expressions and sends them to the server as input data. The server analyzes this data using an emotion engine to determine the user's emotional state. At this time, stress levels and feelings of security are quantified using voice recognition and facial recognition technologies.

[0402] Step 5:

[0403] The server adjusts conversations and responses based on the analyzed emotional data. Specifically, if the user is feeling anxious, it uses a generative AI model to generate additional support messages and delivers them to the device.

[0404] Step 6:

[0405] The terminal sequentially sends the content of the interaction with the user to the server, which then processes the conversation to determine its urgency. The server integrates sentiment data with the conversation content, applies urgency determination logic, and determines the priority.

[0406] Step 7:

[0407] If the situation is deemed highly urgent, the server will instruct the terminal to take immediate action based on rules defined within the system. This includes instructing the dispatch of medical staff as needed and preparing relevant equipment.

[0408] (Application Example 2)

[0409] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0410] Conventional security systems have the challenge of being unable to adequately consider the emotional state of users and respond appropriately to diverse situations. Furthermore, they struggle to provide users with customized, real-time feedback for ensuring their safety. As a result, they are unable to implement optimal security measures and adequately alleviate user anxiety.

[0411] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0412] In this invention, the server includes means for receiving signals transmitted by a user, generation means for providing generated visual and auditory information to an information terminal based on the received signals, and analysis means for analyzing the user's emotional state. This makes it possible to analyze the user's emotions in real time and provide optimal visual and auditory feedback tailored to that situation.

[0413] "Receiving means" refers to devices or systems for receiving signals transmitted by users.

[0414] A "generation means" is a mechanism for providing information terminals with generated visual and auditory content based on received signals.

[0415] A "determination means" is a mechanism that determines priority based on generated visual and auditory information and user information.

[0416] "Instruction means" refers to a means of quickly dispatching personnel or equipment when the determination means determines that the matter has a high priority.

[0417] "Analysis methods" refer to processes and techniques for analyzing the emotional state of users.

[0418] "Adjustment means" refers to means that have the function of adjusting the response content based on the analysis results obtained by the analysis means.

[0419] A "recording system" is a mechanism that sequentially collects dialogue data with users and generates information for use in new responses.

[0420] A "database access means" is a means of obtaining the data necessary to provide a customized response through an information terminal.

[0421] This invention relates to a personal security system that utilizes emotion analysis and was developed with the aim of enhancing user safety. In this system, the user's device, specifically smart glasses, is used as the main interface.

[0422] The server first receives the user's visual and audio data transmitted from the smart glasses. This data is processed to analyze facial expressions and voice tone. Specifically, facial recognition technology is used for facial expressions, and voice analysis technology is used for voice. The hardware used includes the camera and microphone within the smart glasses. This makes it possible to evaluate the user's emotional state in real time.

[0423] Subsequently, the server uses a generative AI model based on the sentiment analysis results to generate a response tailored to the user. At this stage, specific advice and warnings are generated to alleviate the user's anxiety and fear. For example, if the user expresses the sentiment, "I've been feeling uneasy on the street lately," the generative AI will provide advice such as, "Pay attention to your surroundings and be wary of suspicious activity."

[0424] The generated advice and warnings are delivered to the user as audio and visual feedback through smart glasses, and the user can receive them in real time. A key feature of this system is that the responses are personalized to the user and optimized according to the user's emotional state.

[0425] An example of a prompt would be: "Sentiment analysis result: User's sentiment data Voice analysis result: User's voice text Provide the best advice for this user." By inputting this prompt into the generating AI model, a customized response tailored to the user can be obtained.

[0426] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0427] Step 1:

[0428] The server receives visual data (images) and audio data transmitted from the terminal (smart glasses). This input data serves as material for analyzing the user's emotional state. Upon receiving this data, the server sends the visual and audio data to separate analysis processes.

[0429] Step 2:

[0430] The server uses visual data to perform facial recognition technology for expression analysis. Specifically, it extracts facial features from the image and determines the user's emotional state. The output obtained here is an emotion label such as "reassured" or "anxious."

[0431] Step 3:

[0432] Simultaneously, the server applies speech analysis technology to the audio data, evaluating the emotional state by analyzing tone and pitch. This results in the output of speech-based emotion labels such as "calm" or "tense."

[0433] Step 4:

[0434] The server integrates the emotional labels obtained from the visual and auditory data to assess the user's overall emotional state. This integrated emotional data is then used in the next step.

[0435] Step 5:

[0436] The server uses a generative AI model based on integrated sentiment data to create prompt messages. Specifically, it embeds the results of sentiment analysis into the prompt messages and provides them as input to generate optimal feedback for the user.

[0437] Step 6:

[0438] The generative AI model receives prompt messages generated by the server as input and generates advice and warnings to provide to the user. The output of this step is a personalized message tailored to the specific situation.

[0439] Step 7:

[0440] The server sends generated advice and warnings to the device, presenting them to the user as visual and audio feedback. The device plays this back and provides it to the user in real time, thereby improving the user's safety awareness.

[0441] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0442] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0443] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0444] [Third Embodiment]

[0445] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0446] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0447] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0448] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0449] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0451] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0452] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0453] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0454] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0455] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0456] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0457] This invention relates to a system using generative AI for receiving and automatically responding to nurse calls initiated by users in medical facilities. This system can reduce the burden on medical staff and improve user satisfaction.

[0458] First, when a user presses a nurse call button installed in the patient's room, the terminal receives the signal. The terminal sends this information to the server and instructs it to begin responding. The server processes the received data and uses generative AI to generate video and audio of a simulated nurse. This generation is based on past history and user preferences, providing personalized interaction.

[0459] The generated video and audio are sent back to the terminal and displayed on the user's screen. The user can converse with the generated nurse through an interactive interface. This interaction collects information about the user's requests and physical condition.

[0460] The server analyzes the user's conversation and data from connected sensor devices to determine the urgency and priority of the response. For example, if a user complains of chest pain, it is immediately determined to be an "emergency," and instructions are generated to request immediate assistance from medical personnel.

[0461] Instructions are fed back to the user via the terminal. In some cases, physical personnel may be dispatched or medical equipment may be provided. This ensures that the user receives appropriate support without any omissions or delays.

[0462] Furthermore, data collected during everyday conversations can be used to improve future services. For example, the results of users' regular health checks and the frequency with which specific caregiving activities are needed can be analyzed, enabling the provision of more personalized services. The introduction of this system is expected to improve efficiency in medical settings and enhance the quality of medical care provided.

[0463] The following describes the processing flow.

[0464] Step 1:

[0465] The user presses the nurse call button. The terminal detects the button press and receives the signal.

[0466] Step 2:

[0467] The terminal sends the received call signal to the server and generates a request to initiate a response.

[0468] Step 3:

[0469] The server receives the request and retrieves the user's ID, location information, and past history data from the database.

[0470] Step 4:

[0471] Based on the acquired data, the server uses a generation AI to create video and audio of a nurse character.

[0472] Step 5:

[0473] The server transmits the generated video and audio to the terminal, presenting the user with a virtual nurse.

[0474] Step 6:

[0475] The terminal displays the generated video and audio of the nurse to the user and initiates an interactive conversation.

[0476] Step 7:

[0477] The user converses with a nurse on the screen. The terminal continuously transmits the user's input and voice data to the server.

[0478] Step 8:

[0479] The server analyzes the conversation content and sensor data to determine the urgency and priority of the response.

[0480] Step 9:

[0481] The server generates necessary responses and instructions based on its judgment, and issues instructions regarding the dispatch of personnel and equipment.

[0482] Step 10:

[0483] The terminal provides feedback to the user based on instructions from the server, displaying them on the screen and via sound.

[0484] Step 11:

[0485] If necessary, staff will be dispatched to the user's location. The terminal reports the completion of the process to the server.

[0486] Step 12:

[0487] The server records the content of conversations and updates the database to help improve subsequent responses.

[0488] (Example 1)

[0489] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0490] Responding quickly and accurately to inquiries and requests from users in medical facilities is crucial for significantly reducing the burden on healthcare professionals and improving user satisfaction. However, current systems struggle to efficiently manage the enormous number of calls, and the time and resources available for providing appropriate individual responses are limited. This can lead to delays in responding to urgent cases, causing risks and dissatisfaction among users.

[0491] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0492] In this invention, the server includes a device for receiving communication signals transmitted by a user, a generation device for providing visual and auditory information generated using a generation AI model to a terminal device, and an analysis device for determining the urgency of the situation. This makes it possible to provide quickly generated individual responses to user inquiries based on multiple information resources, reducing the workload of medical professionals and improving the accuracy and speed of services to users.

[0493] A "communication signal" is electrical data that a user transmits from their terminal, indicating a call or request.

[0494] "Device" refers to an instrument or machine used to perform a specific function, and in this invention, it refers to a mechanism that receives communication signals and manipulates generated AI models.

[0495] A "generative AI model" is a form of artificial intelligence that constructs visual and auditory information based on past data and personal preferences.

[0496] "Visual and auditory information" refers to images displayed on a screen and sounds emitted from speakers, created by a generative AI model.

[0497] A "terminal device" is a device that serves as an interface with the user, displaying information and playing audio.

[0498] An "analysis device" is a device that has the function of analyzing and determining the urgency level using user information and data collected from sensors.

[0499] A "commanding device" refers to a mechanism that issues commands to workers or equipment to carry out specific actions based on analysis results.

[0500] An "information resource access device" is a device that has the function of accessing a database to generate an adapted response by referring to the user's past history and individual preferences.

[0501] This invention relates to a system for responding quickly and individually to inquiries from users in medical facilities. The following specifically describes embodiments for carrying out this invention.

[0502] This system uses terminals installed in patient rooms to receive communication signals transmitted by users and utilizes a server to process that data. The terminal functions as an interface for user operation, and when a nurse call button is pressed, it sends a signal to the server.

[0503] The server analyzes the received signal to determine the user's room number and the content of the call. Next, it uses a generative AI model to generate personalized visual and auditory information for each user. This information incorporates data reflecting past medical history and the user's specific preferences. Specifically, the server utilizes a cloud-based AI platform for data analysis and video / audio generation.

[0504] For example, if a user presses the nurse call button and requests water, the device sends the voice message to the server. The server, using a generative AI model, generates video and audio responses saying, "We'll bring you water right away," and sends them back to the device. The user can then see this interaction through the device's screen and speaker.

[0505] A concrete example of a prompt message would be, "The patient has pressed the nurse call button. Please generate a response." This allows the server to quickly take an appropriate action via the AI ​​model.

[0506] Furthermore, the server collects data from connected sensor devices to monitor the user's situation and health status. This data helps determine the urgency of a situation and the appropriate response, enabling efficient medical support.

[0507] This system is expected to significantly improve efficiency in healthcare settings and user satisfaction by enabling users to receive prompt and appropriate support and reducing the burden on healthcare professionals.

[0508] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0509] Step 1:

[0510] The user presses the nurse call button on the terminal. This action causes the terminal to generate a communication signal, which is then sent to the server as input data. Specifically, pressing the button causes the terminal to construct a data packet containing information such as the room number and the reason for the call. The output is the transmission of a signal to the server.

[0511] Step 2:

[0512] The server analyzes the signals received from the terminal. The input is a data packet containing the user's room number and requests. The server transforms this data into a prompt for the AI ​​model and inputs it into the generative AI model. The prompt might be, for example, "The patient says they want water. Generate an appropriate response." The output of this process is a template of the visual and auditory information that the generative AI model generates.

[0513] Step 3:

[0514] The server generates visual and auditory information using a generative AI model. The input is a generation template based on the prompts generated in step 2. The AI ​​considers past historical data and user preferences when generating video and audio based on this template. Specifically, the AI ​​model adjusts the tone of the voice and the facial expressions in the video to suit the user. The output is visual and auditory media files sent to the terminal.

[0515] Step 4:

[0516] The terminal receives visual and auditory information transmitted from the server and presents it to the user. The input is a generated media file, and the output is the provision of visual and auditory information to the user. Specifically, the terminal displays images on the screen and plays audio through the speaker, providing information to the user in an intuitive and easy-to-understand manner.

[0517] Step 5:

[0518] The user interacts with a nurse generated through the device. The input for this step is the user's voice and further requests. The output is the user's speech being sent to the server as text data. Specifically, the device's microphone captures the user's voice and converts it to text using speech recognition technology.

[0519] Step 6:

[0520] The server analyzes interaction data from the user. Inputs include text data sent from the terminal and additional information from sensor devices. The server analyzes this data to determine urgency. The results are expressed as work instructions and prioritized response protocols. Outputs are the next instructions and alerts sent to the terminal.

[0521] Step 7:

[0522] The terminal notifies the user of instructions from the server and adjusts physical responses as needed. Inputs are work instructions and alert information from the server. Outputs are notifications of appropriate actions to the user and medical staff. Specifically, the terminal displays emergency messages on the screen and issues alarms to medical staff if necessary.

[0523] (Application Example 1)

[0524] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0525] In industrial settings, there is a need to respond quickly and accurately to abnormal signals and worker requests while maintaining production efficiency and safety. However, conventional systems have struggled to make real-time decisions and respond to these requests. This has resulted in challenges such as disruptions to on-site work, production delays, and even safety risks.

[0526] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0527] In this invention, the server includes a device for receiving signals transmitted by a user, a generation device for providing generated visual media and sound to a terminal based on the received signals, a determination device for determining priority based on the generated visual media and sound and user information, an instruction device for quickly dispatching personnel or equipment when the determination device determines the situation to be important, and a device for receiving abnormal signals from equipment and work stations at an industrial site and responding as a virtual maintenance person using a generated AI model. This enables immediate judgment and response to abnormal situations, thereby improving the efficiency and safety of production at industrial sites.

[0528] A "signal" is a means of transmitting data or requests originating from a user or device.

[0529] A "device" is a machine or electronic device that has a specific function and is used to receive, generate, judge, or direct signals.

[0530] A "generative AI model" is a mathematical or algorithmic model that uses artificial intelligence technology to analyze signals and generate appropriate responses.

[0531] "Visual media" refers to digital content presented in the form of images or videos that contain visual information.

[0532] "Acoustics" refers to information that includes audio data and is used as sound for communication.

[0533] A "terminal" is an electronic device or part of a computer that a user interacts with directly.

[0534] A "decision-making device" is a device or software used to analyze the content of received data and determine its importance and the priority of actions to be taken.

[0535] A "commanding device" is a device that generates or transmits commands to initiate necessary actions or processes based on a judgment.

[0536] An "abnormal signal" is a signal that includes warnings or error messages indicating a condition that is different from the normal state.

[0537] A "virtual maintenance staff" is a maintenance staff system artificially created by a generative AI model, capable of handling problems without human intervention.

[0538] The system that realizes this invention aims to efficiently handle abnormalities in industrial settings. First, a terminal receives a signal transmitted by the user. The terminal sends this signal to a server and starts processing the response.

[0539] The server is equipped with a generative AI model that analyzes abnormal signals and user requests. The generated video and audio are sent back to the terminal as visual and auditory media. This allows users to interact with a virtual maintenance representative through the screen and audio, and quickly receive the necessary information and support.

[0540] The server analyzes the received data, and a determination device is activated to assess the urgency. If it is deemed critical, an instruction device is activated, issuing orders to dispatch personnel and equipment as needed. As a result, an appropriate response is quickly taken at the scene.

[0541] For example, if a worker reports that "the machine is making an unusual noise," the AI ​​model will predict the possibility of an anomaly and suggest appropriate countermeasures. This process is initiated by entering a prompt message into the user's terminal. An example of a prompt message is, "The machine is making an unusual noise. Please tell me how to deal with it."

[0542] The hardware used includes smartphones and head-mounted displays, while the software utilizes TensorFlow and OpenCV for signal processing and data analysis. This allows for the efficient execution of the process from signal reception to processing.

[0543] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0544] Step 1:

[0545] When a user sends a signal, the terminal receives that signal. The input is the call signal from the user, and the output is the digital data of the signal. This allows the terminal to recognize an anomaly or a request.

[0546] Step 2:

[0547] The terminal forwards the received signal to the server. The input is digital data of the signal, and the output is prompt data directed to the server. This data forms the basis for the next processing step.

[0548] Step 3:

[0549] The server inputs the received prompt data into the generating AI model. The input is prompt data, and the output is the analysis results. The generating AI model determines the nature and priority of the anomalies.

[0550] Step 4:

[0551] The server generates visual and auditory media for a virtual maintenance worker based on the analysis results of the generated AI model. The input is the analysis results from the AI ​​model, and the output is the generated visual and auditory media. This is then used as feedback to the user.

[0552] Step 5:

[0553] The generated visual and auditory media are sent to the terminal and presented to the user. The input is the generated visual and auditory media, and the output is video and audio for the user. This allows the user to interact with the system.

[0554] Step 6:

[0555] If the user provides further interaction, the server analyzes the data based on that interaction and generates the necessary instructions. The input is the additional interaction with the user, and the output is further instruction data. If necessary, instructions are given to quickly dispatch personnel and equipment to the site.

[0556] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0557] This invention relates to a generative AI system that combines an emotion engine to automatically and individually respond to nurse calls from users in medical facilities. This system can further improve user satisfaction by recognizing the emotional state of users and providing personalized responses accordingly.

[0558] When a user presses the nurse call button, the terminal receives the signal and sends the information to the server. Based on this information, the server uses generative AI to generate video and audio of a virtual nurse. Because this generation is based on the user's past history and preferences, personalized interaction is possible.

[0559] The server then analyzes the user's emotional state through an emotion engine. This engine uses voice analysis and facial recognition technology to determine whether the user is stressed, at ease, or confused. The analyzed emotional data is then sent to the server and used to adjust the conversation.

[0560] The generated video and audio are sent to the device, initiating interaction with the user. For example, if the user expresses feelings such as "I'm a little anxious, but I don't feel any major problems," the generating AI will respond calmly with something like, "I'll help you feel more at ease. Please let me know if there's anything specific I can do to help."

[0561] The terminal transmits the conversation with the user to the server in real time, and the server determines the urgency level based on the conversation content. Along with the emotional state, appropriate instructions are generated. For example, if the emotional analysis indicates a low level of urgency but the user is feeling anxious, additional counseling via voice call can be provided.

[0562] Another important function of this system is to improve the quality of care and medical treatment by tracking changes in users' emotions and conducting long-term emotional analysis. This allows for the provision of care tailored to each user and enhances their sense of security. In this way, the system is a powerful tool for reducing the burden on healthcare professionals and implementing comprehensive medical care.

[0563] The following describes the processing flow.

[0564] Step 1:

[0565] The user presses the nurse call button. The terminal receives this signal and prepares a request to the server.

[0566] Step 2:

[0567] The terminal sends call information to the server, initiating a processing request along with the user's ID and location information.

[0568] Step 3:

[0569] Based on the received data, the server generates video and audio of a virtual nurse tailored to the user's history and preferences.

[0570] Step 4:

[0571] The server activates the emotion engine, analyzes emotional data from the user's voice and video, and determines the user's emotional state.

[0572] Step 5:

[0573] The server sends the generated screen and emotion data to the terminal, initiating interaction with the user.

[0574] Step 6:

[0575] The terminal uses the received video and audio to display the conversation between the user and the virtual nurse, providing interactive support.

[0576] Step 7:

[0577] As the user converses with the virtual nurse, they express a variety of reactions and emotions. The device transmits the conversation data to the server in real time.

[0578] Step 8:

[0579] The server analyzes the conversation content and emotional data, determines the urgency level, and decides on a flexible response based on the emotions, if necessary.

[0580] Step 9:

[0581] Based on the assessment results, the server generates instruction data to dispatch personnel or provide additional remote support if necessary.

[0582] Step 10:

[0583] Based on instructions from the server, the terminal continuously provides feedback to the user and offers additional support to provide a sense of security.

[0584] Step 11:

[0585] Upon receiving a user response, the terminal reports the response result and sentiment data back to the server, completing the entire process.

[0586] Step 12:

[0587] The server records all conversation and sentiment data, and updates the database to help with future improvements and long-term analysis.

[0588] (Example 2)

[0589] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0590] In healthcare facilities, there is a need for a system that can appropriately respond to calls from users, understand their emotional state, and provide individualized interactions to increase user satisfaction, while also enabling rapid response to emergencies. Conventional systems have the problem of not being able to provide satisfactory service to users because they are not capable of analyzing emotional states or generating individualized responses.

[0591] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0592] In this invention, the server includes a receiving means, a generating means, an analysis means, an adjustment means, and an instruction means. This enables rapid analysis of the user's emotional state and the provision of personalized services. Specifically, it becomes possible to generate a virtual supporter that meets the needs of each user by utilizing a generation AI model and to provide interaction based on emotional analysis.

[0593] A "receiving means" is a function that detects signals transmitted by the user and incorporates them into the system.

[0594] The "generation means" is a function that generates information about a virtual supporter based on the signal information received by the receiving means and supplies this information to the information processing device.

[0595] The "analysis means" is a function that determines the user's emotional state by analyzing data such as the user's voice and video.

[0596] The "adjustment means" is a function that plays the role of generating individually adjusted information based on the results of the analysis means.

[0597] The "determination means" is a function that determines the urgency of the current situation based on generated information and user history information.

[0598] An "instruction mechanism" is a function that issues instructions to promptly take necessary measures based on the results of a determination mechanism.

[0599] A "recording means" is a function that sequentially collects user conversation data and generates information to track changes in emotional state based on this data.

[0600] The "response generation means" is a function that enables the generation of personalized responses based on the user's history data and emotional data.

[0601] This invention relates to a system for responding quickly and appropriately to user calls in medical facilities. The system involves multiple components, including a server, terminals, and users. The server utilizes a generative AI model to generate video and audio of a virtual supporter tailored to each user's needs. The generative AI model used is vendor-independent but implements state-of-the-art machine learning algorithms.

[0602] The terminal is responsible for receiving call signals from the user and transmitting them to the server. The server uses the received signals and past user data to perform voice analysis using an emotion engine and facial recognition technology. This analysis determines the user's emotional state, and personalized interactions are set accordingly. The server can also use general natural language processing techniques for voice analysis and existing image analysis techniques for facial recognition.

[0603] For example, Microsoft Azure's facial recognition API and Google Cloud's natural language processing API can be used. The generated video and audio are sent to the device, and a conversation with the user begins. During this interaction, the device continuously sends the conversation content to the server. The server determines the urgency based on the conversation content and emotional state, and generates appropriate instructions.

[0604] For example, if a user feels anxious, the system uses AI to provide a response such as, "We'll help you feel more at ease. Please let us know if there's anything specific we can do to help." Furthermore, this system tracks the user's emotional changes over the long term, providing data that can be used to improve medical or nursing care services.

[0605] An example of a prompt message would be: "Explain the process for generating an appropriate virtual caregiver based on past history information and sentiment analysis when a patient presses the nurse call button, and for optimizing the interaction with the patient." This allows the system to provide user-optimized medical care and reduce the burden on healthcare professionals.

[0606] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0607] Step 1:

[0608] When a user presses the nurse call button, the terminal receives the signal as input data. This signal is in digital format and contains the user's ID and location information. The terminal forwards this input to the server, which then sends a trigger to prepare for response.

[0609] Step 2:

[0610] The server generates video and audio of a virtual supporter based on signals received from the terminal, utilizing a generative AI model. Specifically, it searches a database for past history information and preferences identified by the user's ID and supplies this as input to the generative AI model. The generative AI model customizes individual interactions based on prompt messages and creates video and audio output.

[0611] Step 3:

[0612] The server sends the generated virtual supporter's video and audio data to the terminal. The terminal receives this data and displays the video and plays the audio for the user. The user experiences this interaction in real time.

[0613] Step 4:

[0614] During user interaction, the device sequentially collects the user's voice and facial expressions and sends them to the server as input data. The server analyzes this data using an emotion engine to determine the user's emotional state. At this time, stress levels and feelings of security are quantified using voice recognition and facial recognition technologies.

[0615] Step 5:

[0616] The server adjusts conversations and responses based on the analyzed emotional data. Specifically, if the user is feeling anxious, it uses a generative AI model to generate additional support messages and delivers them to the device.

[0617] Step 6:

[0618] The terminal sequentially sends the content of the interaction with the user to the server, which then processes the conversation to determine its urgency. The server integrates sentiment data with the conversation content, applies urgency determination logic, and determines the priority.

[0619] Step 7:

[0620] If the situation is deemed highly urgent, the server will instruct the terminal to take immediate action based on rules defined within the system. This includes instructing the dispatch of medical staff as needed and preparing relevant equipment.

[0621] (Application Example 2)

[0622] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0623] Conventional security systems have the challenge of being unable to adequately consider the emotional state of users and respond appropriately to diverse situations. Furthermore, they struggle to provide users with customized, real-time feedback for ensuring their safety. As a result, they are unable to implement optimal security measures and adequately alleviate user anxiety.

[0624] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0625] In this invention, the server includes means for receiving signals transmitted by a user, generation means for providing generated visual and auditory information to an information terminal based on the received signals, and analysis means for analyzing the user's emotional state. This makes it possible to analyze the user's emotions in real time and provide optimal visual and auditory feedback tailored to that situation.

[0626] "Receiving means" refers to devices or systems for receiving signals transmitted by users.

[0627] A "generation means" is a mechanism for providing information terminals with generated visual and auditory content based on received signals.

[0628] A "determination means" is a mechanism that determines priority based on generated visual and auditory information and user information.

[0629] "Instruction means" refers to a means of quickly dispatching personnel or equipment when the determination means determines that the matter has a high priority.

[0630] "Analysis methods" refer to processes and techniques for analyzing the emotional state of users.

[0631] "Adjustment means" refers to means that have the function of adjusting the response content based on the analysis results obtained by the analysis means.

[0632] A "recording system" is a mechanism that sequentially collects dialogue data with users and generates information for use in new responses.

[0633] A "database access means" is a means of obtaining the data necessary to provide a customized response through an information terminal.

[0634] This invention relates to a personal security system that utilizes emotion analysis and was developed with the aim of enhancing user safety. In this system, the user's device, specifically smart glasses, is used as the main interface.

[0635] The server first receives the user's visual and audio data transmitted from the smart glasses. This data is processed to analyze facial expressions and voice tone. Specifically, facial recognition technology is used for facial expressions, and voice analysis technology is used for voice. The hardware used includes the camera and microphone within the smart glasses. This makes it possible to evaluate the user's emotional state in real time.

[0636] Subsequently, the server uses a generative AI model based on the sentiment analysis results to generate a response tailored to the user. At this stage, specific advice and warnings are generated to alleviate the user's anxiety and fear. For example, if the user expresses the sentiment, "I've been feeling uneasy on the street lately," the generative AI will provide advice such as, "Pay attention to your surroundings and be wary of suspicious activity."

[0637] The generated advice and warnings are delivered to the user as audio and visual feedback through smart glasses, and the user can receive them in real time. A key feature of this system is that the responses are personalized to the user and optimized according to the user's emotional state.

[0638] An example of a prompt would be: "Sentiment analysis result: User's sentiment data Voice analysis result: User's voice text Provide the best advice for this user." By inputting this prompt into the generating AI model, a customized response tailored to the user can be obtained.

[0639] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0640] Step 1:

[0641] The server receives visual data (images) and audio data transmitted from the terminal (smart glasses). This input data serves as material for analyzing the user's emotional state. Upon receiving this data, the server sends the visual and audio data to separate analysis processes.

[0642] Step 2:

[0643] The server uses visual data to perform facial recognition technology for expression analysis. Specifically, it extracts facial features from the image and determines the user's emotional state. The output obtained here is an emotion label such as "reassured" or "anxious."

[0644] Step 3:

[0645] Simultaneously, the server applies speech analysis technology to the audio data, evaluating the emotional state by analyzing tone and pitch. This results in the output of speech-based emotion labels such as "calm" or "tense."

[0646] Step 4:

[0647] The server integrates the emotional labels obtained from the visual and auditory data to assess the user's overall emotional state. This integrated emotional data is then used in the next step.

[0648] Step 5:

[0649] The server uses a generative AI model based on integrated sentiment data to create prompt messages. Specifically, it embeds the results of sentiment analysis into the prompt messages and provides them as input to generate optimal feedback for the user.

[0650] Step 6:

[0651] The generative AI model receives prompt messages generated by the server as input and generates advice and warnings to provide to the user. The output of this step is a personalized message tailored to the specific situation.

[0652] Step 7:

[0653] The server sends generated advice and warnings to the device, presenting them to the user as visual and audio feedback. The device plays this back and provides it to the user in real time, thereby improving the user's safety awareness.

[0654] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0655] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0656] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0657] [Fourth Embodiment]

[0658] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0659] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0660] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0661] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0662] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0663] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0664] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0665] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0666] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0667] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0668] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0669] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0670] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0671] This invention relates to a system using generative AI for receiving and automatically responding to nurse calls initiated by users in medical facilities. This system can reduce the burden on medical staff and improve user satisfaction.

[0672] First, when a user presses a nurse call button installed in the patient's room, the terminal receives the signal. The terminal sends this information to the server and instructs it to begin responding. The server processes the received data and uses generative AI to generate video and audio of a simulated nurse. This generation is based on past history and user preferences, providing personalized interaction.

[0673] The generated video and audio are sent back to the terminal and displayed on the user's screen. The user can converse with the generated nurse through an interactive interface. This interaction collects information about the user's requests and physical condition.

[0674] The server analyzes the user's conversation and data from connected sensor devices to determine the urgency and priority of the response. For example, if a user complains of chest pain, it is immediately determined to be an "emergency," and instructions are generated to request immediate assistance from medical personnel.

[0675] Instructions are fed back to the user via the terminal. In some cases, physical personnel may be dispatched or medical equipment may be provided. This ensures that the user receives appropriate support without any omissions or delays.

[0676] Furthermore, data collected during everyday conversations can be used to improve future services. For example, the results of users' regular health checks and the frequency with which specific caregiving activities are needed can be analyzed, enabling the provision of more personalized services. The introduction of this system is expected to improve efficiency in medical settings and enhance the quality of medical care provided.

[0677] The following describes the processing flow.

[0678] Step 1:

[0679] The user presses the nurse call button. The terminal detects the button press and receives the signal.

[0680] Step 2:

[0681] The terminal sends the received call signal to the server and generates a request to initiate a response.

[0682] Step 3:

[0683] The server receives the request and retrieves the user's ID, location information, and past history data from the database.

[0684] Step 4:

[0685] Based on the acquired data, the server uses a generation AI to create video and audio of a nurse character.

[0686] Step 5:

[0687] The server transmits the generated video and audio to the terminal, presenting the user with a virtual nurse.

[0688] Step 6:

[0689] The terminal displays the generated video and audio of the nurse to the user and initiates an interactive conversation.

[0690] Step 7:

[0691] The user converses with a nurse on the screen. The terminal continuously transmits the user's input and voice data to the server.

[0692] Step 8:

[0693] The server analyzes the conversation content and sensor data to determine the urgency and priority of the response.

[0694] Step 9:

[0695] The server generates necessary responses and instructions based on its judgment, and issues instructions regarding the dispatch of personnel and equipment.

[0696] Step 10:

[0697] The terminal provides feedback to the user based on instructions from the server, displaying them on the screen and via sound.

[0698] Step 11:

[0699] If necessary, staff will be dispatched to the user's location. The terminal reports the completion of the process to the server.

[0700] Step 12:

[0701] The server records the content of conversations and updates the database to help improve subsequent responses.

[0702] (Example 1)

[0703] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0704] Responding quickly and accurately to inquiries and requests from users in medical facilities is crucial for significantly reducing the burden on healthcare professionals and improving user satisfaction. However, current systems struggle to efficiently manage the enormous number of calls, and the time and resources available for providing appropriate individual responses are limited. This can lead to delays in responding to urgent cases, causing risks and dissatisfaction among users.

[0705] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0706] In this invention, the server includes a device for receiving communication signals transmitted by a user, a generation device for providing visual and auditory information generated using a generation AI model to a terminal device, and an analysis device for determining the urgency of the situation. This makes it possible to provide quickly generated individual responses to user inquiries based on multiple information resources, reducing the workload of medical professionals and improving the accuracy and speed of services to users.

[0707] A "communication signal" is electrical data that a user transmits from their terminal, indicating a call or request.

[0708] "Device" refers to an instrument or machine used to perform a specific function, and in this invention, it refers to a mechanism that receives communication signals and manipulates generated AI models.

[0709] A "generative AI model" is a form of artificial intelligence that constructs visual and auditory information based on past data and personal preferences.

[0710] "Visual and auditory information" refers to images displayed on a screen and sounds emitted from speakers, created by a generative AI model.

[0711] A "terminal device" is a device that serves as an interface with the user, displaying information and playing audio.

[0712] An "analysis device" is a device that has the function of analyzing and determining the urgency level using user information and data collected from sensors.

[0713] A "commanding device" refers to a mechanism that issues commands to workers or equipment to carry out specific actions based on analysis results.

[0714] An "information resource access device" is a device that has the function of accessing a database to generate an adapted response by referring to the user's past history and individual preferences.

[0715] This invention relates to a system for responding quickly and individually to inquiries from users in medical facilities. The following specifically describes embodiments for carrying out this invention.

[0716] This system uses terminals installed in patient rooms to receive communication signals transmitted by users and utilizes a server to process that data. The terminal functions as an interface for user operation, and when a nurse call button is pressed, it sends a signal to the server.

[0717] The server analyzes the received signal to determine the user's room number and the content of the call. Next, it uses a generative AI model to generate personalized visual and auditory information for each user. This information incorporates data reflecting past medical history and the user's specific preferences. Specifically, the server utilizes a cloud-based AI platform for data analysis and video / audio generation.

[0718] For example, if a user presses the nurse call button and requests water, the device sends the voice message to the server. The server, using a generative AI model, generates video and audio responses saying, "We'll bring you water right away," and sends them back to the device. The user can then see this interaction through the device's screen and speaker.

[0719] A concrete example of a prompt message would be, "The patient has pressed the nurse call button. Please generate a response." This allows the server to quickly take an appropriate action via the AI ​​model.

[0720] Furthermore, the server collects data from connected sensor devices to monitor the user's situation and health status. This data helps determine the urgency of a situation and the appropriate response, enabling efficient medical support.

[0721] This system is expected to significantly improve efficiency in healthcare settings and user satisfaction by enabling users to receive prompt and appropriate support and reducing the burden on healthcare professionals.

[0722] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0723] Step 1:

[0724] The user presses the nurse call button on the terminal. This action causes the terminal to generate a communication signal, which is then sent to the server as input data. Specifically, pressing the button causes the terminal to construct a data packet containing information such as the room number and the reason for the call. The output is the transmission of a signal to the server.

[0725] Step 2:

[0726] The server analyzes the signals received from the terminal. The input is a data packet containing the user's room number and requests. The server transforms this data into a prompt for the AI ​​model and inputs it into the generative AI model. The prompt might be, for example, "The patient says they want water. Generate an appropriate response." The output of this process is a template of the visual and auditory information that the generative AI model generates.

[0727] Step 3:

[0728] The server generates visual and auditory information using a generative AI model. The input is a generation template based on the prompts generated in step 2. The AI ​​considers past historical data and user preferences when generating video and audio based on this template. Specifically, the AI ​​model adjusts the tone of the voice and the facial expressions in the video to suit the user. The output is visual and auditory media files sent to the terminal.

[0729] Step 4:

[0730] The terminal receives visual and auditory information transmitted from the server and presents it to the user. The input is a generated media file, and the output is the provision of visual and auditory information to the user. Specifically, the terminal displays images on the screen and plays audio through the speaker, providing information to the user in an intuitive and easy-to-understand manner.

[0731] Step 5:

[0732] The user interacts with a nurse generated through the device. The input for this step is the user's voice and further requests. The output is the user's speech being sent to the server as text data. Specifically, the device's microphone captures the user's voice and converts it to text using speech recognition technology.

[0733] Step 6:

[0734] The server analyzes interaction data from the user. Inputs include text data sent from the terminal and additional information from sensor devices. The server analyzes this data to determine urgency. The results are expressed as work instructions and prioritized response protocols. Outputs are the next instructions and alerts sent to the terminal.

[0735] Step 7:

[0736] The terminal notifies the user of instructions from the server and adjusts physical responses as needed. Inputs are work instructions and alert information from the server. Outputs are notifications of appropriate actions to the user and medical staff. Specifically, the terminal displays emergency messages on the screen and issues alarms to medical staff if necessary.

[0737] (Application Example 1)

[0738] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0739] In industrial settings, there is a need to respond quickly and accurately to abnormal signals and worker requests while maintaining production efficiency and safety. However, conventional systems have struggled to make real-time decisions and respond to these requests. This has resulted in challenges such as disruptions to on-site work, production delays, and even safety risks.

[0740] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0741] In this invention, the server includes a device for receiving signals transmitted by a user, a generation device for providing generated visual media and sound to a terminal based on the received signals, a determination device for determining priority based on the generated visual media and sound and user information, an instruction device for quickly dispatching personnel or equipment when the determination device determines the situation to be important, and a device for receiving abnormal signals from equipment and work stations at an industrial site and responding as a virtual maintenance person using a generated AI model. This enables immediate judgment and response to abnormal situations, thereby improving the efficiency and safety of production at industrial sites.

[0742] A "signal" is a means of transmitting data or requests originating from a user or device.

[0743] A "device" is a machine or electronic device that has a specific function and is used to receive, generate, judge, or direct signals.

[0744] A "generative AI model" is a mathematical or algorithmic model that uses artificial intelligence technology to analyze signals and generate appropriate responses.

[0745] "Visual media" refers to digital content presented in the form of images or videos that contain visual information.

[0746] "Acoustics" refers to information that includes audio data and is used as sound for communication.

[0747] A "terminal" is an electronic device or part of a computer that a user interacts with directly.

[0748] A "decision-making device" is a device or software used to analyze the content of received data and determine its importance and the priority of actions to be taken.

[0749] A "commanding device" is a device that generates or transmits commands to initiate necessary actions or processes based on a judgment.

[0750] An "abnormal signal" is a signal that includes warnings or error messages indicating a condition that is different from the normal state.

[0751] A "virtual maintenance staff" is a maintenance staff system artificially created by a generative AI model, capable of handling problems without human intervention.

[0752] The system that realizes this invention aims to efficiently handle abnormalities in industrial settings. First, a terminal receives a signal transmitted by the user. The terminal sends this signal to a server and starts processing the response.

[0753] The server is equipped with a generative AI model that analyzes abnormal signals and user requests. The generated video and audio are sent back to the terminal as visual and auditory media. This allows users to interact with a virtual maintenance representative through the screen and audio, and quickly receive the necessary information and support.

[0754] The server analyzes the received data, and a determination device is activated to assess the urgency. If it is deemed critical, an instruction device is activated, issuing orders to dispatch personnel and equipment as needed. As a result, an appropriate response is quickly taken at the scene.

[0755] For example, if a worker reports that "the machine is making an unusual noise," the AI ​​model will predict the possibility of an anomaly and suggest appropriate countermeasures. This process is initiated by entering a prompt message into the user's terminal. An example of a prompt message is, "The machine is making an unusual noise. Please tell me how to deal with it."

[0756] The hardware used includes smartphones and head-mounted displays, while the software utilizes TensorFlow and OpenCV for signal processing and data analysis. This allows for the efficient execution of the process from signal reception to processing.

[0757] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0758] Step 1:

[0759] When a user sends a signal, the terminal receives that signal. The input is the call signal from the user, and the output is the digital data of the signal. This allows the terminal to recognize an anomaly or a request.

[0760] Step 2:

[0761] The terminal forwards the received signal to the server. The input is digital data of the signal, and the output is prompt data directed to the server. This data forms the basis for the next processing step.

[0762] Step 3:

[0763] The server inputs the received prompt data into the generating AI model. The input is prompt data, and the output is the analysis results. The generating AI model determines the nature and priority of the anomalies.

[0764] Step 4:

[0765] The server generates visual and auditory media for a virtual maintenance worker based on the analysis results of the generated AI model. The input is the analysis results from the AI ​​model, and the output is the generated visual and auditory media. This is then used as feedback to the user.

[0766] Step 5:

[0767] The generated visual and auditory media are sent to the terminal and presented to the user. The input is the generated visual and auditory media, and the output is video and audio for the user. This allows the user to interact with the system.

[0768] Step 6:

[0769] If the user provides further interaction, the server analyzes the data based on that interaction and generates the necessary instructions. The input is the additional interaction with the user, and the output is further instruction data. If necessary, instructions are given to quickly dispatch personnel and equipment to the site.

[0770] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0771] This invention relates to a generative AI system that combines an emotion engine to automatically and individually respond to nurse calls from users in medical facilities. This system can further improve user satisfaction by recognizing the emotional state of users and providing personalized responses accordingly.

[0772] When a user presses the nurse call button, the terminal receives the signal and sends the information to the server. Based on this information, the server uses generative AI to generate video and audio of a virtual nurse. Because this generation is based on the user's past history and preferences, personalized interaction is possible.

[0773] The server then analyzes the user's emotional state through an emotion engine. This engine uses voice analysis and facial recognition technology to determine whether the user is stressed, at ease, or confused. The analyzed emotional data is then sent to the server and used to adjust the conversation.

[0774] The generated video and audio are sent to the device, initiating interaction with the user. For example, if the user expresses feelings such as "I'm a little anxious, but I don't feel any major problems," the generating AI will respond calmly with something like, "I'll help you feel more at ease. Please let me know if there's anything specific I can do to help."

[0775] The terminal transmits the conversation with the user to the server in real time, and the server determines the urgency level based on the conversation content. Along with the emotional state, appropriate instructions are generated. For example, if the emotional analysis indicates a low level of urgency but the user is feeling anxious, additional counseling via voice call can be provided.

[0776] Another important function of this system is to improve the quality of care and medical treatment by tracking changes in users' emotions and conducting long-term emotional analysis. This allows for the provision of care tailored to each user and enhances their sense of security. In this way, the system is a powerful tool for reducing the burden on healthcare professionals and implementing comprehensive medical care.

[0777] The following describes the processing flow.

[0778] Step 1:

[0779] The user presses the nurse call button. The terminal receives this signal and prepares a request to the server.

[0780] Step 2:

[0781] The terminal sends call information to the server, initiating a processing request along with the user's ID and location information.

[0782] Step 3:

[0783] Based on the received data, the server generates video and audio of a virtual nurse tailored to the user's history and preferences.

[0784] Step 4:

[0785] The server activates the emotion engine, analyzes emotional data from the user's voice and video, and determines the user's emotional state.

[0786] Step 5:

[0787] The server sends the generated screen and emotion data to the terminal, initiating interaction with the user.

[0788] Step 6:

[0789] The terminal uses the received video and audio to display the conversation between the user and the virtual nurse, providing interactive support.

[0790] Step 7:

[0791] As the user converses with the virtual nurse, they express a variety of reactions and emotions. The device transmits the conversation data to the server in real time.

[0792] Step 8:

[0793] The server analyzes the conversation content and emotional data, determines the urgency level, and decides on a flexible response based on the emotions, if necessary.

[0794] Step 9:

[0795] Based on the assessment results, the server generates instruction data to dispatch personnel or provide additional remote support if necessary.

[0796] Step 10:

[0797] Based on instructions from the server, the terminal continuously provides feedback to the user and offers additional support to provide a sense of security.

[0798] Step 11:

[0799] Upon receiving a user response, the terminal reports the response result and sentiment data back to the server, completing the entire process.

[0800] Step 12:

[0801] The server records all conversation and sentiment data, and updates the database to help with future improvements and long-term analysis.

[0802] (Example 2)

[0803] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0804] In healthcare facilities, there is a need for a system that can appropriately respond to calls from users, understand their emotional state, and provide individualized interactions to increase user satisfaction, while also enabling rapid response to emergencies. Conventional systems have the problem of not being able to provide satisfactory service to users because they are not capable of analyzing emotional states or generating individualized responses.

[0805] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0806] In this invention, the server includes a receiving means, a generating means, an analysis means, an adjustment means, and an instruction means. This enables rapid analysis of the user's emotional state and the provision of personalized services. Specifically, it becomes possible to generate a virtual supporter that meets the needs of each user by utilizing a generation AI model and to provide interaction based on emotional analysis.

[0807] A "receiving means" is a function that detects signals transmitted by the user and incorporates them into the system.

[0808] The "generation means" is a function that generates information about a virtual supporter based on the signal information received by the receiving means and supplies this information to the information processing device.

[0809] The "analysis means" is a function that determines the user's emotional state by analyzing data such as the user's voice and video.

[0810] The "adjustment means" is a function that plays the role of generating individually adjusted information based on the results of the analysis means.

[0811] The "determination means" is a function that determines the urgency of the current situation based on generated information and user history information.

[0812] An "instruction mechanism" is a function that issues instructions to promptly take necessary measures based on the results of a determination mechanism.

[0813] A "recording means" is a function that sequentially collects user conversation data and generates information to track changes in emotional state based on this data.

[0814] The "response generation means" is a function that enables the generation of personalized responses based on the user's history data and emotional data.

[0815] This invention relates to a system for responding quickly and appropriately to user calls in medical facilities. The system involves multiple components, including a server, terminals, and users. The server utilizes a generative AI model to generate video and audio of a virtual supporter tailored to each user's needs. The generative AI model used is vendor-independent but implements state-of-the-art machine learning algorithms.

[0816] The terminal is responsible for receiving call signals from the user and transmitting them to the server. The server uses the received signals and past user data to perform voice analysis using an emotion engine and facial recognition technology. This analysis determines the user's emotional state, and personalized interactions are set accordingly. The server can also use general natural language processing techniques for voice analysis and existing image analysis techniques for facial recognition.

[0817] For example, Microsoft Azure's facial recognition API and Google Cloud's natural language processing API can be used. The generated video and audio are sent to the device, and a conversation with the user begins. During this interaction, the device continuously sends the conversation content to the server. The server determines the urgency based on the conversation content and emotional state, and generates appropriate instructions.

[0818] For example, if a user feels anxious, the system uses AI to provide a response such as, "We'll help you feel more at ease. Please let us know if there's anything specific we can do to help." Furthermore, this system tracks the user's emotional changes over the long term, providing data that can be used to improve medical or nursing care services.

[0819] An example of a prompt message would be: "Explain the process for generating an appropriate virtual caregiver based on past history information and sentiment analysis when a patient presses the nurse call button, and for optimizing the interaction with the patient." This allows the system to provide user-optimized medical care and reduce the burden on healthcare professionals.

[0820] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0821] Step 1:

[0822] When a user presses the nurse call button, the terminal receives the signal as input data. This signal is in digital format and contains the user's ID and location information. The terminal forwards this input to the server, which then sends a trigger to prepare for response.

[0823] Step 2:

[0824] The server generates video and audio of a virtual supporter based on signals received from the terminal, utilizing a generative AI model. Specifically, it searches a database for past history information and preferences identified by the user's ID and supplies this as input to the generative AI model. The generative AI model customizes individual interactions based on prompt messages and creates video and audio output.

[0825] Step 3:

[0826] The server sends the generated virtual supporter's video and audio data to the terminal. The terminal receives this data and displays the video and plays the audio for the user. The user experiences this interaction in real time.

[0827] Step 4:

[0828] During user interaction, the device sequentially collects the user's voice and facial expressions and sends them to the server as input data. The server analyzes this data using an emotion engine to determine the user's emotional state. At this time, stress levels and feelings of security are quantified using voice recognition and facial recognition technologies.

[0829] Step 5:

[0830] The server adjusts conversations and responses based on the analyzed emotional data. Specifically, if the user is feeling anxious, it uses a generative AI model to generate additional support messages and delivers them to the device.

[0831] Step 6:

[0832] The terminal sequentially sends the content of the interaction with the user to the server, which then processes the conversation to determine its urgency. The server integrates sentiment data with the conversation content, applies urgency determination logic, and determines the priority.

[0833] Step 7:

[0834] If the situation is deemed highly urgent, the server will instruct the terminal to take immediate action based on rules defined within the system. This includes instructing the dispatch of medical staff as needed and preparing relevant equipment.

[0835] (Application Example 2)

[0836] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0837] Conventional security systems have the challenge of being unable to adequately consider the emotional state of users and respond appropriately to diverse situations. Furthermore, they struggle to provide users with customized, real-time feedback for ensuring their safety. As a result, they are unable to implement optimal security measures and adequately alleviate user anxiety.

[0838] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0839] In this invention, the server includes means for receiving signals transmitted by a user, generation means for providing generated visual and auditory information to an information terminal based on the received signals, and analysis means for analyzing the user's emotional state. This makes it possible to analyze the user's emotions in real time and provide optimal visual and auditory feedback tailored to that situation.

[0840] "Receiving means" refers to devices or systems for receiving signals transmitted by users.

[0841] A "generation means" is a mechanism for providing information terminals with generated visual and auditory content based on received signals.

[0842] A "determination means" is a mechanism that determines priority based on generated visual and auditory information and user information.

[0843] "Instruction means" refers to a means of quickly dispatching personnel or equipment when the determination means determines that the matter has a high priority.

[0844] "Analysis methods" refer to processes and techniques for analyzing the emotional state of users.

[0845] "Adjustment means" refers to means that have the function of adjusting the response content based on the analysis results obtained by the analysis means.

[0846] A "recording system" is a mechanism that sequentially collects dialogue data with users and generates information for use in new responses.

[0847] A "database access means" is a means of obtaining the data necessary to provide a customized response through an information terminal.

[0848] This invention relates to a personal security system that utilizes emotion analysis and was developed with the aim of enhancing user safety. In this system, the user's device, specifically smart glasses, is used as the main interface.

[0849] The server first receives the user's visual and audio data transmitted from the smart glasses. This data is processed to analyze facial expressions and voice tone. Specifically, facial recognition technology is used for facial expressions, and voice analysis technology is used for voice. The hardware used includes the camera and microphone within the smart glasses. This makes it possible to evaluate the user's emotional state in real time.

[0850] Subsequently, the server uses a generative AI model based on the sentiment analysis results to generate a response tailored to the user. At this stage, specific advice and warnings are generated to alleviate the user's anxiety and fear. For example, if the user expresses the sentiment, "I've been feeling uneasy on the street lately," the generative AI will provide advice such as, "Pay attention to your surroundings and be wary of suspicious activity."

[0851] The generated advice and warnings are delivered to the user as audio and visual feedback through smart glasses, and the user can receive them in real time. A key feature of this system is that the responses are personalized to the user and optimized according to the user's emotional state.

[0852] An example of a prompt would be: "Sentiment analysis result: User's sentiment data Voice analysis result: User's voice text Provide the best advice for this user." By inputting this prompt into the generating AI model, a customized response tailored to the user can be obtained.

[0853] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0854] Step 1:

[0855] The server receives visual data (images) and audio data transmitted from the terminal (smart glasses). This input data serves as material for analyzing the user's emotional state. Upon receiving this data, the server sends the visual and audio data to separate analysis processes.

[0856] Step 2:

[0857] The server uses visual data to perform facial recognition technology for expression analysis. Specifically, it extracts facial features from the image and determines the user's emotional state. The output obtained here is an emotion label such as "reassured" or "anxious."

[0858] Step 3:

[0859] Simultaneously, the server applies speech analysis technology to the audio data, evaluating the emotional state by analyzing tone and pitch. This results in the output of speech-based emotion labels such as "calm" or "tense."

[0860] Step 4:

[0861] The server integrates the emotional labels obtained from the visual and auditory data to assess the user's overall emotional state. This integrated emotional data is then used in the next step.

[0862] Step 5:

[0863] The server uses a generative AI model based on integrated sentiment data to create prompt messages. Specifically, it embeds the results of sentiment analysis into the prompt messages and provides them as input to generate optimal feedback for the user.

[0864] Step 6:

[0865] The generative AI model receives prompt messages generated by the server as input and generates advice and warnings to provide to the user. The output of this step is a personalized message tailored to the specific situation.

[0866] Step 7:

[0867] The server sends generated advice and warnings to the device, presenting them to the user as visual and audio feedback. The device plays this back and provides it to the user in real time, thereby improving the user's safety awareness.

[0868] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0869] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0870] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0871] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0872] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0873] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0874] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0875] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0876] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0877] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0878] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0879] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0880] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0881] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0882] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0883] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0884] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0885] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0886] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0887] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0888] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0889] The following is further disclosed regarding the embodiments described above.

[0890] (Claim 1)

[0891] A receiving means for receiving call signals transmitted by a user,

[0892] A generation means that provides generated video and audio to a user terminal based on the received signal,

[0893] A determination means for determining the priority of urgency based on the generated video and audio and user information,

[0894] If the aforementioned determination means determines that it is an emergency, the instruction means dispatches personnel or equipment promptly,

[0895] A system that includes this.

[0896] (Claim 2)

[0897] The system according to claim 1, comprising recording means for sequentially collecting conversation data with users and generating information for use in subsequent interactions.

[0898] (Claim 3)

[0899] The system according to claim 1, comprising means for accessing a database to generate a customized response based on user preferences.

[0900] "Example 1"

[0901] (Claim 1)

[0902] A device that receives communication signals transmitted by a user,

[0903] A generation device that provides a terminal device with visual and auditory information generated using a generation AI model based on the received signal,

[0904] An analysis device that determines the degree of urgency based on the generated visual and auditory information, user information, and data acquired from multiple sensors,

[0905] An instruction device that, when the aforementioned analytical device determines that an emergency has occurred, promptly dispatches a worker or equipment,

[0906] A system that includes this.

[0907] (Claim 2)

[0908] The system according to claim 1, comprising a recording device that sequentially collects user interaction data and generates information for use in subsequent interactions.

[0909] (Claim 3)

[0910] The system according to claim 1, comprising an information resource access device for generating an adapted response based on the user's past history and individual preferences.

[0911] "Application Example 1"

[0912] (Claim 1)

[0913] A device that receives signals transmitted by users,

[0914] A generating device that provides generated visual media and sound to a terminal based on the received signal,

[0915] A determination device that determines the priority of importance based on the generated visual media and sound and user information,

[0916] An instruction device that promptly dispatches personnel or equipment when the aforementioned determination device determines that the situation is important,

[0917] A device that receives abnormal signals from equipment and work stations in an industrial setting and uses a generated AI model to act as a virtual maintenance worker,

[0918] A system that includes this.

[0919] (Claim 2)

[0920] The system according to claim 1, comprising a recording device that continuously collects user interaction data and generates information for use in subsequent interactions.

[0921] (Claim 3)

[0922] The system according to claim 1, comprising a database access device for generating personalized responses based on user preferences.

[0923] "Example 2 of combining an emotion engine"

[0924] (Claim 1)

[0925] A receiving means for receiving signals transmitted by a user,

[0926] A generation means that provides information about a virtual supporter generated based on a received signal to an information processing device,

[0927] An analytical means for analyzing the emotional state of the receiving user,

[0928] An adjustment means that generates individually adjusted information based on the results of the analysis means,

[0929] A determination means for determining the urgency based on the generated information and the user's history information,

[0930] If the determination means determines that it is an emergency, an instruction means is provided to promptly instruct measures to be taken.

[0931] A system that includes this.

[0932] (Claim 2)

[0933] The system according to claim 1, comprising recording means for sequentially collecting user conversation data and generating data for tracking emotional states.

[0934] (Claim 3)

[0935] The system according to claim 1, comprising a response generation means for generating individually customized responses based on user history and sentiment data.

[0936] "Application example 2 when combining with an emotional engine"

[0937] (Claim 1)

[0938] A means of receiving signals transmitted by users,

[0939] A generation means that provides generated visual and auditory information to an information terminal based on the received signal,

[0940] A determination means for determining priority based on the generated visual and auditory information and user information,

[0941] If the determination means determines that the priority is high, an instruction means for quickly dispatching personnel or equipment,

[0942] Analytical means for analyzing the emotional state of users,

[0943] An adjustment means for adjusting the response content based on the analysis results from the aforementioned analysis means,

[0944] A system that includes this.

[0945] (Claim 2)

[0946] The system according to claim 1, comprising recording means for sequentially collecting user interaction data and generating information for use in new responses.

[0947] (Claim 3)

[0948] The system according to claim 1, comprising means for accessing a database to provide a customized response via an information terminal. [Explanation of symbols]

[0949] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A receiving means for receiving call signals transmitted by a user, A generation means that provides generated video and audio to a user terminal based on the received signal, A determination means for determining the priority of urgency based on the generated video and audio and user information, If the aforementioned determination means determines that it is an emergency, the instruction means dispatches personnel or equipment promptly, A system that includes this.

2. The system according to claim 1, comprising recording means for sequentially collecting conversation data with users and generating information for use in subsequent interactions.

3. The system according to claim 1, comprising means for accessing a database to generate a customized response based on user preferences.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A