system
A system using generative AI and speech-to-text technology addresses loneliness and risk assessment for elderly individuals, enabling real-time health and crime risk detection and prompt service provision.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Elderly individuals living alone often experience loneliness and are at high risk for health problems and crime, with existing systems failing to address these issues effectively.
A system utilizing a communication terminal, generative artificial intelligence, speech-to-text technology, and log recording to assess health and crime risks in real-time, and transfer calls to appropriate service providers as needed.
Reduces loneliness, detects health and crime risks early, and provides prompt appropriate services to elderly individuals.
Smart Images

Figure 2026037131000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] It is common for elderly people living alone to feel lonely on a daily basis. In an effort to alleviate this loneliness, they often call call centers periodically, which increases the burden on the call centers. Another serious issue is that elderly people living alone are at high risk of health problems and crime. There is a need for methods to detect and address these risks early. Therefore, it is necessary to reduce the loneliness of elderly people living alone while assessing their health status and crime risk in real time and taking appropriate measures as needed. [Means for solving the problem]
[0005] To solve this problem, we provide a system that includes a means for a user to communicate using a communication terminal, a means for analyzing the conversation with the user using generative artificial intelligence to evaluate the user's health condition or criminal risk, and a means for transferring the communication to a service provider's operator as needed based on the evaluation. Furthermore, the system includes a means for acquiring the content of the conversation as text data using speech-to-text conversion technology and a means for recording the analysis results and call logs. This makes it possible to reduce loneliness among elderly people living alone, detect health and criminal risks early, and provide appropriate services.
[0006] A "communication terminal" is a device that allows a user to communicate voice and data, and includes telephones, smartphones, computers, and the like.
[0007] "Generative AI" is an AI technology that has the ability to analyze conversations with users and generate appropriate responses.
[0008] "Health Status" refers to information about a user's physical and mental health, used to assess specific symptoms and health risks.
[0009] "Crime risk" is information that indicates the possibility or signs that a user may become involved in a crime.
[0010] A "service provider operator" is a person in charge of a provider that provides a specific service, who takes appropriate measures to address user problems and risks.
[0011] "Speech-to-text technology" is a technology that converts voice data into text data, and is also called voice recognition technology.
[0012] A "log" is data that records information such as conversation content, analysis results, and system operation, and is used for later verification and analysis. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate operator of a service provider as necessary. Specific operations and processing of an embodiment of this system will be described below.
[0035] Specific Embodiments of the System
[0036] A user makes a call using a communication terminal
[0037] A user calls the system using a communication terminal (for example, a telephone or a smartphone). The communication terminal starts a voice call and is connected to the server.
[0038] The server starts the audio service
[0039] When the server receives a call from a user, it activates a voice service using generative artificial intelligence. The server starts a conversation by saying to the user, "Hello, is there anything I can help you with?"
[0040] Conversation initiation and analysis
[0041] The server converts the user's responses into text data in real time using speech-to-text technology. Generative AI then analyzes this text data to assess the user's health status and criminal risk.
[0042] Specific examples
[0043] For example, if a user says, "I've been having headaches lately," the server will analyze the text data containing the keyword "headache" and evaluate it as a health risk. Similarly, if a user says, "I saw a suspicious person," the server will evaluate the crime risk from the keyword "suspicious person."
[0044] Risk assessment and call transfer to an operator
[0045] The server transfers calls to service provider operators if necessary based on the health or crime risk assessed by the generative AI. For example, if the health risk is determined to be high, the call is transferred to a HELPO operator, and if the crime risk is determined to be high, the call is transferred to a security service operator.
[0046] Logging the results
[0047] The server records all conversations, their analysis results, and whether or not the call was transferred in detail as a log, which can be used for future analysis and verification when a problem occurs.
[0048] In this way, the system of the present invention is a system that reduces the user's sense of loneliness, detects health conditions and crime risks early, and takes necessary measures promptly. The objectives of the present invention can be achieved by appropriately arranging communication terminals, generative artificial intelligence, speech-to-text conversion technology, and log recording means to implement the invention.
[0049] The processing flow will be explained below.
[0050] Step 1:
[0051] A user makes a call using a communication terminal.
[0052] The communication terminal initiates a voice call and is connected to the server.
[0053] Step 2:
[0054] The server receives a call from a user.
[0055] The server launches a voice service using generative artificial intelligence and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0056] Step 3:
[0057] The server waits for the user's response.
[0058] The communication terminal picks up the user's voice and transmits it to the server.
[0059] Step 4:
[0060] The server converts the received voice data into text data using voice-to-text conversion technology.
[0061] The server analyzes the text data using generative artificial intelligence.
[0062] Step 5:
[0063] The server assesses the user's health and criminal risk.
[0064] Generative artificial intelligence detects keywords such as "headache," "dizziness," "suspicious person," and "fraud," and determines health or criminal risks based on these.
[0065] Step 6:
[0066] Based on the risk assessment result, the server prepares to transfer the call to an operator of the service provider as necessary.
[0067] If the server is determined to pose a high health risk, it will contact a HELPO operator, and if it is determined to pose a high crime risk, it will contact a security service operator.
[0068] Step 7:
[0069] The server notifies the user that "Important information has been detected and we will connect you to an operator."
[0070] The server then transfers the call to the appropriate operator.
[0071] Step 8:
[0072] The server logs all conversations, their analysis results, and whether or not the call was forwarded.
[0073] The server records logs to allow for future analysis and verification when problems occur.
[0074] Through this procedure, the server, communication terminal, and user cooperate with each other to realize the system of the present invention.
[0075] Example 1
[0076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0077] It is desirable to provide a means to reduce loneliness among elderly people living alone, to detect health conditions and crime risks early, and to take appropriate measures promptly. There is also a need for a system that can effectively analyze the content of conversations with users and transfer calls to the necessary service providers based on the results.
[0078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0079] In this invention, the server includes a means for a user to communicate using an information terminal, a means for analyzing the conversation with the user using generative artificial intelligence to evaluate the user's health condition or criminal risk, a means for transferring the communication to a service provider's operator as needed based on the evaluation, a means for acquiring the user's conversation content as text data using speech-to-text technology, and a means for recording the conversation content and analysis results in a log. This reduces the sense of loneliness felt by elderly people living alone, enables early detection of health conditions and criminal risk, and allows for prompt response. Furthermore, the server can evaluate risk based on the user's speech and quickly transfer the call to an appropriate service provider.
[0080] "Information terminal" refers to any device used for communication, and specifically includes telephones, smartphones, computers, etc.
[0081] "Generative AI" refers to AI technology that analyzes conversations with users and makes decisions and suggestions based on that analysis.
[0082] "Health status" refers to general information about the user's health, such as the user's physical condition and whether or not the user has any illnesses.
[0083] "Crime risk" refers to the possibility that a user will become involved in a crime or the risk that a crime will occur.
[0084] "Service provider" means the operator of a business or organisation that provides health and safety services.
[0085] "Speech-to-text technology" refers to technology that converts voice data into text data in real time.
[0086] "Log" refers to records of conversation content, analysis results, communication transfer history, etc.
[0087] "Conversation content" refers to the content of the voice communication exchanged between the user and the system (or generative artificial intelligence).
[0088] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an operator of an appropriate service provider as necessary. Specific embodiments of the present invention will be described below.
[0089] First, a user calls a specific phone number using an information terminal (e.g., a telephone or smartphone). When the voice call begins, the terminal connects to the server. When the server receives the call from the user, it immediately answers and launches a voice-enabled application using generative artificial intelligence. Specifically, the server speaks to the user, saying, "Hello, is there anything I can help you with?"
[0090] The server then uses speech-to-text technology (e.g., Google® Cloud Speech-to-Text API) to convert the user's speech into text data in real time. Once the user's speech is converted into text data, the text data is stored on the server.
[0091] The server analyzes the text data using generative artificial intelligence (for example, OpenAI's GPT-4 model). This analysis evaluates the user's health condition and criminal risk. For example, if a user says, "I've had a constant headache lately," the server extracts the keyword "headache" and determines that this corresponds to a health risk. Similarly, if a user says, "I saw a suspicious person," the keyword "suspicious person" will evaluate the criminal risk as high.
[0092] Based on the evaluation results, the server transfers the call to a service provider operator as needed. For example, if the call is judged to be a high health risk, the server transfers the call to a health support service operator, and if the call is judged to be a high crime risk, the server transfers the call to a security service operator.
[0093] In addition, the server records detailed logs of all conversations, analysis results, and whether or not a call was forwarded, which can be used for future analysis and verification if a problem occurs.
[0094] Prompt Sentence Examples
[0095] Prompt: "If a user says, 'I've been having a lot of headaches lately,' explain how you can use a generative AI model to assess health risks and route the call to a health support service operator."
[0096] Prompt: "When a user says 'I see a suspicious person,' explain how a generative AI model can be used to assess the crime risk and route the call to a security service operator."
[0097] As described above, the system of the present invention can reduce the user's sense of loneliness, detect health conditions and crime risks early, and take necessary measures promptly.
[0098] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0099] Step 1:
[0100] A user uses an information terminal to call a specified phone number. When the voice call begins, the terminal connects to the server. The input is the user's voice, and the output is a connection with the server. Specifically, the user operates a telephone or smartphone to call the phone number specified by the system.
[0101] Step 2:
[0102] The server receives a call and activates a voice service. The input is an incoming call signal from the user, and the output is a voice-enabled application powered by generative artificial intelligence. Specifically, the server detects the incoming call, activates an auto-answer function, and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0103] Step 3:
[0104] The server uses speech-to-text technology to convert the user's speech into text data in real time. The input is the user's speech, and the output is text data. Specifically, the server uses the Google Cloud Speech-to-Text API to convert the speech into text data.
[0105] Step 4:
[0106] The server analyzes the text data using generative artificial intelligence. The input is text data, and the output is an assessment result regarding the user's health status and crime risk. Specifically, the server uses generative artificial intelligence (e.g., GPT-4) to analyze the text data and perform a risk assessment.
[0107] Step 5:
[0108] The server transfers the call to a service provider's operator as needed based on the risk assessment results. The input is the risk assessment result, and the output is call transfer. As a specific example, if the health risk is determined to be high, the server transfers the call to a health support operator, and if the crime risk is determined to be high, the server transfers the call to a security service operator.
[0109] Step 6:
[0110] The server records all conversation content, analysis results, and whether or not the call was forwarded in detail as a log. The input is the conversation content with the user and the analysis results, and the output is a log file. Specifically, the server processes the conversation content, analysis results, and call history in a recording database.
[0111] Through the above processing steps, the system of the present invention can reduce the user's sense of loneliness, detect health conditions and crime risks early, and take necessary measures promptly.
[0112] (Application example 1)
[0113] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0114] As the elderly population increases, feelings of loneliness, health risks, and crime risks among those living alone are becoming social issues. Elderly people living alone face particular challenges, such as difficulty in responding immediately to health problems or crime risks. As a result, there is a need for a system that can reduce feelings of loneliness while ensuring their safety and peace of mind. To solve this issue, it is important to assess the elderly's situation in real time via voice calls and quickly connect them to appropriate services as needed.
[0115] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0116] In this invention, the server includes: means for a user to use a communication terminal to perform predetermined communications; means for analyzing conversations with the user using generative artificial intelligence to assess the user's health status or criminal risk; means for transferring communications to a service provider operator as needed based on the assessment; means for converting the content of the voice call into text data in real time using speech-to-text technology; means for assessing the risk of exceeding a predetermined threshold based on the analysis results; means for quickly transferring the call to an appropriate operator if the risk exceeds a specific threshold; and means for recording all conversation contents, analysis results, and call transfer history. This allows for rapid detection of potential risks faced by users and enables prompt response.
[0117] A "communication terminal" is a device that a user uses to perform predetermined communications, and includes telephones and smartphones.
[0118] "Generative AI" is an AI technology that analyzes voice and text data to assess a user's health status and crime risk.
[0119] "Speech-to-text technology" is a technology that converts voice data into text data in real time.
[0120] "Evaluation" refers to the act of using generative artificial intelligence to analyze the content of a user's conversations and determine their health status and criminal risk.
[0121] A "service provider operator" is a specialized service person who takes action when a user is assessed as being at high risk for health or crime.
[0122] A "threshold" is a reference value that is set when the evaluation results of the generative artificial intelligence exceed a specific risk level.
[0123] "Call forwarding" is the process of quickly connecting a communication taken from a user to an operator of another service provider.
[0124] A "log" is a detailed record of conversation content, analysis results, call forwarding history, etc.
[0125] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate operator of a service provider as necessary. Specific operations and processing of an embodiment of this system will be described below.
[0126] System Hardware and Software
[0127] The following hardware and software are used to implement this system:
[0128] Hardware:
[0129] Smartphone microphone and speaker
[0130] Networked Servers
[0131] software:
[0132] Python
[0133] speech_recognition library
[0134] The requests library
[0135] Generative AI model
[0136] System operation explanation
[0137] 1. Start a voice call:
[0138] A user starts the application on their smartphone and taps the "emergency call" button, which establishes a real-time voice call between the user and the server.
[0139] 2. Speech recognition and text conversion:
[0140] The voice captured by the smartphone microphone is converted into text data in real time using the speech_recognition library, and the user's speech is then sent to the server as text data.
[0141] 3. Analysis by generative artificial intelligence:
[0142] The server uses a generative artificial intelligence model to analyze the acquired text data. This analysis evaluates the user's health condition and crime risk. For example, if a user says, "I've been having headaches lately," the keyword "headache" is analyzed and evaluated as a health risk. Similarly, if a user says, "I saw a suspicious person," the keyword "suspicious person" is used to evaluate the crime risk.
[0143] 4. Risk Assessment and Call Forwarding:
[0144] Based on the analysis results, the server evaluates whether the risk exceeds a certain threshold. If the risk is high, the generative AI model quickly routes the call to the appropriate service operator based on the risk assessment. For example, if the call is judged to be a high health risk, it will be routed to a medical service operator, and if the call is judged to be a high crime risk, it will be routed to a security service operator.
[0145] 5. Logging:
[0146] The server records all conversations, their analysis results, and call forwarding history in detail, and stores them in protected cloud storage for future analysis and verification in case of problems.
[0147] Specific prompt examples
[0148] Examples:
[0149] If a user says, "I'm scared because I've seen a lot of suspicious people walking down the street lately," the system will detect the phrase "suspicious people" and evaluate them as a crime risk, immediately transferring the call to a security service provider.
[0150] Example prompt sentence:
[0151] Text: "Recently, I've been seeing a lot of suspicious people walking down the street at night and it's scary."
[0152] Prompt: "Analyze this text to assess its health and crime risks."
[0153] In this way, the system provides a practical solution that ensures the safety and security of elderly people living alone, while also allowing for a rapid response.
[0154] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0155] Step 1:
[0156] A user starts an application using a communication terminal and taps the "emergency call" button. This operation causes the communication terminal to establish a voice call with the server. The user's voice is acquired as input, and a call with the server is initiated as output.
[0157] Step 2:
[0158] The server receives the voice data acquired from the communication terminal. The voice data is sent as input and recorded on the server as output.
[0159] Step 3:
[0160] The server uses the speech_recognition library to convert the received voice data into text data in real time. Voice data is used as input and text data is generated as output. Specifically, the server analyzes each phoneme using a speech recognition engine and generates the corresponding string.
[0161] Step 4:
[0162] The server inputs the generated text data into a generative artificial intelligence (AI) model to evaluate the user's health status and crime risk. The text data is used as input, and the health and crime risk assessment results are obtained as output. Specifically, the AI model uses natural language processing technology to analyze keywords and context within the text and conduct a risk assessment.
[0163] Step 5:
[0164] The server determines whether the risk exceeds a certain threshold based on the assessment results. The risk assessment results are used as input, and a judgment of whether the risk is high or low is obtained as output. Specifically, the risk level is determined by comparing it with a predefined threshold.
[0165] Step 6:
[0166] If the server determines that the call is high risk, it transfers the call to an operator at the appropriate service provider. The server uses the risk assessment results and threshold judgment results as inputs, and executes the call transfer as output. Specifically, this includes the process of rerouting the call to a specific phone number.
[0167] Step 7:
[0168] The server records all conversation content, assessment results, and call forwarding history in detailed logs and stores them in protected cloud storage. Assessment results and call history are used as input, and detailed log data is generated as output. Specifically, information such as date and time, conversation content, risk assessment results, and forwarding destinations is written to the log file.
[0169] In this way, the system assesses the situation of elderly people living alone in real time and takes prompt action as needed, providing peace of mind and safety.
[0170] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0171] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate service provider operator as needed. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and generates responses based on those emotions, more personalized responses are possible for users.
[0172] Specific Embodiments of the System
[0173] A user makes a call using a communication terminal
[0174] A user calls the system using a communication terminal (for example, a telephone or a smartphone). The communication terminal starts a voice call and is connected to the server.
[0175] The server starts the audio service
[0176] When the server receives a call from a user, it activates a voice service using generative artificial intelligence. The server starts a conversation by saying to the user, "Hello, is there anything I can help you with?"
[0177] Conversation initiation and analysis
[0178] The server converts the user's responses into text data in real time using speech-to-text technology. The generative AI analyzes this text data and assesses the user's health status and criminal risk. At the same time, the emotion engine recognizes emotions from the user's voice and provides this information to the generative AI.
[0179] Specific examples
[0180] For example, if a user says, "I've been having headaches lately," the server analyzes text data containing the keyword "headache" and evaluates it as a health risk. Meanwhile, the emotion engine detects emotions such as anxiety and stress from the user's voice. Similarly, if a user says, "I saw a suspicious person," the server evaluates the crime risk from the keyword "suspicious person," and the emotion engine recognizes fear and tension.
[0181] Generating responses based on emotion information
[0182] The server uses generative AI to generate an appropriate response based on the user's emotional information provided by the emotion engine. For example, if a user is feeling anxious, the server will respond with "Don't worry, we'll take care of it right away."
[0183] Risk assessment and call transfer to an operator
[0184] The server transfers the call to a service provider operator if necessary based on the health or criminal risk assessed by the generative artificial intelligence. For example, if the health risk is determined to be high, the call is transferred to an appropriate medical service operator, and if the criminal risk is determined to be high, the call is transferred to a security service operator.
[0185] Logging the results
[0186] The server logs all conversations, emotional information, analysis results, and whether or not a call was forwarded. These logs are used for future analysis and verification when problems occur.
[0187] summary
[0188] In this way, the system of the present invention reduces the user's sense of loneliness, detects health conditions and crime risks early, and promptly provides appropriate services. Furthermore, incorporating an emotion engine enables more compassionate and personalized responses to users. The objectives of the present invention can be achieved by appropriately arranging a communication terminal, generative artificial intelligence, speech-to-text technology, an emotion engine, and a means for logging in order to implement the invention.
[0189] The processing flow will be explained below.
[0190] Step 1:
[0191] A user makes a call using a communication terminal.
[0192] The communication terminal initiates a voice call and is connected to the server.
[0193] Step 2:
[0194] The server receives a call from a user.
[0195] The server launches a voice service using generative artificial intelligence and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0196] Step 3:
[0197] The server waits for the user's response.
[0198] The communication terminal picks up the user's voice and transmits it to the server.
[0199] Step 4:
[0200] The server converts the received voice data into text data using voice-to-text conversion technology.
[0201] The server analyzes the text data using generative artificial intelligence.
[0202] Step 5:
[0203] At the same time, the server uses an emotion engine to recognize emotions from the user's voice.
[0204] The emotion engine provides the recognized emotion information to the generative artificial intelligence.
[0205] Step 6:
[0206] The server evaluates the user's health status and crime risk based on text data and emotional information analyzed by generative artificial intelligence.
[0207] For example, if a user says, "I've been having headaches lately," the server detects the keyword "headache," and the emotion engine recognizes the emotion of anxiety.
[0208] Step 7:
[0209] The server uses generative artificial intelligence to generate an appropriate response based on emotional information.
[0210] For example, a user who is feeling anxious can be given a response such as "Don't worry, we'll take care of it right away."
[0211] Step 8:
[0212] Based on the evaluation results, the server transfers the call to an operator at the service provider as necessary.
[0213] If the health risk is determined to be high, the call will be transferred to a medical service operator, and if the crime risk is determined to be high, the call will be transferred to a security service operator.
[0214] Step 9:
[0215] The server performs the call transfer and notifies the user that "Important content has been detected and you will be connected to an operator."
[0216] The server then transfers the call to the appropriate operator.
[0217] Step 10:
[0218] The server records all conversation content, analysis results, emotional information, and whether or not the call was forwarded as a log.
[0219] This record will be used for future analysis and verification when problems occur.
[0220] This specific procedure enables the server, communication terminal, and user to work together, reducing the user's sense of loneliness, detecting health conditions and crime risks early, and taking necessary measures quickly.
[0221] Example 2
[0222] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0223] In modern society, elderly people living alone are increasingly experiencing loneliness, health problems, and crime risks, necessitating rapid and appropriate responses. However, conventional monitoring and warning systems have difficulty accurately grasping a user's emotional state and providing personalized responses based on that information. Furthermore, they lack the ability to adequately assess risk in real time and quickly transfer calls to service providers. To address these challenges, a system is needed that can accurately assess a user's emotions, health risks, and crime risks, and provide appropriate responses.
[0224] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0225] In this invention, the server includes: means for allowing a user to communicate using a communication terminal; means for analyzing a conversation with the user and assessing their health condition or criminal risk using generative artificial intelligence; means for converting the user's voice data into text data using speech-to-text technology; means including an emotion engine for evaluating the user's emotional state based on the analysis results; means for transferring the communication to a service provider operator as needed based on the evaluation; and means for recording the content of the conversation, emotional information, analysis results, and whether or not the call was transferred. This allows for accurate evaluation of the user's emotions and risk state in real time, enabling prompt and appropriate responses.
[0226] "Communication terminal" refers to any device that a user uses to conduct voice or data communications, and specifically includes smartphones and landline phones.
[0227] "Generative AI" refers to AI technology that analyzes user input and generates appropriate responses and analytical results, and includes systems that utilize natural language processing and machine learning.
[0228] "Speech-to-text technology" refers to technology that converts a user's voice into text data in real time, and includes, for example, cloud-based voice recognition services.
[0229] "Emotion engine" refers to technology that analyzes a user's emotional state from their voice or text data and provides that information, and includes systems that utilize emotion recognition algorithms and machine learning.
[0230] "Service provider operators" refer to professional personnel who are on standby to provide a particular service, including those working in fields such as healthcare and security.
[0231] "Database" means a system of record for storing and managing system-generated data, including the means for efficiently storing and accessing data in various formats.
[0232] "Log" refers to historical data that records the processes performed by the system, the data generated, events, etc., and is used for future analysis and troubleshooting.
[0233] This invention relates to a system that reduces loneliness among elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to service provider operators as needed. This system uses generative artificial intelligence to analyze user conversations and utilizes speech-to-text technology and an emotion engine to provide more personalized responses.
[0234] A user calls the system using a communication terminal. The communication terminal can be a landline phone or a smartphone. When the user makes a call, the server receives the call and activates a voice service using generative artificial intelligence. For example, the server may play a standard phrase such as "Hello, is there anything I can help you with?". At this stage, the server activates a voice playback module and provides the user with an initial response.
[0235] The server then receives the user's response and converts it into text data in real time using speech-to-text technology (e.g., Google Cloud Speech-to-Text API). The server then inputs this text data into a generative artificial intelligence (e.g., OpenAI GPT-4) to assess the user's health status and criminal risk. Based on the analysis results, the server then uses an emotion engine (e.g., IBM Watson® Tone Analyzer) to assess the user's emotional state.
[0236] For example, if a user says, "I've been having headaches lately," the server will send text data containing the keyword "headache" to the generative AI to assess whether there is a health risk. At the same time, the emotion engine will detect anxiety and stress from the user's voice. This will complement the user's emotional data, allowing the system to make a more accurate risk assessment.
[0237] The server integrates data from the generative artificial intelligence and emotion engine to generate an appropriate response for the user. For example, if a user feels anxious, the server generates a response such as "Don't worry, we'll take care of it right away," and plays it back to the user using the voice playback module.
[0238] If the server determines that the user faces a high health or criminal risk as a result of the risk assessment, it will transfer the call to a service provider operator as necessary. For example, if the user is assessed as having a high health risk, it will transfer the call to an operator of an appropriate medical service. The server uses the call transfer function to ensure a prompt response.
[0239] The server also records all conversation content, emotional information, analysis results, and whether or not the call was transferred in a database, which can be used for future analysis and troubleshooting.
[0240] As a concrete example, by inputting the following prompt sentence into a generative artificial intelligence, an appropriate response can be obtained.
[0241] Prompt: "The user mentions that they've been having headaches lately. Use this information to generate an appropriate response. They also sound anxious and stressed."
[0242] Example response: "I'm sorry to hear that you're concerned. Your persistent headache is concerning. Please wait a moment while we connect you to medical services."
[0243] As described above, the system of the present invention can accurately evaluate the user's emotions and risk status in real time and provide prompt and appropriate responses. This system can reduce the sense of loneliness felt by elderly people living alone and enable prompt responses to health conditions and crime risks.
[0244] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0245] Step 1:
[0246] A user calls the system using a communication terminal. The input includes the user's voice, and the output is a connection to the server. Specifically, the user calls the system's phone number using a smartphone or landline phone and presses the call button.
[0247] Step 2:
[0248] The server invokes a voice service, which includes as input the user's call connection and as output the playback of a voice message. Specifically, the server detects when the call is answered and invokes the voice playback module to play a voice message saying, "Hello, how can I help you?"
[0249] Step 3:
[0250] The server converts the user's speech into text data. The input includes the user's speech response, and the converted text data is obtained as the output. Specifically, the server uses speech-to-text conversion technology (e.g., Google Cloud Speech-to-Text API) to convert the user's speech into text data in real time.
[0251] Step 4:
[0252] The server analyzes the text data using generative AI. The converted text data is included as input, and the analysis results are obtained as output. Specifically, the server sends the text data to a generative AI (e.g., OpenAI GPT-4), which then evaluates health and crime risks.
[0253] Step 5:
[0254] The server uses an emotion engine to recognize the user's emotion. The input includes the analysis results and voice data, and the output is emotion information. Specifically, the server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to identify emotions from the user's tone of voice and text.
[0255] Step 6:
[0256] The server generates an appropriate response based on the integrated analysis results. The input includes the analysis results and emotional information, and the generated response is obtained as the output. Specifically, the server provides prompts to the generative AI based on the analysis results and emotional information, converts the generated response into speech, and sends it to the user.
[0257] Step 7:
[0258] The server performs risk assessment and transfers the call to an operator if necessary. The input contains the integrated analysis results, and the output is the required call transfer. Specifically, the server checks the risk assessment and, for example, if it determines that the health risk is high, transfers the call to a medical service operator.
[0259] Step 8:
[0260] The server records all conversation content, emotional information, analysis results, and whether or not the call was forwarded as a log. All communication data is included as input, and log data is obtained as output. Specifically, the server records and saves the conversation content, emotional information, analysis results, and whether or not the call was forwarded in a database.
[0261] (Application example 2)
[0262] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0263] There is a need to reduce the sense of loneliness felt by elderly people living alone, and to provide real-time assessments of their health status and crime risk to respond promptly. However, conventional systems are unable to recognize users' emotions and provide personalized responses, and there are also problems with the speed and accuracy of risk assessments. Therefore, a system that improves the quality of life of elderly people living alone and enables appropriate emergency response is needed.
[0264] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0265] In this invention, the server includes means for allowing a user to communicate using a communication terminal, means for analyzing the conversation with the user using a generative artificial intelligence to assess the user's health condition or criminal risk, means for transferring the communication to a service provider operator as needed based on the assessment, means for converting the conversation into text data using speech recognition, and means for recognizing emotions from the user's voice using an emotion engine and providing that information to the generative artificial intelligence. This enables personalized responses that take the user's emotions into consideration, enabling rapid and appropriate risk assessment and emergency response.
[0266] A "communication terminal" is a device that a user uses to make voice calls or data communications, and specifically includes telephones, smartphones, tablets, and the like.
[0267] "Generative AI" is a type of AI model that analyzes user input data and generates appropriate responses and decisions.
[0268] "Health status assessment" is the process of analyzing a user's health-related statements and behaviors and assessing their health risks.
[0269] "Crime risk assessment" is the process of assessing the risk of crime based on user comments and environmental information.
[0270] A "service provider operator" is a professional operator who is on standby to respond to user risks or emergencies, such as a provider of medical services or security services.
[0271] "Speech recognition" is a technology that captures a user's voice as digital data and converts it into text data.
[0272] "Text data" is data expressed as text information generated by voice recognition technology.
[0273] The "emotion engine" is an engine for analyzing and recognizing the user's emotional state from voice and text data.
[0274] "Logging" refers to the act of saving the contents of conversations conducted by the system, risk assessment results, emergency response history, etc.
[0275] This invention aims to develop a system that reduces loneliness among elderly people living alone and assesses their health status and crime risk in real time. The system allows users to communicate using a communication terminal, analyzes the conversation with the user using a generative AI, and, based on the results, forwards the communication to a service provider operator as needed. Furthermore, the system converts the conversation into text data using speech recognition technology, recognizes emotions from the user's voice using an emotion engine, and provides this information to the generative AI.
[0276] Hardware Configuration
[0277] Communication device: A smartphone or telephone used by a user, which allows voice calls.
[0278] Server: A server with high-performance computing power runs generative artificial intelligence, speech recognition technology, and an emotion engine.
[0279] Software Configuration
[0280] Speech Recognition Library: speech_recognition library
[0281] Speech synthesis engine: pyttsx3 library
[0282] Emotion Recognition Module: emotion_recognition library
[0283] Generative AI model: Our own ai_model library
[0284] Emergency Service Connector: EmergencyServiceConnector module
[0285] Process Overview
[0286] 1. Voice input and analysis: A user makes a voice call using a communication device. The voice data is transferred to the server and converted into text data by a voice recognition library.
[0287] 2. Emotion recognition: The emotion recognition module analyzes the user's emotions from the voice data and provides the results to the generative AI.
[0288] 3. Risk assessment: Generative AI analyzes text data to assess health or crime risks.
[0289] 4. Emergency Calls: Based on risk assessment, emergency services will be called as required.
[0290] 5. Logging: Record the conversation content and evaluation results as logs for future analysis.
[0291] Specific examples
[0292] For example, if a user says, "I've been having headaches lately," the speech recognition technology will detect the keyword "headache," and the generative AI will evaluate this as a health risk. At the same time, the emotion engine will detect anxiety in the voice and generate a response to the user saying, "Don't worry, we'll take care of it right away."
[0293] Prompt Sentence Examples
[0294] text
[0295] User: I've been having a lot of headaches lately.
[0296] Generative AI: The keyword "headache" is detected as a high health risk. The emotion engine recognizes the emotion of anxiety.
[0297] User: I saw something suspicious in my neighborhood.
[0298] Generative AI: The keyword "suspicious person" is detected, which indicates a high risk of crime. The emotion engine recognizes the emotion of fear.
[0299] In this way, the present invention enables personalized responses that take into account the user's emotions, and can provide quick and appropriate risk assessments and emergency responses.
[0300] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0301] Step 1:
[0302] A user initiates a voice call using a communication terminal. The communication terminal uses a microphone to collect voice and transmits it to a server. The input data is the user's voice, and the output is the transfer of voice data to the server.
[0303] Step 2:
[0304] The server uses a speech recognition library (speech_recognition) to convert the received voice data into text data. The input is the user's voice data, and the output is the generated text data.
[0305] Step 3:
[0306] The server uses a generative artificial intelligence model (ai_model library) to analyze the text data. As a result of the analysis, a health or crime risk is assessed. The input is the text data, and the output is the risk assessment result.
[0307] Step 4:
[0308] The server uses an emotion engine (emotion_recognition library) to recognize the user's emotions from the voice data. The results are fed back to the generative AI. The input is the voice data, and the output is the recognized emotion data.
[0309] Step 5:
[0310] Generative AI generates appropriate responses based on text data and emotional data. The responses are formatted as text and used as feedback to the user. The input is text data and emotional data, and the output is the generated response text.
[0311] Step 6:
[0312] The server forwards the communication to the service provider's operator via the Emergency Service Connector as needed. For example, if the health risk is assessed as high, it is forwarded to medical services, and if the crime risk is assessed as high, it is forwarded to security services. The input is the risk assessment result, and the output is forwarding the communication to the operator.
[0313] Step 7:
[0314] The server records conversations, risk assessment results, and emergency response history as logs. These logs are useful for future analysis and problem solving. The input is the conversations and assessment results, and the output is the log records.
[0315] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0316] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0317] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0318] [Second embodiment]
[0319] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0320] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0321] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0322] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0323] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0324] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0325] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0326] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0327] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0328] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0329] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0330] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0331] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate operator of a service provider as necessary. Specific operations and processing of an embodiment of this system will be described below.
[0332] Specific Embodiments of the System
[0333] A user makes a call using a communication terminal
[0334] A user calls the system using a communication terminal (for example, a telephone or a smartphone). The communication terminal starts a voice call and is connected to the server.
[0335] The server starts the audio service
[0336] When the server receives a call from a user, it activates a voice service using generative artificial intelligence. The server starts a conversation by saying to the user, "Hello, is there anything I can help you with?"
[0337] Conversation initiation and analysis
[0338] The server converts the user's responses into text data in real time using speech-to-text technology. Generative AI then analyzes this text data to assess the user's health status and criminal risk.
[0339] Specific examples
[0340] For example, if a user says, "I've been having headaches lately," the server will analyze the text data containing the keyword "headache" and evaluate it as a health risk. Similarly, if a user says, "I saw a suspicious person," the server will evaluate the crime risk from the keyword "suspicious person."
[0341] Risk assessment and call transfer to an operator
[0342] The server transfers calls to service provider operators if necessary based on the health or crime risk assessed by the generative AI. For example, if the health risk is determined to be high, the call is transferred to a HELPO operator, and if the crime risk is determined to be high, the call is transferred to a security service operator.
[0343] Logging the results
[0344] The server records all conversations, their analysis results, and whether or not the call was transferred in detail as a log, which can be used for future analysis and verification when a problem occurs.
[0345] In this way, the system of the present invention is a system that reduces the user's sense of loneliness, detects health conditions and crime risks early, and takes necessary measures promptly. The objectives of the present invention can be achieved by appropriately arranging communication terminals, generative artificial intelligence, speech-to-text conversion technology, and log recording means to implement the invention.
[0346] The processing flow will be explained below.
[0347] Step 1:
[0348] A user makes a call using a communication terminal.
[0349] The communication terminal initiates a voice call and is connected to the server.
[0350] Step 2:
[0351] The server receives a call from a user.
[0352] The server launches a voice service using generative artificial intelligence and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0353] Step 3:
[0354] The server waits for the user's response.
[0355] The communication terminal picks up the user's voice and transmits it to the server.
[0356] Step 4:
[0357] The server converts the received voice data into text data using voice-to-text conversion technology.
[0358] The server analyzes the text data using generative artificial intelligence.
[0359] Step 5:
[0360] The server assesses the user's health and criminal risk.
[0361] Generative artificial intelligence detects keywords such as "headache," "dizziness," "suspicious person," and "fraud," and determines health or criminal risks based on these.
[0362] Step 6:
[0363] Based on the risk assessment result, the server prepares to transfer the call to an operator of the service provider as necessary.
[0364] If the server is determined to pose a high health risk, it will contact a HELPO operator, and if it is determined to pose a high crime risk, it will contact a security service operator.
[0365] Step 7:
[0366] The server notifies the user that "Important information has been detected and we will connect you to an operator."
[0367] The server then transfers the call to the appropriate operator.
[0368] Step 8:
[0369] The server logs all conversations, their analysis results, and whether or not the call was forwarded.
[0370] The server records logs to allow for future analysis and verification when problems occur.
[0371] Through this procedure, the server, communication terminal, and user cooperate with each other to realize the system of the present invention.
[0372] Example 1
[0373] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0374] It is desirable to provide a means to reduce loneliness among elderly people living alone, to detect health conditions and crime risks early, and to take appropriate measures promptly. There is also a need for a system that can effectively analyze the content of conversations with users and transfer calls to the necessary service providers based on the results.
[0375] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0376] In this invention, the server includes a means for a user to communicate using an information terminal, a means for analyzing the conversation with the user using generative artificial intelligence to evaluate the user's health condition or criminal risk, a means for transferring the communication to a service provider's operator as needed based on the evaluation, a means for acquiring the user's conversation content as text data using speech-to-text technology, and a means for recording the conversation content and analysis results in a log. This reduces the sense of loneliness felt by elderly people living alone, enables early detection of health conditions and criminal risk, and allows for prompt response. Furthermore, the server can evaluate risk based on the user's speech and quickly transfer the call to an appropriate service provider.
[0377] "Information terminal" refers to any device used for communication, and specifically includes telephones, smartphones, computers, etc.
[0378] "Generative AI" refers to AI technology that analyzes conversations with users and makes decisions and suggestions based on that analysis.
[0379] "Health status" refers to general information about the user's health, such as the user's physical condition and whether or not the user has any illnesses.
[0380] "Crime risk" refers to the possibility that a user will become involved in a crime or the risk that a crime will occur.
[0381] "Service provider" means the operator of a business or organisation that provides health and safety services.
[0382] "Speech-to-text technology" refers to technology that converts voice data into text data in real time.
[0383] "Log" refers to records of conversation content, analysis results, communication transfer history, etc.
[0384] "Conversation content" refers to the content of the voice communication exchanged between the user and the system (or generative artificial intelligence).
[0385] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an operator of an appropriate service provider as necessary. Specific embodiments of the present invention will be described below.
[0386] First, a user calls a specific phone number using an information terminal (e.g., a telephone or smartphone). When the voice call begins, the terminal connects to the server. When the server receives the call from the user, it immediately answers and launches a voice-enabled application using generative artificial intelligence. Specifically, the server speaks to the user, saying, "Hello, is there anything I can help you with?"
[0387] The server then uses speech-to-text technology (e.g., Google Cloud Speech-to-Text API) to convert the user's speech into text data in real time. Once the user's speech is converted into text data, the text data is stored on the server.
[0388] The server analyzes the text data using generative artificial intelligence (for example, OpenAI's GPT-4 model). This analysis evaluates the user's health condition and criminal risk. For example, if a user says, "I've been having headaches lately," the server extracts the keyword "headache" and determines that this corresponds to a health risk. Similarly, if a user says, "I saw a suspicious person," the keyword "suspicious person" will evaluate the criminal risk as high.
[0389] Based on the evaluation results, the server transfers the call to a service provider operator as needed. For example, if the call is judged to be a high health risk, the server transfers the call to a health support service operator, and if the call is judged to be a high crime risk, the server transfers the call to a security service operator.
[0390] In addition, the server records detailed logs of all conversations, analysis results, and whether or not a call was forwarded, which can be used for future analysis and verification if a problem occurs.
[0391] Prompt Sentence Examples
[0392] Prompt: "If a user says, 'I've been having a lot of headaches lately,' explain how you can use a generative AI model to assess health risks and route the call to a health support service operator."
[0393] Prompt: "When a user says 'I see a suspicious person,' explain how a generative AI model can be used to assess the crime risk and route the call to a security service operator."
[0394] As described above, the system of the present invention can reduce the user's sense of loneliness, detect health conditions and crime risks early, and take necessary measures promptly.
[0395] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0396] Step 1:
[0397] A user uses an information terminal to call a specified phone number. When the voice call begins, the terminal connects to the server. The input is the user's voice, and the output is a connection with the server. Specifically, the user operates a telephone or smartphone to call the phone number specified by the system.
[0398] Step 2:
[0399] The server receives a call and activates a voice service. The input is an incoming call signal from the user, and the output is a voice-enabled application powered by generative artificial intelligence. Specifically, the server detects the incoming call, activates an auto-answer function, and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0400] Step 3:
[0401] The server uses speech-to-text technology to convert the user's speech into text data in real time. The input is the user's speech, and the output is text data. Specifically, the server uses the Google Cloud Speech-to-Text API to convert the speech into text data.
[0402] Step 4:
[0403] The server analyzes the text data using generative artificial intelligence. The input is text data, and the output is an assessment result regarding the user's health status and crime risk. Specifically, the server uses generative artificial intelligence (e.g., GPT-4) to analyze the text data and perform a risk assessment.
[0404] Step 5:
[0405] The server transfers the call to a service provider's operator as needed based on the risk assessment results. The input is the risk assessment result, and the output is call transfer. As a specific example, if the health risk is determined to be high, the server transfers the call to a health support operator, and if the crime risk is determined to be high, the server transfers the call to a security service operator.
[0406] Step 6:
[0407] The server records all conversation content, analysis results, and whether or not the call was forwarded in detail as a log. The input is the conversation content with the user and the analysis results, and the output is a log file. Specifically, the server processes the conversation content, analysis results, and call history in a recording database.
[0408] Through the above processing steps, the system of the present invention can reduce the user's sense of loneliness, detect health conditions and crime risks early, and take necessary measures promptly.
[0409] (Application example 1)
[0410] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0411] As the elderly population increases, feelings of loneliness, health risks, and crime risks among those living alone are becoming social issues. Elderly people living alone face particular challenges, such as difficulty in responding immediately to health problems or crime risks. As a result, there is a need for a system that can reduce feelings of loneliness while ensuring their safety and peace of mind. To solve this issue, it is important to assess the elderly's situation in real time via voice calls and quickly connect them to appropriate services as needed.
[0412] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0413] In this invention, the server includes: means for a user to use a communication terminal to perform predetermined communications; means for analyzing conversations with the user using generative artificial intelligence to assess the user's health status or criminal risk; means for transferring communications to a service provider operator as needed based on the assessment; means for converting the content of the voice call into text data in real time using speech-to-text technology; means for assessing the risk of exceeding a predetermined threshold based on the analysis results; means for quickly transferring the call to an appropriate operator if the risk exceeds a specific threshold; and means for recording all conversation contents, analysis results, and call transfer history. This allows for rapid detection of potential risks faced by users and enables prompt response.
[0414] A "communication terminal" is a device that a user uses to perform predetermined communications, and includes telephones and smartphones.
[0415] "Generative AI" is an AI technology that analyzes voice and text data to assess a user's health status and crime risk.
[0416] "Speech-to-text technology" is a technology that converts voice data into text data in real time.
[0417] "Evaluation" refers to the act of using generative artificial intelligence to analyze the content of a user's conversations and determine their health status and criminal risk.
[0418] A "service provider operator" is a specialized service person who takes action when a user is assessed as being at high risk for health or crime.
[0419] A "threshold" is a reference value that is set when the evaluation results of the generative artificial intelligence exceed a specific risk level.
[0420] "Call forwarding" is the process of quickly connecting a communication taken from a user to an operator of another service provider.
[0421] A "log" is a detailed record of conversation content, analysis results, call forwarding history, etc.
[0422] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate operator of a service provider as necessary. Specific operations and processing of an embodiment of this system will be described below.
[0423] System Hardware and Software
[0424] The following hardware and software are used to implement this system:
[0425] Hardware:
[0426] Smartphone microphone and speaker
[0427] Networked Servers
[0428] software:
[0429] Python
[0430] speech_recognition library
[0431] The requests library
[0432] Generative AI model
[0433] System operation explanation
[0434] 1. Start a voice call:
[0435] A user starts the application on their smartphone and taps the "emergency call" button, which establishes a real-time voice call between the user and the server.
[0436] 2. Speech recognition and text conversion:
[0437] The voice captured by the smartphone microphone is converted into text data in real time using the speech_recognition library, and the user's speech is then sent to the server as text data.
[0438] 3. Analysis by generative artificial intelligence:
[0439] The server uses a generative artificial intelligence model to analyze the acquired text data. This analysis evaluates the user's health condition and crime risk. For example, if a user says, "I've been having headaches lately," the keyword "headache" is analyzed and evaluated as a health risk. Similarly, if a user says, "I saw a suspicious person," the keyword "suspicious person" is used to evaluate the crime risk.
[0440] 4. Risk Assessment and Call Forwarding:
[0441] Based on the analysis results, the server evaluates whether the risk exceeds a certain threshold. If the risk is high, the generative AI model quickly routes the call to the appropriate service operator based on the risk assessment. For example, if the call is judged to be a high health risk, it will be routed to a medical service operator, and if the call is judged to be a high crime risk, it will be routed to a security service operator.
[0442] 5. Logging:
[0443] The server records all conversations, their analysis results, and call forwarding history in detail, and stores them in protected cloud storage for future analysis and verification in case of problems.
[0444] Specific prompt examples
[0445] Examples:
[0446] If a user says, "I'm scared because I've seen a lot of suspicious people walking down the street lately," the system will detect the phrase "suspicious people" and evaluate them as a crime risk, immediately transferring the call to a security service provider.
[0447] Example prompt sentence:
[0448] Text: "Recently, I've been seeing a lot of suspicious people walking down the street at night and it's scary."
[0449] Prompt: "Analyze this text to assess its health and crime risks."
[0450] In this way, the system provides a practical solution that ensures the safety and security of elderly people living alone, while also allowing for a rapid response.
[0451] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0452] Step 1:
[0453] A user starts an application using a communication terminal and taps the "emergency call" button. This operation causes the communication terminal to establish a voice call with the server. The user's voice is acquired as input, and a call with the server is initiated as output.
[0454] Step 2:
[0455] The server receives the voice data acquired from the communication terminal. The voice data is sent as input and recorded on the server as output.
[0456] Step 3:
[0457] The server uses the speech_recognition library to convert the received voice data into text data in real time. Voice data is used as input and text data is generated as output. Specifically, the server analyzes each phoneme using a speech recognition engine and generates the corresponding string.
[0458] Step 4:
[0459] The server inputs the generated text data into a generative artificial intelligence (AI) model to evaluate the user's health status and crime risk. The text data is used as input, and the health and crime risk assessment results are obtained as output. Specifically, the AI model uses natural language processing technology to analyze keywords and context within the text and conduct a risk assessment.
[0460] Step 5:
[0461] The server determines whether the risk exceeds a certain threshold based on the assessment results. The risk assessment results are used as input, and a judgment of whether the risk is high or low is obtained as output. Specifically, the risk level is determined by comparing it with a predefined threshold.
[0462] Step 6:
[0463] If the server determines that the call is high risk, it transfers the call to an operator at the appropriate service provider. The server uses the risk assessment results and threshold judgment results as inputs, and executes the call transfer as output. Specifically, this includes the process of rerouting the call to a specific phone number.
[0464] Step 7:
[0465] The server records all conversation content, assessment results, and call forwarding history in detailed logs and stores them in protected cloud storage. Assessment results and call history are used as input, and detailed log data is generated as output. Specifically, information such as date and time, conversation content, risk assessment results, and forwarding destinations is written to the log file.
[0466] In this way, the system assesses the situation of elderly people living alone in real time and takes prompt action as needed, providing peace of mind and safety.
[0467] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0468] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate service provider operator as needed. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and generates responses based on those emotions, more personalized responses are possible for users.
[0469] Specific Embodiments of the System
[0470] A user makes a call using a communication terminal
[0471] A user calls the system using a communication terminal (for example, a telephone or a smartphone). The communication terminal starts a voice call and is connected to the server.
[0472] The server starts the audio service
[0473] When the server receives a call from a user, it activates a voice service using generative artificial intelligence. The server starts a conversation by saying to the user, "Hello, is there anything I can help you with?"
[0474] Conversation initiation and analysis
[0475] The server converts the user's responses into text data in real time using speech-to-text technology. The generative AI analyzes this text data and assesses the user's health status and criminal risk. At the same time, the emotion engine recognizes emotions from the user's voice and provides this information to the generative AI.
[0476] Specific examples
[0477] For example, if a user says, "I've been having headaches lately," the server analyzes text data containing the keyword "headache" and evaluates it as a health risk. Meanwhile, the emotion engine detects emotions such as anxiety and stress from the user's voice. Similarly, if a user says, "I saw a suspicious person," the server evaluates the crime risk from the keyword "suspicious person," and the emotion engine recognizes fear and tension.
[0478] Generating responses based on emotion information
[0479] The server uses generative AI to generate an appropriate response based on the user's emotional information provided by the emotion engine. For example, if a user is feeling anxious, the server will respond with "Don't worry, we'll take care of it right away."
[0480] Risk assessment and call transfer to an operator
[0481] The server transfers the call to a service provider operator if necessary based on the health or criminal risk assessed by the generative artificial intelligence. For example, if the health risk is determined to be high, the call is transferred to an appropriate medical service operator, and if the criminal risk is determined to be high, the call is transferred to a security service operator.
[0482] Logging the results
[0483] The server logs all conversations, emotional information, analysis results, and whether or not a call was forwarded. These logs are used for future analysis and verification when problems occur.
[0484] summary
[0485] In this way, the system of the present invention reduces the user's sense of loneliness, detects health conditions and crime risks early, and promptly provides appropriate services. Furthermore, incorporating an emotion engine enables more compassionate and personalized responses to users. The objectives of the present invention can be achieved by appropriately arranging a communication terminal, generative artificial intelligence, speech-to-text technology, an emotion engine, and a means for logging in order to implement the invention.
[0486] The processing flow will be explained below.
[0487] Step 1:
[0488] A user makes a call using a communication terminal.
[0489] The communication terminal initiates a voice call and is connected to the server.
[0490] Step 2:
[0491] The server receives a call from a user.
[0492] The server launches a voice service using generative artificial intelligence and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0493] Step 3:
[0494] The server waits for the user's response.
[0495] The communication terminal picks up the user's voice and transmits it to the server.
[0496] Step 4:
[0497] The server converts the received voice data into text data using voice-to-text conversion technology.
[0498] The server analyzes the text data using generative artificial intelligence.
[0499] Step 5:
[0500] At the same time, the server uses an emotion engine to recognize emotions from the user's voice.
[0501] The emotion engine provides the recognized emotion information to the generative artificial intelligence.
[0502] Step 6:
[0503] The server evaluates the user's health status and crime risk based on text data and emotional information analyzed by generative artificial intelligence.
[0504] For example, if a user says, "I've been having headaches lately," the server detects the keyword "headache," and the emotion engine recognizes the emotion of anxiety.
[0505] Step 7:
[0506] The server uses generative artificial intelligence to generate an appropriate response based on emotional information.
[0507] For example, a user who is feeling anxious can be given a response such as "Don't worry, we'll take care of it right away."
[0508] Step 8:
[0509] Based on the evaluation results, the server transfers the call to an operator at the service provider as necessary.
[0510] If the health risk is determined to be high, the call will be transferred to a medical service operator, and if the crime risk is determined to be high, the call will be transferred to a security service operator.
[0511] Step 9:
[0512] The server performs the call transfer and notifies the user that "Important content has been detected and you will be connected to an operator."
[0513] The server then transfers the call to the appropriate operator.
[0514] Step 10:
[0515] The server records all conversation content, analysis results, emotional information, and whether or not the call was forwarded as a log.
[0516] This record will be used for future analysis and verification when problems occur.
[0517] This specific procedure enables the server, communication terminal, and user to work together, reducing the user's sense of loneliness, detecting health conditions and crime risks early, and taking necessary measures quickly.
[0518] Example 2
[0519] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0520] In modern society, elderly people living alone are increasingly experiencing loneliness, health problems, and crime risks, necessitating rapid and appropriate responses. However, conventional monitoring and warning systems have difficulty accurately grasping a user's emotional state and providing personalized responses based on that information. Furthermore, they lack the ability to adequately assess risk in real time and quickly transfer calls to service providers. To address these challenges, a system is needed that can accurately assess a user's emotions, health risks, and crime risks, and provide appropriate responses.
[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0522] In this invention, the server includes: means for allowing a user to communicate using a communication terminal; means for analyzing a conversation with the user and assessing their health condition or criminal risk using generative artificial intelligence; means for converting the user's voice data into text data using speech-to-text technology; means including an emotion engine for evaluating the user's emotional state based on the analysis results; means for transferring the communication to a service provider operator as needed based on the evaluation; and means for recording the content of the conversation, emotional information, analysis results, and whether or not the call was transferred. This allows for accurate evaluation of the user's emotions and risk state in real time, enabling prompt and appropriate responses.
[0523] "Communication terminal" refers to any device that a user uses to conduct voice or data communications, and specifically includes smartphones and landline phones.
[0524] "Generative AI" refers to AI technology that analyzes user input and generates appropriate responses and analytical results, and includes systems that utilize natural language processing and machine learning.
[0525] "Speech-to-text technology" refers to technology that converts a user's voice into text data in real time, and includes, for example, cloud-based voice recognition services.
[0526] "Emotion engine" refers to technology that analyzes a user's emotional state from their voice or text data and provides that information, and includes systems that utilize emotion recognition algorithms and machine learning.
[0527] "Service provider operators" refer to professional personnel who are on standby to provide a particular service, including those working in fields such as healthcare and security.
[0528] "Database" means a system of record for storing and managing system-generated data, including the means for efficiently storing and accessing data in various formats.
[0529] "Log" refers to historical data that records the processes performed by the system, the data generated, events, etc., and is used for future analysis and troubleshooting.
[0530] This invention relates to a system that reduces loneliness among elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to service provider operators as needed. This system uses generative artificial intelligence to analyze user conversations and utilizes speech-to-text technology and an emotion engine to provide more personalized responses.
[0531] A user calls the system using a communication terminal. The communication terminal can be a landline phone or a smartphone. When the user makes a call, the server receives the call and activates a voice service using generative artificial intelligence. For example, the server may play a standard phrase such as "Hello, is there anything I can help you with?". At this stage, the server activates a voice playback module and provides the user with an initial response.
[0532] The server then receives the user's response and converts it into text data in real time using speech-to-text technology (e.g., Google Cloud Speech-to-Text API). The server then inputs this text data into a generative artificial intelligence (e.g., OpenAI GPT-4) to assess the user's health status and criminal risk. Based on the analysis results, the server then uses an emotion engine (e.g., IBM Watson Tone Analyzer) to evaluate the user's emotional state.
[0533] For example, if a user says, "I've been having headaches lately," the server will send text data containing the keyword "headache" to the generative AI to assess whether there is a health risk. At the same time, the emotion engine will detect anxiety and stress from the user's voice. This will complement the user's emotional data, allowing the system to make a more accurate risk assessment.
[0534] The server integrates data from the generative artificial intelligence and emotion engine to generate an appropriate response for the user. For example, if a user feels anxious, the server generates a response such as "Don't worry, we'll take care of it right away," and plays it back to the user using the voice playback module.
[0535] If the server determines that the user faces a high health or criminal risk as a result of the risk assessment, it will transfer the call to a service provider operator as necessary. For example, if the user is assessed as having a high health risk, it will transfer the call to an operator of an appropriate medical service. The server uses the call transfer function to ensure a prompt response.
[0536] The server also records all conversation content, emotional information, analysis results, and whether or not the call was transferred in a database, which can be used for future analysis and troubleshooting.
[0537] As a concrete example, by inputting the following prompt sentence into a generative artificial intelligence, an appropriate response can be obtained.
[0538] Prompt: "The user mentions that they've been having headaches lately. Use this information to generate an appropriate response. They also sound anxious and stressed."
[0539] Example response: "I'm sorry to hear that you're concerned. Your persistent headache is concerning. Please wait a moment while we connect you to medical services."
[0540] As described above, the system of the present invention can accurately evaluate the user's emotions and risk status in real time and provide prompt and appropriate responses. This system can reduce the sense of loneliness felt by elderly people living alone and enable prompt responses to health conditions and crime risks.
[0541] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0542] Step 1:
[0543] A user calls the system using a communication terminal. The input includes the user's voice, and the output is a connection to the server. Specifically, the user calls the system's phone number using a smartphone or landline phone and presses the call button.
[0544] Step 2:
[0545] The server invokes a voice service, which includes as input the user's call connection and as output the playback of a voice message. Specifically, the server detects when the call is answered and invokes the voice playback module to play a voice message saying, "Hello, how can I help you?"
[0546] Step 3:
[0547] The server converts the user's speech into text data. The input includes the user's speech response, and the converted text data is obtained as the output. Specifically, the server uses speech-to-text conversion technology (e.g., Google Cloud Speech-to-Text API) to convert the user's speech into text data in real time.
[0548] Step 4:
[0549] The server analyzes the text data using generative AI. The converted text data is included as input, and the analysis results are obtained as output. Specifically, the server sends the text data to a generative AI (e.g., OpenAI GPT-4), which then evaluates health and crime risks.
[0550] Step 5:
[0551] The server uses an emotion engine to recognize the user's emotion. The input includes the analysis results and voice data, and the output is emotion information. Specifically, the server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to identify emotions from the user's tone of voice and text.
[0552] Step 6:
[0553] The server generates an appropriate response based on the integrated analysis results. The input includes the analysis results and emotional information, and the generated response is obtained as the output. Specifically, the server provides prompts to the generative AI based on the analysis results and emotional information, converts the generated response into speech, and sends it to the user.
[0554] Step 7:
[0555] The server performs risk assessment and transfers the call to an operator if necessary. The input contains the integrated analysis results, and the output is the required call transfer. Specifically, the server checks the risk assessment and, for example, if it determines that the health risk is high, transfers the call to a medical service operator.
[0556] Step 8:
[0557] The server records all conversation content, emotional information, analysis results, and whether or not the call was forwarded as a log. All communication data is included as input, and log data is obtained as output. Specifically, the server records and saves the conversation content, emotional information, analysis results, and whether or not the call was forwarded in a database.
[0558] (Application example 2)
[0559] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0560] There is a need to reduce the sense of loneliness felt by elderly people living alone, and to provide real-time assessments of their health status and crime risk to respond promptly. However, conventional systems are unable to recognize users' emotions and provide personalized responses, and there are also problems with the speed and accuracy of risk assessments. Therefore, a system that improves the quality of life of elderly people living alone and enables appropriate emergency response is needed.
[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0562] In this invention, the server includes means for allowing a user to communicate using a communication terminal, means for analyzing the conversation with the user using a generative artificial intelligence to assess the user's health condition or criminal risk, means for transferring the communication to a service provider operator as needed based on the assessment, means for converting the conversation into text data using speech recognition, and means for recognizing emotions from the user's voice using an emotion engine and providing that information to the generative artificial intelligence. This enables personalized responses that take the user's emotions into consideration, enabling rapid and appropriate risk assessment and emergency response.
[0563] A "communication terminal" is a device that a user uses to make voice calls or data communications, and specifically includes telephones, smartphones, tablets, and the like.
[0564] "Generative AI" is a type of AI model that analyzes user input data and generates appropriate responses and decisions.
[0565] "Health status assessment" is the process of analyzing a user's health-related statements and behaviors and assessing their health risks.
[0566] "Crime risk assessment" is the process of assessing the risk of crime based on user comments and environmental information.
[0567] A "service provider operator" is a professional operator who is on standby to respond to user risks or emergencies, such as a provider of medical services or security services.
[0568] "Speech recognition" is a technology that captures a user's voice as digital data and converts it into text data.
[0569] "Text data" is data expressed as text information generated by voice recognition technology.
[0570] The "emotion engine" is an engine for analyzing and recognizing the user's emotional state from voice and text data.
[0571] "Logging" refers to the act of saving the contents of conversations conducted by the system, risk assessment results, emergency response history, etc.
[0572] This invention aims to develop a system that reduces loneliness among elderly people living alone and assesses their health status and crime risk in real time. The system allows users to communicate using a communication terminal, analyzes the conversation with the user using a generative AI, and, based on the results, forwards the communication to a service provider operator as needed. Furthermore, the system converts the conversation into text data using speech recognition technology, recognizes emotions from the user's voice using an emotion engine, and provides this information to the generative AI.
[0573] Hardware Configuration
[0574] Communication device: A smartphone or telephone used by a user, which allows voice calls.
[0575] Server: A server with high-performance computing power runs generative artificial intelligence, speech recognition technology, and an emotion engine.
[0576] Software Configuration
[0577] Speech Recognition Library: speech_recognition library
[0578] Speech synthesis engine: pyttsx3 library
[0579] Emotion Recognition Module: emotion_recognition library
[0580] Generative AI model: Our own ai_model library
[0581] Emergency Service Connector: EmergencyServiceConnector module
[0582] Process Overview
[0583] 1. Voice input and analysis: A user makes a voice call using a communication device. The voice data is transferred to the server and converted into text data by a voice recognition library.
[0584] 2. Emotion recognition: The emotion recognition module analyzes the user's emotions from the voice data and provides the results to the generative AI.
[0585] 3. Risk assessment: Generative AI analyzes text data to assess health or crime risks.
[0586] 4. Emergency Calls: Based on risk assessment, emergency services will be called as required.
[0587] 5. Logging: Record the conversation content and evaluation results as logs for future analysis.
[0588] Specific examples
[0589] For example, if a user says, "I've been having headaches lately," the speech recognition technology will detect the keyword "headache," and the generative AI will evaluate this as a health risk. At the same time, the emotion engine will detect anxiety in the voice and generate a response to the user saying, "Don't worry, we'll take care of it right away."
[0590] Prompt Sentence Examples
[0591] text
[0592] User: I've been having a lot of headaches lately.
[0593] Generative AI: The keyword "headache" is detected as a high health risk. The emotion engine recognizes the emotion of anxiety.
[0594] User: I saw something suspicious in my neighborhood.
[0595] Generative AI: The keyword "suspicious person" is detected, which indicates a high risk of crime. The emotion engine recognizes the emotion of fear.
[0596] In this way, the present invention enables personalized responses that take into account the user's emotions, and can provide quick and appropriate risk assessments and emergency responses.
[0597] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0598] Step 1:
[0599] A user initiates a voice call using a communication terminal. The communication terminal uses a microphone to collect voice and transmits it to a server. The input data is the user's voice, and the output is the transfer of voice data to the server.
[0600] Step 2:
[0601] The server uses a speech recognition library (speech_recognition) to convert the received voice data into text data. The input is the user's voice data, and the output is the generated text data.
[0602] Step 3:
[0603] The server uses a generative artificial intelligence model (ai_model library) to analyze the text data. As a result of the analysis, a health or crime risk is assessed. The input is the text data, and the output is the risk assessment result.
[0604] Step 4:
[0605] The server uses an emotion engine (emotion_recognition library) to recognize the user's emotions from the voice data. The results are fed back to the generative AI. The input is the voice data, and the output is the recognized emotion data.
[0606] Step 5:
[0607] Generative AI generates appropriate responses based on text data and emotional data. The responses are formatted as text and used as feedback to the user. The input is text data and emotional data, and the output is the generated response text.
[0608] Step 6:
[0609] The server forwards the communication to the service provider's operator via the Emergency Service Connector as needed. For example, if the health risk is assessed as high, it is forwarded to medical services, and if the crime risk is assessed as high, it is forwarded to security services. The input is the risk assessment result, and the output is forwarding the communication to the operator.
[0610] Step 7:
[0611] The server records conversations, risk assessment results, and emergency response history as logs. These logs are useful for future analysis and problem solving. The input is the conversations and assessment results, and the output is the log records.
[0612] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0613] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0614] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0615] [Third embodiment]
[0616] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0617] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0618] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0619] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0620] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0621] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0622] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0623] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0624] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0625] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0626] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0627] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0628] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate operator of a service provider as necessary. Specific operations and processing of an embodiment of this system will be described below.
[0629] Specific Embodiments of the System
[0630] A user makes a call using a communication terminal
[0631] A user calls the system using a communication terminal (for example, a telephone or a smartphone). The communication terminal starts a voice call and is connected to the server.
[0632] The server starts the audio service
[0633] When the server receives a call from a user, it activates a voice service using generative artificial intelligence. The server starts a conversation by saying to the user, "Hello, is there anything I can help you with?"
[0634] Conversation initiation and analysis
[0635] The server converts the user's responses into text data in real time using speech-to-text technology. Generative AI then analyzes this text data to assess the user's health status and criminal risk.
[0636] Specific examples
[0637] For example, if a user says, "I've been having headaches lately," the server will analyze the text data containing the keyword "headache" and evaluate it as a health risk. Similarly, if a user says, "I saw a suspicious person," the server will evaluate the crime risk from the keyword "suspicious person."
[0638] Risk assessment and call transfer to an operator
[0639] The server transfers calls to service provider operators if necessary based on the health or crime risk assessed by the generative AI. For example, if the health risk is determined to be high, the call is transferred to a HELPO operator, and if the crime risk is determined to be high, the call is transferred to a security service operator.
[0640] Logging the results
[0641] The server records all conversations, their analysis results, and whether or not the call was transferred in detail as a log, which can be used for future analysis and verification when a problem occurs.
[0642] In this way, the system of the present invention is a system that reduces the user's sense of loneliness, detects health conditions and crime risks early, and takes necessary measures promptly. The objectives of the present invention can be achieved by appropriately arranging communication terminals, generative artificial intelligence, speech-to-text conversion technology, and log recording means to implement the invention.
[0643] The processing flow will be explained below.
[0644] Step 1:
[0645] A user makes a call using a communication terminal.
[0646] The communication terminal initiates a voice call and is connected to the server.
[0647] Step 2:
[0648] The server receives a call from a user.
[0649] The server launches a voice service using generative artificial intelligence and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0650] Step 3:
[0651] The server waits for the user's response.
[0652] The communication terminal picks up the user's voice and transmits it to the server.
[0653] Step 4:
[0654] The server converts the received voice data into text data using voice-to-text conversion technology.
[0655] The server analyzes the text data using generative artificial intelligence.
[0656] Step 5:
[0657] The server assesses the user's health and criminal risk.
[0658] Generative artificial intelligence detects keywords such as "headache," "dizziness," "suspicious person," and "fraud," and determines health or criminal risks based on these.
[0659] Step 6:
[0660] Based on the risk assessment result, the server prepares to transfer the call to an operator of the service provider as necessary.
[0661] If the server is determined to pose a high health risk, it will contact a HELPO operator, and if it is determined to pose a high crime risk, it will contact a security service operator.
[0662] Step 7:
[0663] The server notifies the user that "Important information has been detected and we will connect you to an operator."
[0664] The server then transfers the call to the appropriate operator.
[0665] Step 8:
[0666] The server logs all conversations, their analysis results, and whether or not the call was forwarded.
[0667] The server records logs to allow for future analysis and verification when problems occur.
[0668] Through this procedure, the server, communication terminal, and user cooperate with each other to realize the system of the present invention.
[0669] Example 1
[0670] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0671] It is desirable to provide a means to reduce loneliness among elderly people living alone, to detect health conditions and crime risks early, and to take appropriate measures promptly. There is also a need for a system that can effectively analyze the content of conversations with users and transfer calls to the necessary service providers based on the results.
[0672] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0673] In this invention, the server includes a means for a user to communicate using an information terminal, a means for analyzing the conversation with the user using generative artificial intelligence to evaluate the user's health condition or criminal risk, a means for transferring the communication to a service provider's operator as needed based on the evaluation, a means for acquiring the user's conversation content as text data using speech-to-text technology, and a means for recording the conversation content and analysis results in a log. This reduces the sense of loneliness felt by elderly people living alone, enables early detection of health conditions and criminal risk, and allows for prompt response. Furthermore, the server can evaluate risk based on the user's speech and quickly transfer the call to an appropriate service provider.
[0674] "Information terminal" refers to any device used for communication, and specifically includes telephones, smartphones, computers, etc.
[0675] "Generative AI" refers to AI technology that analyzes conversations with users and makes decisions and suggestions based on that analysis.
[0676] "Health status" refers to general information about the user's health, such as the user's physical condition and whether or not the user has any illnesses.
[0677] "Crime risk" refers to the possibility that a user will become involved in a crime or the risk that a crime will occur.
[0678] "Service provider" means the operator of a business or organisation that provides health and safety services.
[0679] "Speech-to-text technology" refers to technology that converts voice data into text data in real time.
[0680] "Log" refers to records of conversation content, analysis results, communication transfer history, etc.
[0681] "Conversation content" refers to the content of the voice communication exchanged between the user and the system (or generative artificial intelligence).
[0682] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an operator of an appropriate service provider as necessary. Specific embodiments of the present invention will be described below.
[0683] First, a user calls a specific phone number using an information terminal (e.g., a telephone or smartphone). When the voice call begins, the terminal connects to the server. When the server receives the call from the user, it immediately answers and launches a voice-enabled application using generative artificial intelligence. Specifically, the server speaks to the user, saying, "Hello, is there anything I can help you with?"
[0684] The server then uses speech-to-text technology (e.g., Google Cloud Speech-to-Text API) to convert the user's speech into text data in real time. Once the user's speech is converted into text data, the text data is stored on the server.
[0685] The server analyzes the text data using generative artificial intelligence (for example, OpenAI's GPT-4 model). This analysis evaluates the user's health condition and criminal risk. For example, if a user says, "I've been having headaches lately," the server extracts the keyword "headache" and determines that this corresponds to a health risk. Similarly, if a user says, "I saw a suspicious person," the keyword "suspicious person" will evaluate the criminal risk as high.
[0686] Based on the evaluation results, the server transfers the call to a service provider operator as needed. For example, if the call is judged to be a high health risk, the server transfers the call to a health support service operator, and if the call is judged to be a high crime risk, the server transfers the call to a security service operator.
[0687] In addition, the server records detailed logs of all conversations, analysis results, and whether or not a call was forwarded, which can be used for future analysis and verification if a problem occurs.
[0688] Prompt Sentence Examples
[0689] Prompt: "If a user says, 'I've been having a lot of headaches lately,' explain how you can use a generative AI model to assess health risks and route the call to a health support service operator."
[0690] Prompt: "When a user says 'I see a suspicious person,' explain how a generative AI model can be used to assess the crime risk and route the call to a security service operator."
[0691] As described above, the system of the present invention can reduce the user's sense of loneliness, detect health conditions and crime risks early, and take necessary measures promptly.
[0692] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0693] Step 1:
[0694] A user uses an information terminal to call a specified phone number. When the voice call begins, the terminal connects to the server. The input is the user's voice, and the output is a connection with the server. Specifically, the user operates a telephone or smartphone to call the phone number specified by the system.
[0695] Step 2:
[0696] The server receives a call and activates a voice service. The input is an incoming call signal from the user, and the output is a voice-enabled application powered by generative artificial intelligence. Specifically, the server detects the incoming call, activates an auto-answer function, and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0697] Step 3:
[0698] The server uses speech-to-text technology to convert the user's speech into text data in real time. The input is the user's speech, and the output is text data. Specifically, the server uses the Google Cloud Speech-to-Text API to convert the speech into text data.
[0699] Step 4:
[0700] The server analyzes the text data using generative artificial intelligence. The input is text data, and the output is an assessment result regarding the user's health status and crime risk. Specifically, the server uses generative artificial intelligence (e.g., GPT-4) to analyze the text data and perform a risk assessment.
[0701] Step 5:
[0702] The server transfers the call to a service provider's operator as needed based on the risk assessment results. The input is the risk assessment result, and the output is call transfer. As a specific example, if the health risk is determined to be high, the server transfers the call to a health support operator, and if the crime risk is determined to be high, the server transfers the call to a security service operator.
[0703] Step 6:
[0704] The server records all conversation content, analysis results, and whether or not the call was forwarded in detail as a log. The input is the conversation content with the user and the analysis results, and the output is a log file. Specifically, the server processes the conversation content, analysis results, and call history in a recording database.
[0705] Through the above processing steps, the system of the present invention can reduce the user's sense of loneliness, detect health conditions and crime risks early, and take necessary measures promptly.
[0706] (Application example 1)
[0707] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0708] As the elderly population increases, feelings of loneliness, health risks, and crime risks among those living alone are becoming social issues. Elderly people living alone face particular challenges, such as difficulty in responding immediately to health problems or crime risks. As a result, there is a need for a system that can reduce feelings of loneliness while ensuring their safety and peace of mind. To solve this issue, it is important to assess the elderly's situation in real time via voice calls and quickly connect them to appropriate services as needed.
[0709] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0710] In this invention, the server includes: means for a user to use a communication terminal to perform predetermined communications; means for analyzing conversations with the user using generative artificial intelligence to assess the user's health status or criminal risk; means for transferring communications to a service provider operator as needed based on the assessment; means for converting the content of the voice call into text data in real time using speech-to-text technology; means for assessing the risk of exceeding a predetermined threshold based on the analysis results; means for quickly transferring the call to an appropriate operator if the risk exceeds a specific threshold; and means for recording all conversation contents, analysis results, and call transfer history. This allows for rapid detection of potential risks faced by users and enables prompt response.
[0711] A "communication terminal" is a device that a user uses to perform predetermined communications, and includes telephones and smartphones.
[0712] "Generative AI" is an AI technology that analyzes voice and text data to assess a user's health status and crime risk.
[0713] "Speech-to-text technology" is a technology that converts voice data into text data in real time.
[0714] "Evaluation" refers to the act of using generative artificial intelligence to analyze the content of a user's conversations and determine their health status and criminal risk.
[0715] A "service provider operator" is a specialized service person who takes action when a user is assessed as being at high risk for health or crime.
[0716] A "threshold" is a reference value that is set when the evaluation results of the generative artificial intelligence exceed a specific risk level.
[0717] "Call forwarding" is the process of quickly connecting a communication taken from a user to an operator of another service provider.
[0718] A "log" is a detailed record of conversation content, analysis results, call forwarding history, etc.
[0719] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate operator of a service provider as necessary. Specific operations and processing of an embodiment of this system will be described below.
[0720] System Hardware and Software
[0721] The following hardware and software are used to implement this system:
[0722] Hardware:
[0723] Smartphone microphone and speaker
[0724] Networked Servers
[0725] software:
[0726] Python
[0727] speech_recognition library
[0728] The requests library
[0729] Generative AI model
[0730] System operation explanation
[0731] 1. Start a voice call:
[0732] A user starts the application on their smartphone and taps the "emergency call" button, which establishes a real-time voice call between the user and the server.
[0733] 2. Speech recognition and text conversion:
[0734] The voice captured by the smartphone microphone is converted into text data in real time using the speech_recognition library, and the user's speech is then sent to the server as text data.
[0735] 3. Analysis by generative artificial intelligence:
[0736] The server uses a generative artificial intelligence model to analyze the acquired text data. This analysis evaluates the user's health condition and crime risk. For example, if a user says, "I've been having headaches lately," the keyword "headache" is analyzed and evaluated as a health risk. Similarly, if a user says, "I saw a suspicious person," the keyword "suspicious person" is used to evaluate the crime risk.
[0737] 4. Risk Assessment and Call Forwarding:
[0738] Based on the analysis results, the server evaluates whether the risk exceeds a certain threshold. If the risk is high, the generative AI model quickly routes the call to the appropriate service operator based on the risk assessment. For example, if the call is judged to be a high health risk, it will be routed to a medical service operator, and if the call is judged to be a high crime risk, it will be routed to a security service operator.
[0739] 5. Logging:
[0740] The server records all conversations, their analysis results, and call forwarding history in detail, and stores them in protected cloud storage for future analysis and verification in case of problems.
[0741] Specific prompt examples
[0742] Examples:
[0743] If a user says, "I'm scared because I've seen a lot of suspicious people walking down the street lately," the system will detect the phrase "suspicious people" and evaluate them as a crime risk, immediately transferring the call to a security service provider.
[0744] Example prompt sentence:
[0745] Text: "Recently, I've been seeing a lot of suspicious people walking down the street at night and it's scary."
[0746] Prompt: "Analyze this text to assess its health and crime risks."
[0747] In this way, the system provides a practical solution that ensures the safety and security of elderly people living alone, while also allowing for a rapid response.
[0748] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0749] Step 1:
[0750] A user starts an application using a communication terminal and taps the "emergency call" button. This operation causes the communication terminal to establish a voice call with the server. The user's voice is acquired as input, and a call with the server is initiated as output.
[0751] Step 2:
[0752] The server receives the voice data acquired from the communication terminal. The voice data is sent as input and recorded on the server as output.
[0753] Step 3:
[0754] The server uses the speech_recognition library to convert the received voice data into text data in real time. Voice data is used as input and text data is generated as output. Specifically, the server analyzes each phoneme using a speech recognition engine and generates the corresponding string.
[0755] Step 4:
[0756] The server inputs the generated text data into a generative artificial intelligence (AI) model to evaluate the user's health status and crime risk. The text data is used as input, and the health and crime risk assessment results are obtained as output. Specifically, the AI model uses natural language processing technology to analyze keywords and context within the text and conduct a risk assessment.
[0757] Step 5:
[0758] The server determines whether the risk exceeds a certain threshold based on the assessment results. The risk assessment results are used as input, and a judgment of whether the risk is high or low is obtained as output. Specifically, the risk level is determined by comparing it with a predefined threshold.
[0759] Step 6:
[0760] If the server determines that the call is high risk, it transfers the call to an operator at the appropriate service provider. The server uses the risk assessment results and threshold judgment results as inputs, and executes the call transfer as output. Specifically, this includes the process of rerouting the call to a specific phone number.
[0761] Step 7:
[0762] The server records all conversation content, assessment results, and call forwarding history in detailed logs and stores them in protected cloud storage. Assessment results and call history are used as input, and detailed log data is generated as output. Specifically, information such as date and time, conversation content, risk assessment results, and forwarding destinations is written to the log file.
[0763] In this way, the system assesses the situation of elderly people living alone in real time and takes prompt action as needed, providing peace of mind and safety.
[0764] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0765] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate service provider operator as needed. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and generates responses based on those emotions, more personalized responses are possible for users.
[0766] Specific Embodiments of the System
[0767] A user makes a call using a communication terminal
[0768] A user calls the system using a communication terminal (for example, a telephone or a smartphone). The communication terminal starts a voice call and is connected to the server.
[0769] The server starts the audio service
[0770] When the server receives a call from a user, it activates a voice service using generative artificial intelligence. The server starts a conversation by saying to the user, "Hello, is there anything I can help you with?"
[0771] Conversation initiation and analysis
[0772] The server converts the user's responses into text data in real time using speech-to-text technology. The generative AI analyzes this text data and assesses the user's health status and criminal risk. At the same time, the emotion engine recognizes emotions from the user's voice and provides this information to the generative AI.
[0773] Specific examples
[0774] For example, if a user says, "I've been having headaches lately," the server analyzes text data containing the keyword "headache" and evaluates it as a health risk. Meanwhile, the emotion engine detects emotions such as anxiety and stress from the user's voice. Similarly, if a user says, "I saw a suspicious person," the server evaluates the crime risk from the keyword "suspicious person," and the emotion engine recognizes fear and tension.
[0775] Generating responses based on emotion information
[0776] The server uses generative AI to generate an appropriate response based on the user's emotional information provided by the emotion engine. For example, if a user is feeling anxious, the server will respond with "Don't worry, we'll take care of it right away."
[0777] Risk assessment and call transfer to an operator
[0778] The server transfers the call to a service provider operator if necessary based on the health or criminal risk assessed by the generative artificial intelligence. For example, if the health risk is determined to be high, the call is transferred to an appropriate medical service operator, and if the criminal risk is determined to be high, the call is transferred to a security service operator.
[0779] Logging the results
[0780] The server logs all conversations, emotional information, analysis results, and whether or not a call was forwarded. These logs are used for future analysis and verification when problems occur.
[0781] summary
[0782] In this way, the system of the present invention reduces the user's sense of loneliness, detects health conditions and crime risks early, and promptly provides appropriate services. Furthermore, incorporating an emotion engine enables more compassionate and personalized responses to users. The objectives of the present invention can be achieved by appropriately arranging a communication terminal, generative artificial intelligence, speech-to-text technology, an emotion engine, and a means for logging in order to implement the invention.
[0783] The processing flow will be explained below.
[0784] Step 1:
[0785] A user makes a call using a communication terminal.
[0786] The communication terminal initiates a voice call and is connected to the server.
[0787] Step 2:
[0788] The server receives a call from a user.
[0789] The server launches a voice service using generative artificial intelligence and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0790] Step 3:
[0791] The server waits for the user's response.
[0792] The communication terminal picks up the user's voice and transmits it to the server.
[0793] Step 4:
[0794] The server converts the received voice data into text data using voice-to-text conversion technology.
[0795] The server analyzes the text data using generative artificial intelligence.
[0796] Step 5:
[0797] At the same time, the server uses an emotion engine to recognize emotions from the user's voice.
[0798] The emotion engine provides the recognized emotion information to the generative artificial intelligence.
[0799] Step 6:
[0800] The server evaluates the user's health status and crime risk based on text data and emotional information analyzed by generative artificial intelligence.
[0801] For example, if a user says, "I've been having headaches lately," the server detects the keyword "headache," and the emotion engine recognizes the emotion of anxiety.
[0802] Step 7:
[0803] The server uses generative artificial intelligence to generate an appropriate response based on emotional information.
[0804] For example, a user who is feeling anxious can be given a response such as "Don't worry, we'll take care of it right away."
[0805] Step 8:
[0806] Based on the evaluation results, the server transfers the call to an operator at the service provider as necessary.
[0807] If the health risk is determined to be high, the call will be transferred to a medical service operator, and if the crime risk is determined to be high, the call will be transferred to a security service operator.
[0808] Step 9:
[0809] The server performs the call transfer and notifies the user that "Important content has been detected and you will be connected to an operator."
[0810] The server then transfers the call to the appropriate operator.
[0811] Step 10:
[0812] The server records all conversation content, analysis results, emotional information, and whether or not the call was forwarded as a log.
[0813] This record will be used for future analysis and verification when problems occur.
[0814] This specific procedure enables the server, communication terminal, and user to work together, reducing the user's sense of loneliness, detecting health conditions and crime risks early, and taking necessary measures quickly.
[0815] Example 2
[0816] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0817] In modern society, elderly people living alone are increasingly experiencing loneliness, health problems, and crime risks, necessitating rapid and appropriate responses. However, conventional monitoring and warning systems have difficulty accurately grasping a user's emotional state and providing personalized responses based on that information. Furthermore, they lack the ability to adequately assess risk in real time and quickly transfer calls to service providers. To address these challenges, a system is needed that can accurately assess a user's emotions, health risks, and crime risks, and provide appropriate responses.
[0818] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0819] In this invention, the server includes: means for allowing a user to communicate using a communication terminal; means for analyzing a conversation with the user and assessing their health condition or criminal risk using generative artificial intelligence; means for converting the user's voice data into text data using speech-to-text technology; means including an emotion engine for evaluating the user's emotional state based on the analysis results; means for transferring the communication to a service provider operator as needed based on the evaluation; and means for recording the content of the conversation, emotional information, analysis results, and whether or not the call was transferred. This allows for accurate evaluation of the user's emotions and risk state in real time, enabling prompt and appropriate responses.
[0820] "Communication terminal" refers to any device that a user uses to conduct voice or data communications, and specifically includes smartphones and landline phones.
[0821] "Generative AI" refers to AI technology that analyzes user input and generates appropriate responses and analytical results, and includes systems that utilize natural language processing and machine learning.
[0822] "Speech-to-text technology" refers to technology that converts a user's voice into text data in real time, and includes, for example, cloud-based voice recognition services.
[0823] "Emotion engine" refers to technology that analyzes a user's emotional state from their voice or text data and provides that information, and includes systems that utilize emotion recognition algorithms and machine learning.
[0824] "Service provider operators" refer to professional personnel who are on standby to provide a particular service, including those working in fields such as healthcare and security.
[0825] "Database" means a system of record for storing and managing system-generated data, including the means for efficiently storing and accessing data in various formats.
[0826] "Log" refers to historical data that records the processes performed by the system, the data generated, events, etc., and is used for future analysis and troubleshooting.
[0827] This invention relates to a system that reduces loneliness among elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to service provider operators as needed. This system uses generative artificial intelligence to analyze user conversations and utilizes speech-to-text technology and an emotion engine to provide more personalized responses.
[0828] A user calls the system using a communication terminal. The communication terminal can be a landline phone or a smartphone. When the user makes a call, the server receives the call and activates a voice service using generative artificial intelligence. For example, the server may play a standard phrase such as "Hello, is there anything I can help you with?". At this stage, the server activates a voice playback module and provides the user with an initial response.
[0829] The server then receives the user's response and converts it into text data in real time using speech-to-text technology (e.g., Google Cloud Speech-to-Text API). The server then inputs this text data into a generative artificial intelligence (e.g., OpenAI GPT-4) to assess the user's health status and criminal risk. Based on the analysis results, the server then uses an emotion engine (e.g., IBM Watson Tone Analyzer) to evaluate the user's emotional state.
[0830] For example, if a user says, "I've been having headaches lately," the server will send text data containing the keyword "headache" to the generative AI to assess whether there is a health risk. At the same time, the emotion engine will detect anxiety and stress from the user's voice. This will complement the user's emotional data, allowing the system to make a more accurate risk assessment.
[0831] The server integrates data from the generative artificial intelligence and emotion engine to generate an appropriate response for the user. For example, if a user feels anxious, the server generates a response such as "Don't worry, we'll take care of it right away," and plays it back to the user using the voice playback module.
[0832] If the server determines that the user faces a high health or criminal risk as a result of the risk assessment, it will transfer the call to a service provider operator as necessary. For example, if the user is assessed as having a high health risk, it will transfer the call to an operator of an appropriate medical service. The server uses the call transfer function to ensure a prompt response.
[0833] The server also records all conversation content, emotional information, analysis results, and whether or not the call was transferred in a database, which can be used for future analysis and troubleshooting.
[0834] As a concrete example, by inputting the following prompt sentence into a generative artificial intelligence, an appropriate response can be obtained.
[0835] Prompt: "The user mentions that they've been having headaches lately. Use this information to generate an appropriate response. They also sound anxious and stressed."
[0836] Example response: "I'm sorry to hear that you're concerned. Your persistent headache is concerning. Please wait a moment while we connect you to medical services."
[0837] As described above, the system of the present invention can accurately evaluate the user's emotions and risk status in real time and provide prompt and appropriate responses. This system can reduce the sense of loneliness felt by elderly people living alone and enable prompt responses to health conditions and crime risks.
[0838] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0839] Step 1:
[0840] A user calls the system using a communication terminal. The input includes the user's voice, and the output is a connection to the server. Specifically, the user calls the system's phone number using a smartphone or landline phone and presses the call button.
[0841] Step 2:
[0842] The server invokes a voice service, which includes as input the user's call connection and as output the playback of a voice message. Specifically, the server detects when the call is answered and invokes the voice playback module to play a voice message saying, "Hello, how can I help you?"
[0843] Step 3:
[0844] The server converts the user's speech into text data. The input includes the user's speech response, and the converted text data is obtained as the output. Specifically, the server uses speech-to-text conversion technology (e.g., Google Cloud Speech-to-Text API) to convert the user's speech into text data in real time.
[0845] Step 4:
[0846] The server analyzes the text data using generative AI. The converted text data is included as input, and the analysis results are obtained as output. Specifically, the server sends the text data to a generative AI (e.g., OpenAI GPT-4), which then evaluates health and crime risks.
[0847] Step 5:
[0848] The server uses an emotion engine to recognize the user's emotion. The input includes the analysis results and voice data, and the output is emotion information. Specifically, the server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to identify emotions from the user's tone of voice and text.
[0849] Step 6:
[0850] The server generates an appropriate response based on the integrated analysis results. The input includes the analysis results and emotional information, and the generated response is obtained as the output. Specifically, the server provides prompts to the generative AI based on the analysis results and emotional information, converts the generated response into speech, and sends it to the user.
[0851] Step 7:
[0852] The server performs risk assessment and transfers the call to an operator if necessary. The input contains the integrated analysis results, and the output is the required call transfer. Specifically, the server checks the risk assessment and, for example, if it determines that the health risk is high, transfers the call to a medical service operator.
[0853] Step 8:
[0854] The server records all conversation content, emotional information, analysis results, and whether or not the call was forwarded as a log. All communication data is included as input, and log data is obtained as output. Specifically, the server records and saves the conversation content, emotional information, analysis results, and whether or not the call was forwarded in a database.
[0855] (Application example 2)
[0856] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0857] There is a need to reduce the sense of loneliness felt by elderly people living alone, and to provide real-time assessments of their health status and crime risk to respond promptly. However, conventional systems are unable to recognize users' emotions and provide personalized responses, and there are also problems with the speed and accuracy of risk assessments. Therefore, a system that improves the quality of life of elderly people living alone and enables appropriate emergency response is needed.
[0858] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0859] In this invention, the server includes means for allowing a user to communicate using a communication terminal, means for analyzing the conversation with the user using a generative artificial intelligence to assess the user's health condition or criminal risk, means for transferring the communication to a service provider operator as needed based on the assessment, means for converting the conversation into text data using speech recognition, and means for recognizing emotions from the user's voice using an emotion engine and providing that information to the generative artificial intelligence. This enables personalized responses that take the user's emotions into consideration, enabling rapid and appropriate risk assessment and emergency response.
[0860] A "communication terminal" is a device that a user uses to make voice calls or data communications, and specifically includes telephones, smartphones, tablets, and the like.
[0861] "Generative AI" is a type of AI model that analyzes user input data and generates appropriate responses and decisions.
[0862] "Health status assessment" is the process of analyzing a user's health-related statements and behaviors and assessing their health risks.
[0863] "Crime risk assessment" is the process of assessing the risk of crime based on user comments and environmental information.
[0864] A "service provider operator" is a professional operator who is on standby to respond to user risks or emergencies, such as a provider of medical services or security services.
[0865] "Speech recognition" is a technology that captures a user's voice as digital data and converts it into text data.
[0866] "Text data" is data expressed as text information generated by voice recognition technology.
[0867] The "emotion engine" is an engine for analyzing and recognizing the user's emotional state from voice and text data.
[0868] "Logging" refers to the act of saving the contents of conversations conducted by the system, risk assessment results, emergency response history, etc.
[0869] This invention aims to develop a system that reduces loneliness among elderly people living alone and assesses their health status and crime risk in real time. The system allows users to communicate using a communication terminal, analyzes the conversation with the user using a generative AI, and, based on the results, forwards the communication to a service provider operator as needed. Furthermore, the system converts the conversation into text data using speech recognition technology, recognizes emotions from the user's voice using an emotion engine, and provides this information to the generative AI.
[0870] Hardware Configuration
[0871] Communication device: A smartphone or telephone used by a user, which allows voice calls.
[0872] Server: A server with high-performance computing power runs generative artificial intelligence, speech recognition technology, and an emotion engine.
[0873] Software Configuration
[0874] Speech Recognition Library: speech_recognition library
[0875] Speech synthesis engine: pyttsx3 library
[0876] Emotion Recognition Module: emotion_recognition library
[0877] Generative AI model: Our own ai_model library
[0878] Emergency Service Connector: EmergencyServiceConnector module
[0879] Process Overview
[0880] 1. Voice input and analysis: A user makes a voice call using a communication device. The voice data is transferred to the server and converted into text data by a voice recognition library.
[0881] 2. Emotion recognition: The emotion recognition module analyzes the user's emotions from the voice data and provides the results to the generative AI.
[0882] 3. Risk assessment: Generative AI analyzes text data to assess health or crime risks.
[0883] 4. Emergency Calls: Based on risk assessment, emergency services will be called as required.
[0884] 5. Logging: Record the conversation content and evaluation results as logs for future analysis.
[0885] Specific examples
[0886] For example, if a user says, "I've been having headaches lately," the speech recognition technology will detect the keyword "headache," and the generative AI will evaluate this as a health risk. At the same time, the emotion engine will detect anxiety in the voice and generate a response to the user saying, "Don't worry, we'll take care of it right away."
[0887] Prompt Sentence Examples
[0888] text
[0889] User: I've been having a lot of headaches lately.
[0890] Generative AI: The keyword "headache" is detected as a high health risk. The emotion engine recognizes the emotion of anxiety.
[0891] User: I saw something suspicious in my neighborhood.
[0892] Generative AI: The keyword "suspicious person" is detected, which indicates a high risk of crime. The emotion engine recognizes the emotion of fear.
[0893] In this way, the present invention enables personalized responses that take into account the user's emotions, and can provide quick and appropriate risk assessments and emergency responses.
[0894] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0895] Step 1:
[0896] A user initiates a voice call using a communication terminal. The communication terminal uses a microphone to collect voice and transmits it to a server. The input data is the user's voice, and the output is the transfer of voice data to the server.
[0897] Step 2:
[0898] The server uses a speech recognition library (speech_recognition) to convert the received voice data into text data. The input is the user's voice data, and the output is the generated text data.
[0899] Step 3:
[0900] The server uses a generative artificial intelligence model (ai_model library) to analyze the text data. As a result of the analysis, a health or crime risk is assessed. The input is the text data, and the output is the risk assessment result.
[0901] Step 4:
[0902] The server uses an emotion engine (emotion_recognition library) to recognize the user's emotions from the voice data. The results are fed back to the generative AI. The input is the voice data, and the output is the recognized emotion data.
[0903] Step 5:
[0904] Generative AI generates appropriate responses based on text data and emotional data. The responses are formatted as text and used as feedback to the user. The input is text data and emotional data, and the output is the generated response text.
[0905] Step 6:
[0906] The server forwards the communication to the service provider's operator via the Emergency Service Connector as needed. For example, if the health risk is assessed as high, it is forwarded to medical services, and if the crime risk is assessed as high, it is forwarded to security services. The input is the risk assessment result, and the output is forwarding the communication to the operator.
[0907] Step 7:
[0908] The server records conversations, risk assessment results, and emergency response history as logs. These logs are useful for future analysis and problem solving. The input is the conversations and assessment results, and the output is the log records.
[0909] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0910] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0911] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0912] [Fourth embodiment]
[0913] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0914] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0915] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0916] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0917] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0918] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0919] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0920] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0921] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0922] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0923] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0924] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0925] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0926] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate operator of a service provider as necessary. Specific operations and processing of an embodiment of this system will be described below.
[0927] Specific Embodiments of the System
[0928] A user makes a call using a communication terminal
[0929] A user calls the system using a communication terminal (for example, a telephone or a smartphone). The communication terminal starts a voice call and is connected to the server.
[0930] The server starts the audio service
[0931] When the server receives a call from a user, it activates a voice service using generative artificial intelligence. The server starts a conversation by saying to the user, "Hello, is there anything I can help you with?"
[0932] Conversation initiation and analysis
[0933] The server converts the user's responses into text data in real time using speech-to-text technology. Generative AI then analyzes this text data to assess the user's health status and criminal risk.
[0934] Specific examples
[0935] For example, if a user says, "I've been having headaches lately," the server will analyze the text data containing the keyword "headache" and evaluate it as a health risk. Similarly, if a user says, "I saw a suspicious person," the server will evaluate the crime risk from the keyword "suspicious person."
[0936] Risk assessment and call transfer to an operator
[0937] The server transfers calls to service provider operators if necessary based on the health or crime risk assessed by the generative AI. For example, if the health risk is determined to be high, the call is transferred to a HELPO operator, and if the crime risk is determined to be high, the call is transferred to a security service operator.
[0938] Logging the results
[0939] The server records all conversations, their analysis results, and whether or not the call was transferred in detail as a log, which can be used for future analysis and verification when a problem occurs.
[0940] In this way, the system of the present invention is a system that reduces the user's sense of loneliness, detects health conditions and crime risks early, and takes necessary measures promptly. The objectives of the present invention can be achieved by appropriately arranging communication terminals, generative artificial intelligence, speech-to-text conversion technology, and log recording means to implement the invention.
[0941] The processing flow will be explained below.
[0942] Step 1:
[0943] A user makes a call using a communication terminal.
[0944] The communication terminal initiates a voice call and is connected to the server.
[0945] Step 2:
[0946] The server receives a call from a user.
[0947] The server launches a voice service using generative artificial intelligence and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0948] Step 3:
[0949] The server waits for the user's response.
[0950] The communication terminal picks up the user's voice and transmits it to the server.
[0951] Step 4:
[0952] The server converts the received voice data into text data using voice-to-text conversion technology.
[0953] The server analyzes the text data using generative artificial intelligence.
[0954] Step 5:
[0955] The server assesses the user's health and criminal risk.
[0956] Generative artificial intelligence detects keywords such as "headache," "dizziness," "suspicious person," and "fraud," and determines health or criminal risks based on these.
[0957] Step 6:
[0958] Based on the risk assessment result, the server prepares to transfer the call to an operator of the service provider as necessary.
[0959] If the server is determined to pose a high health risk, it will contact a HELPO operator, and if it is determined to pose a high crime risk, it will contact a security service operator.
[0960] Step 7:
[0961] The server notifies the user that "Important information has been detected and we will connect you to an operator."
[0962] The server then transfers the call to the appropriate operator.
[0963] Step 8:
[0964] The server logs all conversations, their analysis results, and whether or not the call was forwarded.
[0965] The server records logs to allow for future analysis and verification when problems occur.
[0966] Through this procedure, the server, communication terminal, and user cooperate with each other to realize the system of the present invention.
[0967] Example 1
[0968] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0969] It is desirable to provide a means to reduce loneliness among elderly people living alone, to detect health conditions and crime risks early, and to take appropriate measures promptly. There is also a need for a system that can effectively analyze the content of conversations with users and transfer calls to the necessary service providers based on the results.
[0970] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0971] In this invention, the server includes a means for a user to communicate using an information terminal, a means for analyzing the conversation with the user using generative artificial intelligence to evaluate the user's health condition or criminal risk, a means for transferring the communication to a service provider's operator as needed based on the evaluation, a means for acquiring the user's conversation content as text data using speech-to-text technology, and a means for recording the conversation content and analysis results in a log. This reduces the sense of loneliness felt by elderly people living alone, enables early detection of health conditions and criminal risk, and allows for prompt response. Furthermore, the server can evaluate risk based on the user's speech and quickly transfer the call to an appropriate service provider.
[0972] "Information terminal" refers to any device used for communication, and specifically includes telephones, smartphones, computers, etc.
[0973] "Generative AI" refers to AI technology that analyzes conversations with users and makes decisions and suggestions based on that analysis.
[0974] "Health status" refers to general information about the user's health, such as the user's physical condition and whether or not the user has any illnesses.
[0975] "Crime risk" refers to the possibility that a user will become involved in a crime or the risk that a crime will occur.
[0976] "Service provider" means the operator of a business or organisation that provides health and safety services.
[0977] "Speech-to-text technology" refers to technology that converts voice data into text data in real time.
[0978] "Log" refers to records of conversation content, analysis results, communication transfer history, etc.
[0979] "Conversation content" refers to the content of the voice communication exchanged between the user and the system (or generative artificial intelligence).
[0980] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an operator of an appropriate service provider as necessary. Specific embodiments of the present invention will be described below.
[0981] First, a user calls a specific phone number using an information terminal (e.g., a telephone or smartphone). When the voice call begins, the terminal connects to the server. When the server receives the call from the user, it immediately answers and launches a voice-enabled application using generative artificial intelligence. Specifically, the server speaks to the user, saying, "Hello, is there anything I can help you with?"
[0982] The server then uses speech-to-text technology (e.g., Google Cloud Speech-to-Text API) to convert the user's speech into text data in real time. Once the user's speech is converted into text data, the text data is stored on the server.
[0983] The server analyzes the text data using generative artificial intelligence (for example, OpenAI's GPT-4 model). This analysis evaluates the user's health condition and criminal risk. For example, if a user says, "I've been having headaches lately," the server extracts the keyword "headache" and determines that this corresponds to a health risk. Similarly, if a user says, "I saw a suspicious person," the keyword "suspicious person" will evaluate the criminal risk as high.
[0984] Based on the evaluation results, the server transfers the call to a service provider operator as needed. For example, if the call is judged to be a high health risk, the server transfers the call to a health support service operator, and if the call is judged to be a high crime risk, the server transfers the call to a security service operator.
[0985] In addition, the server records detailed logs of all conversations, analysis results, and whether or not a call was forwarded, which can be used for future analysis and verification if a problem occurs.
[0986] Prompt Sentence Examples
[0987] Prompt: "If a user says, 'I've been having a lot of headaches lately,' explain how you can use a generative AI model to assess health risks and route the call to a health support service operator."
[0988] Prompt: "When a user says 'I see a suspicious person,' explain how a generative AI model can be used to assess the crime risk and route the call to a security service operator."
[0989] As described above, the system of the present invention can reduce the user's sense of loneliness, detect health conditions and crime risks early, and take necessary measures promptly.
[0990] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0991] Step 1:
[0992] A user uses an information terminal to call a specified phone number. When the voice call begins, the terminal connects to the server. The input is the user's voice, and the output is a connection with the server. Specifically, the user operates a telephone or smartphone to call the phone number specified by the system.
[0993] Step 2:
[0994] The server receives a call and activates a voice service. The input is an incoming call signal from the user, and the output is a voice-enabled application powered by generative artificial intelligence. Specifically, the server detects the incoming call, activates an auto-answer function, and speaks to the user, saying, "Hello, is there anything I can help you with?"
[0995] Step 3:
[0996] The server uses speech-to-text technology to convert the user's speech into text data in real time. The input is the user's speech, and the output is text data. Specifically, the server uses the Google Cloud Speech-to-Text API to convert the speech into text data.
[0997] Step 4:
[0998] The server analyzes the text data using generative artificial intelligence. The input is text data, and the output is an assessment result regarding the user's health status and crime risk. Specifically, the server uses generative artificial intelligence (e.g., GPT-4) to analyze the text data and perform a risk assessment.
[0999] Step 5:
[1000] The server transfers the call to a service provider's operator as needed based on the risk assessment results. The input is the risk assessment result, and the output is call transfer. As a specific example, if the health risk is determined to be high, the server transfers the call to a health support operator, and if the crime risk is determined to be high, the server transfers the call to a security service operator.
[1001] Step 6:
[1002] The server records all conversation content, analysis results, and whether or not the call was forwarded in detail as a log. The input is the conversation content with the user and the analysis results, and the output is a log file. Specifically, the server processes the conversation content, analysis results, and call history in a recording database.
[1003] Through the above processing steps, the system of the present invention can reduce the user's sense of loneliness, detect health conditions and crime risks early, and take necessary measures promptly.
[1004] (Application example 1)
[1005] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1006] As the elderly population increases, feelings of loneliness, health risks, and crime risks among those living alone are becoming social issues. Elderly people living alone face particular challenges, such as difficulty in responding immediately to health problems or crime risks. As a result, there is a need for a system that can reduce feelings of loneliness while ensuring their safety and peace of mind. To solve this issue, it is important to assess the elderly's situation in real time via voice calls and quickly connect them to appropriate services as needed.
[1007] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1008] In this invention, the server includes: means for a user to use a communication terminal to perform predetermined communications; means for analyzing conversations with the user using generative artificial intelligence to assess the user's health status or criminal risk; means for transferring communications to a service provider operator as needed based on the assessment; means for converting the content of the voice call into text data in real time using speech-to-text technology; means for assessing the risk of exceeding a predetermined threshold based on the analysis results; means for quickly transferring the call to an appropriate operator if the risk exceeds a specific threshold; and means for recording all conversation contents, analysis results, and call transfer history. This allows for rapid detection of potential risks faced by users and enables prompt response.
[1009] A "communication terminal" is a device that a user uses to perform predetermined communications, and includes telephones and smartphones.
[1010] "Generative AI" is an AI technology that analyzes voice and text data to assess a user's health status and crime risk.
[1011] "Speech-to-text technology" is a technology that converts voice data into text data in real time.
[1012] "Evaluation" refers to the act of using generative artificial intelligence to analyze the content of a user's conversations and determine their health status and criminal risk.
[1013] A "service provider operator" is a specialized service person who takes action when a user is assessed as being at high risk for health or crime.
[1014] A "threshold" is a reference value that is set when the evaluation results of the generative artificial intelligence exceed a specific risk level.
[1015] "Call forwarding" is the process of quickly connecting a communication taken from a user to an operator of another service provider.
[1016] A "log" is a detailed record of conversation content, analysis results, call forwarding history, etc.
[1017] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate operator of a service provider as necessary. Specific operations and processing of an embodiment of this system will be described below.
[1018] System Hardware and Software
[1019] The following hardware and software are used to implement this system:
[1020] Hardware:
[1021] Smartphone microphone and speaker
[1022] Networked Servers
[1023] software:
[1024] Python
[1025] speech_recognition library
[1026] The requests library
[1027] Generative AI model
[1028] System operation explanation
[1029] 1. Start a voice call:
[1030] A user starts the application on their smartphone and taps the "emergency call" button, which establishes a real-time voice call between the user and the server.
[1031] 2. Speech recognition and text conversion:
[1032] The voice captured by the smartphone microphone is converted into text data in real time using the speech_recognition library, and the user's speech is then sent to the server as text data.
[1033] 3. Analysis by generative artificial intelligence:
[1034] The server uses a generative artificial intelligence model to analyze the acquired text data. This analysis evaluates the user's health condition and crime risk. For example, if a user says, "I've been having headaches lately," the keyword "headache" is analyzed and evaluated as a health risk. Similarly, if a user says, "I saw a suspicious person," the keyword "suspicious person" is used to evaluate the crime risk.
[1035] 4. Risk Assessment and Call Forwarding:
[1036] Based on the analysis results, the server evaluates whether the risk exceeds a certain threshold. If the risk is high, the generative AI model quickly routes the call to the appropriate service operator based on the risk assessment. For example, if the call is judged to be a high health risk, it will be routed to a medical service operator, and if the call is judged to be a high crime risk, it will be routed to a security service operator.
[1037] 5. Logging:
[1038] The server records all conversations, their analysis results, and call forwarding history in detail, and stores them in protected cloud storage for future analysis and verification in case of problems.
[1039] Specific prompt examples
[1040] Examples:
[1041] If a user says, "I'm scared because I've seen a lot of suspicious people walking down the street lately," the system will detect the phrase "suspicious people" and evaluate them as a crime risk, immediately transferring the call to a security service provider.
[1042] Example prompt sentence:
[1043] Text: "Recently, I've been seeing a lot of suspicious people walking down the street at night and it's scary."
[1044] Prompt: "Analyze this text to assess its health and crime risks."
[1045] In this way, the system provides a practical solution that ensures the safety and security of elderly people living alone, while also allowing for a rapid response.
[1046] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1047] Step 1:
[1048] A user starts an application using a communication terminal and taps the "emergency call" button. This operation causes the communication terminal to establish a voice call with the server. The user's voice is acquired as input, and a call with the server is initiated as output.
[1049] Step 2:
[1050] The server receives the voice data acquired from the communication terminal. The voice data is sent as input and recorded on the server as output.
[1051] Step 3:
[1052] The server uses the speech_recognition library to convert the received voice data into text data in real time. Voice data is used as input and text data is generated as output. Specifically, the server analyzes each phoneme using a speech recognition engine and generates the corresponding string.
[1053] Step 4:
[1054] The server inputs the generated text data into a generative artificial intelligence (AI) model to evaluate the user's health status and crime risk. The text data is used as input, and the health and crime risk assessment results are obtained as output. Specifically, the AI model uses natural language processing technology to analyze keywords and context within the text and conduct a risk assessment.
[1055] Step 5:
[1056] The server determines whether the risk exceeds a certain threshold based on the assessment results. The risk assessment results are used as input, and a judgment of whether the risk is high or low is obtained as output. Specifically, the risk level is determined by comparing it with a predefined threshold.
[1057] Step 6:
[1058] If the server determines that the call is high risk, it transfers the call to an operator at the appropriate service provider. The server uses the risk assessment results and threshold judgment results as inputs, and executes the call transfer as output. Specifically, this includes the process of rerouting the call to a specific phone number.
[1059] Step 7:
[1060] The server records all conversation content, assessment results, and call forwarding history in detailed logs and stores them in protected cloud storage. Assessment results and call history are used as input, and detailed log data is generated as output. Specifically, information such as date and time, conversation content, risk assessment results, and forwarding destinations is written to the log file.
[1061] In this way, the system assesses the situation of elderly people living alone in real time and takes prompt action as needed, providing peace of mind and safety.
[1062] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1063] The present invention relates to a system that reduces the sense of loneliness felt by elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to an appropriate service provider operator as needed. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and generates responses based on those emotions, more personalized responses are possible for users.
[1064] Specific Embodiments of the System
[1065] A user makes a call using a communication terminal
[1066] A user calls the system using a communication terminal (for example, a telephone or a smartphone). The communication terminal starts a voice call and is connected to the server.
[1067] The server starts the audio service
[1068] When the server receives a call from a user, it activates a voice service using generative artificial intelligence. The server starts a conversation by saying to the user, "Hello, is there anything I can help you with?"
[1069] Conversation initiation and analysis
[1070] The server converts the user's responses into text data in real time using speech-to-text technology. The generative AI analyzes this text data and assesses the user's health status and criminal risk. At the same time, the emotion engine recognizes emotions from the user's voice and provides this information to the generative AI.
[1071] Specific examples
[1072] For example, if a user says, "I've been having headaches lately," the server analyzes text data containing the keyword "headache" and evaluates it as a health risk. Meanwhile, the emotion engine detects emotions such as anxiety and stress from the user's voice. Similarly, if a user says, "I saw a suspicious person," the server evaluates the crime risk from the keyword "suspicious person," and the emotion engine recognizes fear and tension.
[1073] Generating responses based on emotion information
[1074] The server uses generative AI to generate an appropriate response based on the user's emotional information provided by the emotion engine. For example, if a user is feeling anxious, the server will respond with "Don't worry, we'll take care of it right away."
[1075] Risk assessment and call transfer to an operator
[1076] The server transfers the call to a service provider operator if necessary based on the health or criminal risk assessed by the generative artificial intelligence. For example, if the health risk is determined to be high, the call is transferred to an appropriate medical service operator, and if the criminal risk is determined to be high, the call is transferred to a security service operator.
[1077] Logging the results
[1078] The server logs all conversations, emotional information, analysis results, and whether or not a call was forwarded. These logs are used for future analysis and verification when problems occur.
[1079] summary
[1080] In this way, the system of the present invention reduces the user's sense of loneliness, detects health conditions and crime risks early, and promptly provides appropriate services. Furthermore, incorporating an emotion engine enables more compassionate and personalized responses to users. The objectives of the present invention can be achieved by appropriately arranging a communication terminal, generative artificial intelligence, speech-to-text technology, an emotion engine, and a means for logging in order to implement the invention.
[1081] The processing flow will be explained below.
[1082] Step 1:
[1083] A user makes a call using a communication terminal.
[1084] The communication terminal initiates a voice call and is connected to the server.
[1085] Step 2:
[1086] The server receives a call from a user.
[1087] The server launches a voice service using generative artificial intelligence and speaks to the user, saying, "Hello, is there anything I can help you with?"
[1088] Step 3:
[1089] The server waits for the user's response.
[1090] The communication terminal picks up the user's voice and transmits it to the server.
[1091] Step 4:
[1092] The server converts the received voice data into text data using voice-to-text conversion technology.
[1093] The server analyzes the text data using generative artificial intelligence.
[1094] Step 5:
[1095] At the same time, the server uses an emotion engine to recognize emotions from the user's voice.
[1096] The emotion engine provides the recognized emotion information to the generative artificial intelligence.
[1097] Step 6:
[1098] The server evaluates the user's health status and crime risk based on text data and emotional information analyzed by generative artificial intelligence.
[1099] For example, if a user says, "I've been having headaches lately," the server detects the keyword "headache," and the emotion engine recognizes the emotion of anxiety.
[1100] Step 7:
[1101] The server uses generative artificial intelligence to generate an appropriate response based on emotional information.
[1102] For example, a user who is feeling anxious can be given a response such as "Don't worry, we'll take care of it right away."
[1103] Step 8:
[1104] Based on the evaluation results, the server transfers the call to an operator at the service provider as necessary.
[1105] If the health risk is determined to be high, the call will be transferred to a medical service operator, and if the crime risk is determined to be high, the call will be transferred to a security service operator.
[1106] Step 9:
[1107] The server performs the call transfer and notifies the user that "Important content has been detected and you will be connected to an operator."
[1108] The server then transfers the call to the appropriate operator.
[1109] Step 10:
[1110] The server records all conversation content, analysis results, emotional information, and whether or not the call was forwarded as a log.
[1111] This record will be used for future analysis and verification when problems occur.
[1112] This specific procedure enables the server, communication terminal, and user to work together, reducing the user's sense of loneliness, detecting health conditions and crime risks early, and taking necessary measures quickly.
[1113] Example 2
[1114] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1115] In modern society, elderly people living alone are increasingly experiencing loneliness, health problems, and crime risks, necessitating rapid and appropriate responses. However, conventional monitoring and warning systems have difficulty accurately grasping a user's emotional state and providing personalized responses based on that information. Furthermore, they lack the ability to adequately assess risk in real time and quickly transfer calls to service providers. To address these challenges, a system is needed that can accurately assess a user's emotions, health risks, and crime risks, and provide appropriate responses.
[1116] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1117] In this invention, the server includes: means for allowing a user to communicate using a communication terminal; means for analyzing a conversation with the user and assessing their health condition or criminal risk using generative artificial intelligence; means for converting the user's voice data into text data using speech-to-text technology; means including an emotion engine for evaluating the user's emotional state based on the analysis results; means for transferring the communication to a service provider operator as needed based on the evaluation; and means for recording the content of the conversation, emotional information, analysis results, and whether or not the call was transferred. This allows for accurate evaluation of the user's emotions and risk state in real time, enabling prompt and appropriate responses.
[1118] "Communication terminal" refers to any device that a user uses to conduct voice or data communications, and specifically includes smartphones and landline phones.
[1119] "Generative AI" refers to AI technology that analyzes user input and generates appropriate responses and analytical results, and includes systems that utilize natural language processing and machine learning.
[1120] "Speech-to-text technology" refers to technology that converts a user's voice into text data in real time, and includes, for example, cloud-based voice recognition services.
[1121] "Emotion engine" refers to technology that analyzes a user's emotional state from their voice or text data and provides that information, and includes systems that utilize emotion recognition algorithms and machine learning.
[1122] "Service provider operators" refer to professional personnel who are on standby to provide a particular service, including those working in fields such as healthcare and security.
[1123] "Database" means a system of record for storing and managing system-generated data, including the means for efficiently storing and accessing data in various formats.
[1124] "Log" refers to historical data that records the processes performed by the system, the data generated, events, etc., and is used for future analysis and troubleshooting.
[1125] This invention relates to a system that reduces loneliness among elderly people living alone, assesses their health status and crime risk in real time, and transfers calls to service provider operators as needed. This system uses generative artificial intelligence to analyze user conversations and utilizes speech-to-text technology and an emotion engine to provide more personalized responses.
[1126] A user calls the system using a communication terminal. The communication terminal can be a landline phone or a smartphone. When the user makes a call, the server receives the call and activates a voice service using generative artificial intelligence. For example, the server may play a standard phrase such as "Hello, is there anything I can help you with?". At this stage, the server activates a voice playback module and provides the user with an initial response.
[1127] The server then receives the user's response and converts it into text data in real time using speech-to-text technology (e.g., Google Cloud Speech-to-Text API). The server then inputs this text data into a generative artificial intelligence (e.g., OpenAI GPT-4) to assess the user's health status and criminal risk. Based on the analysis results, the server then uses an emotion engine (e.g., IBM Watson Tone Analyzer) to evaluate the user's emotional state.
[1128] For example, if a user says, "I've been having headaches lately," the server will send text data containing the keyword "headache" to the generative AI to assess whether there is a health risk. At the same time, the emotion engine will detect anxiety and stress from the user's voice. This will complement the user's emotional data, allowing the system to make a more accurate risk assessment.
[1129] The server integrates data from the generative artificial intelligence and emotion engine to generate an appropriate response for the user. For example, if a user feels anxious, the server generates a response such as "Don't worry, we'll take care of it right away," and plays it back to the user using the voice playback module.
[1130] If the server determines that the user faces a high health or criminal risk as a result of the risk assessment, it will transfer the call to a service provider operator as necessary. For example, if the user is assessed as having a high health risk, it will transfer the call to an operator of an appropriate medical service. The server uses the call transfer function to ensure a prompt response.
[1131] The server also records all conversation content, emotional information, analysis results, and whether or not the call was transferred in a database, which can be used for future analysis and troubleshooting.
[1132] As a concrete example, by inputting the following prompt sentence into a generative artificial intelligence, an appropriate response can be obtained.
[1133] Prompt: "The user mentions that they've been having headaches lately. Use this information to generate an appropriate response. They also sound anxious and stressed."
[1134] Example response: "I'm sorry to hear that you're concerned. Your persistent headache is concerning. Please wait a moment while we connect you to medical services."
[1135] As described above, the system of the present invention can accurately evaluate the user's emotions and risk status in real time and provide prompt and appropriate responses. This system can reduce the sense of loneliness felt by elderly people living alone and enable prompt responses to health conditions and crime risks.
[1136] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1137] Step 1:
[1138] A user calls the system using a communication terminal. The input includes the user's voice, and the output is a connection to the server. Specifically, the user calls the system's phone number using a smartphone or landline phone and presses the call button.
[1139] Step 2:
[1140] The server invokes a voice service, which includes as input the user's call connection and as output the playback of a voice message. Specifically, the server detects when the call is answered and invokes the voice playback module to play a voice message saying, "Hello, how can I help you?"
[1141] Step 3:
[1142] The server converts the user's speech into text data. The input includes the user's speech response, and the converted text data is obtained as the output. Specifically, the server uses speech-to-text conversion technology (e.g., Google Cloud Speech-to-Text API) to convert the user's speech into text data in real time.
[1143] Step 4:
[1144] The server analyzes the text data using generative AI. The converted text data is included as input, and the analysis results are obtained as output. Specifically, the server sends the text data to a generative AI (e.g., OpenAI GPT-4), which then evaluates health and crime risks.
[1145] Step 5:
[1146] The server uses an emotion engine to recognize the user's emotion. The input includes the analysis results and voice data, and the output is emotion information. Specifically, the server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to identify emotions from the user's tone of voice and text.
[1147] Step 6:
[1148] The server generates an appropriate response based on the integrated analysis results. The input includes the analysis results and emotional information, and the generated response is obtained as the output. Specifically, the server provides prompts to the generative AI based on the analysis results and emotional information, converts the generated response into speech, and sends it to the user.
[1149] Step 7:
[1150] The server performs risk assessment and transfers the call to an operator if necessary. The input contains the integrated analysis results, and the output is the required call transfer. Specifically, the server checks the risk assessment and, for example, if it determines that the health risk is high, transfers the call to a medical service operator.
[1151] Step 8:
[1152] The server records all conversation content, emotional information, analysis results, and whether or not the call was forwarded as a log. All communication data is included as input, and log data is obtained as output. Specifically, the server records and saves the conversation content, emotional information, analysis results, and whether or not the call was forwarded in a database.
[1153] (Application example 2)
[1154] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1155] There is a need to reduce the sense of loneliness felt by elderly people living alone, and to provide real-time assessments of their health status and crime risk to respond promptly. However, conventional systems are unable to recognize users' emotions and provide personalized responses, and there are also problems with the speed and accuracy of risk assessments. Therefore, a system that improves the quality of life of elderly people living alone and enables appropriate emergency response is needed.
[1156] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1157] In this invention, the server includes means for allowing a user to communicate using a communication terminal, means for analyzing the conversation with the user using a generative artificial intelligence to assess the user's health condition or criminal risk, means for transferring the communication to a service provider operator as needed based on the assessment, means for converting the conversation into text data using speech recognition, and means for recognizing emotions from the user's voice using an emotion engine and providing that information to the generative artificial intelligence. This enables personalized responses that take the user's emotions into consideration, enabling rapid and appropriate risk assessment and emergency response.
[1158] A "communication terminal" is a device that a user uses to make voice calls or data communications, and specifically includes telephones, smartphones, tablets, and the like.
[1159] "Generative AI" is a type of AI model that analyzes user input data and generates appropriate responses and decisions.
[1160] "Health status assessment" is the process of analyzing a user's health-related statements and behaviors and assessing their health risks.
[1161] "Crime risk assessment" is the process of assessing the risk of crime based on user comments and environmental information.
[1162] A "service provider operator" is a professional operator who is on standby to respond to user risks or emergencies, such as a provider of medical services or security services.
[1163] "Speech recognition" is a technology that captures a user's voice as digital data and converts it into text data.
[1164] "Text data" is data expressed as text information generated by voice recognition technology.
[1165] The "emotion engine" is an engine for analyzing and recognizing the user's emotional state from voice and text data.
[1166] "Logging" refers to the act of saving the contents of conversations conducted by the system, risk assessment results, emergency response history, etc.
[1167] This invention aims to develop a system that reduces loneliness among elderly people living alone and assesses their health status and crime risk in real time. The system allows users to communicate using a communication terminal, analyzes the conversation with the user using a generative AI, and, based on the results, forwards the communication to a service provider operator as needed. Furthermore, the system converts the conversation into text data using speech recognition technology, recognizes emotions from the user's voice using an emotion engine, and provides this information to the generative AI.
[1168] Hardware Configuration
[1169] Communication device: A smartphone or telephone used by a user, which allows voice calls.
[1170] Server: A server with high-performance computing power runs generative artificial intelligence, speech recognition technology, and an emotion engine.
[1171] Software Configuration
[1172] Speech Recognition Library: speech_recognition library
[1173] Speech synthesis engine: pyttsx3 library
[1174] Emotion Recognition Module: emotion_recognition library
[1175] Generative AI model: Our own ai_model library
[1176] Emergency Service Connector: EmergencyServiceConnector module
[1177] Process Overview
[1178] 1. Voice input and analysis: A user makes a voice call using a communication device. The voice data is transferred to the server and converted into text data by a voice recognition library.
[1179] 2. Emotion recognition: The emotion recognition module analyzes the user's emotions from the voice data and provides the results to the generative AI.
[1180] 3. Risk assessment: Generative AI analyzes text data to assess health or crime risks.
[1181] 4. Emergency Calls: Based on risk assessment, emergency services will be called as required.
[1182] 5. Logging: Record the conversation content and evaluation results as logs for future analysis.
[1183] Specific examples
[1184] For example, if a user says, "I've been having headaches lately," the speech recognition technology will detect the keyword "headache," and the generative AI will evaluate this as a health risk. At the same time, the emotion engine will detect anxiety in the voice and generate a response to the user saying, "Don't worry, we'll take care of it right away."
[1185] Prompt Sentence Examples
[1186] text
[1187] User: I've been having a lot of headaches lately.
[1188] Generative AI: The keyword "headache" is detected as a high health risk. The emotion engine recognizes the emotion of anxiety.
[1189] User: I saw something suspicious in my neighborhood.
[1190] Generative AI: The keyword "suspicious person" is detected, which indicates a high risk of crime. The emotion engine recognizes the emotion of fear.
[1191] In this way, the present invention enables personalized responses that take into account the user's emotions, and can provide quick and appropriate risk assessments and emergency responses.
[1192] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1193] Step 1:
[1194] A user initiates a voice call using a communication terminal. The communication terminal uses a microphone to collect voice and transmits it to a server. The input data is the user's voice, and the output is the transfer of voice data to the server.
[1195] Step 2:
[1196] The server uses a speech recognition library (speech_recognition) to convert the received voice data into text data. The input is the user's voice data, and the output is the generated text data.
[1197] Step 3:
[1198] The server uses a generative artificial intelligence model (ai_model library) to analyze the text data. As a result of the analysis, a health or crime risk is assessed. The input is the text data, and the output is the risk assessment result.
[1199] Step 4:
[1200] The server uses an emotion engine (emotion_recognition library) to recognize the user's emotions from the voice data. The results are fed back to the generative AI. The input is the voice data, and the output is the recognized emotion data.
[1201] Step 5:
[1202] Generative AI generates appropriate responses based on text data and emotional data. The responses are formatted as text and used as feedback to the user. The input is text data and emotional data, and the output is the generated response text.
[1203] Step 6:
[1204] The server forwards the communication to the service provider's operator via the Emergency Service Connector as needed. For example, if the health risk is assessed as high, it is forwarded to medical services, and if the crime risk is assessed as high, it is forwarded to security services. The input is the risk assessment result, and the output is forwarding the communication to the operator.
[1205] Step 7:
[1206] The server records conversations, risk assessment results, and emergency response history as logs. These logs are useful for future analysis and problem solving. The input is the conversations and assessment results, and the output is the log records.
[1207] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1208] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1209] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1210] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1211] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1212] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1213] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1214] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1215] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1216] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1217] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1218] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1219] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1220] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1221] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1222] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1223] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1224] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1225] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1226] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1227] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1228] The following is further disclosed regarding the above embodiment.
[1229] (Claim 1)
[1230] means for a user to perform predetermined communication using a communication terminal;
[1231] A means for analyzing conversations with users using generative artificial intelligence to assess their health status or criminal risk;
[1232] means for forwarding the communication to a service provider operator as necessary based on said evaluation;
[1233] A system including:
[1234] (Claim 2)
[1235] 10. The system of claim 1, further comprising: means for acquiring the conversation content as text data using speech-to-text technology.
[1236] (Claim 3)
[1237] 10. The system of claim 1, further comprising means for recording analysis results and call logs.
[1238] "Example 1"
[1239] (Claim 1)
[1240] means for a user to perform predetermined communication using an information terminal;
[1241] A means for analyzing conversations with users using generative artificial intelligence to assess their health status or criminal risk;
[1242] means for forwarding the communication to a service provider operator as necessary based on said evaluation;
[1243] A means for acquiring the content of a user's conversation as text data using speech-to-text conversion technology;
[1244] A means for recording a log of the conversation content and analysis results;
[1245] A system including:
[1246] (Claim 2)
[1247] The system of claim 1, further comprising means for the generative artificial intelligence to accurately assess health risks and crime risks based on the content of user statements.
[1248] (Claim 3)
[1249] 10. The system of claim 1, wherein the speech-to-text technology further comprises means for converting user utterances into text data in real time.
[1250] "Application Example 1"
[1251] (Claim 1)
[1252] means for a user to perform predetermined communication using a communication terminal;
[1253] A means for analyzing conversations with users using generative artificial intelligence to assess their health status or criminal risk;
[1254] means for forwarding the communication to a service provider operator as necessary based on said evaluation;
[1255] A means for converting the contents of a voice call into text data in real time using speech-to-text conversion technology;
[1256] A means for assessing the risk of exceeding a predetermined threshold based on the analysis results;
[1257] If the risk exceeds a certain threshold, a means of quickly transferring the call to an appropriate operator;
[1258] A means of recording all conversations, analysis results, and call forwarding history;
[1259] A system including:
[1260] (Claim 2)
[1261] 10. The system of claim 1, further comprising: means for acquiring the conversation content as text data using speech-to-text technology.
[1262] (Claim 3)
[1263] 10. The system of claim 1, further comprising means for recording analysis results and call logs.
[1264] "Example 2: Combining Emotion Engines"
[1265] (Claim 1)
[1266] means for a user to perform predetermined communication using a communication terminal;
[1267] A means for analyzing conversations with users using generative artificial intelligence to assess their health status or criminal risk;
[1268] means for converting the user's voice data into text data using speech-to-text technology;
[1269] means including an emotion engine for assessing an emotional state of the user based on the analysis results;
[1270] means for forwarding the communication to a service provider operator as necessary based on said evaluation;
[1271] A means for recording the content of the conversation, emotional information, analysis results, and whether or not the call is forwarded;
[1272] A system including:
[1273] (Claim 2)
[1274] 10. The system of claim 1, further comprising means for obtaining the text data using speech-to-text technology.
[1275] (Claim 3)
[1276] 10. The system of claim 1, further comprising means for recording analysis results and call logs.
[1277] "Application example 2 when combining emotion engines"
[1278] (Claim 1)
[1279] means for a user to perform predetermined communication using a communication terminal;
[1280] A means for analyzing conversations with users using generative artificial intelligence to assess their health status or criminal risk;
[1281] means for forwarding the communication to a service provider operator as necessary based on said evaluation;
[1282] a means for converting the conversation into text data using speech recognition;
[1283] A means for recognizing emotions from a user's voice using an emotion engine and providing the information to a generative artificial intelligence;
[1284] A system including:
[1285] (Claim 2)
[1286] 10. The system of claim 1, further comprising: means for acquiring the conversation content as text data using speech-to-text technology.
[1287] (Claim 3)
[1288] 10. The system of claim 1, further comprising means for recording analysis results and call logs. [Explanation of symbols]
[1289] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for a user to perform predetermined communication using a communication terminal; A means for analyzing conversations with users using generative artificial intelligence to assess their health status or criminal risk; means for forwarding the communication to a service provider operator as necessary based on said evaluation; A system including:
2. The system according to claim 1 , further comprising means for acquiring the conversation content as text data using speech-to-text technology.
3. The system of claim 1 further comprising means for recording analysis results and call logs.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A