System
A system that collects and summarizes voice data using a generative model to convey the deceased's wishes to bereaved family members, simplifying the creation of legally binding wills by organizing and summarizing voice data into readable formats and providing expert guidance.
Patent Information
- Application Number
- JP2024118193
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Existing methods for conveying the wishes of the deceased, such as wills and video letters, are time-consuming, require literary skill, and lack real-time communication, making it difficult for bereaved family members to understand the deceased's thoughts, and the process of creating a legally binding will is complex without expert support.
A system that collects voice data, converts it into text, summarizes it using a generative model, and provides summarized information through a user interface, while also supporting the creation of a legally binding will with expert guidance.
Enables easy and effective communication of the deceased's wishes to bereaved family members, simplifying the process of creating a legally valid will by organizing and summarizing voice data into readable formats and providing expert assistance.
Smart Images

Figure 2026017411000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Traditionally, the deceased could communicate their wishes to their bereaved family members through "wills," letters, and video letters, but these required a lot of time, literary skill, and technical work, so they were rarely used. Furthermore, there was a lack of ways to share the deceased's thoughts in real time, making it difficult for bereaved family members to learn their wishes. Furthermore, creating a legally binding will required the support of an expert, which was also a time-consuming challenge. [Means for solving the problem]
[0005] The present invention provides a system including means for collecting voice data from a user, means for converting the voice data into text data, means for storing the text data in a database, means for using a generative model to analyze and summarize the stored text data, and means for providing summarized information through a user interface. Furthermore, the system includes means for outputting the summarized information in letter format based on a user request, and means for collecting information necessary to support the creation of a legally binding will and providing expert guidance, thereby enabling the wishes of the deceased to be easily and effectively conveyed to the bereaved.
[0006] "User" refers to any individual or deceased person who uses the System.
[0007] "Voice Data" refers to a collection of digital information, including a user's voice, collected by a device such as a microphone.
[0008] "Text data" refers to audio data converted into text format.
[0009] "Database" refers to software or a system for systematically storing and managing text data and other related information.
[0010] A "generative model" refers to an algorithm or program that uses artificial intelligence to summarize data or generate text.
[0011] "User interface" refers to the interface means through which a user interacts with a system, and typically includes a display and input devices.
[0012] "Letter format" refers to a format in which personal messages and memories are organized into text documents.
[0013] A "legally valid will" is a will that is valid under law.
[0014] "Experts" refer to lawyers and notaries who have the knowledge and qualifications to prepare legally binding wills and to carry out notarization procedures. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and a specific embodiment thereof is described below. This system includes a process in which a user uses a smart speaker to collect voice data, converts the data into text data, organizes and summarizes it, and provides it to the bereaved.
[0037] System Configuration
[0038] 1. Smart Speaker
[0039] This system collects voice data by having the user (deceased person) speak into a smart speaker, which is equipped with a voice recognition engine that instantly converts the voice data into text data.
[0040] 2. Server
[0041] The text data is sent over the Internet to a server, which then stores the received text data in a database. This database contains not only the text data, but also the date and time the data was collected and the user's identification information.
[0042] 3. Generative Model
[0043] The server is equipped with a generative model using artificial intelligence that periodically analyzes the stored text data, extracts important keywords and sentences from the text data, and summarizes the memories and messages of the deceased.
[0044] 4. User Interface
[0045] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server provides information organized and summarized by the generative model in an interactive format through a user interface. It is also possible to output the information in letter format upon request from family members.
[0046] 5. Notary support and expert guidance
[0047] Support for creating a legally valid will is also provided. When a user makes such a request to the server, the server collects the necessary information and directs the user to a lawyer or notary public, thereby simplifying the process of creating a legally valid will.
[0048] Specific examples
[0049] Example 1: Collecting and summarizing a conversation
[0050] User: Speaks into smart speaker: "Today is my wife's wedding anniversary. I'm so grateful to her."
[0051] Server: Receives voice data and the voice recognition engine converts it into text data.
[0052] Generative model: Analyzes and summarizes text data. "Expressing gratitude on wedding anniversary."
[0053] Example 2: Providing information through a user interface
[0054] Survivors: Access the system from a computer and type in "I'd like to know about your memories of our wedding anniversary."
[0055] Server: Uses the generative model to find relevant summary data and displays it through a user interface. "Today is my wife's wedding anniversary. I'm so grateful to her."
[0056] Example 3: Letter-style output
[0057] Bereaved family: Request a letter-style output. "Please create a letter about our wedding anniversary."
[0058] Server: Generates a letter-style sentence using a generative model. "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0059] In this way, a system is realized that can easily and reliably convey the thoughts of the deceased to the bereaved family.
[0060] The processing flow will be explained below.
[0061] Step 1:
[0062] The user speaks to the smart speaker.
[0063] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0064] Smart speakers collect user voice in real time.
[0065] Step 2:
[0066] The smart speaker converts the voice data into text data.
[0067] The speech recognition engine analyzes the voice data and converts it into text.
[0068] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0069] Step 3:
[0070] The smart speaker sends the converted text data to the server.
[0071] The text data is sent to a server via the Internet.
[0072] Step 4:
[0073] The server receives the text data and stores it in a database.
[0074] The received text data is assigned a timestamp and user identification information and registered in a database.
[0075] Example: { "timestamp": "2023-10-12T14:30:00Z", "user": "Deceased Person A", "message": "Today is my wedding anniversary with my wife. I am truly grateful to her."}
[0076] Step 5:
[0077] A generative model in the server analyzes and summarizes the text data in the database.
[0078] A generative model extracts important keywords and sentences from text data and creates a summary.
[0079] Example: { "topic": "Wedding anniversary", "content": "Expressing gratitude for his wife on their wedding anniversary."}
[0080] Step 6:
[0081] The user (surviving family member) accesses the system from a terminal and enters a request.
[0082] Example: "I'd like to know your memories of your wedding anniversary."
[0083] The request is sent over the Internet to a server.
[0084] Step 7:
[0085] The server retrieves the summarized data from the generative model and provides it through a user interface.
[0086] Matching data is retrieved and displayed on the user interface.
[0087] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0088] Step 8:
[0089] The user requests output in letter format.
[0090] Example: "Please write me a letter about our wedding anniversary."
[0091] The request is sent over the Internet to a server.
[0092] Step 9:
[0093] The server's generative model generates a letter-style document based on the request.
[0094] A letter-style text is created based on past text data and sent to the device.
[0095] For example: "Dear wife, every time our wedding anniversary comes around, my heart fills with gratitude for you. You are the treasure of my life."
[0096] The above is the specific processing flow of the "Conclusion" system.
[0097] Example 1
[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0099] In modern society, there are limited ways to accurately convey the wishes of the deceased to their bereaved families, and there is a particular problem of a lack of an efficient system for the process from collecting to transmitting information using voice. Furthermore, the process of creating a legally valid will is complicated, and there is a lack of support methods to simplify the process.
[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0101] In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing and summarizing the stored text data using a generative model, means for providing information summarized by the generative model through a user interface, means for searching for relevant summary data based on a request entered by a user, means for generating the summary data in letter format, and means for collecting information to support the creation of a legally binding will and guiding experts. This allows the wishes of the deceased to be effectively conveyed to the surviving family and simplifies legal procedures.
[0102] "User" means an individual who uses the System to provide voice data and receive subsequent services.
[0103] "Voice data" refers to a digital recording of a voice signal emitted by a user through a device such as a smart speaker.
[0104] "Text data" refers to digital data that has been converted from voice data into text information.
[0105] "Database" refers to an information management system for storing text data and its associated metadata (date and time, identification information, etc.).
[0106] A "generative model" refers to a software model that uses artificial intelligence to analyze text data, extract important keywords and sentences, and summarize them.
[0107] "User interface" refers to the interface through which users and bereaved families access the system and input or retrieve data.
[0108] A "request" is an instruction from a user or family member to the system requesting the provision of specific data or services.
[0109] "Summary data" is shortened text data that contains key points or sentences extracted by a generative model.
[0110] "Letter format" refers to a document in which summary data has been reconstructed into a letter format that is easy for humans to read.
[0111] "Experts" refers to professionals such as lawyers and notaries who have knowledge and qualifications regarding legal procedures and the preparation of wills.
[0112] A "legally valid will" is a will that has legal effect and is officially recognized under the law.
[0113] This invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and a specific embodiment thereof is described below. This system involves a process in which a user collects voice data, converts it into text data, summarizes and organizes it using a generative AI model, and provides it to the bereaved.
[0114] System Configuration
[0115] 1. Collection of audio data
[0116] Users collect voice data by speaking into a voice input device such as a smart speaker, which then converts the voice data into text data in real time using a voice recognition system such as Amazon Alexa or Google Assistant.
[0117] 2. Sending and saving text data
[0118] The text data generated by the voice recognition system is sent over the Internet to a server. The server stores the received text data and its associated metadata (date and time, user identification information, etc.) in a database. This database serves to store and manage the memories and messages of the deceased.
[0119] 3. Analysis and Summarization of Text Data
[0120] The server is equipped with a generative AI model (e.g., OpenAI's GPT-3) that periodically analyzes the text data in the database. The generative model extracts important keywords and sentences from the text data and summarizes and organizes the memories and messages of the deceased.
[0121] 4. Providing information through the user interface
[0122] Family members access the system through devices such as computers or smartphones. Using the user interface, they input specific requests, such as "I want to know about memories of our wedding anniversary." The server uses the generative AI model to search for the relevant summary data and displays it through the user interface.
[0123] An example of a specific prompt is "Show me a message about your wedding anniversary memories."
[0124] 5. Generating Letter Format
[0125] If the family requests a letter-style output, the server sends a specific prompt to the generative AI model, which generates a letter with a message like, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0126] 6. Legal process support
[0127] When a user wishes to create a legally valid will, the server collects the necessary information and directs them to professionals such as lawyers and notaries. Also, when the bereaved family requests the creation of a will through the system, the server organizes and provides the necessary information, helping to simplify the legal procedures.
[0128] In this way, this system can easily and reliably convey the wishes of the deceased to their family members.It is also extremely convenient because it has a function to assist users in legal procedures.
[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0130] Step 1:
[0131] The user speaks to the smart speaker. For example, "Today is my wedding anniversary with my wife. I'm really grateful to her." This voice is input into the smart speaker.
[0132] Step 2:
[0133] Smart speakers use a built-in voice recognition engine (e.g., Google Assistant) to convert the received voice data into text data in real time. This conversion process outputs the voice data as text data.
[0134] Step 3:
[0135] The text data is sent to a server via the Internet. The data sent includes not only the text data but also the date and time the voice was collected and the user's identification information. The server receives this data.
[0136] Step 4:
[0137] The server saves the received text data and its associated metadata in a database. The saving process stores the text data and metadata in the database.
[0138] Step 5:
[0139] The server passes the text data stored in the database to a generative AI model (for example, OpenAI's GPT-3) for analysis. The generative model reads the text data, extracts important keywords and sentences, and generates summary data. This summary data is the output from the generative model.
[0140] Step 6:
[0141] The bereaved family accesses the system from a PC or smartphone and inputs a specific request into the user interface (e.g., "I would like to know about memories of our wedding anniversary.") This request is sent to the server.
[0142] Step 7:
[0143] The server receives the request from the family and uses the generative AI model to search for the relevant summary data. The search process is carried out and the relevant summary data is output.
[0144] Step 8:
[0145] The server displays the retrieved summary data through a user interface, allowing the family members to view the information.
[0146] Step 9:
[0147] If the family requests a letter-style output, the server sends a specific prompt to the generative AI model, which generates a letter-style text, such as, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0148] Step 10:
[0149] If a user wishes to create a legally valid will, the server collects the necessary information and directs them to professionals such as lawyers and notaries, helping them to navigate the legal process more easily.
[0150] Through the above processing steps, the system can accurately convey the wishes of the deceased to the bereaved family and can also provide support in legal procedures.
[0151] (Application example 1)
[0152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0153] In conventional estate sorting, it has been difficult to accurately convey the deceased's memories and messages to the bereaved. Also, the task of organizing and summarizing the deceased's thoughts is time-consuming and laborious, placing a heavy burden on the bereaved. Furthermore, the means to easily retrieve and access the deceased's thoughts were limited, making it difficult for the bereaved to reminisce. The present invention aims to solve these problems.
[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0155] In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing and summarizing the stored text data using a generative model, means for providing information summarized by the generative model through a user interface, and means for delivering the summarized information to a smartphone or head-mounted display. This allows memories and messages of the deceased to be organized and easily conveyed to the bereaved. Furthermore, the bereaved can access memories of the deceased at any time using a smartphone or head-mounted display, thereby reducing the burden on the bereaved.
[0156] "User" refers to an individual who provides voice data using this system.
[0157] "Voice Data" refers to recordings of voice signals collected when a user speaks.
[0158] "Text data" refers to data that has been converted from voice data into text information using voice recognition technology.
[0159] "Database" refers to a digital storage system for storing collected text data.
[0160] A "generative model" refers to an algorithm that uses artificial intelligence techniques to analyze and summarize text data.
[0161] "User interface" refers to the interface on a computer system that allows users to view information and perform operations.
[0162] A "smartphone" refers to a multi-function mobile phone that can make calls and connect to the Internet.
[0163] "Head-mounted display" refers to a display device worn on the head.
[0164] "Distribution" refers to the act of transmitting or providing information via the Internet or other means of communication.
[0165] The present invention is a system for accurately conveying the memories and messages of a deceased person to their bereaved families. A specific embodiment of the system is described below. This system covers the collection of voice data, storing it in a database, summarizing the text data using a generative model, and delivering the information via smartphones and head-mounted displays.
[0166] First, the user speaks to the smart speaker to collect voice data. The smart speaker is equipped with a speech recognition engine that converts the voice data into text data in real time. This speech recognition engine can be Google's speech recognition API or IBM Watson.
[0167] The converted text data is sent to a server via the Internet and stored in a database on the server side, which may use a storage system such as SQL Server or MongoDB.
[0168] The server implements a generative model to periodically analyze and summarize text data. This generative model, for example, the "BART" model based on the Hugging Face "transformers" library, is used. The generative model analyzes the stored text data, extracts important keywords and sentences, and creates a summary. The summarized information is then stored on the server.
[0169] Family members can access the system through a user interface using a smartphone or head-mounted display. The user interface provides information summarized by the generative model. Android and iOS app development frameworks are used as the user interface technology. This system allows family members to access the memories and messages of the deceased at any time.
[0170] The detailed operation of the system will be explained below by showing a specific example.
[0171] Examples:
[0172] If a user wants to collect memories of the deceased using the following prompt, they can speak to the smart speaker, saying, "Tell me your memories of Mother's Day." The smart speaker converts the speech into text and sends it to the server. The server summarizes the text and generates a summary such as, "I always sent flowers on Mother's Day. My mother was always very happy." The bereaved can view this summarized message at any time via their smartphone or head-mounted display.
[0173] This makes it possible to efficiently organize memories of the deceased and properly convey them to the bereaved family.
[0174] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0175] Step 1:
[0176] The user speaks to the smart speaker to provide voice data. The smart speaker receives the user's voice as input, and the built-in voice recognition engine converts the voice data into text data in real time. As a result, the input (voice data) acquired by the smart speaker is output as text data.
[0177] Step 2:
[0178] The smart speaker sends the converted text data over the Internet to a server. The server receives this text data as input and stores it in a database in an appropriate format. The database uses a cloud platform such as Azure SQL Database or Amazon RDS. The text data is stored in the database along with the date, time, and user identification information.
[0179] Step 3:
[0180] The server periodically accesses the database and analyzes the stored text data. At this time, it uses a generative AI model (for example, the BART model implemented using Hugging Face's transformers library) to summarize the text data. The server takes the text data as input and uses the generative model to extract important keywords and sentences. As a result of the analysis, the summarized text is output and stored again in the database.
[0181] Step 4:
[0182] The bereaved family members access the user interface via a smartphone or head-mounted display. The user interface requests information from the server using a specific prompt (e.g., "Tell me your memories of Mother's Day"). The server receives this prompt, searches for relevant summary data from a database, and responds to the user interface. The user interface receives the user's request as input and displays the summarized information as output.
[0183] Step 5:
[0184] The user checks the summarized information through a smartphone or head-mounted display. Specifically, by checking the summary message displayed on the user interface (e.g., "I always sent flowers to my mother on Mother's Day. She was always very happy every time"), the user can reminisce about the deceased. This allows the user to access the memories and messages of the deceased at any time.
[0185] This series of processes creates a system that efficiently organizes the memories and messages of the deceased and conveys them appropriately to the bereaved.
[0186] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0187] This invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and further describes a form that combines an emotion engine that recognizes the user's emotions. This system involves a process in which the user uses a smart speaker to collect voice data, converts that data into text data, organizes and summarizes it, and provides it to the bereaved. In addition, the emotion engine recognizes the user's emotions and provides a more in-depth summary or letter-style output based on that information.
[0188] System Configuration
[0189] 1. Smart Speaker
[0190] This system collects voice data when the user (deceased person) speaks to a smart speaker. The smart speaker is equipped with a voice recognition engine that instantly converts the voice data into text data. An emotion engine is also built into the smart speaker, which recognizes emotions from the user's voice in real time.
[0191] 2. Server
[0192] The text data and emotion information are sent to a server via the Internet. The server receives the data and stores it in a database. The database stores the text data, emotion information, collection date and time, and user identification information.
[0193] 3. Generative Model
[0194] The server is equipped with a generative model using artificial intelligence that periodically analyzes the stored text data and emotional information. The generative model extracts important keywords and sentences from the text data and summarizes the memories and messages of the deceased along with emotional information.
[0195] 4. User Interface
[0196] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server then provides information organized and summarized by the generative model in an interactive format through a user interface. Adding emotional information allows for a more in-depth dialogue. It is also possible to output the information in the form of a letter upon request from family members.
[0197] 5. Notary support and expert guidance
[0198] Support for creating a legally valid will is also provided. When a user makes such a request to the server, the server collects the necessary information and directs the user to a lawyer or notary public, thereby simplifying the process of creating a legally valid will.
[0199] Specific examples
[0200] Example 1: Collecting and summarizing a conversation
[0201] User: Speaks into smart speaker: "Today is my wife's wedding anniversary. I'm so grateful to her."
[0202] Smart speaker: Recognizes voice and the emotion engine identifies emotions such as "gratitude."
[0203] Server: Converts the voice data into text data and emotional information and stores them in a database.
[0204] Generative model: Analyzes and summarizes text data and emotional information. "Expressing gratitude on wedding anniversary."
[0205] Example 2: Providing information through a user interface
[0206] Survivors: Access the system from a computer and type in "I'd like to know about your memories of our wedding anniversary."
[0207] Server: Uses a generative model to search for relevant summary data, adds emotional information, and displays it on the user interface. "Today is my wife's wedding anniversary. I'm really grateful to her."
[0208] Example 3: Letter-style output
[0209] Bereaved family: Request a letter-style output. "Please create a letter about our wedding anniversary."
[0210] Server: Generates letter-style text using a generative model and adds emotional information. "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0211] In this way, by combining the emotion engine, a system is realized that can easily and deeply convey the feelings of the deceased to the bereaved family.
[0212] The processing flow will be explained below.
[0213] Step 1:
[0214] The user speaks to the smart speaker.
[0215] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0216] Smart speakers collect user voice in real time.
[0217] Step 2:
[0218] The smart speaker converts the voice data into text data.
[0219] The speech recognition engine analyzes the voice data and converts it into text.
[0220] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0221] Step 3:
[0222] The smart speaker uses an emotion engine to recognize emotions from voice data.
[0223] The emotion engine analyzes the voice data and identifies emotions such as "gratitude" and "joy."
[0224] Example: Recognizing the emotion of "gratitude."
[0225] Step 4:
[0226] The smart speaker sends the converted text data and emotional information to the server.
[0227] The text data and emotion information are sent to a server via the Internet.
[0228] Step 5:
[0229] The server receives the text data and emotion information and stores them in a database.
[0230] The received text data and emotion information are assigned a timestamp and user identification information and registered in a database.
[0231] Example: { "timestamp": "2023-10-12T14:30:00Z", "user": "Deceased Person A", "message": "Today is my wedding anniversary with my wife. I am truly grateful to her.", "emotion": "Thank you"}
[0232] Step 6:
[0233] A generative model in the server analyzes and summarizes the text data and sentiment information in the database.
[0234] The generative model extracts important keywords and sentences from text data and emotional information, and creates summaries that reflect the emotional information.
[0235] Example: { "topic": "Wedding anniversary", "content": "Expressing gratitude for his wife on their wedding anniversary."}
[0236] Step 7:
[0237] The user (surviving family member) accesses the system from a terminal and enters a request.
[0238] Example: "I'd like to know your memories of your wedding anniversary."
[0239] The request is sent over the Internet to a server.
[0240] Step 8:
[0241] The server retrieves summary data based on the generative model, adds emotional information, and provides it through a user interface.
[0242] Matching data is searched for and displayed on the user interface based on emotion information.
[0243] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0244] Step 9:
[0245] The user requests output in letter format.
[0246] Example: "Please write me a letter about our wedding anniversary."
[0247] The request is sent over the Internet to a server.
[0248] Step 10:
[0249] The server's generative model generates a letter-style document based on the request.
[0250] A letter-style text is created based on past text data and emotional information and sent to the device.
[0251] For example: "Dear wife, every time our wedding anniversary comes around, my heart fills with gratitude for you. You are the treasure of my life."
[0252] The above is the specific processing flow that combines the emotion engine in the "Conclusion" system.
[0253] Example 2
[0254] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0255] There is a problem in accurately conveying the deceased's thoughts to the bereaved. Conventional methods do not accurately convey the feelings and intentions of the deceased, and do not provide deep meaning to the bereaved. In addition, the process of recording and organizing the deceased's thoughts and creating a legally valid will, if necessary, is complicated and time-consuming.
[0256] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0257] In this invention, the server includes means for collecting voice data from a user, means for converting the voice data into text data, means for recognizing the user's emotions from the voice data, means for storing the text data and emotional information in a database, means for analyzing and summarizing the stored text data and emotional information using a generative model, and means for providing information summarized by the generative model through a user interface. This allows the deceased's thoughts, along with their emotional information, to be organized and summarized and conveyed to the bereaved. It can also output the message in letter format and support the creation of a legally valid will, allowing for a deeper understanding of the bereaved and legal response.
[0258] "Voice data" refers to voice signal information collected when a user speaks to a smart speaker or other voice input device.
[0259] "Text data" is character information converted from voice data by a voice recognition engine.
[0260] "Emotion information" is data that is analyzed from voice data by the emotion engine and indicates the user's emotional state.
[0261] A "database" is a collection of information for systematically storing and managing text data and emotional information.
[0262] A "generative model" is a program that uses artificial intelligence to analyze text data and emotional information, extract important keywords and sentences, and generate summaries.
[0263] A "user interface" refers to the screen or operating means that allows bereaved family members to access the system through devices such as computers or smartphones and obtain the necessary information.
[0264] "Letter-format output" is a function that provides information summarized by the generative model in the form of a letter.
[0265] A "legally valid will" is a written document that expresses the final wishes of a deceased person and is legally valid.
[0266] "Expert Guidance" is a support function that encourages users who wish to write a will to contact or consult with legal experts (such as lawyers or notaries).
[0267] The system of the present invention aims to accurately convey the thoughts of the deceased to the bereaved family, and is realized by combining an emotion engine that recognizes the user's emotions. This system includes the following components.
[0268] 1. Smart Speaker
[0269] The first component of this system is the smart speaker. Users speak into the smart speaker to collect their messages and thoughts as voice data. The smart speaker has a built-in voice recognition engine (e.g., Google Speech-to-Text or Amazon Alexa) that instantly converts the voice data into text data. In addition, the built-in emotion engine recognizes the user's emotions from the collected voice in real time.
[0270] Examples:
[0271] A user says to a smart speaker, "Today is my wife's wedding anniversary. I'm so grateful to her."
[0272] 2. Server
[0273] The text data and emotion information are sent to a server via the Internet. The server receives the data and stores it in a database. The database stores the text data, emotion information, collection date and time, and user identification information.
[0274] Examples:
[0275] The server stores the text data (e.g., "Today is my wedding anniversary with my wife. I am truly grateful to her") and emotional information (gratitude) sent from the smart speaker in a database.
[0276] 3. Generative AI Models
[0277] The server is equipped with a generative AI model that periodically analyzes the stored text data and emotional information. Using the generative model (e.g., OpenAI GPT-3), it extracts important keywords and sentences from the text data and summarizes the deceased's memories and messages along with emotional information.
[0278] Examples:
[0279] The generative AI model analyzes the stored text data and emotional information to generate summaries such as "expressing gratitude on the wedding anniversary."
[0280] 4. User Interface
[0281] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server provides information organized and summarized by the generative model through a user interface. Adding emotional information allows for deeper dialogue.
[0282] Examples:
[0283] When a family member accesses the system from a computer and types, "I'd like to know about your memories of our wedding anniversary," the server searches for relevant summary data and displays information such as, "Today is my wedding anniversary with my wife. I'm truly grateful to her."
[0284] 5. Letter-style output
[0285] Upon request of the bereaved family, the information can be output in letter format.
[0286] Examples:
[0287] If a family member requests a letter about their wedding anniversary, the generative AI model will generate a letter that reads, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0288] 6. Support for creating legally binding wills
[0289] When a user requests the server for assistance in creating a legally binding will, the server collects the necessary information and directs the user to a lawyer or notary public.
[0290] Examples:
[0291] When a user requests the creation of a legally binding will, the server collects the appropriate information and guides them through the process and how to contact the appropriate professional.
[0292] This system makes it possible to organize and summarize the feelings of the deceased along with emotional information, and to convey them in depth to the bereaved. It can also output them in letter format and assist in the creation of legally valid wills, making it easier for the bereaved to understand and respond to the situation.
[0293] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0294] Step 1:
[0295] The user speaks to the smart speaker.
[0296] Input: User's voice message (e.g., "Today is my wedding anniversary with my wife. I'm so grateful to her.")
[0297] How it works: Smart speakers use built-in microphones to collect voice data.
[0298] Output: Collected audio data
[0299] Step 2:
[0300] Smart speakers convert voice data into text data and recognize emotions.
[0301] Input: Audio data
[0302] How it works: A voice recognition engine (e.g., Google Speech-to-Text) built into a smart speaker converts voice data into text data. At the same time, an emotion engine analyzes the text data and generates emotion information (e.g., "thank you").
[0303] Output: Text data (e.g., "Today is my wedding anniversary with my wife. I am so grateful to her."), emotional information (e.g., "gratitude")
[0304] Step 3:
[0305] The smart speaker sends text data and emotional information to the server.
[0306] Input: Text data, emotion information
[0307] How it works: Smart speakers send data over the internet to a server, which often encrypts the data before sending it.
[0308] Output: Text data and emotion information sent to the server
[0309] Step 4:
[0310] The server stores the text data and emotion information in a database.
[0311] Input: Text data, emotion information
[0312] Operation: The server analyzes the received data and stores the text data, emotion information, collection date and time, user identification information, etc. in a relational database.
[0313] Output: Information stored in the database
[0314] Step 5:
[0315] A generative AI model implemented on the server analyzes text data and emotional information to generate a summary.
[0316] Input: Text data and emotion information stored in a database
[0317] How it works: A generative AI model (e.g., OpenAI GPT-3) periodically scans the database, analyzes the text data and sentiment information, extracts important keywords and sentences, and generates summaries taking sentiment information into account.
[0318] Output: Summarized information (e.g., "Expressing gratitude for our wedding anniversary.")
[0319] Step 6:
[0320] Family members access the system through a terminal and request information.
[0321] Input: Request from the bereaved family (e.g., "I'd like to know about your memories of your wedding anniversary.")
[0322] How it works: Family members access the system using a computer or smartphone and enter information.
[0323] Output: Input information of bereaved family members
[0324] Step 7:
[0325] The server provides the summarized information through a user interface.
[0326] Input: Family request information
[0327] How it works: The server uses the generative model to search for relevant summary data, adds emotional information, and displays it on the user interface.
[0328] Output: Summary information displayed in a user interface (e.g., "Today is my wife's wedding anniversary. I'm so grateful to her.")
[0329] Step 8:
[0330] The family requests a letter-style output.
[0331] Input: A request for output in the form of a letter (e.g., "Please write me a letter about my wedding anniversary.")
[0332] How it works: The family requests a letter-style output from the system.
[0333] Output: Request information
[0334] Step 9:
[0335] The server uses a generative AI model to generate a letter-style document.
[0336] Input: Request information for output in letter format, database information
[0337] How it works: The generative AI model generates letter-style text based on summarized text data and emotional information.
[0338] Output: A letter-style sentence (e.g., "Dear Wife, Every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life.")
[0339] Step 10:
[0340] The server collects information to support the creation of a legally binding will and provides guidance to experts.
[0341] Input: User request information
[0342] How it works: Based on the user's request, the server gathers the necessary information from the database and directs them to the appropriate lawyer or notary public.
[0343] Output: Guidance information (e.g., contact instructions for a lawyer or notary public, list of required documents)
[0344] (Application example 2)
[0345] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0346] When dealing with customers in brick-and-mortar stores, it is difficult for employees to properly understand customer emotions and provide optimal service based on that. A particular challenge is the lack of a system that can recognize and reflect customer emotions and requests in real time. Therefore, there is a need for a system that can improve customer satisfaction and the quality of employee service.
[0347] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for using an emotion engine to analyze and summarize the saved text data and emotion information, and means for providing information summarized by the generative model and emotion engine in real time via an audiovisual device used by a customer service representative. This makes it possible to properly recognize customer emotions when serving customers in a physical store and provide services based on those emotions quickly and accurately.
[0348] Definitions of important words
[0349] "Voice data" refers to digitized data of the voice uttered by the user.
[0350] "Text data" refers to data obtained by converting voice data into character information.
[0351] A "database" is a digital storage device for systematically storing text data and emotional information.
[0352] A "generative model" is an algorithm or software that uses artificial intelligence to extract and summarize important keywords and sentences from text data.
[0353] "Emotion information" is information that represents the emotional state of a user, as recognized from voice data or text data.
[0354] An "emotion engine" is software that analyzes a user's voice data and text data to extract emotional information.
[0355] "Audiovisual devices" is a general term for devices that provide visual and auditory information, such as smart glasses and head-mounted displays.
[0356] "Real-time" refers to processing or response occurring with almost no delay after an event occurs.
[0357] "Customer service personnel" refers to employees and staff who interact directly with customers and provide services in physical stores.
[0358] MODE FOR CARRYING OUT THE INVENTION
[0359] This invention relates to a system for supporting customer service in brick-and-mortar stores, and aims to increase customer satisfaction by utilizing emotion recognition technology in particular. This system uses audiovisual devices such as smart glasses to enable store staff to grasp customer emotions in real time and provide optimal service.
[0360] System configuration
[0361] 1. Collection of audio data
[0362] Customer service representatives will wear smart glasses or head-mounted displays to communicate with customers. The smart glasses have built-in microphones that collect customer voice data in real time.
[0363] 2. Converting audio data to text data
[0364] The collected voice data is instantly converted into text data using the device's voice recognition engine (e.g., Google Speech-to-Text API), which is temporarily stored on the device and then sent to a server.
[0365] 3. Extraction of Emotional Information
[0366] The server analyzes the received text data and extracts customer emotional information using an emotion engine (e.g., EmotionRecognition software). This emotional information is then stored in a database along with the text data.
[0367] 4. Data Analysis and Summarization
[0368] A generative AI model implemented on the server analyzes the stored text data and sentiment information, extracts important keywords and sentences, and summarizes them.
[0369] 5. Real-time information provision
[0370] The server then provides the customer service representative with real-time information summarized by the generative model via an audiovisual device, and the employee's smart glasses display the customer's emotional state and appropriate response.
[0371] Specific examples
[0372] Program processing
[0373] 1. A customer in the store says, "I've been feeling stressed lately and I want a product to help me relax."
[0374] 2. Voice is collected by the microphone in the smart glasses and instantly converted into text data.
[0375] 3. On the server side, the text data is analyzed by an emotion engine, and the emotion "stress" is extracted.
[0376] 4. The generative model analyzes the text data and generates a summary: "We suggest a relaxing aroma set."
[0377] 5. This generated summary information is displayed on an audiovisual device, and the employee can recommend a product to the customer by saying, "We suggest an aroma set that will have a relaxing effect."
[0378] Prompt Sentence Examples
[0379] "A customer might say, 'I've been feeling stressed lately and I want a product to help me relax.' Analyze their emotions with an emotion recognition engine and generate a dialogue that suggests relevant products."
[0380] This system makes it possible to understand customer sentiment in real time and quickly make suggestions based on that sentiment, thereby significantly improving customer satisfaction.
[0381] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0382] Program processing steps
[0383] Step 1:
[0384] Audio data collection
[0385] A user (customer service representative) wears smart glasses and interacts with customers in a physical store. When a customer speaks, the microphone built into the smart glasses collects the voice.
[0386] Input: Customer voice.
[0387] Output: Digitized audio data.
[0388] What it does: The microphone in the smart glasses activates and records what the customer says in real time.
[0389] Step 2:
[0390] Converting audio data to text data
[0391] The device (smart glasses) converts the collected voice data into text data using an internal voice recognition engine, which uses the Google Speech-to-Text API.
[0392] Input: Digitized audio data.
[0393] Output: Text data.
[0394] How it works: The voice recognition engine processes the voice data and converts it into corresponding text data, which is then temporarily stored in the smart glasses.
[0395] Step 3:
[0396] Sending text data
[0397] The terminal (smart glasses) sends the converted text data to a server via the Internet.
[0398] Input: Text data.
[0399] Output: The text data sent to the server.
[0400] How it works: The smart glasses use Wi-Fi or mobile networks to upload text data to a server.
[0401] Step 4:
[0402] Extracting Emotional Information
[0403] The server analyzes the received text data using an emotion engine (e.g., EmotionRecognition software) to extract customer emotional information.
[0404] Input: The text data sent to the server.
[0405] Output: Emotion information (e.g., "joy", "sad", "stress", etc.).
[0406] What it does: The emotion engine on the server scans the text data and identifies relevant emotions.
[0407] Step 5:
[0408] Data storage
[0409] The server stores the converted text data and the extracted emotion information in a database.
[0410] Input: Text data, emotion information.
[0411] Output: Information stored in a database.
[0412] Specific operation: The server's database management system stores text data and emotional information in an orderly manner.
[0413] Step 6:
[0414] Data analysis and summary
[0415] A generative AI model (e.g., OpenAI GPT-3) implemented on the server analyzes the stored text data and emotional information, extracts important keywords and sentences, and generates a summary.
[0416] Input: Text data and emotion information stored in the database.
[0417] Output: Summarized information.
[0418] How it works: The generative AI model analyzes the information in the database, extracts the most important parts, and generates a summary.
[0419] Step 7:
[0420] Real-time provision
[0421] The server provides information summarized by the generative model to the audiovisual device (smart glasses) in real time.
[0422] Input: Summarized information.
[0423] Output: Summary information displayed or spoken on the smart glasses.
[0424] How it works: The server uses Wi-Fi or mobile networks to send the summarized text to the smart glasses, which then displays it on the customer service representative's display or notifies them via voice.
[0425] Step 8:
[0426] Service delivery based on feedback
[0427] Customer service representatives will suggest the most suitable services and products to customers based on the summary information and suggestions displayed on the smart glasses.
[0428] Input: Summary information displayed on smart glasses.
[0429] Output: Services and offers to customers.
[0430] Specific action: The customer service representative checks the summary information and makes appropriate suggestions or asks questions to the customer.
[0431] By implementing the above steps, it is possible to provide services that are sensitive to the customer's emotions when dealing with customers in physical stores.
[0432] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0433] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0434] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0435] [Second embodiment]
[0436] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0437] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0438] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0439] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0440] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0441] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0442] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0443] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0444] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0445] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0446] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0447] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0448] The present invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and a specific embodiment thereof is described below. This system includes a process in which a user uses a smart speaker to collect voice data, converts the data into text data, organizes and summarizes it, and provides it to the bereaved.
[0449] System Configuration
[0450] 1. Smart Speaker
[0451] This system collects voice data by having the user (deceased person) speak into a smart speaker, which is equipped with a voice recognition engine that instantly converts the voice data into text data.
[0452] 2. Server
[0453] The text data is sent over the Internet to a server, which then stores the received text data in a database. This database contains not only the text data, but also the date and time the data was collected and the user's identification information.
[0454] 3. Generative Model
[0455] The server is equipped with a generative model using artificial intelligence that periodically analyzes the stored text data, extracts important keywords and sentences from the text data, and summarizes the memories and messages of the deceased.
[0456] 4. User Interface
[0457] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server provides information organized and summarized by the generative model in an interactive format through a user interface. It is also possible to output the information in letter format upon request from family members.
[0458] 5. Notary support and expert guidance
[0459] Support for creating a legally valid will is also provided. When a user makes such a request to the server, the server collects the necessary information and directs the user to a lawyer or notary public, thereby simplifying the process of creating a legally valid will.
[0460] Specific examples
[0461] Example 1: Collecting and summarizing a conversation
[0462] User: Speaks into smart speaker: "Today is my wife's wedding anniversary. I'm so grateful to her."
[0463] Server: Receives voice data and the voice recognition engine converts it into text data.
[0464] Generative model: Analyzes and summarizes text data. "Expressing gratitude on wedding anniversary."
[0465] Example 2: Providing information through a user interface
[0466] Survivors: Access the system from a computer and type in "I'd like to know about your memories of our wedding anniversary."
[0467] Server: Uses the generative model to find relevant summary data and displays it through a user interface. "Today is my wife's wedding anniversary. I'm so grateful to her."
[0468] Example 3: Letter-style output
[0469] Bereaved family: Request a letter-style output. "Please create a letter about our wedding anniversary."
[0470] Server: Generates a letter-style sentence using a generative model. "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0471] In this way, a system is realized that can easily and reliably convey the thoughts of the deceased to the bereaved family.
[0472] The processing flow will be explained below.
[0473] Step 1:
[0474] The user speaks to the smart speaker.
[0475] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0476] Smart speakers collect user voice in real time.
[0477] Step 2:
[0478] The smart speaker converts the voice data into text data.
[0479] The speech recognition engine analyzes the voice data and converts it into text.
[0480] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0481] Step 3:
[0482] The smart speaker sends the converted text data to the server.
[0483] The text data is sent to a server via the Internet.
[0484] Step 4:
[0485] The server receives the text data and stores it in a database.
[0486] The received text data is assigned a timestamp and user identification information and registered in a database.
[0487] Example: { "timestamp": "2023-10-12T14:30:00Z", "user": "Deceased Person A", "message": "Today is my wedding anniversary with my wife. I am truly grateful to her."}
[0488] Step 5:
[0489] A generative model in the server analyzes and summarizes the text data in the database.
[0490] A generative model extracts important keywords and sentences from text data and creates a summary.
[0491] Example: { "topic": "Wedding anniversary", "content": "Expressing gratitude for his wife on their wedding anniversary."}
[0492] Step 6:
[0493] The user (surviving family member) accesses the system from a terminal and enters a request.
[0494] Example: "I'd like to know your memories of your wedding anniversary."
[0495] The request is sent over the Internet to a server.
[0496] Step 7:
[0497] The server retrieves the summarized data from the generative model and provides it through a user interface.
[0498] Matching data is retrieved and displayed on the user interface.
[0499] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0500] Step 8:
[0501] The user requests output in letter format.
[0502] Example: "Please write me a letter about our wedding anniversary."
[0503] The request is sent over the Internet to a server.
[0504] Step 9:
[0505] The server's generative model generates a letter-style document based on the request.
[0506] A letter-style text is created based on past text data and sent to the device.
[0507] For example: "Dear wife, every time our wedding anniversary comes around, my heart fills with gratitude for you. You are the treasure of my life."
[0508] The above is the specific processing flow of the "Conclusion" system.
[0509] Example 1
[0510] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0511] In modern society, there are limited ways to accurately convey the wishes of the deceased to their bereaved families, and there is a particular problem of a lack of an efficient system for the process from collecting to transmitting information using voice. Furthermore, the process of creating a legally valid will is complicated, and there is a lack of support methods to simplify the process.
[0512] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0513] In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing and summarizing the stored text data using a generative model, means for providing information summarized by the generative model through a user interface, means for searching for relevant summary data based on a request entered by a user, means for generating the summary data in letter format, and means for collecting information to support the creation of a legally binding will and guiding experts. This allows the wishes of the deceased to be effectively conveyed to the surviving family and simplifies legal procedures.
[0514] "User" means an individual who uses the System to provide voice data and receive subsequent services.
[0515] "Voice data" refers to a digital recording of a voice signal emitted by a user through a device such as a smart speaker.
[0516] "Text data" refers to digital data that has been converted from voice data into text information.
[0517] "Database" refers to an information management system for storing text data and its associated metadata (date and time, identification information, etc.).
[0518] A "generative model" refers to a software model that uses artificial intelligence to analyze text data, extract important keywords and sentences, and summarize them.
[0519] "User interface" refers to the interface through which users and bereaved families access the system and input or retrieve data.
[0520] A "request" is an instruction from a user or family member to the system requesting the provision of specific data or services.
[0521] "Summary data" is shortened text data that contains key points or sentences extracted by a generative model.
[0522] "Letter format" refers to a document in which summary data has been reconstructed into a letter format that is easy for humans to read.
[0523] "Experts" refers to professionals such as lawyers and notaries who have knowledge and qualifications regarding legal procedures and the preparation of wills.
[0524] A "legally valid will" is a will that has legal effect and is officially recognized under the law.
[0525] This invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and a specific embodiment thereof is described below. This system involves a process in which a user collects voice data, converts it into text data, summarizes and organizes it using a generative AI model, and provides it to the bereaved.
[0526] System Configuration
[0527] 1. Collection of audio data
[0528] Users collect voice data by speaking into a voice input device such as a smart speaker, which then converts the voice data into text data in real time using a voice recognition system such as Amazon Alexa or Google Assistant.
[0529] 2. Sending and saving text data
[0530] The text data generated by the voice recognition system is sent over the Internet to a server. The server stores the received text data and its associated metadata (date and time, user identification information, etc.) in a database. This database serves to store and manage the memories and messages of the deceased.
[0531] 3. Analysis and Summarization of Text Data
[0532] The server is equipped with a generative AI model (e.g., OpenAI's GPT-3) that periodically analyzes the text data in the database. The generative model extracts important keywords and sentences from the text data and summarizes and organizes the memories and messages of the deceased.
[0533] 4. Providing information through the user interface
[0534] Family members access the system through devices such as computers or smartphones. Using the user interface, they input specific requests, such as "I want to know about memories of our wedding anniversary." The server uses the generative AI model to search for the relevant summary data and displays it through the user interface.
[0535] An example of a specific prompt is "Show me a message about your wedding anniversary memories."
[0536] 5. Generating Letter Format
[0537] If the family requests a letter-style output, the server sends a specific prompt to the generative AI model, which generates a letter with a message like, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0538] 6. Legal process support
[0539] When a user wishes to create a legally valid will, the server collects the necessary information and directs them to professionals such as lawyers and notaries. Also, when the bereaved family requests the creation of a will through the system, the server organizes and provides the necessary information, helping to simplify the legal procedures.
[0540] In this way, this system can easily and reliably convey the wishes of the deceased to their family members.It is also extremely convenient because it has a function to assist users in legal procedures.
[0541] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0542] Step 1:
[0543] The user speaks to the smart speaker. For example, "Today is my wedding anniversary with my wife. I'm really grateful to her." This voice is input into the smart speaker.
[0544] Step 2:
[0545] Smart speakers use a built-in voice recognition engine (such as Google Assistant) to convert the received voice data into text data in real time, and this conversion process outputs the voice data as text data.
[0546] Step 3:
[0547] The text data is sent to a server via the Internet. The data sent includes not only the text data but also the date and time the voice was collected and the user's identification information. The server receives this data.
[0548] Step 4:
[0549] The server saves the received text data and its associated metadata in a database. The saving process stores the text data and metadata in the database.
[0550] Step 5:
[0551] The server passes the text data stored in the database to a generative AI model (for example, OpenAI's GPT-3) for analysis. The generative model reads the text data, extracts important keywords and sentences, and generates summary data. This summary data is the output from the generative model.
[0552] Step 6:
[0553] The bereaved family accesses the system from a PC or smartphone and inputs a specific request into the user interface (e.g., "I would like to know about memories of our wedding anniversary.") This request is sent to the server.
[0554] Step 7:
[0555] The server receives the request from the family and uses the generative AI model to search for the relevant summary data. The search process is carried out and the relevant summary data is output.
[0556] Step 8:
[0557] The server displays the retrieved summary data through a user interface, allowing the family members to view the information.
[0558] Step 9:
[0559] If the family requests a letter-style output, the server sends a specific prompt to the generative AI model, which generates a letter-style text, such as, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0560] Step 10:
[0561] If a user wishes to create a legally valid will, the server collects the necessary information and directs them to professionals such as lawyers and notaries, helping them to navigate the legal process more easily.
[0562] Through the above processing steps, the system can accurately convey the wishes of the deceased to the bereaved family and can also provide support in legal procedures.
[0563] (Application example 1)
[0564] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0565] In conventional estate sorting, it has been difficult to accurately convey the deceased's memories and messages to the bereaved. Also, the task of organizing and summarizing the deceased's thoughts is time-consuming and laborious, placing a heavy burden on the bereaved. Furthermore, the means to easily retrieve and access the deceased's thoughts were limited, making it difficult for the bereaved to reminisce. The present invention aims to solve these problems.
[0566] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0567] In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing and summarizing the stored text data using a generative model, means for providing information summarized by the generative model through a user interface, and means for delivering the summarized information to a smartphone or head-mounted display. This allows memories and messages of the deceased to be organized and easily conveyed to the bereaved. Furthermore, the bereaved can access memories of the deceased at any time using a smartphone or head-mounted display, thereby reducing the burden on the bereaved.
[0568] "User" refers to an individual who provides voice data using this system.
[0569] "Voice Data" refers to recordings of voice signals collected when a user speaks.
[0570] "Text data" refers to data that has been converted from voice data into text information using voice recognition technology.
[0571] "Database" refers to a digital storage system for storing collected text data.
[0572] A "generative model" refers to an algorithm that uses artificial intelligence techniques to analyze and summarize text data.
[0573] "User interface" refers to the interface on a computer system that allows users to view information and perform operations.
[0574] A "smartphone" refers to a multi-function mobile phone that can make calls and connect to the Internet.
[0575] "Head-mounted display" refers to a display device worn on the head.
[0576] "Distribution" refers to the act of transmitting or providing information via the Internet or other means of communication.
[0577] The present invention is a system for accurately conveying the memories and messages of a deceased person to their bereaved families. A specific embodiment of the system is described below. This system covers the collection of voice data, storing it in a database, summarizing the text data using a generative model, and delivering the information via smartphones and head-mounted displays.
[0578] First, the user speaks to the smart speaker to collect voice data. The smart speaker is equipped with a speech recognition engine that converts the voice data into text data in real time. This speech recognition engine can be Google's speech recognition API or IBM Watson.
[0579] The converted text data is sent to a server via the Internet and stored in a database on the server side, which may use a storage system such as SQL Server or MongoDB.
[0580] The server implements a generative model to periodically analyze and summarize text data. This generative model, for example, the "BART" model based on the Hugging Face "transformers" library, is used. The generative model analyzes the stored text data, extracts important keywords and sentences, and creates a summary. The summarized information is then stored on the server.
[0581] Family members can access the system through a user interface using a smartphone or head-mounted display. The user interface provides information summarized by the generative model. Android and iOS app development frameworks are used as the user interface technology. This system allows family members to access the memories and messages of the deceased at any time.
[0582] The detailed operation of the system will be explained below by showing a specific example.
[0583] Examples:
[0584] If a user wants to collect memories of the deceased using the following prompt, they can speak to the smart speaker, saying, "Tell me your memories of Mother's Day." The smart speaker converts the speech into text and sends it to the server. The server summarizes the text and generates a summary such as, "I always sent flowers on Mother's Day. My mother was always very happy." The bereaved can view this summarized message at any time via their smartphone or head-mounted display.
[0585] This makes it possible to efficiently organize memories of the deceased and properly convey them to the bereaved family.
[0586] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0587] Step 1:
[0588] The user speaks to the smart speaker to provide voice data. The smart speaker receives the user's voice as input, and the built-in voice recognition engine converts the voice data into text data in real time. As a result, the input (voice data) acquired by the smart speaker is output as text data.
[0589] Step 2:
[0590] The smart speaker sends the converted text data over the Internet to a server. The server receives this text data as input and stores it in a database in an appropriate format. The database uses a cloud platform such as Azure SQL Database or Amazon RDS. The text data is stored in the database along with the date, time, and user identification information.
[0591] Step 3:
[0592] The server periodically accesses the database and analyzes the stored text data. At this time, it uses a generative AI model (for example, the BART model implemented using Hugging Face's transformers library) to summarize the text data. The server takes the text data as input and uses the generative model to extract important keywords and sentences. As a result of the analysis, the summarized text is output and stored again in the database.
[0593] Step 4:
[0594] The bereaved family members access the user interface via a smartphone or head-mounted display. The user interface requests information from the server using a specific prompt (e.g., "Tell me your memories of Mother's Day"). The server receives this prompt, searches for relevant summary data from a database, and responds to the user interface. The user interface receives the user's request as input and displays the summarized information as output.
[0595] Step 5:
[0596] The user checks the summarized information through a smartphone or head-mounted display. Specifically, by checking the summary message displayed on the user interface (e.g., "I always sent flowers to my mother on Mother's Day. She was always very happy every time"), the user can reminisce about the deceased. This allows the user to access the memories and messages of the deceased at any time.
[0597] This series of processes creates a system that efficiently organizes the memories and messages of the deceased and conveys them appropriately to the bereaved.
[0598] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0599] This invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and further describes a form that combines an emotion engine that recognizes the user's emotions. This system involves a process in which the user uses a smart speaker to collect voice data, converts that data into text data, organizes and summarizes it, and provides it to the bereaved. In addition, the emotion engine recognizes the user's emotions and provides a more in-depth summary or letter-style output based on that information.
[0600] System Configuration
[0601] 1. Smart Speaker
[0602] This system collects voice data when the user (deceased person) speaks to a smart speaker. The smart speaker is equipped with a voice recognition engine that instantly converts the voice data into text data. An emotion engine is also built into the smart speaker, which recognizes emotions from the user's voice in real time.
[0603] 2. Server
[0604] The text data and emotion information are sent to a server via the Internet. The server receives the data and stores it in a database. The database stores the text data, emotion information, collection date and time, and user identification information.
[0605] 3. Generative Model
[0606] The server is equipped with a generative model using artificial intelligence that periodically analyzes the stored text data and emotional information. The generative model extracts important keywords and sentences from the text data and summarizes the memories and messages of the deceased along with emotional information.
[0607] 4. User Interface
[0608] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server then provides information organized and summarized by the generative model in an interactive format through a user interface. Adding emotional information allows for a more in-depth dialogue. It is also possible to output the information in the form of a letter upon request from the family members.
[0609] 5. Notary support and expert guidance
[0610] Support for creating a legally valid will is also provided. When a user makes such a request to the server, the server collects the necessary information and directs the user to a lawyer or notary public, thereby simplifying the process of creating a legally valid will.
[0611] Specific examples
[0612] Example 1: Collecting and summarizing a conversation
[0613] User: Speaks into smart speaker: "Today is my wife's wedding anniversary. I'm so grateful to her."
[0614] Smart speaker: Recognizes voice and the emotion engine identifies emotions such as "gratitude."
[0615] Server: Converts the voice data into text data and emotional information and stores them in a database.
[0616] Generative model: Analyzes and summarizes text data and emotional information. "Expressing gratitude on wedding anniversary."
[0617] Example 2: Providing information through a user interface
[0618] Survivors: Access the system from a computer and type in "I'd like to know about your memories of our wedding anniversary."
[0619] Server: Uses a generative model to search for relevant summary data, adds emotional information, and displays it on the user interface. "Today is my wedding anniversary with my wife. I'm really grateful to her."
[0620] Example 3: Letter-style output
[0621] Bereaved family: Request a letter-style output. "Please create a letter about our wedding anniversary."
[0622] Server: Generates letter-style text using a generative model and adds emotional information. "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0623] In this way, by combining the emotion engine, a system is realized that can easily and deeply convey the feelings of the deceased to the bereaved family.
[0624] The processing flow will be explained below.
[0625] Step 1:
[0626] The user speaks to the smart speaker.
[0627] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0628] Smart speakers collect user voice in real time.
[0629] Step 2:
[0630] The smart speaker converts the voice data into text data.
[0631] The speech recognition engine analyzes the voice data and converts it into text.
[0632] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0633] Step 3:
[0634] The smart speaker uses an emotion engine to recognize emotions from voice data.
[0635] The emotion engine analyzes the voice data and identifies emotions such as "gratitude" and "joy."
[0636] Example: Recognizing the emotion of "gratitude."
[0637] Step 4:
[0638] The smart speaker sends the converted text data and emotional information to the server.
[0639] The text data and emotion information are sent to a server via the Internet.
[0640] Step 5:
[0641] The server receives the text data and emotion information and stores them in a database.
[0642] The received text data and emotion information are assigned a timestamp and user identification information and registered in a database.
[0643] Example: { "timestamp": "2023-10-12T14:30:00Z", "user": "Deceased Person A", "message": "Today is my wedding anniversary with my wife. I am truly grateful to her.", "emotion": "Thank you"}
[0644] Step 6:
[0645] A generative model in the server analyzes and summarizes the text data and sentiment information in the database.
[0646] The generative model extracts important keywords and sentences from text data and emotional information, and creates summaries that reflect the emotional information.
[0647] Example: { "topic": "Wedding anniversary", "content": "Expressing gratitude for his wife on their wedding anniversary."}
[0648] Step 7:
[0649] The user (surviving family member) accesses the system from a terminal and enters a request.
[0650] Example: "I'd like to know your memories of your wedding anniversary."
[0651] The request is sent over the Internet to a server.
[0652] Step 8:
[0653] The server retrieves summary data based on the generative model, adds emotional information, and provides it through a user interface.
[0654] Matching data is searched for and displayed on the user interface based on emotion information.
[0655] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0656] Step 9:
[0657] The user requests output in letter format.
[0658] Example: "Please write me a letter about our wedding anniversary."
[0659] The request is sent over the Internet to a server.
[0660] Step 10:
[0661] The server's generative model generates a letter-style document based on the request.
[0662] A letter-style text is created based on past text data and emotional information and sent to the device.
[0663] For example: "Dear wife, every time our wedding anniversary comes around, my heart fills with gratitude for you. You are the treasure of my life."
[0664] The above is the specific processing flow that combines the emotion engine in the "Conclusion" system.
[0665] Example 2
[0666] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0667] There is a problem in accurately conveying the deceased's thoughts to the bereaved. Conventional methods do not accurately convey the feelings and intentions of the deceased, and do not provide deep meaning to the bereaved. In addition, the process of recording and organizing the deceased's thoughts and creating a legally valid will, if necessary, is complicated and time-consuming.
[0668] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0669] In this invention, the server includes means for collecting voice data from a user, means for converting the voice data into text data, means for recognizing the user's emotions from the voice data, means for storing the text data and emotional information in a database, means for analyzing and summarizing the stored text data and emotional information using a generative model, and means for providing information summarized by the generative model through a user interface. This allows the deceased's thoughts, along with their emotional information, to be organized and summarized and conveyed to the bereaved. It can also output the message in letter format and support the creation of a legally valid will, allowing for a deeper understanding of the bereaved and legal response.
[0670] "Voice data" refers to voice signal information collected when a user speaks to a smart speaker or other voice input device.
[0671] "Text data" is character information converted from voice data by a voice recognition engine.
[0672] "Emotion information" is data that is analyzed from voice data by the emotion engine and indicates the user's emotional state.
[0673] A "database" is a collection of information for systematically storing and managing text data and emotional information.
[0674] A "generative model" is a program that uses artificial intelligence to analyze text data and emotional information, extract important keywords and sentences, and generate summaries.
[0675] A "user interface" refers to the screen or operating means that allows bereaved family members to access the system through devices such as computers or smartphones and obtain the necessary information.
[0676] "Letter-format output" is a function that provides information summarized by the generative model in the form of a letter.
[0677] A "legally valid will" is a written document that expresses the final wishes of a deceased person and is legally valid.
[0678] "Expert Guidance" is a support function that encourages users who wish to write a will to contact or consult with legal experts (such as lawyers or notaries).
[0679] The system of the present invention aims to accurately convey the thoughts of the deceased to the bereaved family, and is realized by combining an emotion engine that recognizes the user's emotions. This system includes the following components.
[0680] 1. Smart Speaker
[0681] The first component of this system is the smart speaker. Users speak into the smart speaker to collect their messages and thoughts as voice data. The smart speaker has a built-in voice recognition engine (e.g., Google Speech-to-Text or Amazon Alexa) that instantly converts the voice data into text data. In addition, a built-in emotion engine recognizes the user's emotions in real time from the collected voice.
[0682] Examples:
[0683] A user says to a smart speaker, "Today is my wife's wedding anniversary. I'm so grateful to her."
[0684] 2. Server
[0685] The text data and emotion information are sent to a server via the Internet. The server receives the data and stores it in a database. The database stores the text data, emotion information, collection date and time, and user identification information.
[0686] Examples:
[0687] The server stores the text data (e.g., "Today is my wedding anniversary with my wife. I am truly grateful to her") and emotional information (gratitude) sent from the smart speaker in a database.
[0688] 3. Generative AI Models
[0689] The server is equipped with a generative AI model that periodically analyzes the stored text data and emotional information. Using the generative model (e.g., OpenAI GPT-3), it extracts important keywords and sentences from the text data and summarizes the deceased's memories and messages along with emotional information.
[0690] Examples:
[0691] The generative AI model analyzes the stored text data and emotional information to generate summaries such as "expressing gratitude on the wedding anniversary."
[0692] 4. User Interface
[0693] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server provides information organized and summarized by the generative model through a user interface. Adding emotional information allows for deeper dialogue.
[0694] Examples:
[0695] When a family member accesses the system from their computer and types, "I'd like to know about your memories of our wedding anniversary," the server searches for relevant summary data and displays information such as, "Today is my wedding anniversary with my wife. I'm truly grateful to her."
[0696] 5. Letter-style output
[0697] Upon request of the bereaved family, the information can be output in letter format.
[0698] Examples:
[0699] If a family member requests a letter about their wedding anniversary, the generative AI model will generate a letter such as, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0700] 6. Support for creating legally binding wills
[0701] When a user requests the server for assistance in creating a legally binding will, the server collects the necessary information and directs the user to a lawyer or notary public.
[0702] Examples:
[0703] When a user requests the creation of a legally binding will, the server collects the appropriate information and guides them through the process and contacting the appropriate professional.
[0704] This system makes it possible to organize and summarize the feelings of the deceased along with emotional information, and to convey them in depth to the bereaved. It can also output them in letter format and assist in the creation of legally valid wills, making it easier for the bereaved to understand and respond to the situation.
[0705] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0706] Step 1:
[0707] The user speaks to the smart speaker.
[0708] Input: User's voice message (e.g., "Today is my wedding anniversary with my wife. I'm so grateful to her.")
[0709] How it works: Smart speakers use built-in microphones to collect voice data.
[0710] Output: Collected audio data
[0711] Step 2:
[0712] Smart speakers convert voice data into text data and recognize emotions.
[0713] Input: Audio data
[0714] How it works: A voice recognition engine (e.g., Google Speech-to-Text) built into a smart speaker converts voice data into text data. At the same time, an emotion engine analyzes the text data and generates emotion information (e.g., "thank you").
[0715] Output: Text data (e.g., "Today is my wedding anniversary with my wife. I am so grateful to her."), emotional information (e.g., "gratitude")
[0716] Step 3:
[0717] The smart speaker sends text data and emotional information to the server.
[0718] Input: Text data, emotion information
[0719] How it works: Smart speakers send data over the internet to a server, which often encrypts the data before sending it.
[0720] Output: Text data and emotion information sent to the server
[0721] Step 4:
[0722] The server stores the text data and emotion information in a database.
[0723] Input: Text data, emotion information
[0724] Operation: The server analyzes the received data and stores the text data, emotion information, collection date and time, user identification information, etc. in a relational database.
[0725] Output: Information stored in the database
[0726] Step 5:
[0727] A generative AI model implemented on the server analyzes text data and emotional information to generate a summary.
[0728] Input: Text data and emotion information stored in a database
[0729] How it works: A generative AI model (e.g., OpenAI GPT-3) periodically scans the database, analyzes the text data and sentiment information, extracts important keywords and sentences, and generates summaries taking sentiment information into account.
[0730] Output: Summarized information (e.g., "Expressing gratitude for our wedding anniversary.")
[0731] Step 6:
[0732] Family members access the system through a terminal and request information.
[0733] Input: Request from the bereaved family (e.g., "I'd like to know about your memories of your wedding anniversary.")
[0734] How it works: Family members access the system using a computer or smartphone and enter information.
[0735] Output: Input information of bereaved family members
[0736] Step 7:
[0737] The server provides the summarized information through a user interface.
[0738] Input: Family request information
[0739] How it works: The server uses the generative model to search for relevant summary data, adds emotional information, and displays it on the user interface.
[0740] Output: Summary information displayed in a user interface (e.g., "Today is my wife's wedding anniversary. I'm so grateful to her.")
[0741] Step 8:
[0742] The family requests a letter-style output.
[0743] Input: A request for output in the form of a letter (e.g., "Please write me a letter about my wedding anniversary.")
[0744] How it works: The family requests a letter-style output from the system.
[0745] Output: Request information
[0746] Step 9:
[0747] The server uses a generative AI model to generate a letter-style document.
[0748] Input: Request information for output in letter format, database information
[0749] How it works: The generative AI model generates letter-style text based on summarized text data and emotional information.
[0750] Output: A letter-style sentence (e.g., "Dear Wife, Every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life.")
[0751] Step 10:
[0752] The server collects information to support the creation of a legally binding will and provides guidance to experts.
[0753] Input: User request information
[0754] How it works: Based on the user's request, the server gathers the necessary information from the database and directs them to the appropriate lawyer or notary public.
[0755] Output: Guidance information (e.g., contact instructions for a lawyer or notary public, list of required documents)
[0756] (Application example 2)
[0757] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0758] When dealing with customers in brick-and-mortar stores, it is difficult for employees to properly understand customer emotions and provide optimal service based on that. A particular challenge is the lack of a system that can recognize and reflect customer emotions and requests in real time. Therefore, there is a need for a system that can improve customer satisfaction and the quality of employee service.
[0759] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for using an emotion engine to analyze and summarize the saved text data and emotion information, and means for providing information summarized by the generative model and emotion engine in real time via an audiovisual device used by a customer service representative. This makes it possible to properly recognize customer emotions when serving customers in a physical store and provide services based on those emotions quickly and accurately.
[0760] Definitions of important words
[0761] "Voice data" refers to digitized data of the voice uttered by the user.
[0762] "Text data" refers to data obtained by converting voice data into character information.
[0763] A "database" is a digital storage device for systematically storing text data and emotional information.
[0764] A "generative model" is an algorithm or software that uses artificial intelligence to extract and summarize important keywords and sentences from text data.
[0765] "Emotion information" is information that represents the emotional state of a user, as recognized from voice data or text data.
[0766] An "emotion engine" is software that analyzes a user's voice data and text data to extract emotional information.
[0767] "Audiovisual devices" is a general term for devices that provide visual and auditory information, such as smart glasses and head-mounted displays.
[0768] "Real-time" refers to processing and response occurring with almost no delay after an event occurs.
[0769] "Customer service personnel" refers to employees and staff who interact directly with customers and provide services in physical stores.
[0770] MODE FOR CARRYING OUT THE INVENTION
[0771] This invention relates to a system for supporting customer service in brick-and-mortar stores, and aims to increase customer satisfaction by utilizing emotion recognition technology in particular. This system uses audiovisual devices such as smart glasses to enable store staff to grasp customer emotions in real time and provide optimal service.
[0772] System configuration
[0773] 1. Collection of audio data
[0774] Customer service representatives will wear smart glasses or head-mounted displays to communicate with customers. The smart glasses have built-in microphones that collect customer voice data in real time.
[0775] 2. Converting audio data to text data
[0776] The collected voice data is instantly converted into text data using the device's voice recognition engine (e.g., Google Speech-to-Text API), which is temporarily stored on the device and then sent to a server.
[0777] 3. Extraction of Emotional Information
[0778] The server analyzes the received text data and extracts customer emotional information using an emotion engine (e.g., EmotionRecognition software). This emotional information is then stored in a database along with the text data.
[0779] 4. Data Analysis and Summarization
[0780] A generative AI model implemented on the server analyzes the stored text data and sentiment information, extracts important keywords and sentences, and summarizes them.
[0781] 5. Real-time information provision
[0782] The server then provides the customer service representative with real-time information summarized by the generative model via an audiovisual device, and the employee's smart glasses display the customer's emotional state and appropriate response.
[0783] Specific examples
[0784] Program processing
[0785] 1. A customer in the store says, "I've been feeling stressed lately and I want a product to help me relax."
[0786] 2. Voice is collected by the microphone in the smart glasses and instantly converted into text data.
[0787] 3. On the server side, the text data is analyzed by an emotion engine, and the emotion "stress" is extracted.
[0788] 4. The generative model analyzes the text data and generates a summary: "We suggest a relaxing aroma set."
[0789] 5. This generated summary information is displayed on an audiovisual device, and the employee can recommend a product to the customer by saying, "We suggest an aroma set that will have a relaxing effect."
[0790] Prompt Sentence Examples
[0791] "A customer might say, 'I've been feeling stressed lately and I want a product to help me relax.' Analyze their emotions with an emotion recognition engine and generate a dialogue that suggests relevant products."
[0792] This system makes it possible to understand customer sentiment in real time and quickly make suggestions based on that sentiment, thereby significantly improving customer satisfaction.
[0793] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0794] Program processing steps
[0795] Step 1:
[0796] Audio data collection
[0797] A user (customer service representative) wears smart glasses and interacts with customers in a physical store. When a customer speaks, the microphone built into the smart glasses collects the voice.
[0798] Input: Customer voice.
[0799] Output: Digitized audio data.
[0800] What it does: The microphone in the smart glasses activates and records what the customer says in real time.
[0801] Step 2:
[0802] Converting audio data to text data
[0803] The device (smart glasses) converts the collected voice data into text data using an internal voice recognition engine, which uses the Google Speech-to-Text API.
[0804] Input: Digitized audio data.
[0805] Output: Text data.
[0806] How it works: The voice recognition engine processes the voice data and converts it into corresponding text data, which is then temporarily stored in the smart glasses.
[0807] Step 3:
[0808] Sending text data
[0809] The terminal (smart glasses) sends the converted text data to a server via the Internet.
[0810] Input: Text data.
[0811] Output: The text data sent to the server.
[0812] How it works: The smart glasses use Wi-Fi or mobile networks to upload text data to a server.
[0813] Step 4:
[0814] Extracting Emotional Information
[0815] The server analyzes the received text data using an emotion engine (e.g., EmotionRecognition software) to extract customer emotional information.
[0816] Input: The text data sent to the server.
[0817] Output: Emotion information (e.g., "joy", "sad", "stress", etc.).
[0818] What it does: The emotion engine on the server scans the text data and identifies relevant emotions.
[0819] Step 5:
[0820] Data storage
[0821] The server stores the converted text data and the extracted emotion information in a database.
[0822] Input: Text data, emotion information.
[0823] Output: Information stored in a database.
[0824] Specific operation: The server's database management system stores text data and emotional information in an orderly manner.
[0825] Step 6:
[0826] Data analysis and summary
[0827] A generative AI model (e.g., OpenAI GPT-3) implemented on the server analyzes the stored text data and emotional information, extracts important keywords and sentences, and generates a summary.
[0828] Input: Text data and emotion information stored in the database.
[0829] Output: Summarized information.
[0830] How it works: The generative AI model analyzes the information in the database, extracts the most important parts, and generates a summary.
[0831] Step 7:
[0832] Real-time provision
[0833] The server provides information summarized by the generative model to the audiovisual device (smart glasses) in real time.
[0834] Input: Summarized information.
[0835] Output: Summary information displayed or spoken on the smart glasses.
[0836] How it works: The server uses Wi-Fi or mobile networks to send the summarized text to the smart glasses, which then displays it on the customer service representative's display or notifies them via voice.
[0837] Step 8:
[0838] Service delivery based on feedback
[0839] Customer service representatives will suggest the most suitable services and products to customers based on the summary information and suggestions displayed on the smart glasses.
[0840] Input: Summary information displayed on smart glasses.
[0841] Output: Services and offers to customers.
[0842] Specific action: The customer service representative checks the summary information and makes appropriate suggestions or asks questions to the customer.
[0843] By implementing the above steps, it is possible to provide services that are sensitive to the customer's emotions when dealing with customers in physical stores.
[0844] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0845] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0846] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0847] [Third embodiment]
[0848] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0849] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0850] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0851] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0852] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0853] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0854] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0855] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0856] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0857] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0858] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0859] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0860] The present invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and a specific embodiment thereof is described below. This system includes a process in which a user uses a smart speaker to collect voice data, converts the data into text data, organizes and summarizes it, and provides it to the bereaved.
[0861] System Configuration
[0862] 1. Smart Speaker
[0863] This system collects voice data by having the user (deceased person) speak into a smart speaker, which is equipped with a voice recognition engine that instantly converts the voice data into text data.
[0864] 2. Server
[0865] The text data is sent over the Internet to a server, which then stores the received text data in a database. This database contains not only the text data, but also the date and time the data was collected and the user's identification information.
[0866] 3. Generative Model
[0867] The server is equipped with a generative model using artificial intelligence that periodically analyzes the stored text data, extracts important keywords and sentences from the text data, and summarizes the memories and messages of the deceased.
[0868] 4. User Interface
[0869] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server provides information organized and summarized by the generative model in an interactive format through a user interface. It is also possible to output the information in letter format upon request from family members.
[0870] 5. Notary support and expert guidance
[0871] Support for creating a legally valid will is also provided. When a user makes such a request to the server, the server collects the necessary information and directs the user to a lawyer or notary public, thereby simplifying the process of creating a legally valid will.
[0872] Specific examples
[0873] Example 1: Collecting and summarizing a conversation
[0874] User: Speaks into smart speaker: "Today is my wife's wedding anniversary. I'm so grateful to her."
[0875] Server: Receives voice data and the voice recognition engine converts it into text data.
[0876] Generative model: Analyzes and summarizes text data. "Expressing gratitude on wedding anniversary."
[0877] Example 2: Providing information through a user interface
[0878] Survivors: Access the system from a computer and type in "I'd like to know about your memories of our wedding anniversary."
[0879] Server: Uses the generative model to find relevant summary data and displays it through a user interface. "Today is my wife's wedding anniversary. I'm so grateful to her."
[0880] Example 3: Letter-style output
[0881] Bereaved family: Request a letter-style output. "Please create a letter about our wedding anniversary."
[0882] Server: Generates a letter-style sentence using a generative model. "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0883] In this way, a system is realized that can easily and reliably convey the thoughts of the deceased to the bereaved family.
[0884] The processing flow will be explained below.
[0885] Step 1:
[0886] The user speaks to the smart speaker.
[0887] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0888] Smart speakers collect user voice in real time.
[0889] Step 2:
[0890] The smart speaker converts the voice data into text data.
[0891] The speech recognition engine analyzes the voice data and converts it into text.
[0892] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0893] Step 3:
[0894] The smart speaker sends the converted text data to the server.
[0895] The text data is sent to a server via the Internet.
[0896] Step 4:
[0897] The server receives the text data and stores it in a database.
[0898] The received text data is assigned a timestamp and user identification information and registered in a database.
[0899] Example: { "timestamp": "2023-10-12T14:30:00Z", "user": "Deceased Person A", "message": "Today is my wedding anniversary with my wife. I am truly grateful to her."}
[0900] Step 5:
[0901] A generative model in the server analyzes and summarizes the text data in the database.
[0902] A generative model extracts important keywords and sentences from text data and creates a summary.
[0903] Example: { "topic": "Wedding anniversary", "content": "Expressing gratitude for his wife on their wedding anniversary."}
[0904] Step 6:
[0905] The user (surviving family member) accesses the system from a terminal and enters a request.
[0906] Example: "I'd like to know your memories of your wedding anniversary."
[0907] The request is sent over the Internet to a server.
[0908] Step 7:
[0909] The server retrieves the summarized data from the generative model and provides it through a user interface.
[0910] Matching data is retrieved and displayed on the user interface.
[0911] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[0912] Step 8:
[0913] The user requests output in letter format.
[0914] Example: "Please write me a letter about our wedding anniversary."
[0915] The request is sent over the Internet to a server.
[0916] Step 9:
[0917] The server's generative model generates a letter-style document based on the request.
[0918] A letter-style text is created based on past text data and sent to the device.
[0919] For example: "Dear wife, every time our wedding anniversary comes around, my heart fills with gratitude for you. You are the treasure of my life."
[0920] The above is the specific processing flow of the "Conclusion" system.
[0921] Example 1
[0922] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0923] In modern society, there are limited ways to accurately convey the wishes of the deceased to their bereaved families, and there is a particular problem of a lack of an efficient system for the process from collecting to transmitting information using voice. Furthermore, the process of creating a legally valid will is complicated, and there is a lack of support methods to simplify the process.
[0924] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0925] In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing and summarizing the stored text data using a generative model, means for providing information summarized by the generative model through a user interface, means for searching for relevant summary data based on a request entered by a user, means for generating the summary data in letter format, and means for collecting information to support the creation of a legally binding will and guiding experts. This allows the wishes of the deceased to be effectively conveyed to the surviving family and simplifies legal procedures.
[0926] "User" means an individual who uses the System to provide voice data and receive subsequent services.
[0927] "Voice data" refers to a digital recording of a voice signal emitted by a user through a device such as a smart speaker.
[0928] "Text data" refers to digital data that has been converted from voice data into text information.
[0929] "Database" refers to an information management system for storing text data and its associated metadata (date and time, identification information, etc.).
[0930] A "generative model" refers to a software model that uses artificial intelligence to analyze text data, extract important keywords and sentences, and summarize them.
[0931] "User interface" refers to the interface through which users and bereaved families access the system and input or retrieve data.
[0932] A "request" is an instruction from a user or family member to the system requesting the provision of specific data or services.
[0933] "Summary data" is shortened text data that contains key points or sentences extracted by a generative model.
[0934] "Letter format" refers to a document in which summary data has been reconstructed into a letter format that is easy for humans to read.
[0935] "Experts" refers to professionals such as lawyers and notaries who have knowledge and qualifications regarding legal procedures and the preparation of wills.
[0936] A "legally valid will" is a will that has legal effect and is officially recognized under the law.
[0937] This invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and a specific embodiment thereof is described below. This system involves a process in which a user collects voice data, converts it into text data, summarizes and organizes it using a generative AI model, and provides it to the bereaved.
[0938] System Configuration
[0939] 1. Collection of audio data
[0940] Users collect voice data by speaking into a voice input device such as a smart speaker, which then converts the voice data into text data in real time using a voice recognition system such as Amazon Alexa or Google Assistant.
[0941] 2. Sending and saving text data
[0942] The text data generated by the voice recognition system is sent over the Internet to a server. The server stores the received text data and its associated metadata (date and time, user identification information, etc.) in a database. This database serves to store and manage the memories and messages of the deceased.
[0943] 3. Analysis and Summarization of Text Data
[0944] The server is equipped with a generative AI model (e.g., OpenAI's GPT-3) that periodically analyzes the text data in the database. The generative model extracts important keywords and sentences from the text data and summarizes and organizes the memories and messages of the deceased.
[0945] 4. Providing information through the user interface
[0946] Family members access the system through devices such as computers or smartphones. Using the user interface, they input specific requests, such as "I want to know about memories of our wedding anniversary." The server uses the generative AI model to search for the relevant summary data and displays it through the user interface.
[0947] An example of a specific prompt is "Show me a message about your wedding anniversary memories."
[0948] 5. Generating Letter Format
[0949] If the family requests a letter-style output, the server sends a specific prompt to the generative AI model, which generates a letter with a message like, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0950] 6. Legal process support
[0951] When a user wishes to create a legally valid will, the server collects the necessary information and directs them to professionals such as lawyers and notaries. Also, when the bereaved family requests the creation of a will through the system, the server organizes and provides the necessary information, helping to simplify the legal procedures.
[0952] In this way, this system can easily and reliably convey the wishes of the deceased to their family members.It is also extremely convenient because it has a function to assist users in legal procedures.
[0953] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0954] Step 1:
[0955] The user speaks to the smart speaker. For example, "Today is my wedding anniversary with my wife. I'm really grateful to her." This voice is input into the smart speaker.
[0956] Step 2:
[0957] Smart speakers use a built-in voice recognition engine (such as Google Assistant) to convert the received voice data into text data in real time, and this conversion process outputs the voice data as text data.
[0958] Step 3:
[0959] The text data is sent to a server via the Internet. The data sent includes not only the text data but also the date and time the voice was collected and the user's identification information. The server receives this data.
[0960] Step 4:
[0961] The server saves the received text data and its associated metadata in a database. The saving process stores the text data and metadata in the database.
[0962] Step 5:
[0963] The server passes the text data stored in the database to a generative AI model (for example, OpenAI's GPT-3) for analysis. The generative model reads the text data, extracts important keywords and sentences, and generates summary data. This summary data is the output from the generative model.
[0964] Step 6:
[0965] The bereaved family accesses the system from a PC or smartphone and inputs a specific request into the user interface (e.g., "I would like to know about memories of our wedding anniversary.") This request is sent to the server.
[0966] Step 7:
[0967] The server receives the request from the family and uses the generative AI model to search for the relevant summary data. The search process is carried out and the relevant summary data is output.
[0968] Step 8:
[0969] The server displays the retrieved summary data through a user interface, allowing the family members to view the information.
[0970] Step 9:
[0971] If the family requests a letter-style output, the server sends a specific prompt to the generative AI model, which generates a letter-style text, such as, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[0972] Step 10:
[0973] If a user wishes to create a legally valid will, the server collects the necessary information and directs them to professionals such as lawyers and notaries, helping them to navigate the legal process more easily.
[0974] Through the above processing steps, the system can accurately convey the wishes of the deceased to the bereaved family and can also provide support in legal procedures.
[0975] (Application example 1)
[0976] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0977] In conventional estate sorting, it has been difficult to accurately convey the deceased's memories and messages to the bereaved. Also, the task of organizing and summarizing the deceased's thoughts is time-consuming and laborious, placing a heavy burden on the bereaved. Furthermore, the means to easily retrieve and access the deceased's thoughts were limited, making it difficult for the bereaved to reminisce. The present invention aims to solve these problems.
[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0979] In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing and summarizing the stored text data using a generative model, means for providing information summarized by the generative model through a user interface, and means for delivering the summarized information to a smartphone or head-mounted display. This allows memories and messages of the deceased to be organized and easily conveyed to the bereaved. Furthermore, the bereaved can access memories of the deceased at any time using a smartphone or head-mounted display, thereby reducing the burden on the bereaved.
[0980] "User" refers to an individual who provides voice data using this system.
[0981] "Voice Data" refers to recordings of voice signals collected when a user speaks.
[0982] "Text data" refers to data that has been converted from voice data into text information using voice recognition technology.
[0983] "Database" refers to a digital storage system for storing collected text data.
[0984] A "generative model" refers to an algorithm that uses artificial intelligence techniques to analyze and summarize text data.
[0985] "User interface" refers to the interface on a computer system that allows users to view information and perform operations.
[0986] A "smartphone" refers to a multi-function mobile phone that can make calls and connect to the Internet.
[0987] "Head-mounted display" refers to a display device worn on the head.
[0988] "Distribution" refers to the act of transmitting or providing information via the Internet or other means of communication.
[0989] The present invention is a system for accurately conveying the memories and messages of a deceased person to their bereaved families. A specific embodiment of the system is described below. This system covers the collection of voice data, storing it in a database, summarizing the text data using a generative model, and delivering the information via smartphones and head-mounted displays.
[0990] First, the user speaks to the smart speaker to collect voice data. The smart speaker is equipped with a speech recognition engine that converts the voice data into text data in real time. This speech recognition engine can be Google's speech recognition API or IBM Watson.
[0991] The converted text data is sent to a server via the Internet and stored in a database on the server side, which may use a storage system such as SQL Server or MongoDB.
[0992] The server implements a generative model to periodically analyze and summarize text data. This generative model, for example, the "BART" model based on the Hugging Face "transformers" library, is used. The generative model analyzes the stored text data, extracts important keywords and sentences, and creates a summary. The summarized information is then stored on the server.
[0993] Family members can access the system through a user interface using a smartphone or head-mounted display. The user interface provides information summarized by the generative model. Android and iOS app development frameworks are used as the user interface technology. This system allows family members to access the memories and messages of the deceased at any time.
[0994] The detailed operation of the system will be explained below by showing a specific example.
[0995] Examples:
[0996] If a user wants to collect memories of the deceased using the following prompt, they can speak to the smart speaker, saying, "Tell me your memories of Mother's Day." The smart speaker converts the speech into text and sends it to the server. The server summarizes the text and generates a summary such as, "I always sent flowers on Mother's Day. My mother was always very happy." The bereaved can view this summarized message at any time via their smartphone or head-mounted display.
[0997] This makes it possible to efficiently organize memories of the deceased and properly convey them to the bereaved family.
[0998] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0999] Step 1:
[1000] The user speaks to the smart speaker to provide voice data. The smart speaker receives the user's voice as input, and the built-in voice recognition engine converts the voice data into text data in real time. As a result, the input (voice data) acquired by the smart speaker is output as text data.
[1001] Step 2:
[1002] The smart speaker sends the converted text data over the Internet to a server. The server receives this text data as input and stores it in a database in an appropriate format. The database uses a cloud platform such as Azure SQL Database or Amazon RDS. The text data is stored in the database along with the date, time, and user identification information.
[1003] Step 3:
[1004] The server periodically accesses the database and analyzes the stored text data. At this time, it uses a generative AI model (for example, the BART model implemented using Hugging Face's transformers library) to summarize the text data. The server takes the text data as input and uses the generative model to extract important keywords and sentences. As a result of the analysis, the summarized text is output and stored again in the database.
[1005] Step 4:
[1006] The bereaved family members access the user interface via a smartphone or head-mounted display. The user interface requests information from the server using a specific prompt (e.g., "Tell me your memories of Mother's Day"). The server receives this prompt, searches for relevant summary data from a database, and responds to the user interface. The user interface receives the user's request as input and displays the summarized information as output.
[1007] Step 5:
[1008] The user checks the summarized information through a smartphone or head-mounted display. Specifically, by checking the summary message displayed on the user interface (e.g., "I always sent flowers to my mother on Mother's Day. She was always very happy every time"), the user can reminisce about the deceased. This allows the user to access the memories and messages of the deceased at any time.
[1009] This series of processes creates a system that efficiently organizes the memories and messages of the deceased and conveys them appropriately to the bereaved.
[1010] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1011] This invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and further describes a form that combines an emotion engine that recognizes the user's emotions. This system involves a process in which the user uses a smart speaker to collect voice data, converts that data into text data, organizes and summarizes it, and provides it to the bereaved. In addition, the emotion engine recognizes the user's emotions and provides a more in-depth summary or letter-style output based on that information.
[1012] System Configuration
[1013] 1. Smart Speaker
[1014] This system collects voice data when the user (deceased person) speaks to a smart speaker. The smart speaker is equipped with a voice recognition engine that instantly converts the voice data into text data. An emotion engine is also built into the smart speaker, which recognizes emotions from the user's voice in real time.
[1015] 2. Server
[1016] The text data and emotion information are sent to a server via the Internet. The server receives the data and stores it in a database. The database stores the text data, emotion information, collection date and time, and user identification information.
[1017] 3. Generative Model
[1018] The server is equipped with a generative model using artificial intelligence that periodically analyzes the stored text data and emotional information. The generative model extracts important keywords and sentences from the text data and summarizes the memories and messages of the deceased along with emotional information.
[1019] 4. User Interface
[1020] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server then provides information organized and summarized by the generative model in an interactive format through a user interface. Adding emotional information allows for a more in-depth dialogue. It is also possible to output the information in the form of a letter upon request from the family members.
[1021] 5. Notary support and expert guidance
[1022] Support for creating a legally valid will is also provided. When a user makes such a request to the server, the server collects the necessary information and directs the user to a lawyer or notary public, thereby simplifying the process of creating a legally valid will.
[1023] Specific examples
[1024] Example 1: Collecting and summarizing a conversation
[1025] User: Speaks into smart speaker: "Today is my wife's wedding anniversary. I'm so grateful to her."
[1026] Smart speaker: Recognizes voice and the emotion engine identifies emotions such as "gratitude."
[1027] Server: Converts the voice data into text data and emotional information and stores them in a database.
[1028] Generative model: Analyzes and summarizes text data and emotional information. "Expressing gratitude on wedding anniversary."
[1029] Example 2: Providing information through a user interface
[1030] Survivors: Access the system from a computer and type in "I'd like to know about your memories of our wedding anniversary."
[1031] Server: Uses a generative model to search for relevant summary data, adds emotional information, and displays it on the user interface. "Today is my wedding anniversary with my wife. I'm really grateful to her."
[1032] Example 3: Letter-style output
[1033] Bereaved family: Request a letter-style output. "Please create a letter about our wedding anniversary."
[1034] Server: Generates letter-style text using a generative model and adds emotional information. "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[1035] In this way, by combining the emotion engine, a system is realized that can easily and deeply convey the feelings of the deceased to the bereaved family.
[1036] The processing flow will be explained below.
[1037] Step 1:
[1038] The user speaks to the smart speaker.
[1039] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1040] Smart speakers collect user voice in real time.
[1041] Step 2:
[1042] The smart speaker converts the voice data into text data.
[1043] The speech recognition engine analyzes the voice data and converts it into text.
[1044] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1045] Step 3:
[1046] The smart speaker uses an emotion engine to recognize emotions from voice data.
[1047] The emotion engine analyzes the voice data and identifies emotions such as "gratitude" and "joy."
[1048] Example: Recognizing the emotion of "gratitude."
[1049] Step 4:
[1050] The smart speaker sends the converted text data and emotional information to the server.
[1051] The text data and emotion information are sent to a server via the Internet.
[1052] Step 5:
[1053] The server receives the text data and emotion information and stores them in a database.
[1054] The received text data and emotion information are assigned a timestamp and user identification information and registered in a database.
[1055] Example: { "timestamp": "2023-10-12T14:30:00Z", "user": "Deceased Person A", "message": "Today is my wedding anniversary with my wife. I am truly grateful to her.", "emotion": "Thank you"}
[1056] Step 6:
[1057] A generative model in the server analyzes and summarizes the text data and sentiment information in the database.
[1058] The generative model extracts important keywords and sentences from text data and emotional information, and creates summaries that reflect the emotional information.
[1059] Example: { "topic": "Wedding anniversary", "content": "Expressing gratitude for his wife on their wedding anniversary."}
[1060] Step 7:
[1061] The user (surviving family member) accesses the system from a terminal and enters a request.
[1062] Example: "I'd like to know your memories of your wedding anniversary."
[1063] The request is sent over the Internet to a server.
[1064] Step 8:
[1065] The server retrieves summary data based on the generative model, adds emotional information, and provides it through a user interface.
[1066] Matching data is searched for and displayed on the user interface based on emotion information.
[1067] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1068] Step 9:
[1069] The user requests output in letter format.
[1070] Example: "Please write me a letter about our wedding anniversary."
[1071] The request is sent over the Internet to a server.
[1072] Step 10:
[1073] The server's generative model generates a letter-style document based on the request.
[1074] A letter-style text is created based on past text data and emotional information and sent to the device.
[1075] For example: "Dear wife, every time our wedding anniversary comes around, my heart fills with gratitude for you. You are the treasure of my life."
[1076] The above is the specific processing flow that combines the emotion engine in the "Conclusion" system.
[1077] Example 2
[1078] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1079] There is a problem in accurately conveying the deceased's thoughts to the bereaved. Conventional methods do not accurately convey the feelings and intentions of the deceased, and do not provide deep meaning to the bereaved. In addition, the process of recording and organizing the deceased's thoughts and creating a legally valid will, if necessary, is complicated and time-consuming.
[1080] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1081] In this invention, the server includes means for collecting voice data from a user, means for converting the voice data into text data, means for recognizing the user's emotions from the voice data, means for storing the text data and emotional information in a database, means for analyzing and summarizing the stored text data and emotional information using a generative model, and means for providing information summarized by the generative model through a user interface. This allows the deceased's thoughts, along with their emotional information, to be organized and summarized and conveyed to the bereaved. It can also output the message in letter format and support the creation of a legally valid will, allowing for a deeper understanding of the bereaved and legal response.
[1082] "Voice data" refers to voice signal information collected when a user speaks to a smart speaker or other voice input device.
[1083] "Text data" is character information converted from voice data by a voice recognition engine.
[1084] "Emotion information" is data that is analyzed from voice data by the emotion engine and indicates the user's emotional state.
[1085] A "database" is a collection of information for systematically storing and managing text data and emotional information.
[1086] A "generative model" is a program that uses artificial intelligence to analyze text data and emotional information, extract important keywords and sentences, and generate summaries.
[1087] A "user interface" refers to the screen or operating means that allows bereaved family members to access the system through devices such as computers or smartphones and obtain the necessary information.
[1088] "Letter-format output" is a function that provides information summarized by the generative model in the form of a letter.
[1089] A "legally valid will" is a written document that expresses the final wishes of a deceased person and is legally valid.
[1090] "Expert Guidance" is a support function that encourages users who wish to write a will to contact or consult with legal experts (such as lawyers or notaries).
[1091] The system of the present invention aims to accurately convey the thoughts of the deceased to the bereaved family, and is realized by combining an emotion engine that recognizes the user's emotions. This system includes the following components.
[1092] 1. Smart Speaker
[1093] The first component of this system is the smart speaker. Users speak into the smart speaker to collect their messages and thoughts as voice data. The smart speaker has a built-in voice recognition engine (e.g., Google Speech-to-Text or Amazon Alexa) that instantly converts the voice data into text data. In addition, a built-in emotion engine recognizes the user's emotions in real time from the collected voice.
[1094] Examples:
[1095] A user says to a smart speaker, "Today is my wife's wedding anniversary. I'm so grateful to her."
[1096] 2. Server
[1097] The text data and emotion information are sent to a server via the Internet. The server receives the data and stores it in a database. The database stores the text data, emotion information, collection date and time, and user identification information.
[1098] Examples:
[1099] The server stores the text data (e.g., "Today is my wedding anniversary with my wife. I am truly grateful to her") and emotional information (gratitude) sent from the smart speaker in a database.
[1100] 3. Generative AI Models
[1101] The server is equipped with a generative AI model that periodically analyzes the stored text data and emotional information. Using the generative model (e.g., OpenAI GPT-3), it extracts important keywords and sentences from the text data and summarizes the deceased's memories and messages along with emotional information.
[1102] Examples:
[1103] The generative AI model analyzes the stored text data and emotional information to generate summaries such as "expressing gratitude on the wedding anniversary."
[1104] 4. User Interface
[1105] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server provides information organized and summarized by the generative model through a user interface. Adding emotional information allows for deeper dialogue.
[1106] Examples:
[1107] When a family member accesses the system from their computer and types, "I'd like to know about your memories of our wedding anniversary," the server searches for relevant summary data and displays information such as, "Today is my wedding anniversary with my wife. I'm truly grateful to her."
[1108] 5. Letter-style output
[1109] Upon request of the bereaved family, the information can be output in letter format.
[1110] Examples:
[1111] If a family member requests a letter about their wedding anniversary, the generative AI model will generate a letter such as, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[1112] 6. Support for creating legally binding wills
[1113] When a user requests the server for assistance in creating a legally binding will, the server collects the necessary information and directs the user to a lawyer or notary public.
[1114] Examples:
[1115] When a user requests the creation of a legally binding will, the server collects the appropriate information and guides them through the process and contacting the appropriate professional.
[1116] This system makes it possible to organize and summarize the feelings of the deceased along with emotional information, and to convey them in depth to the bereaved. It can also output them in letter format and assist in the creation of legally valid wills, making it easier for the bereaved to understand and respond to the situation.
[1117] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1118] Step 1:
[1119] The user speaks to the smart speaker.
[1120] Input: User's voice message (e.g., "Today is my wedding anniversary with my wife. I'm so grateful to her.")
[1121] How it works: Smart speakers use built-in microphones to collect voice data.
[1122] Output: Collected audio data
[1123] Step 2:
[1124] Smart speakers convert voice data into text data and recognize emotions.
[1125] Input: Audio data
[1126] How it works: A voice recognition engine (e.g., Google Speech-to-Text) built into a smart speaker converts voice data into text data. At the same time, an emotion engine analyzes the text data and generates emotion information (e.g., "thank you").
[1127] Output: Text data (e.g., "Today is my wedding anniversary with my wife. I am so grateful to her."), emotional information (e.g., "gratitude")
[1128] Step 3:
[1129] The smart speaker sends text data and emotional information to the server.
[1130] Input: Text data, emotion information
[1131] How it works: Smart speakers send data over the internet to a server, which often encrypts the data before sending it.
[1132] Output: Text data and emotion information sent to the server
[1133] Step 4:
[1134] The server stores the text data and emotion information in a database.
[1135] Input: Text data, emotion information
[1136] Operation: The server analyzes the received data and stores the text data, emotion information, collection date and time, user identification information, etc. in a relational database.
[1137] Output: Information stored in the database
[1138] Step 5:
[1139] A generative AI model implemented on the server analyzes text data and emotional information to generate a summary.
[1140] Input: Text data and emotion information stored in a database
[1141] How it works: A generative AI model (e.g., OpenAI GPT-3) periodically scans the database, analyzes the text data and sentiment information, extracts important keywords and sentences, and generates summaries taking sentiment information into account.
[1142] Output: Summarized information (e.g., "Expressing gratitude for our wedding anniversary.")
[1143] Step 6:
[1144] Family members access the system through a terminal and request information.
[1145] Input: Request from the bereaved family (e.g., "I'd like to know about your memories of your wedding anniversary.")
[1146] How it works: Family members access the system using a computer or smartphone and enter information.
[1147] Output: Input information of bereaved family members
[1148] Step 7:
[1149] The server provides the summarized information through a user interface.
[1150] Input: Family request information
[1151] How it works: The server uses the generative model to search for relevant summary data, adds emotional information, and displays it on the user interface.
[1152] Output: Summary information displayed in a user interface (e.g., "Today is my wife's wedding anniversary. I'm so grateful to her.")
[1153] Step 8:
[1154] The family requests a letter-style output.
[1155] Input: A request for output in the form of a letter (e.g., "Please write me a letter about my wedding anniversary.")
[1156] How it works: The family requests a letter-style output from the system.
[1157] Output: Request information
[1158] Step 9:
[1159] The server uses a generative AI model to generate a letter-style document.
[1160] Input: Request information for output in letter format, database information
[1161] How it works: The generative AI model generates letter-style text based on summarized text data and emotional information.
[1162] Output: A letter-style sentence (e.g., "Dear Wife, Every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life.")
[1163] Step 10:
[1164] The server collects information to support the creation of a legally binding will and provides guidance to experts.
[1165] Input: User request information
[1166] How it works: Based on the user's request, the server gathers the necessary information from the database and directs them to the appropriate lawyer or notary public.
[1167] Output: Guidance information (e.g., contact instructions for a lawyer or notary public, list of required documents)
[1168] (Application example 2)
[1169] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1170] When dealing with customers in brick-and-mortar stores, it is difficult for employees to properly understand customer emotions and provide optimal service based on that. A particular challenge is the lack of a system that can recognize and reflect customer emotions and requests in real time. Therefore, there is a need for a system that can improve customer satisfaction and the quality of employee service.
[1171] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for using an emotion engine to analyze and summarize the saved text data and emotion information, and means for providing information summarized by the generative model and emotion engine in real time via an audiovisual device used by a customer service representative. This makes it possible to properly recognize customer emotions when serving customers in a physical store and provide services based on those emotions quickly and accurately.
[1172] Definitions of important words
[1173] "Voice data" refers to digitized data of the voice uttered by the user.
[1174] "Text data" refers to data obtained by converting voice data into character information.
[1175] A "database" is a digital storage device for systematically storing text data and emotional information.
[1176] A "generative model" is an algorithm or software that uses artificial intelligence to extract and summarize important keywords and sentences from text data.
[1177] "Emotion information" is information that represents the emotional state of a user, as recognized from voice data or text data.
[1178] An "emotion engine" is software that analyzes a user's voice data and text data to extract emotional information.
[1179] "Audiovisual devices" is a general term for devices that provide visual and auditory information, such as smart glasses and head-mounted displays.
[1180] "Real-time" refers to processing and response occurring with almost no delay after an event occurs.
[1181] "Customer service personnel" refers to employees and staff who interact directly with customers and provide services in physical stores.
[1182] MODE FOR CARRYING OUT THE INVENTION
[1183] This invention relates to a system for supporting customer service in brick-and-mortar stores, and aims to increase customer satisfaction by utilizing emotion recognition technology in particular. This system uses audiovisual devices such as smart glasses to enable store staff to grasp customer emotions in real time and provide optimal service.
[1184] System configuration
[1185] 1. Collection of audio data
[1186] Customer service representatives will wear smart glasses or head-mounted displays to communicate with customers. The smart glasses have built-in microphones that collect customer voice data in real time.
[1187] 2. Converting audio data to text data
[1188] The collected voice data is instantly converted into text data using the device's voice recognition engine (e.g., Google Speech-to-Text API), which is temporarily stored on the device and then sent to a server.
[1189] 3. Extraction of Emotional Information
[1190] The server analyzes the received text data and extracts customer emotional information using an emotion engine (e.g., EmotionRecognition software). This emotional information is then stored in a database along with the text data.
[1191] 4. Data Analysis and Summarization
[1192] A generative AI model implemented on the server analyzes the stored text data and sentiment information, extracts important keywords and sentences, and summarizes them.
[1193] 5. Real-time information provision
[1194] The server then provides the customer service representative with real-time information summarized by the generative model via an audiovisual device, and the employee's smart glasses display the customer's emotional state and appropriate response.
[1195] Specific examples
[1196] Program processing
[1197] 1. A customer in the store says, "I've been feeling stressed lately and I want a product to help me relax."
[1198] 2. Voice is collected by the microphone in the smart glasses and instantly converted into text data.
[1199] 3. On the server side, the text data is analyzed by an emotion engine, and the emotion "stress" is extracted.
[1200] 4. The generative model analyzes the text data and generates a summary: "We suggest a relaxing aroma set."
[1201] 5. This generated summary information is displayed on an audiovisual device, and the employee can recommend a product to the customer by saying, "We suggest an aroma set that will have a relaxing effect."
[1202] Prompt Sentence Examples
[1203] "A customer might say, 'I've been feeling stressed lately and I want a product to help me relax.' Analyze their emotions with an emotion recognition engine and generate a dialogue that suggests relevant products."
[1204] This system makes it possible to understand customer sentiment in real time and quickly make suggestions based on that sentiment, thereby significantly improving customer satisfaction.
[1205] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1206] Program processing steps
[1207] Step 1:
[1208] Audio data collection
[1209] A user (customer service representative) wears smart glasses and interacts with customers in a physical store. When a customer speaks, the microphone built into the smart glasses collects the voice.
[1210] Input: Customer voice.
[1211] Output: Digitized audio data.
[1212] What it does: The microphone in the smart glasses activates and records what the customer says in real time.
[1213] Step 2:
[1214] Converting audio data to text data
[1215] The device (smart glasses) converts the collected voice data into text data using an internal voice recognition engine, which uses the Google Speech-to-Text API.
[1216] Input: Digitized audio data.
[1217] Output: Text data.
[1218] How it works: The voice recognition engine processes the voice data and converts it into corresponding text data, which is then temporarily stored in the smart glasses.
[1219] Step 3:
[1220] Sending text data
[1221] The terminal (smart glasses) sends the converted text data to a server via the Internet.
[1222] Input: Text data.
[1223] Output: The text data sent to the server.
[1224] How it works: The smart glasses use Wi-Fi or mobile networks to upload text data to a server.
[1225] Step 4:
[1226] Extracting Emotional Information
[1227] The server analyzes the received text data using an emotion engine (e.g., EmotionRecognition software) to extract customer emotional information.
[1228] Input: The text data sent to the server.
[1229] Output: Emotion information (e.g., "joy", "sad", "stress", etc.).
[1230] What it does: The emotion engine on the server scans the text data and identifies relevant emotions.
[1231] Step 5:
[1232] Data storage
[1233] The server stores the converted text data and the extracted emotion information in a database.
[1234] Input: Text data, emotion information.
[1235] Output: Information stored in a database.
[1236] Specific operation: The server's database management system stores text data and emotional information in an orderly manner.
[1237] Step 6:
[1238] Data analysis and summary
[1239] A generative AI model (e.g., OpenAI GPT-3) implemented on the server analyzes the stored text data and emotional information, extracts important keywords and sentences, and generates a summary.
[1240] Input: Text data and emotion information stored in the database.
[1241] Output: Summarized information.
[1242] How it works: The generative AI model analyzes the information in the database, extracts the most important parts, and generates a summary.
[1243] Step 7:
[1244] Real-time provision
[1245] The server provides information summarized by the generative model to the audiovisual device (smart glasses) in real time.
[1246] Input: Summarized information.
[1247] Output: Summary information displayed or spoken on the smart glasses.
[1248] How it works: The server uses Wi-Fi or mobile networks to send the summarized text to the smart glasses, which then displays it on the customer service representative's display or notifies them via voice.
[1249] Step 8:
[1250] Service delivery based on feedback
[1251] Customer service representatives will suggest the most suitable services and products to customers based on the summary information and suggestions displayed on the smart glasses.
[1252] Input: Summary information displayed on smart glasses.
[1253] Output: Services and offers to customers.
[1254] Specific action: The customer service representative checks the summary information and makes appropriate suggestions or asks questions to the customer.
[1255] By implementing the above steps, it is possible to provide services that are sensitive to the customer's emotions when dealing with customers in physical stores.
[1256] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1257] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1258] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1259] [Fourth embodiment]
[1260] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1261] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1262] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1263] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1264] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1265] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1266] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1267] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1268] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1269] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1270] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1271] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1272] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1273] The present invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and a specific embodiment thereof is described below. This system includes a process in which a user uses a smart speaker to collect voice data, converts the data into text data, organizes and summarizes it, and provides it to the bereaved.
[1274] System Configuration
[1275] 1. Smart Speaker
[1276] This system collects voice data by having the user (deceased person) speak into a smart speaker, which is equipped with a voice recognition engine that instantly converts the voice data into text data.
[1277] 2. Server
[1278] The text data is sent over the Internet to a server, which then stores the received text data in a database. This database contains not only the text data, but also the date and time the data was collected and the user's identification information.
[1279] 3. Generative Model
[1280] The server is equipped with a generative model using artificial intelligence that periodically analyzes the stored text data, extracts important keywords and sentences from the text data, and summarizes the memories and messages of the deceased.
[1281] 4. User Interface
[1282] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server provides information organized and summarized by the generative model in an interactive format through a user interface. It is also possible to output the information in letter format upon request from family members.
[1283] 5. Notary support and expert guidance
[1284] Support for creating a legally valid will is also provided. When a user makes such a request to the server, the server collects the necessary information and directs the user to a lawyer or notary public, thereby simplifying the process of creating a legally valid will.
[1285] Specific examples
[1286] Example 1: Collecting and summarizing a conversation
[1287] User: Speaks into smart speaker: "Today is my wife's wedding anniversary. I'm so grateful to her."
[1288] Server: Receives voice data and the voice recognition engine converts it into text data.
[1289] Generative model: Analyzes and summarizes text data. "Expressing gratitude on wedding anniversary."
[1290] Example 2: Providing information through a user interface
[1291] Survivors: Access the system from a computer and type in "I'd like to know about your memories of our wedding anniversary."
[1292] Server: Uses the generative model to find relevant summary data and displays it through a user interface. "Today is my wife's wedding anniversary. I'm so grateful to her."
[1293] Example 3: Letter-style output
[1294] Bereaved family: Request a letter-style output. "Please create a letter about our wedding anniversary."
[1295] Server: Generates a letter-style sentence using a generative model. "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[1296] In this way, a system is realized that can easily and reliably convey the thoughts of the deceased to the bereaved family.
[1297] The processing flow will be explained below.
[1298] Step 1:
[1299] The user speaks to the smart speaker.
[1300] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1301] Smart speakers collect user voice in real time.
[1302] Step 2:
[1303] The smart speaker converts the voice data into text data.
[1304] The speech recognition engine analyzes the voice data and converts it into text.
[1305] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1306] Step 3:
[1307] The smart speaker sends the converted text data to the server.
[1308] The text data is sent to a server via the Internet.
[1309] Step 4:
[1310] The server receives the text data and stores it in a database.
[1311] The received text data is assigned a timestamp and user identification information and registered in a database.
[1312] Example: { "timestamp": "2023-10-12T14:30:00Z", "user": "Deceased Person A", "message": "Today is my wedding anniversary with my wife. I am truly grateful to her."}
[1313] Step 5:
[1314] A generative model in the server analyzes and summarizes the text data in the database.
[1315] A generative model extracts important keywords and sentences from text data and creates a summary.
[1316] Example: { "topic": "Wedding anniversary", "content": "Expressing gratitude for his wife on their wedding anniversary."}
[1317] Step 6:
[1318] The user (surviving family member) accesses the system from a terminal and enters a request.
[1319] Example: "I'd like to know your memories of your wedding anniversary."
[1320] The request is sent over the Internet to a server.
[1321] Step 7:
[1322] The server retrieves the summarized data from the generative model and provides it through a user interface.
[1323] Matching data is retrieved and displayed on the user interface.
[1324] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1325] Step 8:
[1326] The user requests output in letter format.
[1327] Example: "Please write me a letter about our wedding anniversary."
[1328] The request is sent over the Internet to a server.
[1329] Step 9:
[1330] The server's generative model generates a letter-style document based on the request.
[1331] A letter-style text is created based on past text data and sent to the device.
[1332] For example: "Dear wife, every time our wedding anniversary comes around, my heart fills with gratitude for you. You are the treasure of my life."
[1333] The above is the specific processing flow of the "Conclusion" system.
[1334] Example 1
[1335] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1336] In modern society, there are limited ways to accurately convey the wishes of the deceased to their bereaved families, and there is a particular problem of a lack of an efficient system for the process from collecting to transmitting information using voice. Furthermore, the process of creating a legally valid will is complicated, and there is a lack of support methods to simplify the process.
[1337] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1338] In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing and summarizing the stored text data using a generative model, means for providing information summarized by the generative model through a user interface, means for searching for relevant summary data based on a request entered by a user, means for generating the summary data in letter format, and means for collecting information to support the creation of a legally binding will and guiding experts. This allows the wishes of the deceased to be effectively conveyed to the surviving family and simplifies legal procedures.
[1339] "User" means an individual who uses the System to provide voice data and receive subsequent services.
[1340] "Voice data" refers to a digital recording of a voice signal emitted by a user through a device such as a smart speaker.
[1341] "Text data" refers to digital data that has been converted from voice data into text information.
[1342] "Database" refers to an information management system for storing text data and its associated metadata (date and time, identification information, etc.).
[1343] A "generative model" refers to a software model that uses artificial intelligence to analyze text data, extract important keywords and sentences, and summarize them.
[1344] "User interface" refers to the interface through which users and bereaved families access the system and input or retrieve data.
[1345] A "request" is an instruction from a user or family member to the system requesting the provision of specific data or services.
[1346] "Summary data" is shortened text data that contains key points or sentences extracted by a generative model.
[1347] "Letter format" refers to a document in which summary data has been reconstructed into a letter format that is easy for humans to read.
[1348] "Experts" refers to professionals such as lawyers and notaries who have knowledge and qualifications regarding legal procedures and the preparation of wills.
[1349] A "legally valid will" is a will that has legal effect and is officially recognized under the law.
[1350] This invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and a specific embodiment thereof is described below. This system involves a process in which a user collects voice data, converts it into text data, summarizes and organizes it using a generative AI model, and provides it to the bereaved.
[1351] System Configuration
[1352] 1. Collection of audio data
[1353] Users collect voice data by speaking into a voice input device such as a smart speaker, which then converts the voice data into text data in real time using a voice recognition system such as Amazon Alexa or Google Assistant.
[1354] 2. Sending and saving text data
[1355] The text data generated by the voice recognition system is sent over the Internet to a server. The server stores the received text data and its associated metadata (date and time, user identification information, etc.) in a database. This database serves to store and manage the memories and messages of the deceased.
[1356] 3. Analysis and Summarization of Text Data
[1357] The server is equipped with a generative AI model (e.g., OpenAI's GPT-3) that periodically analyzes the text data in the database. The generative model extracts important keywords and sentences from the text data and summarizes and organizes the memories and messages of the deceased.
[1358] 4. Providing information through the user interface
[1359] Family members access the system through devices such as computers or smartphones. Using the user interface, they input specific requests, such as "I want to know about memories of our wedding anniversary." The server uses the generative AI model to search for the relevant summary data and displays it through the user interface.
[1360] An example of a specific prompt is "Show me a message about your wedding anniversary memories."
[1361] 5. Generating Letter Format
[1362] If the family requests a letter-style output, the server sends a specific prompt to the generative AI model, which generates a letter with a message like, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[1363] 6. Legal process support
[1364] When a user wishes to create a legally valid will, the server collects the necessary information and directs them to professionals such as lawyers and notaries. Also, when the bereaved family requests the creation of a will through the system, the server organizes and provides the necessary information, helping to simplify the legal procedures.
[1365] In this way, this system can easily and reliably convey the wishes of the deceased to their family members.It is also extremely convenient because it has a function to assist users in legal procedures.
[1366] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1367] Step 1:
[1368] The user speaks to the smart speaker. For example, "Today is my wedding anniversary with my wife. I'm really grateful to her." This voice is input into the smart speaker.
[1369] Step 2:
[1370] Smart speakers use a built-in voice recognition engine (such as Google Assistant) to convert the received voice data into text data in real time, and this conversion process outputs the voice data as text data.
[1371] Step 3:
[1372] The text data is sent to a server via the Internet. The data sent includes not only the text data but also the date and time the voice was collected and the user's identification information. The server receives this data.
[1373] Step 4:
[1374] The server saves the received text data and its associated metadata in a database. The saving process stores the text data and metadata in the database.
[1375] Step 5:
[1376] The server passes the text data stored in the database to a generative AI model (for example, OpenAI's GPT-3) for analysis. The generative model reads the text data, extracts important keywords and sentences, and generates summary data. This summary data is the output from the generative model.
[1377] Step 6:
[1378] The bereaved family accesses the system from a PC or smartphone and inputs a specific request into the user interface (e.g., "I would like to know about memories of our wedding anniversary.") This request is sent to the server.
[1379] Step 7:
[1380] The server receives the request from the family and uses the generative AI model to search for the relevant summary data. The search process is carried out and the relevant summary data is output.
[1381] Step 8:
[1382] The server displays the retrieved summary data through a user interface, allowing the family members to view the information.
[1383] Step 9:
[1384] If the family requests a letter-style output, the server sends a specific prompt to the generative AI model, which generates a letter-style text, such as, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[1385] Step 10:
[1386] If a user wishes to create a legally valid will, the server collects the necessary information and directs them to professionals such as lawyers and notaries, helping them to navigate the legal process more easily.
[1387] Through the above processing steps, the system can accurately convey the wishes of the deceased to the bereaved family and can also provide support in legal procedures.
[1388] (Application example 1)
[1389] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1390] In conventional estate sorting, it has been difficult to accurately convey the deceased's memories and messages to the bereaved. Also, the task of organizing and summarizing the deceased's thoughts is time-consuming and laborious, placing a heavy burden on the bereaved. Furthermore, the means to easily retrieve and access the deceased's thoughts were limited, making it difficult for the bereaved to reminisce. The present invention aims to solve these problems.
[1391] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1392] In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for storing the text data in a database, means for analyzing and summarizing the stored text data using a generative model, means for providing information summarized by the generative model through a user interface, and means for delivering the summarized information to a smartphone or head-mounted display. This allows memories and messages of the deceased to be organized and easily conveyed to the bereaved. Furthermore, the bereaved can access memories of the deceased at any time using a smartphone or head-mounted display, thereby reducing the burden on the bereaved.
[1393] "User" refers to an individual who provides voice data using this system.
[1394] "Voice Data" refers to recordings of voice signals collected when a user speaks.
[1395] "Text data" refers to data that has been converted from voice data into text information using voice recognition technology.
[1396] "Database" refers to a digital storage system for storing collected text data.
[1397] A "generative model" refers to an algorithm that uses artificial intelligence techniques to analyze and summarize text data.
[1398] "User interface" refers to the interface on a computer system that allows users to view information and perform operations.
[1399] A "smartphone" refers to a multi-function mobile phone that can make calls and connect to the Internet.
[1400] "Head-mounted display" refers to a display device worn on the head.
[1401] "Distribution" refers to the act of transmitting or providing information via the Internet or other means of communication.
[1402] The present invention is a system for accurately conveying the memories and messages of the deceased to the bereaved, and specific embodiments thereof are described below. This system covers the collection of voice data, storing it in a database, summarizing the text data using a generative model, and distributing the information via smartphones and head-mounted displays.
[1403] First, the user speaks to the smart speaker to collect voice data. The smart speaker is equipped with a speech recognition engine that converts the voice data into text data in real time. This speech recognition engine can be Google's speech recognition API or IBM Watson.
[1404] The converted text data is sent to a server via the Internet and stored in a database on the server side, which may use a storage system such as SQL Server or MongoDB.
[1405] The server implements a generative model to periodically analyze and summarize text data. This generative model, for example, the "BART" model based on the Hugging Face "transformers" library, is used. The generative model analyzes the stored text data, extracts important keywords and sentences, and creates a summary. The summarized information is then stored on the server.
[1406] Family members can access the system through a user interface using a smartphone or head-mounted display. The user interface provides information summarized by the generative model. Android and iOS app development frameworks are used as the user interface technology. This system allows family members to access the memories and messages of the deceased at any time.
[1407] The detailed operation of the system will be explained below by showing a specific example.
[1408] Examples:
[1409] If a user wants to collect memories of the deceased using the following prompt, they can speak to the smart speaker, saying, "Tell me your memories of Mother's Day." The smart speaker converts the speech into text and sends it to the server. The server summarizes the text and generates a summary such as, "I always sent flowers on Mother's Day. My mother was always very happy." The bereaved can view this summarized message at any time via their smartphone or head-mounted display.
[1410] This makes it possible to efficiently organize memories of the deceased and properly convey them to the bereaved family.
[1411] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1412] Step 1:
[1413] The user speaks to the smart speaker to provide voice data. The smart speaker receives the user's voice as input, and the built-in voice recognition engine converts the voice data into text data in real time. As a result, the input (voice data) acquired by the smart speaker is output as text data.
[1414] Step 2:
[1415] The smart speaker sends the converted text data over the Internet to a server. The server receives this text data as input and stores it in a database in an appropriate format. The database uses a cloud platform such as Azure SQL Database or Amazon RDS. The text data is stored in the database along with the date, time, and user identification information.
[1416] Step 3:
[1417] The server periodically accesses the database and analyzes the stored text data. At this time, it uses a generative AI model (for example, the BART model implemented using Hugging Face's transformers library) to summarize the text data. The server takes the text data as input and uses the generative model to extract important keywords and sentences. As a result of the analysis, the summarized text is output and stored again in the database.
[1418] Step 4:
[1419] The bereaved family members access the user interface via a smartphone or head-mounted display. The user interface requests information from the server using a specific prompt (e.g., "Tell me your memories of Mother's Day"). The server receives this prompt, searches for relevant summary data from a database, and responds to the user interface. The user interface receives the user's request as input and displays the summarized information as output.
[1420] Step 5:
[1421] The user checks the summarized information through a smartphone or head-mounted display. Specifically, by checking the summary message displayed on the user interface (e.g., "I always sent flowers to my mother on Mother's Day. She was always very happy every time"), the user can reminisce about the deceased. This allows the user to access the memories and messages of the deceased at any time.
[1422] This series of processes creates a system that efficiently organizes the memories and messages of the deceased and conveys them appropriately to the bereaved.
[1423] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1424] This invention relates to a system for accurately conveying the thoughts of the deceased to the bereaved, and further describes a form that combines an emotion engine that recognizes the user's emotions. This system involves a process in which the user uses a smart speaker to collect voice data, converts that data into text data, organizes and summarizes it, and provides it to the bereaved. In addition, the emotion engine recognizes the user's emotions and provides a more in-depth summary or letter-style output based on that information.
[1425] System Configuration
[1426] 1. Smart Speaker
[1427] This system collects voice data when the user (deceased person) speaks to a smart speaker. The smart speaker is equipped with a voice recognition engine that instantly converts the voice data into text data. An emotion engine is also built into the smart speaker, which recognizes emotions from the user's voice in real time.
[1428] 2. Server
[1429] The text data and emotion information are sent to a server via the Internet. The server receives the data and stores it in a database. The database stores the text data, emotion information, collection date and time, and user identification information.
[1430] 3. Generative Model
[1431] The server is equipped with a generative model using artificial intelligence that periodically analyzes the stored text data and emotional information. The generative model extracts important keywords and sentences from the text data and summarizes the memories and messages of the deceased along with emotional information.
[1432] 4. User Interface
[1433] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server then provides information organized and summarized by the generative model in an interactive format through a user interface. Adding emotional information allows for a more in-depth dialogue. It is also possible to output the information in the form of a letter upon request from family members.
[1434] 5. Notary support and expert guidance
[1435] Support for creating a legally valid will is also provided. When a user makes such a request to the server, the server collects the necessary information and directs the user to a lawyer or notary public, thereby simplifying the process of creating a legally valid will.
[1436] Specific examples
[1437] Example 1: Collecting and summarizing a conversation
[1438] User: Speaks into smart speaker: "Today is my wife's wedding anniversary. I'm so grateful to her."
[1439] Smart speaker: Recognizes voice and the emotion engine identifies emotions such as "gratitude."
[1440] Server: Converts the voice data into text data and emotional information and stores them in a database.
[1441] Generative model: Analyzes and summarizes text data and emotional information. "Expressing gratitude on wedding anniversary."
[1442] Example 2: Providing information through a user interface
[1443] Survivors: Access the system from a computer and type in "I'd like to know about your memories of our wedding anniversary."
[1444] Server: Uses a generative model to search for relevant summary data, adds emotional information, and displays it on the user interface. "Today is my wife's wedding anniversary. I'm really grateful to her."
[1445] Example 3: Letter-style output
[1446] Bereaved family: Request a letter-style output. "Please create a letter about our wedding anniversary."
[1447] Server: Generates letter-style text using a generative model and adds emotional information. "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[1448] In this way, by combining the emotion engine, a system is realized that can easily and deeply convey the feelings of the deceased to the bereaved family.
[1449] The processing flow will be explained below.
[1450] Step 1:
[1451] The user speaks to the smart speaker.
[1452] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1453] Smart speakers collect user voice in real time.
[1454] Step 2:
[1455] The smart speaker converts the voice data into text data.
[1456] The speech recognition engine analyzes the voice data and converts it into text.
[1457] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1458] Step 3:
[1459] The smart speaker uses an emotion engine to recognize emotions from voice data.
[1460] The emotion engine analyzes the voice data and identifies emotions such as "gratitude" and "joy."
[1461] Example: Recognizing the emotion of "gratitude."
[1462] Step 4:
[1463] The smart speaker sends the converted text data and emotional information to the server.
[1464] The text data and emotion information are sent to a server via the Internet.
[1465] Step 5:
[1466] The server receives the text data and emotion information and stores them in a database.
[1467] The received text data and emotion information are assigned a timestamp and user identification information and registered in a database.
[1468] Example: { "timestamp": "2023-10-12T14:30:00Z", "user": "Deceased Person A", "message": "Today is my wedding anniversary with my wife. I am truly grateful to her.", "emotion": "Thank you"}
[1469] Step 6:
[1470] A generative model in the server analyzes and summarizes the text data and sentiment information in the database.
[1471] The generative model extracts important keywords and sentences from text data and emotional information, and creates summaries that reflect the emotional information.
[1472] Example: { "topic": "Wedding anniversary", "content": "Expressing gratitude for his wife on their wedding anniversary."}
[1473] Step 7:
[1474] The user (surviving family member) accesses the system from a terminal and enters a request.
[1475] Example: "I'd like to know your memories of your wedding anniversary."
[1476] The request is sent over the Internet to a server.
[1477] Step 8:
[1478] The server retrieves summary data based on the generative model, adds emotional information, and provides it through a user interface.
[1479] Matching data is searched for and displayed on the user interface based on emotion information.
[1480] For example: "Today is my wife's wedding anniversary. I'm so grateful for her."
[1481] Step 9:
[1482] The user requests output in letter format.
[1483] Example: "Please write me a letter about our wedding anniversary."
[1484] The request is sent over the Internet to a server.
[1485] Step 10:
[1486] The server's generative model generates a letter-style document based on the request.
[1487] A letter-style text is created based on past text data and emotional information and sent to the device.
[1488] For example: "Dear wife, every time our wedding anniversary comes around, my heart fills with gratitude for you. You are the treasure of my life."
[1489] The above is the specific processing flow that combines the emotion engine in the "Conclusion" system.
[1490] Example 2
[1491] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1492] There is a problem in accurately conveying the deceased's thoughts to the bereaved. Conventional methods do not accurately convey the feelings and intentions of the deceased, and do not provide deep meaning to the bereaved. In addition, the process of recording and organizing the deceased's thoughts and creating a legally valid will, if necessary, is complicated and time-consuming.
[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1494] In this invention, the server includes means for collecting voice data from a user, means for converting the voice data into text data, means for recognizing the user's emotions from the voice data, means for storing the text data and emotional information in a database, means for analyzing and summarizing the stored text data and emotional information using a generative model, and means for providing information summarized by the generative model through a user interface. This allows the deceased's thoughts, along with their emotional information, to be organized and summarized and conveyed to the bereaved. It can also output the message in letter format and support the creation of a legally valid will, allowing for a deeper understanding of the bereaved and legal response.
[1495] "Voice data" refers to voice signal information collected when a user speaks to a smart speaker or other voice input device.
[1496] "Text data" is character information converted from voice data by a voice recognition engine.
[1497] "Emotion information" is data that is analyzed from voice data by the emotion engine and indicates the user's emotional state.
[1498] A "database" is a collection of information for systematically storing and managing text data and emotional information.
[1499] A "generative model" is a program that uses artificial intelligence to analyze text data and emotional information, extract important keywords and sentences, and generate summaries.
[1500] A "user interface" refers to the screen or operating means that allows bereaved family members to access the system through devices such as computers or smartphones and obtain the necessary information.
[1501] "Letter-format output" is a function that provides information summarized by the generative model in the form of a letter.
[1502] A "legally valid will" is a written document that expresses the final wishes of a deceased person and is legally valid.
[1503] "Expert Guidance" is a support function that encourages users who wish to write a will to contact or consult with legal experts (such as lawyers or notaries).
[1504] The system of the present invention aims to accurately convey the thoughts of the deceased to the bereaved family, and is realized by combining an emotion engine that recognizes the user's emotions. This system includes the following components.
[1505] 1. Smart Speaker
[1506] The first component of this system is the smart speaker. Users speak into the smart speaker to collect their messages and thoughts as voice data. The smart speaker has a built-in voice recognition engine (e.g., Google Speech-to-Text or Amazon Alexa) that instantly converts the voice data into text data. In addition, the built-in emotion engine recognizes the user's emotions from the collected voice in real time.
[1507] Examples:
[1508] A user says to a smart speaker, "Today is my wife's wedding anniversary. I'm so grateful to her."
[1509] 2. Server
[1510] The text data and emotion information are sent to a server via the Internet. The server receives the data and stores it in a database. The database stores the text data, emotion information, collection date and time, and user identification information.
[1511] Examples:
[1512] The server stores the text data (e.g., "Today is my wedding anniversary with my wife. I am truly grateful to her") and emotional information (gratitude) sent from the smart speaker in a database.
[1513] 3. Generative AI Models
[1514] The server is equipped with a generative AI model that periodically analyzes the stored text data and emotional information. Using the generative model (e.g., OpenAI GPT-3), it extracts important keywords and sentences from the text data and summarizes the deceased's memories and messages along with emotional information.
[1515] Examples:
[1516] The generative AI model analyzes the stored text data and emotional information to generate summaries such as "expressing gratitude on the wedding anniversary."
[1517] 4. User Interface
[1518] Family members can access the system through devices such as PCs or smartphones and request information about the deceased's memories and messages. The server provides information organized and summarized by the generative model through a user interface. Adding emotional information allows for a more in-depth dialogue.
[1519] Examples:
[1520] When a family member accesses the system from a computer and types, "I'd like to know about your memories of our wedding anniversary," the server searches for relevant summary data and displays information such as, "Today is my wedding anniversary with my wife. I'm truly grateful to her."
[1521] 5. Letter-style output
[1522] Upon request of the bereaved family, the information can be output in letter format.
[1523] Examples:
[1524] If a family member requests a letter about their wedding anniversary, the generative AI model will generate a letter such as, "Dear wife, every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life."
[1525] 6. Support for creating legally binding wills
[1526] When a user requests the server for assistance in creating a legally binding will, the server collects the necessary information and directs the user to a lawyer or notary public.
[1527] Examples:
[1528] When a user requests the creation of a legally binding will, the server collects the appropriate information and guides them through the process and contacting the appropriate professional.
[1529] This system makes it possible to organize and summarize the feelings of the deceased along with emotional information, and to convey them in depth to the bereaved. It can also output them in letter format and assist in the creation of legally valid wills, making it easier for the bereaved to understand and respond to the situation.
[1530] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1531] Step 1:
[1532] The user speaks to the smart speaker.
[1533] Input: User's voice message (e.g., "Today is my wedding anniversary with my wife. I'm so grateful to her.")
[1534] How it works: Smart speakers use built-in microphones to collect voice data.
[1535] Output: Collected audio data
[1536] Step 2:
[1537] Smart speakers convert voice data into text data and recognize emotions.
[1538] Input: Audio data
[1539] How it works: A voice recognition engine (e.g., Google Speech-to-Text) built into a smart speaker converts voice data into text data. At the same time, an emotion engine analyzes the text data and generates emotion information (e.g., "thank you").
[1540] Output: Text data (e.g., "Today is my wedding anniversary with my wife. I am so grateful to her."), emotional information (e.g., "gratitude")
[1541] Step 3:
[1542] The smart speaker sends text data and emotional information to the server.
[1543] Input: Text data, emotion information
[1544] How it works: Smart speakers send data over the internet to a server, which often encrypts the data before sending it.
[1545] Output: Text data and emotion information sent to the server
[1546] Step 4:
[1547] The server stores the text data and emotion information in a database.
[1548] Input: Text data, emotion information
[1549] Operation: The server analyzes the received data and stores the text data, emotion information, collection date and time, user identification information, etc. in a relational database.
[1550] Output: Information stored in the database
[1551] Step 5:
[1552] A generative AI model implemented on the server analyzes text data and emotional information to generate a summary.
[1553] Input: Text data and emotion information stored in a database
[1554] How it works: A generative AI model (e.g., OpenAI GPT-3) periodically scans the database, analyzes the text data and sentiment information, extracts important keywords and sentences, and generates summaries taking sentiment information into account.
[1555] Output: Summarized information (e.g., "Expressing gratitude for our wedding anniversary.")
[1556] Step 6:
[1557] Family members access the system through a terminal and request information.
[1558] Input: Request from the bereaved family (e.g., "I'd like to know about your memories of your wedding anniversary.")
[1559] How it works: Family members access the system using a computer or smartphone and enter information.
[1560] Output: Input information of bereaved family members
[1561] Step 7:
[1562] The server provides the summarized information through a user interface.
[1563] Input: Family request information
[1564] How it works: The server uses the generative model to search for relevant summary data, adds emotional information, and displays it on the user interface.
[1565] Output: Summary information displayed in a user interface (e.g., "Today is my wife's wedding anniversary. I'm so grateful to her.")
[1566] Step 8:
[1567] The family requests a letter-style output.
[1568] Input: A request for output in the form of a letter (e.g., "Please write me a letter about my wedding anniversary.")
[1569] How it works: The family requests a letter-style output from the system.
[1570] Output: Request information
[1571] Step 9:
[1572] The server uses a generative AI model to generate a letter-style document.
[1573] Input: Request information for output in letter format, database information
[1574] How it works: The generative AI model generates letter-style text based on summarized text data and emotional information.
[1575] Output: A letter-style sentence (e.g., "Dear Wife, Every time our wedding anniversary comes around, my heart is filled with gratitude for you. You are the treasure of my life.")
[1576] Step 10:
[1577] The server collects information to support the creation of a legally binding will and provides guidance to experts.
[1578] Input: User request information
[1579] How it works: Based on the user's request, the server gathers the necessary information from the database and directs them to the appropriate lawyer or notary public.
[1580] Output: Guidance information (e.g., contact instructions for a lawyer or notary public, list of required documents)
[1581] (Application example 2)
[1582] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1583] When dealing with customers in brick-and-mortar stores, it is difficult for employees to properly understand customer emotions and provide optimal service based on that. A particular challenge is the lack of a system that can recognize and reflect customer emotions and requests in real time. Therefore, there is a need for a system that can improve customer satisfaction and the quality of employee service.
[1584] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data from users, means for converting the voice data into text data, means for using an emotion engine to analyze and summarize the saved text data and emotion information, and means for providing information summarized by the generative model and emotion engine in real time via an audiovisual device used by a customer service representative. This makes it possible to properly recognize customer emotions when serving customers in a physical store and provide services based on those emotions quickly and accurately.
[1585] Definitions of important words
[1586] "Voice data" refers to digitized data of the voice uttered by the user.
[1587] "Text data" refers to data obtained by converting voice data into character information.
[1588] A "database" is a digital storage device for systematically storing text data and emotional information.
[1589] A "generative model" is an algorithm or software that uses artificial intelligence to extract and summarize important keywords and sentences from text data.
[1590] "Emotion information" is information that represents the emotional state of a user, as recognized from voice data or text data.
[1591] An "emotion engine" is software that analyzes a user's voice data and text data to extract emotional information.
[1592] "Audiovisual devices" is a general term for devices that provide visual and auditory information, such as smart glasses and head-mounted displays.
[1593] "Real-time" refers to processing and response occurring with almost no delay after an event occurs.
[1594] "Customer service personnel" refers to employees and staff who interact directly with customers and provide services in physical stores.
[1595] MODE FOR CARRYING OUT THE INVENTION
[1596] This invention relates to a system for supporting customer service in brick-and-mortar stores, and aims to increase customer satisfaction by utilizing emotion recognition technology in particular. This system uses audiovisual devices such as smart glasses to enable store staff to grasp customer emotions in real time and provide optimal service.
[1597] System configuration
[1598] 1. Collection of audio data
[1599] Customer service representatives will wear smart glasses or head-mounted displays to communicate with customers. The smart glasses have built-in microphones that collect customer voice data in real time.
[1600] 2. Converting audio data to text data
[1601] The collected voice data is instantly converted into text data using the device's voice recognition engine (e.g., Google Speech-to-Text API), which is temporarily stored on the device and then sent to a server.
[1602] 3. Extraction of Emotional Information
[1603] The server analyzes the received text data and extracts customer emotional information using an emotion engine (e.g., EmotionRecognition software). This emotional information is then stored in a database along with the text data.
[1604] 4. Data Analysis and Summarization
[1605] A generative AI model implemented on the server analyzes the stored text data and sentiment information, extracts important keywords and sentences, and summarizes them.
[1606] 5. Real-time information provision
[1607] The server then provides the customer service representative with real-time information summarized by the generative model via an audiovisual device, and the employee's smart glasses display the customer's emotional state and appropriate response.
[1608] Specific examples
[1609] Program processing
[1610] 1. A customer in the store says, "I've been feeling stressed lately and I want a product to help me relax."
[1611] 2. Voice is collected by the microphone in the smart glasses and instantly converted into text data.
[1612] 3. On the server side, the text data is analyzed by an emotion engine, and the emotion "stress" is extracted.
[1613] 4. The generative model analyzes the text data and generates a summary: "We suggest a relaxing aroma set."
[1614] 5. This generated summary information is displayed on an audiovisual device, and the employee can recommend a product to the customer by saying, "We suggest an aroma set that will have a relaxing effect."
[1615] Prompt Sentence Examples
[1616] "A customer might say, 'I've been feeling stressed lately and I want a product to help me relax.' Analyze their emotions with an emotion recognition engine and generate a dialogue that suggests relevant products."
[1617] This system makes it possible to understand customer sentiment in real time and quickly make suggestions based on that sentiment, thereby significantly improving customer satisfaction.
[1618] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1619] Program processing steps
[1620] Step 1:
[1621] Audio data collection
[1622] A user (customer service representative) wears smart glasses and interacts with customers in a physical store. When a customer speaks, the microphone built into the smart glasses collects the voice.
[1623] Input: Customer voice.
[1624] Output: Digitized audio data.
[1625] What it does: The microphone in the smart glasses activates and records what the customer says in real time.
[1626] Step 2:
[1627] Converting audio data to text data
[1628] The device (smart glasses) converts the collected voice data into text data using an internal voice recognition engine, which uses the Google Speech-to-Text API.
[1629] Input: Digitized audio data.
[1630] Output: Text data.
[1631] How it works: The voice recognition engine processes the voice data and converts it into corresponding text data, which is then temporarily stored in the smart glasses.
[1632] Step 3:
[1633] Sending text data
[1634] The terminal (smart glasses) sends the converted text data to a server via the Internet.
[1635] Input: Text data.
[1636] Output: The text data sent to the server.
[1637] How it works: The smart glasses use Wi-Fi or mobile networks to upload text data to a server.
[1638] Step 4:
[1639] Extracting Emotional Information
[1640] The server analyzes the received text data using an emotion engine (e.g., EmotionRecognition software) to extract customer emotional information.
[1641] Input: The text data sent to the server.
[1642] Output: Emotion information (e.g., "joy", "sad", "stress", etc.).
[1643] What it does: The emotion engine on the server scans the text data and identifies relevant emotions.
[1644] Step 5:
[1645] Data storage
[1646] The server stores the converted text data and the extracted emotion information in a database.
[1647] Input: Text data, emotion information.
[1648] Output: Information stored in a database.
[1649] Specific operation: The server's database management system stores text data and emotional information in an orderly manner.
[1650] Step 6:
[1651] Data analysis and summary
[1652] A generative AI model (e.g., OpenAI GPT-3) implemented on the server analyzes the stored text data and emotional information, extracts important keywords and sentences, and generates a summary.
[1653] Input: Text data and emotion information stored in the database.
[1654] Output: Summarized information.
[1655] How it works: The generative AI model analyzes the information in the database, extracts the most important parts, and generates a summary.
[1656] Step 7:
[1657] Real-time provision
[1658] The server provides information summarized by the generative model to the audiovisual device (smart glasses) in real time.
[1659] Input: Summarized information.
[1660] Output: Summary information displayed or spoken on the smart glasses.
[1661] How it works: The server uses Wi-Fi or mobile networks to send the summarized text to the smart glasses, which then displays it on the customer service representative's display or notifies them via voice.
[1662] Step 8:
[1663] Service delivery based on feedback
[1664] Customer service representatives will suggest the most suitable services and products to customers based on the summary information and suggestions displayed on the smart glasses.
[1665] Input: Summary information displayed on smart glasses.
[1666] Output: Services and offers to customers.
[1667] Specific action: The customer service representative checks the summary information and makes appropriate suggestions or asks questions to the customer.
[1668] By implementing the above steps, it is possible to provide services that are sensitive to the customer's emotions when dealing with customers in physical stores.
[1669] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1670] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1671] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1672] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1673] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1674] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1675] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1676] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1677] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1678] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1679] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1680] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1681] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1682] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1683] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1684] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1685] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1686] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1687] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1688] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1689] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1690] The following is further disclosed regarding the above embodiment.
[1691] (Claim 1)
[1692] A means for collecting voice data from a user;
[1693] means for converting the voice data into text data;
[1694] means for storing the text data in a database;
[1695] means for using a generative model to analyze and summarize the stored text data;
[1696] a means for providing information summarized by the generative model through a user interface;
[1697] A system including:
[1698] (Claim 2)
[1699] 10. The system of claim 1, further comprising means for outputting the summarized information in letter form upon request from the user.
[1700] (Claim 3)
[1701] 10. The system of claim 1, further comprising means for collecting information and providing expert guidance necessary to support the preparation of a legally binding will.
[1702] "Example 1"
[1703] (Claim 1)
[1704] A means for collecting voice data from a user;
[1705] means for converting the voice data into text data;
[1706] means for storing the text data in a database;
[1707] means for using a generative model to analyze and summarize the stored text data;
[1708] a means for providing information summarized by the generative model through a user interface;
[1709] a means for retrieving relevant summary data based on a request input by a user;
[1710] means for generating summary data in letter form;
[1711] A means of gathering information and guiding professionals to assist in the preparation of legally binding wills;
[1712] A system including:
[1713] (Claim 2)
[1714] 10. The system of claim 1, further comprising means for outputting the summarized information in letter form.
[1715] (Claim 3)
[1716] 10. The system of claim 1, further comprising means for collecting information and guiding an expert necessary to support the preparation of a legally binding will.
[1717] "Application Example 1"
[1718] (Claim 1)
[1719] A means for collecting voice data from a user;
[1720] means for converting the voice data into text data;
[1721] means for storing the text data in a database;
[1722] means for using a generative model to analyze and summarize the stored text data;
[1723] a means for providing information summarized by the generative model through a user interface;
[1724] means for delivering the summarized information to a smartphone or a head-mounted display;
[1725] A system including:
[1726] (Claim 2)
[1727] 10. The system of claim 1, further comprising means for outputting the summarized information in letter form upon request from the user.
[1728] (Claim 3)
[1729] 10. The system of claim 1, further comprising means for collecting information and providing expert guidance necessary to support the preparation of a legally binding will.
[1730] "Example 2: Combining Emotion Engines"
[1731] (Claim 1)
[1732] A means for collecting voice data from a user;
[1733] means for converting the voice data into text data;
[1734] means for recognizing a user's emotion from the voice data;
[1735] means for storing the text data and emotion information in a database;
[1736] means for using a generative model to analyze and summarize the stored text data and sentiment information;
[1737] a means for providing information summarized by the generative model through a user interface;
[1738] A system including:
[1739] (Claim 2)
[1740] 10. The system of claim 1, further comprising means for outputting the summarized information in letter form.
[1741] (Claim 3)
[1742] 10. The system of claim 1, further comprising means for collecting information and providing expert guidance necessary to support the preparation of a legally binding will.
[1743] "Application example 2 when combining emotion engines"
[1744] Rewriting of claims
[1745] (Claim 1)
[1746] A means for collecting voice data from a user;
[1747] means for converting the voice data into text data;
[1748] means for storing the text data in a database;
[1749] means for using a generative model to analyze and summarize the stored text data;
[1750] means for using an emotion engine to analyze and summarize the text data and emotion information;
[1751] a means for providing, in real time, the information summarized by said generative model and emotion engine via an audiovisual device used by a customer service representative;
[1752] A system including:
[1753] (Claim 2)
[1754] 10. The system of claim 1, further comprising means for the customer service representative to appropriately provide feedback of the summarized information to the customer using an audiovisual device.
[1755] (Claim 3)
[1756] 10. The system of claim 1, further comprising means for outputting the summarized information in letter form upon request from the user.
[1757] (Claim 4)
[1758] 10. The system of claim 1, further comprising means for collecting information and providing expert guidance necessary to support the preparation of a legally binding will. [Explanation of symbols]
[1759] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for collecting voice data from a user; means for converting the voice data into text data; means for storing the text data in a database; means for using a generative model to analyze and summarize the stored text data; a means for providing information summarized by the generative model through a user interface; A system including:
2. 10. The system of claim 1, further comprising means for outputting the summarized information in letter form upon request from the user.
3. 10. The system of claim 1, further comprising means for gathering information and providing expert guidance necessary to support the preparation of a legally binding will.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A