System

A system for recording and recalling personal memories and emotions by converting voice to text, analyzing sentiment, and managing data for easy retrieval addresses the challenge of inefficient memory recording and playback.

JP2026019836APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121584
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

There is a lack of means to easily record and review personal memories and emotions, with existing systems failing to efficiently convert voice input to text, analyze emotions, and manage data for easy retrieval.

Method used

A system that accepts voice or text input, converts it to text, performs sentiment analysis, stores the data in a database, and allows retrieval based on specific conditions for playback in text or voice.

Benefits of technology

Enables users to efficiently record and easily recall past memories and emotions, enhancing personal reflection and family bonding through effective data management and retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019836000001_ABST
    Figure 2026019836000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving an input by voice or sentence; means for converting voice data into text data; means for performing an emotion analysis on the input text data; means for storing the text data subjected to the emotion analysis in a database; means for acquiring the stored data based on a specific condition; and means for reproducing the acquired data by text or voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern life, there is a lack of means to record past memories and emotions and review them at a later date. In particular, there is no concrete and easy way to look back on personal emotions and memories, and many people forget important memories. Furthermore, there is a need for a system that allows people to easily recall past experiences and emotions in order to record personal growth and deepen family bonds. [Means for solving the problem]

[0005] The present invention provides a system including means for accepting voice or text input, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for storing the sentiment-analyzed text data in a database, means for retrieving the stored data based on specific conditions, and means for playing back the retrieved data in text or voice, thereby enabling users to efficiently record their past memories and emotions and easily look back on them in the future.

[0006] "Means for accepting voice or text input" refers to the functionality of an interface device or software that allows a user to input memories or emotions into the system in voice or text form.

[0007] The "means for converting voice data into text data" refers to a voice recognition engine or software function for converting data input by voice by the user into text format.

[0008] The "means for performing emotion analysis on input text data" refers to the function of an analysis engine or algorithm for analyzing the emotions contained in the input text data and tagging those emotions.

[0009] The "means for storing emotion-analyzed text data in a database" refers to software or system functionality for storing emotion-analyzed text data in a database together with date, time, and emotion tags.

[0010] "Means for retrieving stored data based on specific conditions" refers to the functionality of a search engine or algorithm for searching and retrieving appropriate data from a database based on a user request.

[0011] "Means for reproducing the captured data as text or audio" refers to software or system functionality for displaying the captured text data to the user or reproducing it as audio using a speech synthesis engine. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0020] [First embodiment]

[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0033] The present invention is a system that allows users to record their memories and emotions and easily look back on them later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when necessary. Specific embodiments envisioned for implementing the present invention will now be described.

[0034] 1. Data Entry

[0035] Users can input memories and events by voice or text. The interface for accepting this input is a device such as a smartphone or computer. If voice input is selected, the device uses a voice recognition engine to convert the voice into text. If text input is selected, the voice is received as text data.

[0036] 2. Emotion analysis

[0037] The server performs sentiment analysis on the input text data. The sentiment analysis engine analyzes the emotions contained in the text data and identifies the type of emotion (e.g., joy, sadness, surprise, etc.). This sentiment tag is useful for later data search.

[0038] 3. Data storage

[0039] The server stores the text data in a database along with the analyzed emotion data, along with the date and time of entry, allowing data to be searched for based on a specific date or time.

[0040] 4. Data Recall

[0041] When a user wants to reminisce about past memories based on a specific date or emotion, they send a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0042] 5. Data playback

[0043] The device receives the data retrieved from the server and presents it to the user. The user can read the data in text format or play it back in audio format. In the audio format, the text data is sent to a speech synthesis engine, which plays it back as audio data.

[0044] Specific examples

[0045] 1. The user speaks "Memories of summer vacation 2015."

[0046] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[0047] 3. The server sends the received text data to the emotion analysis engine and obtains the emotion tag "fun."

[0048] 4. The server stores the emotion tag and text data in the database with the date summer 2015.

[0049] 5. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0050] 6. The server retrieves the day's data from the database and sends it to the terminal.

[0051] 7. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[0052] Thus, the present invention is a system that allows users to efficiently record past memories and emotions and easily look back on them when necessary.

[0053] The processing flow will be explained below.

[0054] Step 1:

[0055] Users input memories and events using voice or text. For voice input, users speak into the device, and for text input, users input using a keyboard or touch panel.

[0056] Step 2:

[0057] The terminal receives the user's voice input and sends it to the voice recognition engine. The voice recognition engine converts the voice data into text data and obtains the text data. If the input is text, this step is skipped and the process proceeds to the next step.

[0058] Step 3:

[0059] The device sends the acquired text data to the server, where information such as the input date and time is also added to the text data.

[0060] Step 4:

[0061] The server analyzes the received text data using a sentiment analysis engine, which identifies the emotions in the text data and generates corresponding sentiment tags (e.g., "happy," "sad," etc.).

[0062] Step 5:

[0063] The server stores the emotion-analyzed text data and emotion tags in a database, along with the input date and time.

[0064] Step 6:

[0065] Years later, if a user wants to relive past memories based on a specific date or emotion, they can use their device to enter a request, for example, "Tell me what happened in the summer of 2015."

[0066] Step 7:

[0067] The device sends the user's request to the server, which includes search criteria such as a specific date or emotion tag.

[0068] Step 8:

[0069] The server searches and retrieves the relevant data from the database, selecting stored data based on specific dates and emotion tags.

[0070] Step 9:

[0071] The server sends the acquired data to the device, which includes text data and emotion tags.

[0072] Step 10:

[0073] The device presents the acquired data to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine, which converts it into audio data and plays it back.

[0074] The above is the specific processing flow of the program in the present invention.

[0075] Example 1

[0076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0077] In today's world, systems that allow users to record and easily recall their daily memories and emotions are extremely important. However, conventional systems have struggled to integrate a series of operations, such as converting voice input to text, analyzing emotions, and storing, searching, and playing back data. As a result, it has been difficult for users to efficiently search and play back past memories, and data storage and management has been cumbersome. Therefore, there is a need for a system that can accept voice and text input, perform emotion analysis, and easily store, search, and play back data.

[0078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0079] In this invention, the server includes means for accepting input by voice or text, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for saving the sentiment-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing the retrieved data in text or voice, means for a user to send a request using a terminal and search for data based on the conditions, and means for the terminal to retrieve data based on the request sent to the server and present it to the user. This allows the user to input by voice or text and easily save, search, and play sentiment-analyzed data.

[0080] "Voice input" refers to the voice data collected through a microphone and input into the system from what the user says.

[0081] "Text input" refers to a method in which a user directly inputs text using a keyboard or touch screen.

[0082] A "speech recognition engine" is software or algorithms used to analyze collected voice data and convert it into corresponding text data.

[0083] An "emotion analysis engine" is software or an algorithm that identifies the type and intensity of emotions from input text data and assigns emotion tags.

[0084] A "database" is a system or software for efficiently storing, managing, and retrieving structured data.

[0085] "Storing data" refers to the act of recording the input text data and its accompanying information in a database.

[0086] "Data search" refers to the operation a user performs to retrieve stored data based on specific criteria (e.g., date, emotion tag).

[0087] A "speech synthesis engine" is software or algorithm that analyzes text data and generates and plays back speech that resembles a human voice.

[0088] A "request" is an instruction from a user requesting the system to perform some operation.

[0089] A "terminal" is a device operated by a user (e.g., a smartphone, a computer).

[0090] "Server" means a central computer system that processes and manages data.

[0091] A "user" is a person who uses the system.

[0092] The present invention is a system that allows users to record their memories and emotions and easily review them later. The system includes a set of means for accepting voice or text input, storing it, and playing it back when needed.

[0093] Data Entry

[0094] Users can input memories and events using devices such as smartphones or computers. They can choose to input via voice or text. In the case of voice input, the device uses a speech recognition engine (e.g., automatic speech recognition software) to convert the speech into text. In the case of text input, the speech is accepted as text data as is.

[0095] Voice Recognition

[0096] If voice input is selected, the device will use automatic speech recognition software to convert speech to text. For example, if a user speaks "Memories of summer vacation 2015," the device will convert this speech to text.

[0097] Emotion analysis

[0098] The server passes the received text data to an emotion analysis engine (e.g., natural language processing software) to analyze the emotions contained in the text. As a result of the analysis, an emotion tag (e.g., "joy," "sadness," etc.) is assigned.

[0099] Data storage

[0100] The server stores the text data, including the analysis results, in a database (e.g., a relational database system), along with the input date and time.

[0101] Recalling Data

[0102] When a user wants to reminisce about past memories, they use their device to send a request based on a specific condition to the server, for example, "Tell me what happened in the summer of 2015."

[0103] Data Search and Playback

[0104] The server searches a database based on the request and retrieves the relevant data. The server then sends the retrieved data to the device. The device plays the retrieved data to the user in text or audio format. In the audio format, the text data is sent to a speech synthesis engine (e.g., text-to-speech software) and played as audio data.

[0105] Specific examples

[0106] Below are some examples of specific prompt sentences.

[0107] 1. The user speaks "Memories of summer vacation 2015."

[0108] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[0109] 3. The server sends the received text data to the emotion analysis engine and obtains the emotion tag "fun."

[0110] 4. The server stores the emotion tag and text data in the database with the date summer 2015.

[0111] 5. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0112] 6. The server retrieves the day's data from the database and sends it to the terminal.

[0113] 7. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[0114] This system allows users to efficiently record past memories and emotions and easily look back on them when needed. In addition to the convenience of voice and text input, it also enables effective data retrieval through emotion analysis.

[0115] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0116] Step 1: Data entry

[0117] Users can input memories and events using devices such as smartphones or computers. Here, users can choose between voice input or text input. For voice input, users use the device's microphone and tap the voice input button on the app to speak. For text input, users enter text using a keyboard or touchscreen.

[0118] Input: Audio or text data

[0119] Output: Converted text data (in the case of speech input) or raw text data (in the case of sentence input)

[0120] Step 2: Voice Recognition

[0121] When voice data is input, the terminal uses a voice recognition engine (e.g., automatic speech recognition software) to convert the voice into text. The converted text data is displayed on the screen for the user to confirm. If necessary, the user can correct this text.

[0122] Input: Audio data

[0123] Output: Converted text data

[0124] Step 3: Send text data

[0125] The terminal sends the converted text data or directly entered text to the server, where the data is converted into JSON format and sent securely to the server using the HTTPS protocol.

[0126] Input: Text data (entered by the user)

[0127] Output: Text data sent to the server

[0128] Step 4: Sentiment Analysis

[0129] The server passes the received text data to a sentiment analysis engine (e.g., natural language processing software) to analyze the sentiment contained in the text. Keywords and sentence structure within the text data are used for the analysis, and sentiment is tagged as the analysis result.

[0130] Input: Text data sent to the server

[0131] Output: Text data with emotion tags added

[0132] Step 5: Save your data

[0133] The server stores the emotion-tagged text data in a database (e.g., a relational database system), along with the input date and time and the user ID.

[0134] Input: Text data with emotion tags added

[0135] Output: Records stored in a database

[0136] Step 6: Calling the data

[0137] When a user wants to reminisce about past memories, they use their device to send a request based on specific criteria to the server, such as "Tell me what happened in the summer of 2015."

[0138] Input: Request data sent from the user's device (conditional search query)

[0139] Output: Request data sent to the server

[0140] Step 7: Search for data

[0141] The server searches the database based on the request received and extracts the relevant data, using SQL queries to extract records that match the criteria.

[0142] Input: Request data sent to the server

[0143] Output: Extracted data that matches the conditions

[0144] Step 8: Replaying the Data

[0145] The device presents the data retrieved from the server to the user. This data can be displayed in text format or played back as audio. In the latter case, the text data is sent to a speech synthesis engine (e.g., text-to-speech software) and played back as audio data.

[0146] Input: Extracted data that matches the conditions

[0147] Output: Text or audio data presented to the user

[0148] (Application example 1)

[0149] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0150] The problem is that there is currently no system that records the memories and emotions experienced by customers when they visit a store, efficiently manages this information, and allows them to offer special services the next time they visit. In particular, there is a need for a system that aims to improve customer experience in physical stores, which can analyze emotions from voice and text input and use the data to improve services the next time.

[0151] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0152] In this invention, the server includes means for accepting voice or text input, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for saving the sentiment-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing back the retrieved data in text or voice, and means for recording a customer's voice or text input when they visit the store and retrieving and playing back data for providing special service the next time they visit. This makes it possible to effectively record and manage customer memories and emotions and to provide special service based on that data the next time they visit.

[0153] "Voice or text input" is the process by which a user provides memories or emotions to a system as a recording or text.

[0154] "Converting voice data into text data" refers to a technique for analyzing recorded voice information and expressing its contents as character string data.

[0155] "Performing emotion analysis on input text data" refers to the process of analyzing the character string data to identify the emotion contained therein and determining the type of emotion (for example, joy, sadness, surprise, etc.).

[0156] "Saving emotion-analyzed text data in a database" refers to recording character string data in a storage device together with information on the analyzed emotion.

[0157] "Retrieving stored data based on specific conditions" refers to the operation of searching and retrieving recorded data according to conditions such as a specific date and time or type of emotion.

[0158] "Reproducing retrieved data as text or voice" means displaying the retrieved data as text or reproducing it as voice using voice synthesis technology.

[0159] "Recording customer voice or text input when they visit the store, and acquiring and playing back data to provide special services the next time they visit" refers to the process of recording the memories and emotions entered by customers when they visit the store, and proposing and providing special experiences and services based on that data the next time they visit.

[0160] This invention is a system for improving customer experience in physical stores. It records the memories and emotions experienced by customers when they visit the store, and provides special services based on that data the next time they visit.

[0161] The system accepts voice or text input, analyzes the data emotionally, and stores it in a database. It then retrieves the stored data based on specific conditions and plays it back in text or voice format, providing personalized service to customers.

[0162] Hardware and Software Used

[0163] Hardware: Tablets, smartphones

[0164] Software: Flask (web application framework), SQLite (database), SpeechRecognition (speech recognition library), text2emotion (emotion analysis library)

[0165] Step Description

[0166] 1. Voice or text input:

[0167] Users use a tablet or smartphone to input the events and emotions they experienced in real time using voice or text.

[0168] In the case of voice input, the terminal uses a voice recognition engine to convert the voice data into text data.

[0169] 2. Emotion analysis:

[0170] The server performs sentiment analysis on the input text data.

[0171] Use the text2emotion library to analyze emotions contained in text data and identify the type of emotion.

[0172] 3. Data storage:

[0173] The server stores the text data together with the analyzed emotion data in a database.

[0174] The data is tagged with the date and time it was entered and an emotion tag.

[0175] 4. Data retrieval:

[0176] The next time the user visits the store, they can send a request to search for past memories and emotions via the terminal.

[0177] The server searches and acquires the relevant data from the database.

[0178] 5. Data Regeneration:

[0179] The terminal presents the acquired text data to the user.

[0180] If the user so desires, the text data is sent to a speech synthesis engine and reproduced as voice data.

[0181] Specific examples

[0182] Below are some examples of specific prompt sentences.

[0183] Example of an input prompt:

[0184] "Please record your memories from the travel destinations you visited in the fall of 2019."

[0185] "Please tell us about a fun memory you had the last time you visited us."

[0186] Example of a retrieval prompt:

[0187] "Tell me about your memories of the last time you visited us."

[0188] "Show me the record of your most inspiring visit."

[0189] In this way, by implementing this invention, it is possible to effectively record and manage customer memories and emotions, and provide special services based on that data the next time the customer visits the store. This is an effective means of improving the customer experience in physical stores and increasing customer satisfaction.

[0190] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0191] Step 1:

[0192] Input - Users use a tablet or smartphone to input their experiences and memories using voice or text.

[0193] Operation - The device displays an interface that accepts input from the user. If the user speaks, the device activates a speech recognition engine.

[0194] Output - In the case of voice input, voice data is obtained, and in the case of text input, text data is obtained.

[0195] Step 2:

[0196] Input - Sends speech data to the speech recognition engine.

[0197] Operation - The device uses the speech recognition library (SpeechRecognition) to convert voice data into text data, analyze the voice data, and represent it as string data.

[0198] Output - The audio data is converted to text data.

[0199] Step 3:

[0200] Input - Sends text data to the server.

[0201] Operation - The server runs the emotion analysis engine (text2emotion) on the text data to analyze the emotions in the text data, thereby identifying the type of emotion (e.g., joy, sadness, surprise, etc.) from the input text data.

[0202] Output - Emotion tags are generated for the text data.

[0203] Step 4:

[0204] Input - Receives text data with emotion tags.

[0205] Operation - The server saves the analyzed emotion data and text data in a database, along with the input date and time and emotion tag.

[0206] Output - Text data, sentiment tags, and date / time stored in a database.

[0207] Step 5:

[0208] Input - When a user returns to the store, they request a search for data based on specific criteria (e.g., date and time or type of emotion).

[0209] Operation - The device accepts the user's request and sends it to the server, which then accesses the database and searches for and retrieves the relevant data based on the specified criteria.

[0210] Output - The text data and sentiment tags retrieved based on the criteria.

[0211] Step 6:

[0212] Input - Data retrieved from the server.

[0213] Operation - The device presents the acquired text data and emotion tags to the user. If the user wishes, the text is sent to a speech synthesis engine and played back as speech.

[0214] Output - The text data presented to the user or the audio data played.

[0215] This allows the system to record the customer's memories and emotions and provide personalized service the next time they visit.

[0216] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0217] The present invention is a system that allows users to record their memories and emotions and easily look back on them later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when necessary, and further combines it with an emotion engine that recognizes the user's emotions. Specific embodiments envisioned for implementing the present invention will now be described.

[0218] 1. Data Entry

[0219] Users can input memories and events by voice or text. The interface for accepting this input is a device such as a smartphone or computer. If voice input is selected, the device uses a voice recognition engine to convert the voice into text. If text input is selected, the voice is received as text data.

[0220] 2. Emotion recognition

[0221] As the device receives voice input, the emotion engine analyzes the user's tone and speech patterns to recognize the user's emotions in real time, and this emotion information is sent to the server along with the text data.

[0222] 3. Emotion analysis

[0223] The server analyzes the input text data using an emotion analysis engine. Using the emotion information received from the emotion engine, the server tags the text data with emotions. These emotion tags are useful for later data searches.

[0224] 4. Data storage

[0225] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input, making it possible to search for data based on a specific date or time.

[0226] 5. Data Recall

[0227] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0228] 6. Data playback

[0229] The device receives the data retrieved from the server and presents it to the user. The user can read the data in text format or play it back in audio format. In the audio format, the text data is sent to a speech synthesis engine, which plays it back as audio data.

[0230] Specific examples

[0231] 1. The user speaks "Memories of summer vacation 2015."

[0232] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[0233] 3. At the same time, the emotion engine analyzes the user's voice tone and generates emotional information such as "feeling happy."

[0234] 4. The server sends the received text data and emotional information to the emotion analysis engine, and obtains the emotion tag "fun."

[0235] 5. The server stores the emotion tag and text data in the database with a date of summer 2015.

[0236] 6. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0237] 7. The server retrieves the day's data from the database and sends it to the terminal.

[0238] 8. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[0239] In this way, the present invention is a system that allows users to efficiently record past memories and emotions and easily look back on them when necessary. By combining it with an emotion engine, it becomes possible to record and replay even deeper emotional experiences.

[0240] The processing flow will be explained below.

[0241] Step 1:

[0242] Users can input memories and events by voice or text. Users can speak into the device or input text using a keyboard or touch panel.

[0243] Step 2:

[0244] The device receives voice input and first sends it to a speech recognition engine, which converts the voice data into text data. If the input is text, this step is skipped.

[0245] Step 3:

[0246] The device sends the converted text data to the emotion engine, and at the same time, sends the input voice data to the emotion engine, which analyzes the user's voice tone and speech patterns.

[0247] Step 4:

[0248] The emotion engine analyzes the voice tone and speech patterns to obtain the user's emotion information (e.g., "happy," "sad," etc.). The emotion engine then adds this emotion information to the text data.

[0249] Step 5:

[0250] The device sends text data and emotion information to the server, including information such as the input date and time.

[0251] Step 6:

[0252] The server sends the received text data and emotional information to the emotion analysis engine for further detailed emotion analysis, which then tags the emotions in the text data.

[0253] Step 7:

[0254] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input.

[0255] Step 8:

[0256] Years later, if a user wants to relive past memories based on a specific date or emotion, they can use their device to enter a request, for example, "Tell me what happened in the summer of 2015."

[0257] Step 9:

[0258] The device sends the user's request to the server, which includes search criteria such as a specific date or emotion tag.

[0259] Step 10:

[0260] The server searches and retrieves the relevant data from the database, selecting stored data based on specific dates and emotion tags.

[0261] Step 11:

[0262] The server sends the acquired data to the device, which includes text data and emotion tags.

[0263] Step 12:

[0264] The device presents the acquired data to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine, which converts it into audio data and plays it back.

[0265] The above is the specific processing flow of the program in the present invention.

[0266] Example 2

[0267] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0268] In recent years, the importance of recording personal memories and emotions and looking back on them has increased. However, existing systems have difficulty efficiently managing voice and text input and recording detailed information, including emotions. In particular, there is a lack of systems that allow users to easily search and play back memories associated with emotions later. To solve these problems, there is a need for a system that can handle voice and text input, recognize and analyze emotions, and store and play back data in a single step.

[0269] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0270] In this invention, the server includes means for accepting voice or text input, means for converting the accepted voice data into text data using a voice recognition engine, means for utilizing an emotion recognition engine that generates emotional information in real time based on the text data and the user's tone of voice, means for emotionally analyzing the generated emotional information together with the text data and assigning an emotion tag, means for storing the emotionally analyzed text data and emotional information in a database together with date information, means for searching and retrieving the stored data under specific conditions based on a user's request, and means for converting the retrieved data into voice using a text or voice synthesis engine and playing it back. This allows users to efficiently record individual memories and emotions and easily look back on them in a form linked to their emotions.

[0271] "Voice or text input" refers to the act of a user providing information to a system in the form of voice or text.

[0272] A "voice recognition engine" refers to software or hardware for converting voice data into text data.

[0273] "Text data" refers to data that is stored and processed as textual information.

[0274] An "emotion recognition engine" refers to software or hardware that analyzes a user's voice tone, speech patterns, etc., and generates emotional information.

[0275] "Emotion information" refers to data that indicates the user's emotional state according to an emotion recognition engine.

[0276] "Sentiment analysis" refers to the process of identifying and classifying emotions within data based on text data and emotional information.

[0277] An "emotion tag" refers to a label that indicates a specific emotional state, assigned through emotion analysis.

[0278] "Database" refers to a system for structured storage of text data, emotional information, and other related information.

[0279] "Speech synthesis engine" refers to software or hardware for converting text data into speech data.

[0280] "User request" refers to the operation or instruction given by a user to the system to search for or retrieve specific data.

[0281] This invention is a system that efficiently records a user's memories and emotions and allows them to be easily reviewed later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when needed. In addition, by combining it with an emotion engine that recognizes the user's emotions, it becomes possible to record and play back richer emotional experiences.

[0282] Data Entry

[0283] Users can input memories and events using devices such as smartphones or computers, either by voice or text. If voice input is selected, the device uses a voice recognition engine (e.g., a general voice recognition engine) to convert the voice into text data. If text input is selected, the voice is received as text data.

[0284] emotion recognition

[0285] The device simultaneously sends the input voice data to an emotion recognition engine (a common emotion recognition engine) and generates emotion information by analyzing the user's voice tone and speech pattern. This emotion information is then sent to the server along with the text data.

[0286] Emotion analysis

[0287] The server receives the text data and emotional information and analyzes it using an emotion analysis engine. Based on the emotional information received from the emotion recognition engine, it assigns emotional tags to the text data. These emotional tags are useful for later searches.

[0288] Data storage

[0289] The server stores the emotion-analyzed text data and emotion tags in a database along with date information, allowing users to search for data based on specific dates or emotions.

[0290] Recalling Data

[0291] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0292] Data playback

[0293] The device receives the data retrieved from the server and presents it to the user, who can read it in text format or play it back in audio format using a speech synthesis engine (a common speech synthesis engine).

[0294] Specific examples

[0295] For example, if a user speaks "Memories of summer vacation 2015," the following happens:

[0296] 1. The user uses a smartphone to speak "Memories of summer vacation 2015."

[0297] 2. The device converts the voice into text, creating text data called "Memories of summer vacation 2015."

[0298] 3. At the same time, the emotion recognition engine analyzes the tone of the voice and generates emotional information such as "happy."

[0299] 4. The server sends the received text data and emotional information to the emotion analysis engine, and obtains the emotion tag "fun."

[0300] 5. The server stores this text data and emotion tag in a database along with the date in the summer of 2015.

[0301] 6. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0302] 7. The server searches the database for the relevant data and sends it to the terminal.

[0303] 8. The device converts the received data into audio format and plays back to the user, "The summer of 2015 was fun."

[0304] In this way, this system allows users to look back on their past memories and emotions more easily and in a richer way. By combining it with an emotion engine, it is possible to play back emotionally rich memories rather than just recording them.

[0305] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0306] Step 1:

[0307] Users can use a device such as a smartphone or computer to input memories and events using voice or text.

[0308] Input: User voice or text data

[0309] Output: Audio data or raw text data

[0310] Specific behavior:

[0311] The user speaks into the microphone on their smartphone about their "memories of summer vacation in 2015."

[0312] The device acquires the voice data and prepares it to be sent to the voice recognition engine.

[0313] Step 2:

[0314] When a voice input is received, the device uses a voice recognition engine to convert the voice data into text data. When a sentence is input, the device receives the text data as is.

[0315] Input: Audio data or text data

[0316] Output: Text data

[0317] Specific behavior:

[0318] The terminal calls a general voice recognition engine and converts the voice data into text data.

[0319] The converted text data "Memories of summer vacation 2015" is obtained.

[0320] Step 3:

[0321] The device analyzes the user's voice tone and speech patterns and generates emotional information using an emotion recognition engine, which is then sent to the server along with the text data.

[0322] Input: Audio data, text data

[0323] Output: Emotional information, text data

[0324] Specific behavior:

[0325] The device uses a general emotion recognition engine to generate emotional information such as "it looks fun" from the voice data.

[0326] The generated emotion information is attached to text data and transmitted to a server.

[0327] Step 4:

[0328] The server performs emotion analysis based on the received text data and emotion information, and assigns emotion tags to the text data.

[0329] Input: Text data, emotion information

[0330] Output: Text data with emotion tags

[0331] Specific behavior:

[0332] The server sends the text data and the emotional information "it looks fun" to the emotion analysis engine.

[0333] The server receives the emotion tag "fun" from the emotion analysis engine and assigns it to the text data.

[0334] Step 5:

[0335] The server stores the emotion-tagged text data and emotion information together with date information in a database.

[0336] Input: Emotion-tagged text data, emotion information, date information

[0337] Output: Save to database

[0338] Specific behavior:

[0339] The server stores the text data "Memories of summer vacation 2015" and the emotion tag "It was fun" in a database along with date information.

[0340] Step 6:

[0341] When a user wants to look back on past memories based on a specific date or emotion, the user inputs a request using the terminal.

[0342] Input: Request (e.g. "Tell me what happened in the summer of 2015")

[0343] Output: A request to retrieve the corresponding data

[0344] Specific behavior:

[0345] A few years later, a user uses a smartphone app to enter a request: "Tell me what happened in the summer of 2015."

[0346] The terminal sends this request to the server.

[0347] Step 7:

[0348] The server searches for and retrieves data corresponding to the request from the database.

[0349] Input: Request, Database

[0350] Output: Acquisition of relevant data

[0351] Specific behavior:

[0352] The server searches the database for and retrieves data tagged with "Memories of summer vacation 2015" and the emotion "It was fun."

[0353] The server transmits the acquired data to the terminal.

[0354] Step 8:

[0355] The terminal presents the received data to the user, who can read it in text format or play it in audio format.

[0356] Input: Retrieved data

[0357] Output: Text display or audio playback

[0358] Specific behavior:

[0359] The device displays the received data in text format, telling the user, "The summer of 2015 was fun."

[0360] The device uses a speech synthesis engine to convert the text data into audio data, which is then played through the speaker.

[0361] (Application example 2)

[0362] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0363] In today's busy daily lives, users often miss opportunities to record their memories and emotions. Furthermore, the time and effort required to review recorded memories and emotions makes it difficult to effectively utilize the accumulated data. Furthermore, the lack of relevant information based on the user's past emotional experiences makes it difficult to provide appropriate information tailored to the user's needs.

[0364] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting input by voice or text, means for converting voice data into text data, means for performing emotion analysis on the input text data, means for saving the emotion-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing the retrieved data in text or voice, and means for recommending related information based on the saved emotion data. This not only enables users to efficiently record past memories and emotions and easily look back on them when necessary, but also enables recommendations of related information based on past emotional experiences.

[0365] "Means for accepting voice or text input" refers to an interface for receiving input from a user in the form of voice or text and processing it appropriately.

[0366] The "means for converting voice data into text data" is a function for converting voice data into text data using voice recognition technology.

[0367] The "means for performing sentiment analysis on input text data" is an engine that analyzes the content of the text data and determines the polarity and intensity of sentiment from that content.

[0368] The "means for saving emotion-analyzed text data in a database" is a function for recording and managing the analyzed emotion information and text data in a database.

[0369] "Means for retrieving stored data based on specific conditions" refers to a function that searches for and retrieves necessary information from a database based on a user request or set conditions.

[0370] "Means for reproducing the acquired data in text or audio" refers to a function for presenting the extracted data in a format that is easy for the user to understand, and includes text display and audio reproduction.

[0371] The "means for recommending related information based on stored emotion data" is a function that suggests related items and information to the user based on the recorded emotion data of the user.

[0372] The present invention provides a system that allows users to record emotions and memories felt during their virtual store experience and easily review them later. This system is implemented using devices such as smartphones and head-mounted displays.

[0373] 1. Data Entry

[0374] Users can input their experiences and emotions in the virtual store by voice. The interface for this input is a device such as a smartphone or head-mounted display. If voice input is selected, the device uses a voice recognition engine to convert the voice into text.

[0375] 2. Emotion recognition

[0376] As the device receives voice input, the emotion analysis engine analyzes the user's tone and speech patterns to recognize the user's emotions in real time, and this emotion information is sent to the server along with the text data.

[0377] 3. Emotion analysis

[0378] The server analyzes the input text data using a sentiment analysis engine. Using the emotional information received from the sentiment analysis engine, the server tags the text data with emotions. These emotional tags are useful for later data searches.

[0379] 4. Data storage

[0380] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input, making it possible to search for data based on a specific date or time.

[0381] 5. Data Recall

[0382] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0383] 6. Data playback

[0384] The device receives the data from the server and presents it to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine and played back as audio data.

[0385] 7. Recommendation of related information

[0386] Based on the stored emotional data, the system recommends information related to past emotional experiences to the user. For example, the system has the function of recommending new items based on items that the user found enjoyable in the past.

[0387] Hardware and software used

[0388] Voice input devices: smartphones, head-mounted displays

[0389] Speech recognition engine: speech_recognition

[0390] Sentiment Analysis Engine: TextBlob Library

[0391] Database: SQLite

[0392] Server: General web server software (e.g. Flask, Django)

[0393] Specific examples

[0394] For example, suppose a user purchases new running shoes in a virtual store on October 1, 2023, and wants to record the excitement they felt. The user speaks into a microphone, saying, "I'm so happy I bought my new running shoes." The system recognizes this speech as text and records a positive sentiment polarity (e.g., 0.8). This data, along with the date, is stored in a database, and if the user later searches for "happy events in October 2023," the record will appear.

[0395] Prompt Sentence Examples

[0396] I was very happy to purchase new running shoes on October 1, 2023. My emotional polarity is 0.8.

[0397] In this way, users can efficiently record their past emotions and memories, enriching their shopping experience within the virtual store.

[0398] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0399] Step 1: User provides voice or text input

[0400] Specific behavior:

[0401] Users use devices such as smartphones or head-mounted displays to input their experiences and emotions in the virtual store by voice or text. For voice input, they use the device's microphone input function, and for text input, they use an on-screen input form.

[0402] Input: User voice or text data

[0403] Output: Raw audio or text data

[0404] Step 2: The device converts the audio data into text data.

[0405] Specific behavior:

[0406] The device uses a speech recognition engine (e.g., the speech_recognition library) to convert the user's voice input into text data. This process involves inputting the voice data into a speech recognition model, which then outputs the speech content as a string of characters.

[0407] Input: User's voice data

[0408] Output: Data converted to text

[0409] Step 3: The device sends the entered text to the server

[0410] Specific behavior:

[0411] The terminal transmits the converted text data and its meta information (e.g., input date and time) to the server via a network communication protocol (e.g., HTTP).

[0412] Input: Text data, input date and time

[0413] Output: Text data and meta information sent to the server

[0414] Step 4: The server inputs the text data into the sentiment analysis engine

[0415] Specific behavior:

[0416] The server inputs the received text data into a sentiment analysis engine (e.g., TextBlob library) to analyze the sentiment polarity (positive, negative, neutral) of each text. The analysis results include sentiment polarity values ​​and sentiment-related tags.

[0417] Input: Text data

[0418] Output: Sentiment polarity value and sentiment tag

[0419] Step 5: The server stores the analysis results in a database

[0420] Specific behavior:

[0421] The server stores the sentiment polarity values, text data, and meta information (e.g., input date and time) in a database (e.g., SQLite). During this storage process, each data item is inserted into the appropriate database table.

[0422] Input: Sentiment polarity value, text data, input date and time

[0423] Output: Sentiment analysis data stored in a database

[0424] Step 6: User requests historical emotion data based on specific criteria

[0425] Specific behavior:

[0426] A user uses a device to request historical emotion data based on a specific date and emotion. The request is sent to the server as a search query containing the specified criteria (e.g., specific date, emotion tag).

[0427] Input: Search criteria (e.g. date, emotion tag)

[0428] Output: Search request sent to the server

[0429] Step 7: The server searches and retrieves the relevant data from the database.

[0430] Specific behavior:

[0431] The server searches and retrieves the corresponding emotion data from the database based on the received search request. In this process, it filters the data that matches the search criteria and generates the data to be sent back to the user.

[0432] Input: Search criteria (e.g. date, emotion tag)

[0433] Output: Emotion data as search results

[0434] Step 8: Play back the data retrieved by the device as text or audio

[0435] Specific behavior:

[0436] The device presents the emotion data received from the server to the user, who can read the data in text format or play it back in audio format using a speech synthesis engine (e.g., a TTS engine).

[0437] Input: Emotion data as search results

[0438] Output: Text display or audio playback data

[0439] Step 9: The server generates a prompt to recommend related information and presents it to the user.

[0440] Specific behavior:

[0441] The server generates prompts to recommend related information based on the stored emotional data. For example, it can recommend new items based on items that users have previously found enjoyable. The generated prompts are input into a generative AI model and presented to the user as recommendations.

[0442] Input: Emotional data, past emotional experiences

[0443] Output: Prompt statement with relevant information

[0444] Prompt Sentence Examples

[0445] I was very happy to purchase new running shoes on October 1, 2023. My emotional polarity is 0.8.

[0446] This prompt allows the user to reflect on past experiences while simultaneously receiving new, relevant information.

[0447] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0448] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0449] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0450] [Second embodiment]

[0451] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0452] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0453] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0454] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0455] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0456] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0457] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0458] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0459] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0460] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0461] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0462] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0463] The present invention is a system that allows users to record their memories and emotions and easily look back on them later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when necessary. Specific embodiments envisioned for implementing the present invention will now be described.

[0464] 1. Data Entry

[0465] Users can input memories and events by voice or text. The interface for accepting this input is a device such as a smartphone or computer. If voice input is selected, the device uses a voice recognition engine to convert the voice into text. If text input is selected, the voice is received as text data.

[0466] 2. Emotion analysis

[0467] The server performs sentiment analysis on the input text data. The sentiment analysis engine analyzes the emotions contained in the text data and identifies the type of emotion (e.g., joy, sadness, surprise, etc.). This sentiment tag is useful for later data search.

[0468] 3. Data storage

[0469] The server stores the text data in a database along with the analyzed emotion data, along with the date and time of entry, allowing data to be searched for based on a specific date or time.

[0470] 4. Data Recall

[0471] When a user wants to reminisce about past memories based on a specific date or emotion, they send a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0472] 5. Data playback

[0473] The device receives the data retrieved from the server and presents it to the user. The user can read the data in text format or play it back in audio format. In the audio format, the text data is sent to a speech synthesis engine, which plays it back as audio data.

[0474] Specific examples

[0475] 1. The user speaks "Memories of summer vacation 2015."

[0476] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[0477] 3. The server sends the received text data to the emotion analysis engine and obtains the emotion tag "fun."

[0478] 4. The server stores the emotion tag and text data in the database with the date summer 2015.

[0479] 5. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0480] 6. The server retrieves the day's data from the database and sends it to the terminal.

[0481] 7. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[0482] Thus, the present invention is a system that allows users to efficiently record past memories and emotions and easily look back on them when necessary.

[0483] The processing flow will be explained below.

[0484] Step 1:

[0485] Users input memories and events using voice or text. For voice input, users speak into the device, and for text input, users input using a keyboard or touch panel.

[0486] Step 2:

[0487] The terminal receives the user's voice input and sends it to the voice recognition engine. The voice recognition engine converts the voice data into text data and obtains the text data. If the input is text, this step is skipped and the process proceeds to the next step.

[0488] Step 3:

[0489] The device sends the acquired text data to the server, where information such as the input date and time is also added to the text data.

[0490] Step 4:

[0491] The server analyzes the received text data using a sentiment analysis engine, which identifies the emotions in the text data and generates corresponding sentiment tags (e.g., "happy," "sad," etc.).

[0492] Step 5:

[0493] The server stores the emotion-analyzed text data and emotion tags in a database, along with the input date and time.

[0494] Step 6:

[0495] Years later, if a user wants to relive past memories based on a specific date or emotion, they can use their device to enter a request, for example, "Tell me what happened in the summer of 2015."

[0496] Step 7:

[0497] The device sends the user's request to the server, which includes search criteria such as a specific date or emotion tag.

[0498] Step 8:

[0499] The server searches and retrieves the relevant data from the database, selecting stored data based on specific dates and emotion tags.

[0500] Step 9:

[0501] The server sends the acquired data to the device, which includes text data and emotion tags.

[0502] Step 10:

[0503] The device presents the acquired data to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine, which converts it into audio data and plays it back.

[0504] The above is the specific processing flow of the program in the present invention.

[0505] Example 1

[0506] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0507] In today's world, systems that allow users to record and easily recall their daily memories and emotions are extremely important. However, conventional systems have struggled to integrate a series of operations, such as converting voice input to text, analyzing emotions, and storing, searching, and playing back data. As a result, it has been difficult for users to efficiently search and play back past memories, and data storage and management has been cumbersome. Therefore, there is a need for a system that can accept voice and text input, perform emotion analysis, and easily store, search, and play back data.

[0508] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0509] In this invention, the server includes means for accepting input by voice or text, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for saving the sentiment-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing the retrieved data in text or voice, means for a user to send a request using a terminal and search for data based on the conditions, and means for the terminal to retrieve data based on the request sent to the server and present it to the user. This allows the user to input by voice or text and easily save, search, and play sentiment-analyzed data.

[0510] "Voice input" refers to the voice data collected through a microphone and input into the system from what the user says.

[0511] "Text input" refers to a method in which a user directly inputs text using a keyboard or touch screen.

[0512] A "speech recognition engine" is software or algorithms used to analyze collected voice data and convert it into corresponding text data.

[0513] An "emotion analysis engine" is software or an algorithm that identifies the type and intensity of emotions from input text data and assigns emotion tags.

[0514] A "database" is a system or software for efficiently storing, managing, and retrieving structured data.

[0515] "Storing data" refers to the act of recording the input text data and its accompanying information in a database.

[0516] "Data search" refers to the operation a user performs to retrieve stored data based on specific criteria (e.g., date, emotion tag).

[0517] A "speech synthesis engine" is software or algorithm that analyzes text data and generates and plays back speech that resembles a human voice.

[0518] A "request" is an instruction from a user requesting the system to perform some operation.

[0519] A "terminal" is a device operated by a user (e.g., a smartphone, a computer).

[0520] "Server" means a central computer system that processes and manages data.

[0521] A "user" is a person who uses the system.

[0522] The present invention is a system that allows users to record their memories and emotions and easily review them later. The system includes a set of means for accepting voice or text input, storing it, and playing it back when needed.

[0523] Data Entry

[0524] Users can input memories and events using devices such as smartphones or computers. They can choose to input via voice or text. In the case of voice input, the device uses a speech recognition engine (e.g., automatic speech recognition software) to convert the speech into text. In the case of text input, the speech is accepted as text data as is.

[0525] Voice Recognition

[0526] If voice input is selected, the device will use automatic speech recognition software to convert speech to text. For example, if a user speaks "Memories of summer vacation 2015," the device will convert this speech to text.

[0527] Emotion analysis

[0528] The server passes the received text data to an emotion analysis engine (e.g., natural language processing software) to analyze the emotions contained in the text. As a result of the analysis, an emotion tag (e.g., "joy," "sadness," etc.) is assigned.

[0529] Data storage

[0530] The server stores the text data, including the analysis results, in a database (e.g., a relational database system), along with the input date and time.

[0531] Recalling Data

[0532] When a user wants to reminisce about past memories, they use their device to send a request based on a specific condition to the server, for example, "Tell me what happened in the summer of 2015."

[0533] Data Search and Playback

[0534] The server searches a database based on the request and retrieves the relevant data. The server then sends the retrieved data to the device. The device plays the retrieved data to the user in text or audio format. In the audio format, the text data is sent to a speech synthesis engine (e.g., text-to-speech software) and played as audio data.

[0535] Specific examples

[0536] Below are some examples of specific prompt sentences.

[0537] 1. The user speaks "Memories of summer vacation 2015."

[0538] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[0539] 3. The server sends the received text data to the emotion analysis engine and obtains the emotion tag "fun."

[0540] 4. The server stores the emotion tag and text data in the database with the date summer 2015.

[0541] 5. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0542] 6. The server retrieves the day's data from the database and sends it to the terminal.

[0543] 7. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[0544] This system allows users to efficiently record past memories and emotions and easily look back on them when needed. In addition to the convenience of voice and text input, it also enables effective data retrieval through emotion analysis.

[0545] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0546] Step 1: Data entry

[0547] Users can input memories and events using devices such as smartphones or computers. Here, users can choose between voice input or text input. For voice input, users use the device's microphone and tap the voice input button on the app to speak. For text input, users enter text using a keyboard or touchscreen.

[0548] Input: Audio or text data

[0549] Output: Converted text data (in the case of speech input) or raw text data (in the case of sentence input)

[0550] Step 2: Voice Recognition

[0551] When voice data is input, the terminal uses a voice recognition engine (e.g., automatic speech recognition software) to convert the voice into text. The converted text data is displayed on the screen for the user to confirm. If necessary, the user can correct this text.

[0552] Input: Audio data

[0553] Output: Converted text data

[0554] Step 3: Send text data

[0555] The terminal sends the converted text data or directly entered text to the server, where the data is converted into JSON format and sent securely to the server using the HTTPS protocol.

[0556] Input: Text data (entered by the user)

[0557] Output: Text data sent to the server

[0558] Step 4: Sentiment Analysis

[0559] The server passes the received text data to a sentiment analysis engine (e.g., natural language processing software) to analyze the sentiment contained in the text. Keywords and sentence structure within the text data are used for the analysis, and sentiment is tagged as the analysis result.

[0560] Input: Text data sent to the server

[0561] Output: Text data with emotion tags added

[0562] Step 5: Save your data

[0563] The server stores the emotion-tagged text data in a database (e.g., a relational database system), along with the input date and time and the user ID.

[0564] Input: Text data with emotion tags added

[0565] Output: Records stored in a database

[0566] Step 6: Calling the data

[0567] When a user wants to reminisce about past memories, they use their device to send a request based on specific criteria to the server, such as "Tell me what happened in the summer of 2015."

[0568] Input: Request data sent from the user's device (conditional search query)

[0569] Output: Request data sent to the server

[0570] Step 7: Search for data

[0571] The server searches the database based on the request received and extracts the relevant data, using SQL queries to extract records that match the criteria.

[0572] Input: Request data sent to the server

[0573] Output: Extracted data that matches the conditions

[0574] Step 8: Replaying the Data

[0575] The device presents the data retrieved from the server to the user. This data can be displayed in text format or played back as audio. In the latter case, the text data is sent to a speech synthesis engine (e.g., text-to-speech software) and played back as audio data.

[0576] Input: Extracted data that matches the conditions

[0577] Output: Text or audio data presented to the user

[0578] (Application example 1)

[0579] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0580] The problem is that there is currently no system that records the memories and emotions experienced by customers when they visit a store, efficiently manages this information, and allows them to offer special services the next time they visit. In particular, there is a need for a system that aims to improve customer experience in physical stores, which can analyze emotions from voice and text input and use the data to improve services the next time.

[0581] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0582] In this invention, the server includes means for accepting voice or text input, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for saving the sentiment-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing back the retrieved data in text or voice, and means for recording a customer's voice or text input when they visit the store and retrieving and playing back data for providing special service the next time they visit. This makes it possible to effectively record and manage customer memories and emotions and to provide special service based on that data the next time they visit.

[0583] "Voice or text input" is the process by which a user provides memories or emotions to a system as a recording or text.

[0584] "Converting voice data into text data" refers to a technique for analyzing recorded voice information and expressing its contents as character string data.

[0585] "Performing emotion analysis on input text data" refers to the process of analyzing the character string data to identify the emotion contained therein and determining the type of emotion (for example, joy, sadness, surprise, etc.).

[0586] "Saving emotion-analyzed text data in a database" refers to recording character string data in a storage device together with information on the analyzed emotion.

[0587] "Retrieving stored data based on specific conditions" refers to the operation of searching and retrieving recorded data according to conditions such as a specific date and time or type of emotion.

[0588] "Reproducing retrieved data as text or voice" means displaying the retrieved data as text or reproducing it as voice using voice synthesis technology.

[0589] "Recording customer voice or text input when they visit the store, and acquiring and playing back data to provide special services the next time they visit" refers to the process of recording the memories and emotions entered by customers when they visit the store, and proposing and providing special experiences and services based on that data the next time they visit.

[0590] This invention is a system for improving customer experience in physical stores. It records the memories and emotions experienced by customers when they visit the store, and provides special services based on that data the next time they visit.

[0591] The system accepts voice or text input, analyzes the data emotionally, and stores it in a database. It then retrieves the stored data based on specific conditions and plays it back in text or voice format, providing personalized service to customers.

[0592] Hardware and Software Used

[0593] Hardware: Tablets, smartphones

[0594] Software: Flask (web application framework), SQLite (database), SpeechRecognition (speech recognition library), text2emotion (emotion analysis library)

[0595] Step Description

[0596] 1. Voice or text input:

[0597] Users use a tablet or smartphone to input the events and emotions they experienced in real time using voice or text.

[0598] In the case of voice input, the terminal uses a voice recognition engine to convert the voice data into text data.

[0599] 2. Emotion analysis:

[0600] The server performs sentiment analysis on the input text data.

[0601] Use the text2emotion library to analyze emotions contained in text data and identify the type of emotion.

[0602] 3. Data storage:

[0603] The server stores the text data together with the analyzed emotion data in a database.

[0604] The data is tagged with the date and time it was entered and an emotion tag.

[0605] 4. Data retrieval:

[0606] The next time the user visits the store, they can send a request to search for past memories and emotions via the terminal.

[0607] The server searches and acquires the relevant data from the database.

[0608] 5. Data Regeneration:

[0609] The terminal presents the acquired text data to the user.

[0610] If the user so desires, the text data is sent to a speech synthesis engine and reproduced as voice data.

[0611] Specific examples

[0612] Below are some examples of specific prompt sentences.

[0613] Example of an input prompt:

[0614] "Please record your memories from the travel destinations you visited in the fall of 2019."

[0615] "Please tell us about a fun memory you had the last time you visited us."

[0616] Example of a retrieval prompt:

[0617] "Tell me about your memories of the last time you visited us."

[0618] "Show me the record of your most inspiring visit."

[0619] In this way, by implementing this invention, it is possible to effectively record and manage customer memories and emotions, and provide special services based on that data the next time the customer visits the store. This is an effective means of improving the customer experience in physical stores and increasing customer satisfaction.

[0620] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0621] Step 1:

[0622] Input - Users use a tablet or smartphone to input their experiences and memories using voice or text.

[0623] Operation - The device displays an interface that accepts input from the user. If the user speaks, the device activates a speech recognition engine.

[0624] Output - In the case of voice input, voice data is obtained, and in the case of text input, text data is obtained.

[0625] Step 2:

[0626] Input - Sends speech data to the speech recognition engine.

[0627] Operation - The device uses the speech recognition library (SpeechRecognition) to convert voice data into text data, analyze the voice data, and represent it as string data.

[0628] Output - The audio data is converted to text data.

[0629] Step 3:

[0630] Input - Sends text data to the server.

[0631] Operation - The server runs the emotion analysis engine (text2emotion) on the text data to analyze the emotions in the text data, thereby identifying the type of emotion (e.g., joy, sadness, surprise, etc.) from the input text data.

[0632] Output - Emotion tags are generated for the text data.

[0633] Step 4:

[0634] Input - Receives text data with emotion tags.

[0635] Operation - The server saves the analyzed emotion data and text data in a database, along with the input date and time and emotion tag.

[0636] Output - Text data, sentiment tags, and date / time stored in a database.

[0637] Step 5:

[0638] Input - When a user returns to the store, they request a search for data based on specific criteria (e.g., date and time or type of emotion).

[0639] Operation - The device accepts the user's request and sends it to the server, which then accesses the database and searches for and retrieves the relevant data based on the specified criteria.

[0640] Output - The text data and sentiment tags retrieved based on the criteria.

[0641] Step 6:

[0642] Input - Data retrieved from the server.

[0643] Operation - The device presents the acquired text data and emotion tags to the user. If the user wishes, the text is sent to a speech synthesis engine and played back as speech.

[0644] Output - The text data presented to the user or the audio data played.

[0645] This allows the system to record the customer's memories and emotions and provide personalized service the next time they visit.

[0646] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0647] The present invention is a system that allows users to record their memories and emotions and easily look back on them later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when necessary, and further combines it with an emotion engine that recognizes the user's emotions. Specific embodiments envisioned for implementing the present invention will now be described.

[0648] 1. Data Entry

[0649] Users can input memories and events by voice or text. The interface for accepting this input is a device such as a smartphone or computer. If voice input is selected, the device uses a voice recognition engine to convert the voice into text. If text input is selected, the voice is received as text data.

[0650] 2. Emotion recognition

[0651] As the device receives voice input, the emotion engine analyzes the user's tone and speech patterns to recognize the user's emotions in real time, and this emotion information is sent to the server along with the text data.

[0652] 3. Emotion analysis

[0653] The server analyzes the input text data using an emotion analysis engine. Using the emotion information received from the emotion engine, the server tags the text data with emotions. These emotion tags are useful for later data searches.

[0654] 4. Data storage

[0655] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input, making it possible to search for data based on a specific date or time.

[0656] 5. Data Recall

[0657] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0658] 6. Data playback

[0659] The device receives the data retrieved from the server and presents it to the user. The user can read the data in text format or play it back in audio format. In the audio format, the text data is sent to a speech synthesis engine, which plays it back as audio data.

[0660] Specific examples

[0661] 1. The user speaks "Memories of summer vacation 2015."

[0662] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[0663] 3. At the same time, the emotion engine analyzes the user's voice tone and generates emotional information such as "feeling happy."

[0664] 4. The server sends the received text data and emotional information to the emotion analysis engine, and obtains the emotion tag "fun."

[0665] 5. The server stores the emotion tag and text data in the database with a date of summer 2015.

[0666] 6. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0667] 7. The server retrieves the day's data from the database and sends it to the terminal.

[0668] 8. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[0669] In this way, the present invention is a system that allows users to efficiently record past memories and emotions and easily look back on them when necessary. By combining it with an emotion engine, it becomes possible to record and replay even deeper emotional experiences.

[0670] The processing flow will be explained below.

[0671] Step 1:

[0672] Users can input memories and events by voice or text. Users can speak into the device or input text using a keyboard or touch panel.

[0673] Step 2:

[0674] The device receives voice input and first sends it to a speech recognition engine, which converts the voice data into text data. If the input is text, this step is skipped.

[0675] Step 3:

[0676] The device sends the converted text data to the emotion engine, and at the same time, sends the input voice data to the emotion engine, which analyzes the user's voice tone and speech patterns.

[0677] Step 4:

[0678] The emotion engine analyzes the voice tone and speech patterns to obtain the user's emotion information (e.g., "happy," "sad," etc.). The emotion engine then adds this emotion information to the text data.

[0679] Step 5:

[0680] The device sends text data and emotion information to the server, including information such as the input date and time.

[0681] Step 6:

[0682] The server sends the received text data and emotional information to the emotion analysis engine for further detailed emotion analysis, which then tags the emotions in the text data.

[0683] Step 7:

[0684] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input.

[0685] Step 8:

[0686] Years later, if a user wants to relive past memories based on a specific date or emotion, they can use their device to enter a request, for example, "Tell me what happened in the summer of 2015."

[0687] Step 9:

[0688] The device sends the user's request to the server, which includes search criteria such as a specific date or emotion tag.

[0689] Step 10:

[0690] The server searches and retrieves the relevant data from the database, selecting stored data based on specific dates and emotion tags.

[0691] Step 11:

[0692] The server sends the acquired data to the device, which includes text data and emotion tags.

[0693] Step 12:

[0694] The device presents the acquired data to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine, which converts it into audio data and plays it back.

[0695] The above is the specific processing flow of the program in the present invention.

[0696] Example 2

[0697] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0698] In recent years, the importance of recording personal memories and emotions and looking back on them has increased. However, existing systems have difficulty efficiently managing voice and text input and recording detailed information, including emotions. In particular, there is a lack of systems that allow users to easily search and play back memories associated with emotions later. To solve these problems, there is a need for a system that can handle voice and text input, recognize and analyze emotions, and store and play back data in a single step.

[0699] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0700] In this invention, the server includes means for accepting voice or text input, means for converting the accepted voice data into text data using a voice recognition engine, means for utilizing an emotion recognition engine that generates emotional information in real time based on the text data and the user's tone of voice, means for emotionally analyzing the generated emotional information together with the text data and assigning an emotion tag, means for storing the emotionally analyzed text data and emotional information in a database together with date information, means for searching and retrieving the stored data under specific conditions based on a user's request, and means for converting the retrieved data into voice using a text or voice synthesis engine and playing it back. This allows users to efficiently record individual memories and emotions and easily look back on them in a form linked to their emotions.

[0701] "Voice or text input" refers to the act of a user providing information to a system in the form of voice or text.

[0702] A "voice recognition engine" refers to software or hardware for converting voice data into text data.

[0703] "Text data" refers to data that is stored and processed as textual information.

[0704] An "emotion recognition engine" refers to software or hardware that analyzes a user's voice tone, speech patterns, etc., and generates emotional information.

[0705] "Emotion information" refers to data that indicates the user's emotional state according to an emotion recognition engine.

[0706] "Sentiment analysis" refers to the process of identifying and classifying emotions within data based on text data and emotional information.

[0707] An "emotion tag" refers to a label that indicates a specific emotional state, assigned through emotion analysis.

[0708] "Database" refers to a system for structured storage of text data, emotional information, and other related information.

[0709] "Speech synthesis engine" refers to software or hardware for converting text data into speech data.

[0710] "User request" refers to the operation or instruction given by a user to the system to search for or retrieve specific data.

[0711] This invention is a system that efficiently records a user's memories and emotions and allows them to be easily reviewed later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when needed. In addition, by combining it with an emotion engine that recognizes the user's emotions, it becomes possible to record and play back richer emotional experiences.

[0712] Data Entry

[0713] Users can input memories and events using devices such as smartphones or computers, either by voice or text. If voice input is selected, the device uses a voice recognition engine (e.g., a general voice recognition engine) to convert the voice into text data. If text input is selected, the voice is received as text data.

[0714] emotion recognition

[0715] The device simultaneously sends the input voice data to an emotion recognition engine (a common emotion recognition engine) and generates emotion information by analyzing the user's voice tone and speech pattern. This emotion information is then sent to the server along with the text data.

[0716] Emotion analysis

[0717] The server receives the text data and emotional information and analyzes it using an emotion analysis engine. Based on the emotional information received from the emotion recognition engine, it assigns emotional tags to the text data. These emotional tags are useful for later searches.

[0718] Data storage

[0719] The server stores the emotion-analyzed text data and emotion tags in a database along with date information, allowing users to search for data based on specific dates or emotions.

[0720] Recalling Data

[0721] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0722] Data playback

[0723] The device receives the data retrieved from the server and presents it to the user, who can read it in text format or play it back in audio format using a speech synthesis engine (a common speech synthesis engine).

[0724] Specific examples

[0725] For example, if a user speaks "Memories of summer vacation 2015," the following happens:

[0726] 1. The user uses a smartphone to speak "Memories of summer vacation 2015."

[0727] 2. The device converts the voice into text, creating text data called "Memories of summer vacation 2015."

[0728] 3. At the same time, the emotion recognition engine analyzes the tone of the voice and generates emotional information such as "happy."

[0729] 4. The server sends the received text data and emotional information to the emotion analysis engine, and obtains the emotion tag "fun."

[0730] 5. The server stores this text data and emotion tag in a database along with the date in the summer of 2015.

[0731] 6. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0732] 7. The server searches the database for the relevant data and sends it to the terminal.

[0733] 8. The device converts the received data into audio format and plays back to the user, "The summer of 2015 was fun."

[0734] In this way, this system allows users to look back on their past memories and emotions more easily and in a richer way. By combining it with an emotion engine, it is possible to play back emotionally rich memories rather than just recording them.

[0735] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0736] Step 1:

[0737] Users can use a device such as a smartphone or computer to input memories and events using voice or text.

[0738] Input: User voice or text data

[0739] Output: Audio data or raw text data

[0740] Specific behavior:

[0741] The user speaks into the microphone on their smartphone about their "memories of summer vacation in 2015."

[0742] The device acquires the voice data and prepares it to be sent to the voice recognition engine.

[0743] Step 2:

[0744] When a voice input is received, the device uses a voice recognition engine to convert the voice data into text data. When a sentence is input, the device receives the text data as is.

[0745] Input: Audio data or text data

[0746] Output: Text data

[0747] Specific behavior:

[0748] The terminal calls a general voice recognition engine and converts the voice data into text data.

[0749] The converted text data "Memories of summer vacation 2015" is obtained.

[0750] Step 3:

[0751] The device analyzes the user's voice tone and speech patterns and generates emotional information using an emotion recognition engine, which is then sent to the server along with the text data.

[0752] Input: Audio data, text data

[0753] Output: Emotional information, text data

[0754] Specific behavior:

[0755] The device uses a general emotion recognition engine to generate emotional information such as "it looks fun" from the voice data.

[0756] The generated emotion information is attached to text data and transmitted to a server.

[0757] Step 4:

[0758] The server performs emotion analysis based on the received text data and emotion information, and assigns emotion tags to the text data.

[0759] Input: Text data, emotion information

[0760] Output: Text data with emotion tags

[0761] Specific behavior:

[0762] The server sends the text data and the emotional information "it looks fun" to the emotion analysis engine.

[0763] The server receives the emotion tag "fun" from the emotion analysis engine and assigns it to the text data.

[0764] Step 5:

[0765] The server stores the emotion-tagged text data and emotion information together with date information in a database.

[0766] Input: Emotion-tagged text data, emotion information, date information

[0767] Output: Save to database

[0768] Specific behavior:

[0769] The server stores the text data "Memories of summer vacation 2015" and the emotion tag "It was fun" in a database along with date information.

[0770] Step 6:

[0771] When a user wants to look back on past memories based on a specific date or emotion, the user inputs a request using the terminal.

[0772] Input: Request (e.g. "Tell me what happened in the summer of 2015")

[0773] Output: A request to retrieve the corresponding data

[0774] Specific behavior:

[0775] A few years later, a user uses a smartphone app to enter a request: "Tell me what happened in the summer of 2015."

[0776] The terminal sends this request to the server.

[0777] Step 7:

[0778] The server searches for and retrieves data corresponding to the request from the database.

[0779] Input: Request, Database

[0780] Output: Acquisition of relevant data

[0781] Specific behavior:

[0782] The server searches the database for and retrieves data tagged with "Memories of summer vacation 2015" and the emotion "It was fun."

[0783] The server transmits the acquired data to the terminal.

[0784] Step 8:

[0785] The terminal presents the received data to the user, who can read it in text format or play it in audio format.

[0786] Input: Retrieved data

[0787] Output: Text display or audio playback

[0788] Specific behavior:

[0789] The device displays the received data in text format, telling the user, "The summer of 2015 was fun."

[0790] The device uses a speech synthesis engine to convert the text data into audio data, which is then played through the speaker.

[0791] (Application example 2)

[0792] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0793] In today's busy daily lives, users often miss opportunities to record their memories and emotions. Furthermore, the time and effort required to review recorded memories and emotions makes it difficult to effectively utilize the accumulated data. Furthermore, the lack of relevant information based on the user's past emotional experiences makes it difficult to provide appropriate information tailored to the user's needs.

[0794] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting input by voice or text, means for converting voice data into text data, means for performing emotion analysis on the input text data, means for saving the emotion-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing the retrieved data in text or voice, and means for recommending related information based on the saved emotion data. This not only enables users to efficiently record past memories and emotions and easily look back on them when necessary, but also enables recommendations of related information based on past emotional experiences.

[0795] "Means for accepting voice or text input" refers to an interface for receiving input from a user in the form of voice or text and processing it appropriately.

[0796] The "means for converting voice data into text data" is a function for converting voice data into text data using voice recognition technology.

[0797] The "means for performing sentiment analysis on input text data" is an engine that analyzes the content of the text data and determines the polarity and intensity of sentiment from that content.

[0798] The "means for saving emotion-analyzed text data in a database" is a function for recording and managing the analyzed emotion information and text data in a database.

[0799] "Means for retrieving stored data based on specific conditions" refers to a function that searches for and retrieves necessary information from a database based on a user request or set conditions.

[0800] "Means for reproducing the acquired data in text or audio" refers to a function for presenting the extracted data in a format that is easy for the user to understand, and includes text display and audio reproduction.

[0801] The "means for recommending related information based on stored emotion data" is a function that suggests related items and information to the user based on the recorded emotion data of the user.

[0802] The present invention provides a system that allows users to record emotions and memories felt during their virtual store experience and easily review them later. This system is implemented using devices such as smartphones and head-mounted displays.

[0803] 1. Data Entry

[0804] Users can input their experiences and emotions in the virtual store by voice. The interface for this input is a device such as a smartphone or head-mounted display. If voice input is selected, the device uses a voice recognition engine to convert the voice into text.

[0805] 2. Emotion recognition

[0806] As the device receives voice input, the emotion analysis engine analyzes the user's tone and speech patterns to recognize the user's emotions in real time, and this emotion information is sent to the server along with the text data.

[0807] 3. Emotion analysis

[0808] The server analyzes the input text data using a sentiment analysis engine. Using the emotional information received from the sentiment analysis engine, the server tags the text data with emotions. These emotional tags are useful for later data searches.

[0809] 4. Data storage

[0810] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input, making it possible to search for data based on a specific date or time.

[0811] 5. Data Recall

[0812] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0813] 6. Data playback

[0814] The device receives the data from the server and presents it to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine and played back as audio data.

[0815] 7. Recommendation of related information

[0816] Based on the stored emotional data, the system recommends information related to past emotional experiences to the user. For example, the system has the function of recommending new items based on items that the user found enjoyable in the past.

[0817] Hardware and software used

[0818] Voice input devices: smartphones, head-mounted displays

[0819] Speech recognition engine: speech_recognition

[0820] Sentiment Analysis Engine: TextBlob Library

[0821] Database: SQLite

[0822] Server: General web server software (e.g. Flask, Django)

[0823] Specific examples

[0824] For example, suppose a user purchases new running shoes in a virtual store on October 1, 2023, and wants to record the excitement they felt. The user speaks into a microphone, saying, "I'm so happy I bought my new running shoes." The system recognizes this speech as text and records a positive sentiment polarity (e.g., 0.8). This data, along with the date, is stored in a database, and if the user later searches for "happy events in October 2023," the record will appear.

[0825] Prompt Sentence Examples

[0826] I was very happy to purchase new running shoes on October 1, 2023. My emotional polarity is 0.8.

[0827] In this way, users can efficiently record their past emotions and memories, enriching their shopping experience within the virtual store.

[0828] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0829] Step 1: User provides voice or text input

[0830] Specific behavior:

[0831] Users use devices such as smartphones or head-mounted displays to input their experiences and emotions in the virtual store by voice or text. For voice input, they use the device's microphone input function, and for text input, they use an on-screen input form.

[0832] Input: User voice or text data

[0833] Output: Raw audio or text data

[0834] Step 2: The device converts the audio data into text data.

[0835] Specific behavior:

[0836] The device uses a speech recognition engine (e.g., the speech_recognition library) to convert the user's voice input into text data. This process involves inputting the voice data into a speech recognition model, which then outputs the speech content as a string of characters.

[0837] Input: User's voice data

[0838] Output: Data converted to text

[0839] Step 3: The device sends the entered text to the server

[0840] Specific behavior:

[0841] The terminal transmits the converted text data and its meta information (e.g., input date and time) to the server via a network communication protocol (e.g., HTTP).

[0842] Input: Text data, input date and time

[0843] Output: Text data and meta information sent to the server

[0844] Step 4: The server inputs the text data into the sentiment analysis engine

[0845] Specific behavior:

[0846] The server inputs the received text data into a sentiment analysis engine (e.g., TextBlob library) to analyze the sentiment polarity (positive, negative, neutral) of each text. The analysis results include sentiment polarity values ​​and sentiment-related tags.

[0847] Input: Text data

[0848] Output: Sentiment polarity value and sentiment tag

[0849] Step 5: The server stores the analysis results in a database

[0850] Specific behavior:

[0851] The server stores the sentiment polarity values, text data, and meta information (e.g., input date and time) in a database (e.g., SQLite). During this storage process, each data item is inserted into the appropriate database table.

[0852] Input: Sentiment polarity value, text data, input date and time

[0853] Output: Sentiment analysis data stored in a database

[0854] Step 6: User requests historical emotion data based on specific criteria

[0855] Specific behavior:

[0856] A user uses a device to request historical emotion data based on a specific date and emotion. The request is sent to the server as a search query containing the specified criteria (e.g., specific date, emotion tag).

[0857] Input: Search criteria (e.g. date, emotion tag)

[0858] Output: Search request sent to the server

[0859] Step 7: The server searches and retrieves the relevant data from the database.

[0860] Specific behavior:

[0861] The server searches and retrieves the corresponding emotion data from the database based on the received search request. In this process, it filters the data that matches the search criteria and generates the data to be sent back to the user.

[0862] Input: Search criteria (e.g. date, emotion tag)

[0863] Output: Emotion data as search results

[0864] Step 8: Play back the data retrieved by the device as text or audio

[0865] Specific behavior:

[0866] The device presents the emotion data received from the server to the user, who can read the data in text format or play it back in audio format using a speech synthesis engine (e.g., a TTS engine).

[0867] Input: Emotion data as search results

[0868] Output: Text display or audio playback data

[0869] Step 9: The server generates a prompt to recommend related information and presents it to the user.

[0870] Specific behavior:

[0871] The server generates prompts to recommend related information based on the stored emotional data. For example, it can recommend new items based on items that users have previously found enjoyable. The generated prompts are input into a generative AI model and presented to the user as recommendations.

[0872] Input: Emotional data, past emotional experiences

[0873] Output: Prompt statement with relevant information

[0874] Prompt Sentence Examples

[0875] I was very happy to purchase new running shoes on October 1, 2023. My emotional polarity is 0.8.

[0876] This prompt allows the user to reflect on past experiences while simultaneously receiving new, relevant information.

[0877] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0878] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0879] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0880] [Third embodiment]

[0881] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0882] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0883] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0884] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0885] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0886] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0887] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0888] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0889] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0890] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0891] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0892] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0893] The present invention is a system that allows users to record their memories and emotions and easily look back on them later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when necessary. Specific embodiments envisioned for implementing the present invention will now be described.

[0894] 1. Data Entry

[0895] Users can input memories and events by voice or text. The interface for accepting this input is a device such as a smartphone or computer. If voice input is selected, the device uses a voice recognition engine to convert the voice into text. If text input is selected, the voice is received as text data.

[0896] 2. Emotion analysis

[0897] The server performs sentiment analysis on the input text data. The sentiment analysis engine analyzes the emotions contained in the text data and identifies the type of emotion (e.g., joy, sadness, surprise, etc.). This sentiment tag is useful for later data search.

[0898] 3. Data storage

[0899] The server stores the text data in a database along with the analyzed emotion data, along with the date and time of entry, allowing data to be searched for based on a specific date or time.

[0900] 4. Data Recall

[0901] When a user wants to reminisce about past memories based on a specific date or emotion, they send a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[0902] 5. Data playback

[0903] The device receives the data retrieved from the server and presents it to the user. The user can read the data in text format or play it back in audio format. In the audio format, the text data is sent to a speech synthesis engine, which plays it back as audio data.

[0904] Specific examples

[0905] 1. The user speaks "Memories of summer vacation 2015."

[0906] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[0907] 3. The server sends the received text data to the emotion analysis engine and obtains the emotion tag "fun."

[0908] 4. The server stores the emotion tag and text data in the database with the date summer 2015.

[0909] 5. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0910] 6. The server retrieves the day's data from the database and sends it to the terminal.

[0911] 7. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[0912] Thus, the present invention is a system that allows users to efficiently record past memories and emotions and easily look back on them when necessary.

[0913] The processing flow will be explained below.

[0914] Step 1:

[0915] Users input memories and events using voice or text. For voice input, users speak into the device, and for text input, users input using a keyboard or touch panel.

[0916] Step 2:

[0917] The terminal receives the user's voice input and sends it to the voice recognition engine. The voice recognition engine converts the voice data into text data and obtains the text data. If the input is text, this step is skipped and the process proceeds to the next step.

[0918] Step 3:

[0919] The device sends the acquired text data to the server, where information such as the input date and time is also added to the text data.

[0920] Step 4:

[0921] The server analyzes the received text data using a sentiment analysis engine, which identifies the emotions in the text data and generates corresponding sentiment tags (e.g., "happy," "sad," etc.).

[0922] Step 5:

[0923] The server stores the emotion-analyzed text data and emotion tags in a database, along with the input date and time.

[0924] Step 6:

[0925] Years later, if a user wants to relive past memories based on a specific date or emotion, they can use their device to enter a request, for example, "Tell me what happened in the summer of 2015."

[0926] Step 7:

[0927] The device sends the user's request to the server, which includes search criteria such as a specific date or emotion tag.

[0928] Step 8:

[0929] The server searches and retrieves the relevant data from the database, selecting stored data based on specific dates and emotion tags.

[0930] Step 9:

[0931] The server sends the acquired data to the device, which includes text data and emotion tags.

[0932] Step 10:

[0933] The device presents the acquired data to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine, which converts it into audio data and plays it back.

[0934] The above is the specific processing flow of the program in the present invention.

[0935] Example 1

[0936] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0937] In today's world, systems that allow users to record and easily recall their daily memories and emotions are extremely important. However, conventional systems have struggled to integrate a series of operations, such as converting voice input to text, analyzing emotions, and storing, searching, and playing back data. As a result, it has been difficult for users to efficiently search and play back past memories, and data storage and management has been cumbersome. Therefore, there is a need for a system that can accept voice and text input, perform emotion analysis, and easily store, search, and play back data.

[0938] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0939] In this invention, the server includes means for accepting input by voice or text, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for saving the sentiment-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing the retrieved data in text or voice, means for a user to send a request using a terminal and search for data based on the conditions, and means for the terminal to retrieve data based on the request sent to the server and present it to the user. This allows the user to input by voice or text and easily save, search, and play sentiment-analyzed data.

[0940] "Voice input" refers to the voice data collected through a microphone and input into the system from what the user says.

[0941] "Text input" refers to a method in which a user directly inputs text using a keyboard or touch screen.

[0942] A "speech recognition engine" is software or algorithms used to analyze collected voice data and convert it into corresponding text data.

[0943] An "emotion analysis engine" is software or an algorithm that identifies the type and intensity of emotions from input text data and assigns emotion tags.

[0944] A "database" is a system or software for efficiently storing, managing, and retrieving structured data.

[0945] "Storing data" refers to the act of recording the input text data and its accompanying information in a database.

[0946] "Data search" refers to the operation a user performs to retrieve stored data based on specific criteria (e.g., date, emotion tag).

[0947] A "speech synthesis engine" is software or algorithm that analyzes text data and generates and plays back speech that resembles a human voice.

[0948] A "request" is an instruction from a user requesting the system to perform some operation.

[0949] A "terminal" is a device operated by a user (e.g., a smartphone, a computer).

[0950] "Server" means a central computer system that processes and manages data.

[0951] A "user" is a person who uses the system.

[0952] The present invention is a system that allows users to record their memories and emotions and easily review them later. The system includes a set of means for accepting voice or text input, storing it, and playing it back when needed.

[0953] Data Entry

[0954] Users can input memories and events using devices such as smartphones or computers. They can choose to input via voice or text. In the case of voice input, the device uses a speech recognition engine (e.g., automatic speech recognition software) to convert the speech into text. In the case of text input, the speech is accepted as text data as is.

[0955] Voice Recognition

[0956] If voice input is selected, the device will use automatic speech recognition software to convert speech to text. For example, if a user speaks "Memories of summer vacation 2015," the device will convert this speech to text.

[0957] Emotion analysis

[0958] The server passes the received text data to an emotion analysis engine (e.g., natural language processing software) to analyze the emotions contained in the text. As a result of the analysis, an emotion tag (e.g., "joy," "sadness," etc.) is assigned.

[0959] Data storage

[0960] The server stores the text data, including the analysis results, in a database (e.g., a relational database system), along with the input date and time.

[0961] Recalling Data

[0962] When a user wants to reminisce about past memories, they use their device to send a request based on a specific condition to the server, for example, "Tell me what happened in the summer of 2015."

[0963] Data Search and Playback

[0964] The server searches a database based on the request and retrieves the relevant data. The server then sends the retrieved data to the device. The device plays the retrieved data to the user in text or audio format. In the audio format, the text data is sent to a speech synthesis engine (e.g., text-to-speech software) and played as audio data.

[0965] Specific examples

[0966] Below are some examples of specific prompt sentences.

[0967] 1. The user speaks "Memories of summer vacation 2015."

[0968] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[0969] 3. The server sends the received text data to the emotion analysis engine and obtains the emotion tag "fun."

[0970] 4. The server stores the emotion tag and text data in the database with the date summer 2015.

[0971] 5. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[0972] 6. The server retrieves the day's data from the database and sends it to the terminal.

[0973] 7. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[0974] This system allows users to efficiently record past memories and emotions and easily look back on them when needed. In addition to the convenience of voice and text input, it also enables effective data retrieval through emotion analysis.

[0975] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0976] Step 1: Data entry

[0977] Users can input memories and events using devices such as smartphones or computers. Here, users can choose between voice input or text input. For voice input, users use the device's microphone and tap the voice input button on the app to speak. For text input, users enter text using a keyboard or touchscreen.

[0978] Input: Audio or text data

[0979] Output: Converted text data (in the case of speech input) or raw text data (in the case of sentence input)

[0980] Step 2: Voice Recognition

[0981] When voice data is input, the terminal uses a voice recognition engine (e.g., automatic speech recognition software) to convert the voice into text. The converted text data is displayed on the screen for the user to confirm. If necessary, the user can correct this text.

[0982] Input: Audio data

[0983] Output: Converted text data

[0984] Step 3: Send text data

[0985] The terminal sends the converted text data or directly entered text to the server, where the data is converted into JSON format and sent securely to the server using the HTTPS protocol.

[0986] Input: Text data (entered by the user)

[0987] Output: Text data sent to the server

[0988] Step 4: Sentiment Analysis

[0989] The server passes the received text data to a sentiment analysis engine (e.g., natural language processing software) to analyze the sentiment contained in the text. Keywords and sentence structure within the text data are used for the analysis, and sentiment is tagged as the analysis result.

[0990] Input: Text data sent to the server

[0991] Output: Text data with emotion tags added

[0992] Step 5: Save your data

[0993] The server stores the emotion-tagged text data in a database (e.g., a relational database system), along with the input date and time and the user ID.

[0994] Input: Text data with emotion tags added

[0995] Output: Records stored in a database

[0996] Step 6: Calling the data

[0997] When a user wants to reminisce about past memories, they use their device to send a request based on specific criteria to the server, such as "Tell me what happened in the summer of 2015."

[0998] Input: Request data sent from the user's device (conditional search query)

[0999] Output: Request data sent to the server

[1000] Step 7: Search for data

[1001] The server searches the database based on the request received and extracts the relevant data, using SQL queries to extract records that match the criteria.

[1002] Input: Request data sent to the server

[1003] Output: Extracted data that matches the conditions

[1004] Step 8: Replaying the Data

[1005] The device presents the data retrieved from the server to the user. This data can be displayed in text format or played back as audio. In the latter case, the text data is sent to a speech synthesis engine (e.g., text-to-speech software) and played back as audio data.

[1006] Input: Extracted data that matches the conditions

[1007] Output: Text or audio data presented to the user

[1008] (Application example 1)

[1009] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1010] The problem is that there is currently no system that records the memories and emotions experienced by customers when they visit a store, efficiently manages this information, and allows them to offer special services the next time they visit. In particular, there is a need for a system that aims to improve customer experience in physical stores, which can analyze emotions from voice and text input and use the data to improve services the next time.

[1011] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1012] In this invention, the server includes means for accepting voice or text input, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for saving the sentiment-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing back the retrieved data in text or voice, and means for recording a customer's voice or text input when they visit the store and retrieving and playing back data for providing special service the next time they visit. This makes it possible to effectively record and manage customer memories and emotions and to provide special service based on that data the next time they visit.

[1013] "Voice or text input" is the process by which a user provides memories or emotions to a system as a recording or text.

[1014] "Converting voice data into text data" refers to a technique for analyzing recorded voice information and expressing its contents as character string data.

[1015] "Performing emotion analysis on input text data" refers to the process of analyzing the character string data to identify the emotion contained therein and determining the type of emotion (for example, joy, sadness, surprise, etc.).

[1016] "Saving emotion-analyzed text data in a database" refers to recording character string data in a storage device together with information on the analyzed emotion.

[1017] "Retrieving stored data based on specific conditions" refers to the operation of searching and retrieving recorded data according to conditions such as a specific date and time or type of emotion.

[1018] "Reproducing retrieved data as text or voice" means displaying the retrieved data as text or reproducing it as voice using voice synthesis technology.

[1019] "Recording customer voice or text input when they visit the store, and acquiring and playing back data to provide special services the next time they visit" refers to the process of recording the memories and emotions entered by customers when they visit the store, and proposing and providing special experiences and services based on that data the next time they visit.

[1020] This invention is a system for improving customer experience in physical stores. It records the memories and emotions experienced by customers when they visit the store, and provides special services based on that data the next time they visit.

[1021] The system accepts voice or text input, analyzes the data emotionally, and stores it in a database. It then retrieves the stored data based on specific conditions and plays it back in text or voice format, providing personalized service to customers.

[1022] Hardware and Software Used

[1023] Hardware: Tablets, smartphones

[1024] Software: Flask (web application framework), SQLite (database), SpeechRecognition (speech recognition library), text2emotion (emotion analysis library)

[1025] Step Description

[1026] 1. Voice or text input:

[1027] Users use a tablet or smartphone to input the events and emotions they experienced in real time using voice or text.

[1028] In the case of voice input, the terminal uses a voice recognition engine to convert the voice data into text data.

[1029] 2. Emotion analysis:

[1030] The server performs sentiment analysis on the input text data.

[1031] Use the text2emotion library to analyze emotions contained in text data and identify the type of emotion.

[1032] 3. Data storage:

[1033] The server stores the text data together with the analyzed emotion data in a database.

[1034] The data is tagged with the date and time it was entered and an emotion tag.

[1035] 4. Data retrieval:

[1036] The next time the user visits the store, they can send a request to search for past memories and emotions via the terminal.

[1037] The server searches and acquires the relevant data from the database.

[1038] 5. Data Regeneration:

[1039] The terminal presents the acquired text data to the user.

[1040] If the user so desires, the text data is sent to a speech synthesis engine and reproduced as voice data.

[1041] Specific examples

[1042] Below are some examples of specific prompt sentences.

[1043] Example of an input prompt:

[1044] "Please record your memories from the travel destinations you visited in the fall of 2019."

[1045] "Please tell us about a fun memory you had the last time you visited us."

[1046] Example of a retrieval prompt:

[1047] "Tell me about your memories of the last time you visited us."

[1048] "Show me the record of your most inspiring visit."

[1049] In this way, by implementing this invention, it is possible to effectively record and manage customer memories and emotions, and provide special services based on that data the next time the customer visits the store. This is an effective means of improving the customer experience in physical stores and increasing customer satisfaction.

[1050] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1051] Step 1:

[1052] Input - Users use a tablet or smartphone to input their experiences and memories using voice or text.

[1053] Operation - The device displays an interface that accepts input from the user. If the user speaks, the device activates a speech recognition engine.

[1054] Output - In the case of voice input, voice data is obtained, and in the case of text input, text data is obtained.

[1055] Step 2:

[1056] Input - Sends speech data to the speech recognition engine.

[1057] Operation - The device uses the speech recognition library (SpeechRecognition) to convert voice data into text data, analyze the voice data, and represent it as string data.

[1058] Output - The audio data is converted to text data.

[1059] Step 3:

[1060] Input - Sends text data to the server.

[1061] Operation - The server runs the emotion analysis engine (text2emotion) on the text data to analyze the emotions in the text data, thereby identifying the type of emotion (e.g., joy, sadness, surprise, etc.) from the input text data.

[1062] Output - Emotion tags are generated for the text data.

[1063] Step 4:

[1064] Input - Receives text data with emotion tags.

[1065] Operation - The server saves the analyzed emotion data and text data in a database, along with the input date and time and emotion tag.

[1066] Output - Text data, sentiment tags, and date / time stored in a database.

[1067] Step 5:

[1068] Input - When a user returns to the store, they request a search for data based on specific criteria (e.g., date and time or type of emotion).

[1069] Operation - The device accepts the user's request and sends it to the server, which then accesses the database and searches for and retrieves the relevant data based on the specified criteria.

[1070] Output - The text data and sentiment tags retrieved based on the criteria.

[1071] Step 6:

[1072] Input - Data retrieved from the server.

[1073] Operation - The device presents the acquired text data and emotion tags to the user. If the user wishes, the text is sent to a speech synthesis engine and played back as speech.

[1074] Output - The text data presented to the user or the audio data played.

[1075] This allows the system to record the customer's memories and emotions and provide personalized service the next time they visit.

[1076] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1077] The present invention is a system that allows users to record their memories and emotions and easily look back on them later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when necessary, and further combines it with an emotion engine that recognizes the user's emotions. Specific embodiments envisioned for implementing the present invention will now be described.

[1078] 1. Data Entry

[1079] Users can input memories and events by voice or text. The interface for accepting this input is a device such as a smartphone or computer. If voice input is selected, the device uses a voice recognition engine to convert the voice into text. If text input is selected, the voice is received as text data.

[1080] 2. Emotion recognition

[1081] As the device receives voice input, the emotion engine analyzes the user's tone and speech patterns to recognize the user's emotions in real time, and this emotion information is sent to the server along with the text data.

[1082] 3. Emotion analysis

[1083] The server analyzes the input text data using an emotion analysis engine. Using the emotion information received from the emotion engine, the server tags the text data with emotions. These emotion tags are useful for later data searches.

[1084] 4. Data storage

[1085] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input, making it possible to search for data based on a specific date or time.

[1086] 5. Data Recall

[1087] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[1088] 6. Data playback

[1089] The device receives the data retrieved from the server and presents it to the user. The user can read the data in text format or play it back in audio format. In the audio format, the text data is sent to a speech synthesis engine, which plays it back as audio data.

[1090] Specific examples

[1091] 1. The user speaks "Memories of summer vacation 2015."

[1092] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[1093] 3. At the same time, the emotion engine analyzes the user's voice tone and generates emotional information such as "feeling happy."

[1094] 4. The server sends the received text data and emotional information to the emotion analysis engine, and obtains the emotion tag "fun."

[1095] 5. The server stores the emotion tag and text data in the database with a date of summer 2015.

[1096] 6. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[1097] 7. The server retrieves the day's data from the database and sends it to the terminal.

[1098] 8. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[1099] In this way, the present invention is a system that allows users to efficiently record past memories and emotions and easily look back on them when necessary. By combining it with an emotion engine, it becomes possible to record and replay even deeper emotional experiences.

[1100] The processing flow will be explained below.

[1101] Step 1:

[1102] Users can input memories and events by voice or text. Users can speak into the device or input text using a keyboard or touch panel.

[1103] Step 2:

[1104] The device receives voice input and first sends it to a speech recognition engine, which converts the voice data into text data. If the input is text, this step is skipped.

[1105] Step 3:

[1106] The device sends the converted text data to the emotion engine, and at the same time, sends the input voice data to the emotion engine, which analyzes the user's voice tone and speech patterns.

[1107] Step 4:

[1108] The emotion engine analyzes the voice tone and speech patterns to obtain the user's emotion information (e.g., "happy," "sad," etc.). The emotion engine then adds this emotion information to the text data.

[1109] Step 5:

[1110] The device sends text data and emotion information to the server, including information such as the input date and time.

[1111] Step 6:

[1112] The server sends the received text data and emotional information to the emotion analysis engine for further detailed emotion analysis, which then tags the emotions in the text data.

[1113] Step 7:

[1114] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input.

[1115] Step 8:

[1116] Years later, if a user wants to relive past memories based on a specific date or emotion, they can use their device to enter a request, for example, "Tell me what happened in the summer of 2015."

[1117] Step 9:

[1118] The device sends the user's request to the server, which includes search criteria such as a specific date or emotion tag.

[1119] Step 10:

[1120] The server searches and retrieves the relevant data from the database, selecting stored data based on specific dates and emotion tags.

[1121] Step 11:

[1122] The server sends the acquired data to the device, which includes text data and emotion tags.

[1123] Step 12:

[1124] The device presents the acquired data to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine, which converts it into audio data and plays it back.

[1125] The above is the specific processing flow of the program in the present invention.

[1126] Example 2

[1127] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1128] In recent years, the importance of recording personal memories and emotions and looking back on them has increased. However, existing systems have difficulty efficiently managing voice and text input and recording detailed information, including emotions. In particular, there is a lack of systems that allow users to easily search and play back memories associated with emotions later. To solve these problems, there is a need for a system that can handle voice and text input, recognize and analyze emotions, and store and play back data in a single step.

[1129] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1130] In this invention, the server includes means for accepting voice or text input, means for converting the accepted voice data into text data using a voice recognition engine, means for utilizing an emotion recognition engine that generates emotional information in real time based on the text data and the user's tone of voice, means for emotionally analyzing the generated emotional information together with the text data and assigning an emotion tag, means for storing the emotionally analyzed text data and emotional information in a database together with date information, means for searching and retrieving the stored data under specific conditions based on a user's request, and means for converting the retrieved data into voice using a text or voice synthesis engine and playing it back. This allows users to efficiently record individual memories and emotions and easily look back on them in a form linked to their emotions.

[1131] "Voice or text input" refers to the act of a user providing information to a system in the form of voice or text.

[1132] A "voice recognition engine" refers to software or hardware for converting voice data into text data.

[1133] "Text data" refers to data that is stored and processed as textual information.

[1134] An "emotion recognition engine" refers to software or hardware that analyzes a user's voice tone, speech patterns, etc., and generates emotional information.

[1135] "Emotion information" refers to data that indicates the user's emotional state according to an emotion recognition engine.

[1136] "Sentiment analysis" refers to the process of identifying and classifying emotions within data based on text data and emotional information.

[1137] An "emotion tag" refers to a label that indicates a specific emotional state, assigned through emotion analysis.

[1138] "Database" refers to a system for structured storage of text data, emotional information, and other related information.

[1139] "Speech synthesis engine" refers to software or hardware for converting text data into speech data.

[1140] "User request" refers to the operation or instruction given by a user to the system to search for or retrieve specific data.

[1141] This invention is a system that efficiently records a user's memories and emotions and allows them to be easily reviewed later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when needed. In addition, by combining it with an emotion engine that recognizes the user's emotions, it becomes possible to record and play back richer emotional experiences.

[1142] Data Entry

[1143] Users can input memories and events using devices such as smartphones or computers, either by voice or text. If voice input is selected, the device uses a voice recognition engine (e.g., a general voice recognition engine) to convert the voice into text data. If text input is selected, the voice is received as text data.

[1144] emotion recognition

[1145] The device simultaneously sends the input voice data to an emotion recognition engine (a common emotion recognition engine) and generates emotion information by analyzing the user's voice tone and speech pattern. This emotion information is then sent to the server along with the text data.

[1146] Emotion analysis

[1147] The server receives the text data and emotional information and analyzes it using an emotion analysis engine. Based on the emotional information received from the emotion recognition engine, it assigns emotional tags to the text data. These emotional tags are useful for later searches.

[1148] Data storage

[1149] The server stores the emotion-analyzed text data and emotion tags in a database along with date information, allowing users to search for data based on specific dates or emotions.

[1150] Recalling Data

[1151] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[1152] Data playback

[1153] The device receives the data retrieved from the server and presents it to the user, who can read it in text format or play it back in audio format using a speech synthesis engine (a common speech synthesis engine).

[1154] Specific examples

[1155] For example, if a user speaks "Memories of summer vacation 2015," the following happens:

[1156] 1. The user uses a smartphone to speak "Memories of summer vacation 2015."

[1157] 2. The device converts the voice into text, creating text data called "Memories of summer vacation 2015."

[1158] 3. At the same time, the emotion recognition engine analyzes the tone of the voice and generates emotional information such as "happy."

[1159] 4. The server sends the received text data and emotional information to the emotion analysis engine, and obtains the emotion tag "fun."

[1160] 5. The server stores this text data and emotion tag in a database along with the date in the summer of 2015.

[1161] 6. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[1162] 7. The server searches the database for the relevant data and sends it to the terminal.

[1163] 8. The device converts the received data into audio format and plays back to the user, "The summer of 2015 was fun."

[1164] In this way, this system allows users to look back on their past memories and emotions more easily and in a richer way. By combining it with an emotion engine, it is possible to play back emotionally rich memories rather than just recording them.

[1165] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1166] Step 1:

[1167] Users can use a device such as a smartphone or computer to input memories and events using voice or text.

[1168] Input: User voice or text data

[1169] Output: Audio data or raw text data

[1170] Specific behavior:

[1171] The user speaks into the microphone on their smartphone about their "memories of summer vacation in 2015."

[1172] The device acquires the voice data and prepares it to be sent to the voice recognition engine.

[1173] Step 2:

[1174] When a voice input is received, the device uses a voice recognition engine to convert the voice data into text data. When a sentence is input, the device receives the text data as is.

[1175] Input: Audio data or text data

[1176] Output: Text data

[1177] Specific behavior:

[1178] The terminal calls a general voice recognition engine and converts the voice data into text data.

[1179] The converted text data "Memories of summer vacation 2015" is obtained.

[1180] Step 3:

[1181] The device analyzes the user's voice tone and speech patterns and generates emotional information using an emotion recognition engine, which is then sent to the server along with the text data.

[1182] Input: Audio data, text data

[1183] Output: Emotional information, text data

[1184] Specific behavior:

[1185] The device uses a general emotion recognition engine to generate emotional information such as "it looks fun" from the voice data.

[1186] The generated emotion information is attached to text data and transmitted to a server.

[1187] Step 4:

[1188] The server performs emotion analysis based on the received text data and emotion information, and assigns emotion tags to the text data.

[1189] Input: Text data, emotion information

[1190] Output: Text data with emotion tags

[1191] Specific behavior:

[1192] The server sends the text data and the emotional information "it looks fun" to the emotion analysis engine.

[1193] The server receives the emotion tag "fun" from the emotion analysis engine and assigns it to the text data.

[1194] Step 5:

[1195] The server stores the emotion-tagged text data and emotion information together with date information in a database.

[1196] Input: Emotion-tagged text data, emotion information, date information

[1197] Output: Save to database

[1198] Specific behavior:

[1199] The server stores the text data "Memories of summer vacation 2015" and the emotion tag "It was fun" in a database along with date information.

[1200] Step 6:

[1201] When a user wants to look back on past memories based on a specific date or emotion, the user inputs a request using the terminal.

[1202] Input: Request (e.g. "Tell me what happened in the summer of 2015")

[1203] Output: A request to retrieve the corresponding data

[1204] Specific behavior:

[1205] A few years later, a user uses a smartphone app to enter a request: "Tell me what happened in the summer of 2015."

[1206] The terminal sends this request to the server.

[1207] Step 7:

[1208] The server searches for and retrieves data corresponding to the request from the database.

[1209] Input: Request, Database

[1210] Output: Acquisition of relevant data

[1211] Specific behavior:

[1212] The server searches the database for and retrieves data tagged with "Memories of summer vacation 2015" and the emotion "It was fun."

[1213] The server transmits the acquired data to the terminal.

[1214] Step 8:

[1215] The terminal presents the received data to the user, who can read it in text format or play it in audio format.

[1216] Input: Retrieved data

[1217] Output: Text display or audio playback

[1218] Specific behavior:

[1219] The device displays the received data in text format, telling the user, "The summer of 2015 was fun."

[1220] The device uses a speech synthesis engine to convert the text data into audio data, which is then played through the speaker.

[1221] (Application example 2)

[1222] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1223] In today's busy daily lives, users often miss opportunities to record their memories and emotions. Furthermore, the time and effort required to review recorded memories and emotions makes it difficult to effectively utilize the accumulated data. Furthermore, the lack of relevant information based on the user's past emotional experiences makes it difficult to provide appropriate information tailored to the user's needs.

[1224] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting input by voice or text, means for converting voice data into text data, means for performing emotion analysis on the input text data, means for saving the emotion-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing the retrieved data in text or voice, and means for recommending related information based on the saved emotion data. This not only enables users to efficiently record past memories and emotions and easily look back on them when necessary, but also enables recommendations of related information based on past emotional experiences.

[1225] "Means for accepting voice or text input" refers to an interface for receiving input from a user in the form of voice or text and processing it appropriately.

[1226] The "means for converting voice data into text data" is a function for converting voice data into text data using voice recognition technology.

[1227] The "means for performing sentiment analysis on input text data" is an engine that analyzes the content of the text data and determines the polarity and intensity of sentiment from that content.

[1228] The "means for saving emotion-analyzed text data in a database" is a function for recording and managing the analyzed emotion information and text data in a database.

[1229] "Means for retrieving stored data based on specific conditions" refers to a function that searches for and retrieves necessary information from a database based on a user request or set conditions.

[1230] "Means for reproducing the acquired data in text or audio" refers to a function for presenting the extracted data in a format that is easy for the user to understand, and includes text display and audio reproduction.

[1231] The "means for recommending related information based on stored emotion data" is a function that suggests related items and information to the user based on the recorded emotion data of the user.

[1232] The present invention provides a system that allows users to record emotions and memories felt during their virtual store experience and easily review them later. This system is implemented using devices such as smartphones and head-mounted displays.

[1233] 1. Data Entry

[1234] Users can input their experiences and emotions in the virtual store by voice. The interface for this input is a device such as a smartphone or head-mounted display. If voice input is selected, the device uses a voice recognition engine to convert the voice into text.

[1235] 2. Emotion recognition

[1236] As the device receives voice input, the emotion analysis engine analyzes the user's tone and speech patterns to recognize the user's emotions in real time, and this emotion information is sent to the server along with the text data.

[1237] 3. Emotion analysis

[1238] The server analyzes the input text data using a sentiment analysis engine. Using the emotional information received from the sentiment analysis engine, the server tags the text data with emotions. These emotional tags are useful for later data searches.

[1239] 4. Data storage

[1240] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input, making it possible to search for data based on a specific date or time.

[1241] 5. Data Recall

[1242] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[1243] 6. Data playback

[1244] The device receives the data from the server and presents it to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine and played back as audio data.

[1245] 7. Recommendation of related information

[1246] Based on the stored emotional data, the system recommends information related to past emotional experiences to the user. For example, the system has the function of recommending new items based on items that the user found enjoyable in the past.

[1247] Hardware and software used

[1248] Voice input devices: smartphones, head-mounted displays

[1249] Speech recognition engine: speech_recognition

[1250] Sentiment Analysis Engine: TextBlob Library

[1251] Database: SQLite

[1252] Server: General web server software (e.g. Flask, Django)

[1253] Specific examples

[1254] For example, suppose a user purchases new running shoes in a virtual store on October 1, 2023, and wants to record the excitement they felt. The user speaks into a microphone, saying, "I'm so happy I bought my new running shoes." The system recognizes this speech as text and records a positive sentiment polarity (e.g., 0.8). This data, along with the date, is stored in a database, and if the user later searches for "happy events in October 2023," the record will appear.

[1255] Prompt Sentence Examples

[1256] I was very happy to purchase new running shoes on October 1, 2023. My emotional polarity is 0.8.

[1257] In this way, users can efficiently record their past emotions and memories, enriching their shopping experience within the virtual store.

[1258] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1259] Step 1: User provides voice or text input

[1260] Specific behavior:

[1261] Users use devices such as smartphones or head-mounted displays to input their experiences and emotions in the virtual store by voice or text. For voice input, they use the device's microphone input function, and for text input, they use an on-screen input form.

[1262] Input: User voice or text data

[1263] Output: Raw audio or text data

[1264] Step 2: The device converts the audio data into text data.

[1265] Specific behavior:

[1266] The device uses a speech recognition engine (e.g., the speech_recognition library) to convert the user's voice input into text data. This process involves inputting the voice data into a speech recognition model, which then outputs the speech content as a string of characters.

[1267] Input: User's voice data

[1268] Output: Data converted to text

[1269] Step 3: The device sends the entered text to the server

[1270] Specific behavior:

[1271] The terminal transmits the converted text data and its meta information (e.g., input date and time) to the server via a network communication protocol (e.g., HTTP).

[1272] Input: Text data, input date and time

[1273] Output: Text data and meta information sent to the server

[1274] Step 4: The server inputs the text data into the sentiment analysis engine

[1275] Specific behavior:

[1276] The server inputs the received text data into a sentiment analysis engine (e.g., TextBlob library) to analyze the sentiment polarity (positive, negative, neutral) of each text. The analysis results include sentiment polarity values ​​and sentiment-related tags.

[1277] Input: Text data

[1278] Output: Sentiment polarity value and sentiment tag

[1279] Step 5: The server stores the analysis results in a database

[1280] Specific behavior:

[1281] The server stores the sentiment polarity values, text data, and meta information (e.g., input date and time) in a database (e.g., SQLite). During this storage process, each data item is inserted into the appropriate database table.

[1282] Input: Sentiment polarity value, text data, input date and time

[1283] Output: Sentiment analysis data stored in a database

[1284] Step 6: User requests historical emotion data based on specific criteria

[1285] Specific behavior:

[1286] A user uses a device to request historical emotion data based on a specific date and emotion. The request is sent to the server as a search query containing the specified criteria (e.g., specific date, emotion tag).

[1287] Input: Search criteria (e.g. date, emotion tag)

[1288] Output: Search request sent to the server

[1289] Step 7: The server searches and retrieves the relevant data from the database.

[1290] Specific behavior:

[1291] The server searches and retrieves the corresponding emotion data from the database based on the received search request. In this process, it filters the data that matches the search criteria and generates the data to be sent back to the user.

[1292] Input: Search criteria (e.g. date, emotion tag)

[1293] Output: Emotion data as search results

[1294] Step 8: Play back the data retrieved by the device as text or audio

[1295] Specific behavior:

[1296] The device presents the emotion data received from the server to the user, who can read the data in text format or play it back in audio format using a speech synthesis engine (e.g., a TTS engine).

[1297] Input: Emotion data as search results

[1298] Output: Text display or audio playback data

[1299] Step 9: The server generates a prompt to recommend related information and presents it to the user.

[1300] Specific behavior:

[1301] The server generates prompts to recommend related information based on the stored emotional data. For example, it can recommend new items based on items that users have previously found enjoyable. The generated prompts are input into a generative AI model and presented to the user as recommendations.

[1302] Input: Emotional data, past emotional experiences

[1303] Output: Prompt statement with relevant information

[1304] Prompt Sentence Examples

[1305] I was very happy to purchase new running shoes on October 1, 2023. My emotional polarity is 0.8.

[1306] This prompt allows the user to reflect on past experiences while simultaneously receiving new, relevant information.

[1307] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1308] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1309] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1310] [Fourth embodiment]

[1311] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1312] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1313] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1314] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1315] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1316] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1317] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1318] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1319] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1320] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1321] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1322] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1323] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1324] The present invention is a system that allows users to record their memories and emotions and easily look back on them later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when necessary. Specific embodiments envisioned for implementing the present invention will now be described.

[1325] 1. Data Entry

[1326] Users can input memories and events by voice or text. The interface for accepting this input is a device such as a smartphone or computer. If voice input is selected, the device uses a voice recognition engine to convert the voice into text. If text input is selected, the voice is received as text data.

[1327] 2. Emotion analysis

[1328] The server performs sentiment analysis on the input text data. The sentiment analysis engine analyzes the emotions contained in the text data and identifies the type of emotion (e.g., joy, sadness, surprise, etc.). This sentiment tag is useful for later data search.

[1329] 3. Data storage

[1330] The server stores the text data in a database along with the analyzed emotion data, along with the date and time of entry, allowing data to be searched for based on a specific date or time.

[1331] 4. Data Recall

[1332] When a user wants to reminisce about past memories based on a specific date or emotion, they send a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[1333] 5. Data playback

[1334] The device receives the data retrieved from the server and presents it to the user. The user can read the data in text format or play it back in audio format. In the audio format, the text data is sent to a speech synthesis engine, which plays it back as audio data.

[1335] Specific examples

[1336] 1. The user speaks "Memories of summer vacation 2015."

[1337] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[1338] 3. The server sends the received text data to the emotion analysis engine and obtains the emotion tag "fun."

[1339] 4. The server stores the emotion tag and text data in the database with the date summer 2015.

[1340] 5. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[1341] 6. The server retrieves the day's data from the database and sends it to the terminal.

[1342] 7. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[1343] Thus, the present invention is a system that allows users to efficiently record past memories and emotions and easily look back on them when necessary.

[1344] The processing flow will be explained below.

[1345] Step 1:

[1346] Users input memories and events using voice or text. For voice input, users speak into the device, and for text input, users input using a keyboard or touch panel.

[1347] Step 2:

[1348] The terminal receives the user's voice input and sends it to the voice recognition engine. The voice recognition engine converts the voice data into text data and obtains the text data. If the input is text, this step is skipped and the process proceeds to the next step.

[1349] Step 3:

[1350] The device sends the acquired text data to the server, where information such as the input date and time is also added to the text data.

[1351] Step 4:

[1352] The server analyzes the received text data using a sentiment analysis engine, which identifies the emotions in the text data and generates corresponding sentiment tags (e.g., "happy," "sad," etc.).

[1353] Step 5:

[1354] The server stores the emotion-analyzed text data and emotion tags in a database, along with the input date and time.

[1355] Step 6:

[1356] Years later, if a user wants to relive past memories based on a specific date or emotion, they can use their device to enter a request, for example, "Tell me what happened in the summer of 2015."

[1357] Step 7:

[1358] The device sends the user's request to the server, which includes search criteria such as a specific date or emotion tag.

[1359] Step 8:

[1360] The server searches and retrieves the relevant data from the database, selecting stored data based on specific dates and emotion tags.

[1361] Step 9:

[1362] The server sends the acquired data to the device, which includes text data and emotion tags.

[1363] Step 10:

[1364] The device presents the acquired data to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine, which converts it into audio data and plays it back.

[1365] The above is the specific processing flow of the program in the present invention.

[1366] Example 1

[1367] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1368] In today's world, systems that allow users to record and easily recall their daily memories and emotions are extremely important. However, conventional systems have struggled to integrate a series of operations, such as converting voice input to text, analyzing emotions, and storing, searching, and playing back data. As a result, it has been difficult for users to efficiently search and play back past memories, and data storage and management has been cumbersome. Therefore, there is a need for a system that can accept voice and text input, perform emotion analysis, and easily store, search, and play back data.

[1369] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1370] In this invention, the server includes means for accepting input by voice or text, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for saving the sentiment-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing the retrieved data in text or voice, means for a user to send a request using a terminal and search for data based on the conditions, and means for the terminal to retrieve data based on the request sent to the server and present it to the user. This allows the user to input by voice or text and easily save, search, and play sentiment-analyzed data.

[1371] "Voice input" refers to the voice data collected through a microphone and input into the system from what the user says.

[1372] "Text input" refers to a method in which a user directly inputs text using a keyboard or touch screen.

[1373] A "speech recognition engine" is software or algorithms used to analyze collected voice data and convert it into corresponding text data.

[1374] An "emotion analysis engine" is software or an algorithm that identifies the type and intensity of emotions from input text data and assigns emotion tags.

[1375] A "database" is a system or software for efficiently storing, managing, and retrieving structured data.

[1376] "Storing data" refers to the act of recording the input text data and its accompanying information in a database.

[1377] "Data search" refers to the operation a user performs to retrieve stored data based on specific criteria (e.g., date, emotion tag).

[1378] A "speech synthesis engine" is software or algorithm that analyzes text data and generates and plays back speech that resembles a human voice.

[1379] A "request" is an instruction from a user requesting the system to perform some operation.

[1380] A "terminal" is a device operated by a user (e.g., a smartphone, a computer).

[1381] "Server" means a central computer system that processes and manages data.

[1382] A "user" is a person who uses the system.

[1383] The present invention is a system that allows users to record their memories and emotions and easily review them later. The system includes a set of means for accepting voice or text input, storing it, and playing it back when needed.

[1384] Data Entry

[1385] Users can input memories and events using devices such as smartphones or computers. They can choose to input via voice or text. In the case of voice input, the device uses a speech recognition engine (e.g., automatic speech recognition software) to convert the speech into text. In the case of text input, the speech is accepted as text data as is.

[1386] Voice Recognition

[1387] If voice input is selected, the device will use automatic speech recognition software to convert speech to text. For example, if a user speaks "Memories of summer vacation 2015," the device will convert this speech to text.

[1388] Emotion analysis

[1389] The server passes the received text data to an emotion analysis engine (e.g., natural language processing software) to analyze the emotions contained in the text. As a result of the analysis, an emotion tag (e.g., "joy," "sadness," etc.) is assigned.

[1390] Data storage

[1391] The server stores the text data, including the analysis results, in a database (e.g., a relational database system), along with the input date and time.

[1392] Recalling Data

[1393] When a user wants to reminisce about past memories, they use their device to send a request based on a specific condition to the server, for example, "Tell me what happened in the summer of 2015."

[1394] Data Search and Playback

[1395] The server searches a database based on the request and retrieves the relevant data. The server then sends the retrieved data to the device. The device plays the retrieved data to the user in text or audio format. In the audio format, the text data is sent to a speech synthesis engine (e.g., text-to-speech software) and played as audio data.

[1396] Specific examples

[1397] Below are some examples of specific prompt sentences.

[1398] 1. The user speaks "Memories of summer vacation 2015."

[1399] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[1400] 3. The server sends the received text data to the emotion analysis engine and obtains the emotion tag "fun."

[1401] 4. The server stores the emotion tag and text data in the database with the date summer 2015.

[1402] 5. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[1403] 6. The server retrieves the day's data from the database and sends it to the terminal.

[1404] 7. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[1405] This system allows users to efficiently record past memories and emotions and easily look back on them when needed. In addition to the convenience of voice and text input, it also enables effective data retrieval through emotion analysis.

[1406] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1407] Step 1: Data entry

[1408] Users can input memories and events using devices such as smartphones or computers. Here, users can choose between voice input or text input. For voice input, users use the device's microphone and tap the voice input button on the app to speak. For text input, users enter text using a keyboard or touchscreen.

[1409] Input: Audio or text data

[1410] Output: Converted text data (in the case of speech input) or raw text data (in the case of sentence input)

[1411] Step 2: Voice Recognition

[1412] When voice data is input, the terminal uses a voice recognition engine (e.g., automatic speech recognition software) to convert the voice into text. The converted text data is displayed on the screen for the user to confirm. If necessary, the user can correct this text.

[1413] Input: Audio data

[1414] Output: Converted text data

[1415] Step 3: Send text data

[1416] The terminal sends the converted text data or directly entered text to the server, where the data is converted into JSON format and sent securely to the server using the HTTPS protocol.

[1417] Input: Text data (entered by the user)

[1418] Output: Text data sent to the server

[1419] Step 4: Sentiment Analysis

[1420] The server passes the received text data to a sentiment analysis engine (e.g., natural language processing software) to analyze the sentiment contained in the text. Keywords and sentence structure within the text data are used for the analysis, and sentiment is tagged as the analysis result.

[1421] Input: Text data sent to the server

[1422] Output: Text data with emotion tags added

[1423] Step 5: Save your data

[1424] The server stores the emotion-tagged text data in a database (e.g., a relational database system), along with the input date and time and the user ID.

[1425] Input: Text data with emotion tags added

[1426] Output: Records stored in a database

[1427] Step 6: Calling the data

[1428] When a user wants to reminisce about past memories, they use their device to send a request based on specific criteria to the server, such as "Tell me what happened in the summer of 2015."

[1429] Input: Request data sent from the user's device (conditional search query)

[1430] Output: Request data sent to the server

[1431] Step 7: Search for data

[1432] The server searches the database based on the request received and extracts the relevant data, using SQL queries to extract records that match the criteria.

[1433] Input: Request data sent to the server

[1434] Output: Extracted data that matches the conditions

[1435] Step 8: Replaying the Data

[1436] The device presents the data retrieved from the server to the user. This data can be displayed in text format or played back as audio. In the latter case, the text data is sent to a speech synthesis engine (e.g., text-to-speech software) and played back as audio data.

[1437] Input: Extracted data that matches the conditions

[1438] Output: Text or audio data presented to the user

[1439] (Application example 1)

[1440] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1441] The problem is that there is currently no system that records the memories and emotions experienced by customers when they visit a store, efficiently manages this information, and allows them to offer special services the next time they visit. In particular, there is a need for a system that aims to improve customer experience in physical stores, which can analyze emotions from voice and text input and use the data to improve services the next time.

[1442] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1443] In this invention, the server includes means for accepting voice or text input, means for converting voice data into text data, means for performing sentiment analysis on the input text data, means for saving the sentiment-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing back the retrieved data in text or voice, and means for recording a customer's voice or text input when they visit the store and retrieving and playing back data for providing special service the next time they visit. This makes it possible to effectively record and manage customer memories and emotions and to provide special service based on that data the next time they visit.

[1444] "Voice or text input" is the process by which a user provides memories or emotions to a system as a recording or text.

[1445] "Converting voice data into text data" refers to a technique for analyzing recorded voice information and expressing its contents as character string data.

[1446] "Performing emotion analysis on input text data" refers to the process of analyzing the character string data to identify the emotion contained therein and determining the type of emotion (for example, joy, sadness, surprise, etc.).

[1447] "Saving emotion-analyzed text data in a database" refers to recording character string data in a storage device together with information on the analyzed emotion.

[1448] "Retrieving stored data based on specific conditions" refers to the operation of searching and retrieving recorded data according to conditions such as a specific date and time or type of emotion.

[1449] "Reproducing retrieved data as text or voice" means displaying the retrieved data as text or reproducing it as voice using voice synthesis technology.

[1450] "Recording customer voice or text input when they visit the store, and acquiring and playing back data to provide special services the next time they visit" refers to the process of recording the memories and emotions entered by customers when they visit the store, and proposing and providing special experiences and services based on that data the next time they visit.

[1451] This invention is a system for improving customer experience in physical stores. It records the memories and emotions experienced by customers when they visit the store, and provides special services based on that data the next time they visit.

[1452] The system accepts voice or text input, analyzes the data emotionally, and stores it in a database. It then retrieves the stored data based on specific conditions and plays it back in text or voice format, providing personalized service to customers.

[1453] Hardware and Software Used

[1454] Hardware: Tablets, smartphones

[1455] Software: Flask (web application framework), SQLite (database), SpeechRecognition (speech recognition library), text2emotion (emotion analysis library)

[1456] Step Description

[1457] 1. Voice or text input:

[1458] Users use a tablet or smartphone to input the events and emotions they experienced in real time using voice or text.

[1459] In the case of voice input, the terminal uses a voice recognition engine to convert the voice data into text data.

[1460] 2. Emotion analysis:

[1461] The server performs sentiment analysis on the input text data.

[1462] Use the text2emotion library to analyze emotions contained in text data and identify the type of emotion.

[1463] 3. Data storage:

[1464] The server stores the text data together with the analyzed emotion data in a database.

[1465] The data is tagged with the date and time it was entered and an emotion tag.

[1466] 4. Data retrieval:

[1467] The next time the user visits the store, they can send a request to search for past memories and emotions via the terminal.

[1468] The server searches and acquires the relevant data from the database.

[1469] 5. Data Regeneration:

[1470] The terminal presents the acquired text data to the user.

[1471] If the user so desires, the text data is sent to a speech synthesis engine and reproduced as voice data.

[1472] Specific examples

[1473] Below are some examples of specific prompt sentences.

[1474] Example of an input prompt:

[1475] "Please record your memories from the travel destinations you visited in the fall of 2019."

[1476] "Please tell us about a fun memory you had the last time you visited us."

[1477] Example of a retrieval prompt:

[1478] "Tell me about your memories of the last time you visited us."

[1479] "Show me the record of your most inspiring visit."

[1480] In this way, by implementing this invention, it is possible to effectively record and manage customer memories and emotions, and provide special services based on that data the next time the customer visits the store. This is an effective means of improving the customer experience in physical stores and increasing customer satisfaction.

[1481] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1482] Step 1:

[1483] Input - Users use a tablet or smartphone to input their experiences and memories using voice or text.

[1484] Operation - The device displays an interface that accepts input from the user. If the user speaks, the device activates a speech recognition engine.

[1485] Output - In the case of voice input, voice data is obtained, and in the case of text input, text data is obtained.

[1486] Step 2:

[1487] Input - Sends speech data to the speech recognition engine.

[1488] Operation - The device uses the speech recognition library (SpeechRecognition) to convert voice data into text data, analyze the voice data, and represent it as string data.

[1489] Output - The audio data is converted to text data.

[1490] Step 3:

[1491] Input - Sends text data to the server.

[1492] Operation - The server runs the emotion analysis engine (text2emotion) on the text data to analyze the emotions in the text data, thereby identifying the type of emotion (e.g., joy, sadness, surprise, etc.) from the input text data.

[1493] Output - Emotion tags are generated for the text data.

[1494] Step 4:

[1495] Input - Receives text data with emotion tags.

[1496] Operation - The server saves the analyzed emotion data and text data in a database, along with the input date and time and emotion tag.

[1497] Output - Text data, sentiment tags, and date / time stored in a database.

[1498] Step 5:

[1499] Input - When a user returns to the store, they request a search for data based on specific criteria (e.g., date and time or type of emotion).

[1500] Operation - The device accepts the user's request and sends it to the server, which then accesses the database and searches for and retrieves the relevant data based on the specified criteria.

[1501] Output - The text data and sentiment tags retrieved based on the criteria.

[1502] Step 6:

[1503] Input - Data retrieved from the server.

[1504] Operation - The device presents the acquired text data and emotion tags to the user. If the user wishes, the text is sent to a speech synthesis engine and played back as speech.

[1505] Output - The text data presented to the user or the audio data played.

[1506] This allows the system to record the customer's memories and emotions and provide personalized service the next time they visit.

[1507] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1508] The present invention is a system that allows users to record their memories and emotions and easily look back on them later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when necessary, and further combines it with an emotion engine that recognizes the user's emotions. Specific embodiments envisioned for implementing the present invention will now be described.

[1509] 1. Data Entry

[1510] Users can input memories and events by voice or text. The interface for accepting this input is a device such as a smartphone or computer. If voice input is selected, the device uses a voice recognition engine to convert the voice into text. If text input is selected, the voice is received as text data.

[1511] 2. Emotion recognition

[1512] As the device receives voice input, the emotion engine analyzes the user's tone and speech patterns to recognize the user's emotions in real time, and this emotion information is sent to the server along with the text data.

[1513] 3. Emotion analysis

[1514] The server analyzes the input text data using an emotion analysis engine. Using the emotion information received from the emotion engine, the server tags the text data with emotions. These emotion tags are useful for later data searches.

[1515] 4. Data storage

[1516] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input, making it possible to search for data based on a specific date or time.

[1517] 5. Data Recall

[1518] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[1519] 6. Data playback

[1520] The device receives the data retrieved from the server and presents it to the user. The user can read the data in text format or play it back in audio format. In the audio format, the text data is sent to a speech synthesis engine, which plays it back as audio data.

[1521] Specific examples

[1522] 1. The user speaks "Memories of summer vacation 2015."

[1523] 2. The device converts the voice into text and sends it to the server as "Memories of Summer Vacation 2015."

[1524] 3. At the same time, the emotion engine analyzes the user's voice tone and generates emotional information such as "feeling happy."

[1525] 4. The server sends the received text data and emotional information to the emotion analysis engine, and obtains the emotion tag "fun."

[1526] 5. The server stores the emotion tag and text data in the database with a date of summer 2015.

[1527] 6. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[1528] 7. The server retrieves the day's data from the database and sends it to the terminal.

[1529] 8. The device plays the received data in audio format, telling the user, "The summer of 2015 was fun."

[1530] In this way, the present invention is a system that allows users to efficiently record past memories and emotions and easily look back on them when necessary. By combining it with an emotion engine, it becomes possible to record and replay even deeper emotional experiences.

[1531] The processing flow will be explained below.

[1532] Step 1:

[1533] Users can input memories and events by voice or text. Users can speak into the device or input text using a keyboard or touch panel.

[1534] Step 2:

[1535] The device receives voice input and first sends it to a speech recognition engine, which converts the voice data into text data. If the input is text, this step is skipped.

[1536] Step 3:

[1537] The device sends the converted text data to the emotion engine, and at the same time, sends the input voice data to the emotion engine, which analyzes the user's voice tone and speech patterns.

[1538] Step 4:

[1539] The emotion engine analyzes the voice tone and speech patterns to obtain the user's emotion information (e.g., "happy," "sad," etc.). The emotion engine then adds this emotion information to the text data.

[1540] Step 5:

[1541] The device sends text data and emotion information to the server, including information such as the input date and time.

[1542] Step 6:

[1543] The server sends the received text data and emotional information to the emotion analysis engine for further detailed emotion analysis, which then tags the emotions in the text data.

[1544] Step 7:

[1545] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input.

[1546] Step 8:

[1547] Years later, if a user wants to relive past memories based on a specific date or emotion, they can use their device to enter a request, for example, "Tell me what happened in the summer of 2015."

[1548] Step 9:

[1549] The device sends the user's request to the server, which includes search criteria such as a specific date or emotion tag.

[1550] Step 10:

[1551] The server searches and retrieves the relevant data from the database, selecting stored data based on specific dates and emotion tags.

[1552] Step 11:

[1553] The server sends the acquired data to the device, which includes text data and emotion tags.

[1554] Step 12:

[1555] The device presents the acquired data to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine, which converts it into audio data and plays it back.

[1556] The above is the specific processing flow of the program in the present invention.

[1557] Example 2

[1558] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1559] In recent years, the importance of recording personal memories and emotions and looking back on them has increased. However, existing systems have difficulty efficiently managing voice and text input and recording detailed information, including emotions. In particular, there is a lack of systems that allow users to easily search and play back memories associated with emotions later. To solve these problems, there is a need for a system that can handle voice and text input, recognize and analyze emotions, and store and play back data in a single step.

[1560] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1561] In this invention, the server includes means for accepting voice or text input, means for converting the accepted voice data into text data using a voice recognition engine, means for utilizing an emotion recognition engine that generates emotional information in real time based on the text data and the user's tone of voice, means for emotionally analyzing the generated emotional information together with the text data and assigning an emotion tag, means for storing the emotionally analyzed text data and emotional information in a database together with date information, means for searching and retrieving the stored data under specific conditions based on a user's request, and means for converting the retrieved data into voice using a text or voice synthesis engine and playing it back. This allows users to efficiently record individual memories and emotions and easily look back on them in a form linked to their emotions.

[1562] "Voice or text input" refers to the act of a user providing information to a system in the form of voice or text.

[1563] A "voice recognition engine" refers to software or hardware for converting voice data into text data.

[1564] "Text data" refers to data that is stored and processed as textual information.

[1565] An "emotion recognition engine" refers to software or hardware that analyzes a user's voice tone, speech patterns, etc., and generates emotional information.

[1566] "Emotion information" refers to data that indicates the user's emotional state according to an emotion recognition engine.

[1567] "Sentiment analysis" refers to the process of identifying and classifying emotions within data based on text data and emotional information.

[1568] An "emotion tag" refers to a label that indicates a specific emotional state, assigned through emotion analysis.

[1569] "Database" refers to a system for structured storage of text data, emotional information, and other related information.

[1570] "Speech synthesis engine" refers to software or hardware for converting text data into speech data.

[1571] "User request" refers to the operation or instruction given by a user to the system to search for or retrieve specific data.

[1572] This invention is a system that efficiently records a user's memories and emotions and allows them to be easily reviewed later. This system includes a series of means for accepting voice or text input, saving it, and playing it back when needed. In addition, by combining it with an emotion engine that recognizes the user's emotions, it becomes possible to record and play back richer emotional experiences.

[1573] Data Entry

[1574] Users can input memories and events using devices such as smartphones or computers, either by voice or text. If voice input is selected, the device uses a voice recognition engine (e.g., a general voice recognition engine) to convert the voice into text data. If text input is selected, the voice is received as text data.

[1575] emotion recognition

[1576] The device simultaneously sends the input voice data to an emotion recognition engine (a common emotion recognition engine) and generates emotion information by analyzing the user's voice tone and speech pattern. This emotion information is then sent to the server along with the text data.

[1577] Emotion analysis

[1578] The server receives the text data and emotional information and analyzes it using an emotion analysis engine. Based on the emotional information received from the emotion recognition engine, it assigns emotional tags to the text data. These emotional tags are useful for later searches.

[1579] Data storage

[1580] The server stores the emotion-analyzed text data and emotion tags in a database along with date information, allowing users to search for data based on specific dates or emotions.

[1581] Recalling Data

[1582] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[1583] Data playback

[1584] The device receives the data retrieved from the server and presents it to the user, who can read it in text format or play it back in audio format using a speech synthesis engine (a common speech synthesis engine).

[1585] Specific examples

[1586] For example, if a user speaks "Memories of summer vacation 2015," the following happens:

[1587] 1. The user uses a smartphone to speak "Memories of summer vacation 2015."

[1588] 2. The device converts the voice into text, creating text data called "Memories of summer vacation 2015."

[1589] 3. At the same time, the emotion recognition engine analyzes the tone of the voice and generates emotional information such as "happy."

[1590] 4. The server sends the received text data and emotional information to the emotion analysis engine, and obtains the emotion tag "fun."

[1591] 5. The server stores this text data and emotion tag in a database along with the date in the summer of 2015.

[1592] 6. A few years later, the user requests on their device, "Tell me what happened in the summer of 2015."

[1593] 7. The server searches the database for the relevant data and sends it to the terminal.

[1594] 8. The device converts the received data into audio format and plays back to the user, "The summer of 2015 was fun."

[1595] In this way, this system allows users to look back on their past memories and emotions more easily and in a richer way. By combining it with an emotion engine, it is possible to play back emotionally rich memories rather than just recording them.

[1596] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1597] Step 1:

[1598] Users can use a device such as a smartphone or computer to input memories and events using voice or text.

[1599] Input: User voice or text data

[1600] Output: Audio data or raw text data

[1601] Specific behavior:

[1602] The user speaks into the microphone on their smartphone about their "memories of summer vacation in 2015."

[1603] The device acquires the voice data and prepares it to be sent to the voice recognition engine.

[1604] Step 2:

[1605] When a voice input is received, the device uses a voice recognition engine to convert the voice data into text data. When a sentence is input, the device receives the text data as is.

[1606] Input: Audio data or text data

[1607] Output: Text data

[1608] Specific behavior:

[1609] The terminal calls a general voice recognition engine and converts the voice data into text data.

[1610] The converted text data "Memories of summer vacation 2015" is obtained.

[1611] Step 3:

[1612] The device analyzes the user's voice tone and speech patterns and generates emotional information using an emotion recognition engine, which is then sent to the server along with the text data.

[1613] Input: Audio data, text data

[1614] Output: Emotional information, text data

[1615] Specific behavior:

[1616] The device uses a general emotion recognition engine to generate emotional information such as "it looks fun" from the voice data.

[1617] The generated emotion information is attached to text data and transmitted to a server.

[1618] Step 4:

[1619] The server performs emotion analysis based on the received text data and emotion information, and assigns emotion tags to the text data.

[1620] Input: Text data, emotion information

[1621] Output: Text data with emotion tags

[1622] Specific behavior:

[1623] The server sends the text data and the emotional information "it looks fun" to the emotion analysis engine.

[1624] The server receives the emotion tag "fun" from the emotion analysis engine and assigns it to the text data.

[1625] Step 5:

[1626] The server stores the emotion-tagged text data and emotion information together with date information in a database.

[1627] Input: Emotion-tagged text data, emotion information, date information

[1628] Output: Save to database

[1629] Specific behavior:

[1630] The server stores the text data "Memories of summer vacation 2015" and the emotion tag "It was fun" in a database along with date information.

[1631] Step 6:

[1632] When a user wants to look back on past memories based on a specific date or emotion, the user inputs a request using the terminal.

[1633] Input: Request (e.g. "Tell me what happened in the summer of 2015")

[1634] Output: A request to retrieve the corresponding data

[1635] Specific behavior:

[1636] A few years later, a user uses a smartphone app to enter a request: "Tell me what happened in the summer of 2015."

[1637] The terminal sends this request to the server.

[1638] Step 7:

[1639] The server searches for and retrieves data corresponding to the request from the database.

[1640] Input: Request, Database

[1641] Output: Acquisition of relevant data

[1642] Specific behavior:

[1643] The server searches the database for and retrieves data tagged with "Memories of summer vacation 2015" and the emotion "It was fun."

[1644] The server transmits the acquired data to the terminal.

[1645] Step 8:

[1646] The terminal presents the received data to the user, who can read it in text format or play it in audio format.

[1647] Input: Retrieved data

[1648] Output: Text display or audio playback

[1649] Specific behavior:

[1650] The device displays the received data in text format, telling the user, "The summer of 2015 was fun."

[1651] The device uses a speech synthesis engine to convert the text data into audio data, which is then played through the speaker.

[1652] (Application example 2)

[1653] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1654] In today's busy daily lives, users often miss opportunities to record their memories and emotions. Furthermore, the time and effort required to review recorded memories and emotions makes it difficult to effectively utilize the accumulated data. Furthermore, the lack of relevant information based on the user's past emotional experiences makes it difficult to provide appropriate information tailored to the user's needs.

[1655] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting input by voice or text, means for converting voice data into text data, means for performing emotion analysis on the input text data, means for saving the emotion-analyzed text data in a database, means for retrieving the saved data based on specific conditions, means for playing the retrieved data in text or voice, and means for recommending related information based on the saved emotion data. This not only enables users to efficiently record past memories and emotions and easily look back on them when necessary, but also enables recommendations of related information based on past emotional experiences.

[1656] "Means for accepting voice or text input" refers to an interface for receiving input from a user in the form of voice or text and processing it appropriately.

[1657] The "means for converting voice data into text data" is a function for converting voice data into text data using voice recognition technology.

[1658] The "means for performing sentiment analysis on input text data" is an engine that analyzes the content of the text data and determines the polarity and intensity of sentiment from that content.

[1659] The "means for saving emotion-analyzed text data in a database" is a function for recording and managing the analyzed emotion information and text data in a database.

[1660] "Means for retrieving stored data based on specific conditions" refers to a function that searches for and retrieves necessary information from a database based on a user request or set conditions.

[1661] "Means for reproducing the acquired data in text or audio" refers to a function for presenting the extracted data in a format that is easy for the user to understand, and includes text display and audio reproduction.

[1662] The "means for recommending related information based on stored emotion data" is a function that suggests related items and information to the user based on the recorded emotion data of the user.

[1663] The present invention provides a system that allows users to record emotions and memories felt during their virtual store experience and easily review them later. This system is implemented using devices such as smartphones and head-mounted displays.

[1664] 1. Data Entry

[1665] Users can input their experiences and emotions in the virtual store by voice. The interface for this input is a device such as a smartphone or head-mounted display. If voice input is selected, the device uses a voice recognition engine to convert the voice into text.

[1666] 2. Emotion recognition

[1667] As the device receives voice input, the emotion analysis engine analyzes the user's tone and speech patterns to recognize the user's emotions in real time, and this emotion information is sent to the server along with the text data.

[1668] 3. Emotion analysis

[1669] The server analyzes the input text data using a sentiment analysis engine. Using the emotional information received from the sentiment analysis engine, the server tags the text data with emotions. These emotional tags are useful for later data searches.

[1670] 4. Data storage

[1671] The server stores the emotion-analyzed text data and emotion tags in a database, along with the date and time of input, making it possible to search for data based on a specific date or time.

[1672] 5. Data Recall

[1673] When a user wants to reminisce about past memories based on a specific date or emotion, they input a request using their device, which is sent to the server, which searches and retrieves the relevant data from the database.

[1674] 6. Data playback

[1675] The device receives the data from the server and presents it to the user. The user can read the data in text format or play it back as audio. In the audio format, the text data is sent to a speech synthesis engine and played back as audio data.

[1676] 7. Recommendation of related information

[1677] Based on the stored emotional data, the system recommends information related to past emotional experiences to the user. For example, the system has the function of recommending new items based on items that the user found enjoyable in the past.

[1678] Hardware and software used

[1679] Voice input devices: smartphones, head-mounted displays

[1680] Speech recognition engine: speech_recognition

[1681] Sentiment Analysis Engine: TextBlob Library

[1682] Database: SQLite

[1683] Server: General web server software (e.g. Flask, Django)

[1684] Specific examples

[1685] For example, suppose a user purchases new running shoes in a virtual store on October 1, 2023, and wants to record the excitement they felt. The user speaks into a microphone, saying, "I'm so happy I bought my new running shoes." The system recognizes this speech as text and records a positive sentiment polarity (e.g., 0.8). This data, along with the date, is stored in a database, and if the user later searches for "happy events in October 2023," the record will appear.

[1686] Prompt Sentence Examples

[1687] I was very happy to purchase new running shoes on October 1, 2023. My emotional polarity is 0.8.

[1688] In this way, users can efficiently record their past emotions and memories, enriching their shopping experience within the virtual store.

[1689] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1690] Step 1: User provides voice or text input

[1691] Specific behavior:

[1692] Users use devices such as smartphones or head-mounted displays to input their experiences and emotions in the virtual store by voice or text. For voice input, they use the device's microphone input function, and for text input, they use an on-screen input form.

[1693] Input: User voice or text data

[1694] Output: Raw audio or text data

[1695] Step 2: The device converts the audio data into text data.

[1696] Specific behavior:

[1697] The device uses a speech recognition engine (e.g., the speech_recognition library) to convert the user's voice input into text data. This process involves inputting the voice data into a speech recognition model, which then outputs the speech content as a string of characters.

[1698] Input: User's voice data

[1699] Output: Data converted to text

[1700] Step 3: The device sends the entered text to the server

[1701] Specific behavior:

[1702] The terminal transmits the converted text data and its meta information (e.g., input date and time) to the server via a network communication protocol (e.g., HTTP).

[1703] Input: Text data, input date and time

[1704] Output: Text data and meta information sent to the server

[1705] Step 4: The server inputs the text data into the sentiment analysis engine

[1706] Specific behavior:

[1707] The server inputs the received text data into a sentiment analysis engine (e.g., TextBlob library) to analyze the sentiment polarity (positive, negative, neutral) of each text. The analysis results include sentiment polarity values ​​and sentiment-related tags.

[1708] Input: Text data

[1709] Output: Sentiment polarity value and sentiment tag

[1710] Step 5: The server stores the analysis results in a database

[1711] Specific behavior:

[1712] The server stores the sentiment polarity values, text data, and meta information (e.g., input date and time) in a database (e.g., SQLite). During this storage process, each data item is inserted into the appropriate database table.

[1713] Input: Sentiment polarity value, text data, input date and time

[1714] Output: Sentiment analysis data stored in a database

[1715] Step 6: User requests historical emotion data based on specific criteria

[1716] Specific behavior:

[1717] A user uses a device to request historical emotion data based on a specific date and emotion. The request is sent to the server as a search query containing the specified criteria (e.g., specific date, emotion tag).

[1718] Input: Search criteria (e.g. date, emotion tag)

[1719] Output: Search request sent to the server

[1720] Step 7: The server searches and retrieves the relevant data from the database.

[1721] Specific behavior:

[1722] The server searches and retrieves the corresponding emotion data from the database based on the received search request. In this process, it filters the data that matches the search criteria and generates the data to be sent back to the user.

[1723] Input: Search criteria (e.g. date, emotion tag)

[1724] Output: Emotion data as search results

[1725] Step 8: Play back the data retrieved by the device as text or audio

[1726] Specific behavior:

[1727] The device presents the emotion data received from the server to the user, who can read the data in text format or play it back in audio format using a speech synthesis engine (e.g., a TTS engine).

[1728] Input: Emotion data as search results

[1729] Output: Text display or audio playback data

[1730] Step 9: The server generates a prompt to recommend related information and presents it to the user.

[1731] Specific behavior:

[1732] The server generates prompts to recommend related information based on the stored emotional data. For example, it can recommend new items based on items that users have previously found enjoyable. The generated prompts are input into a generative AI model and presented to the user as recommendations.

[1733] Input: Emotional data, past emotional experiences

[1734] Output: Prompt statement with relevant information

[1735] Prompt Sentence Examples

[1736] I was very happy to purchase new running shoes on October 1, 2023. My emotional polarity is 0.8.

[1737] This prompt allows the user to reflect on past experiences while simultaneously receiving new, relevant information.

[1738] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1739] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1740] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1741] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1742] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1743] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1744] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1745] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1746] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1747] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1748] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1749] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1750] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1751] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1752] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1753] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1754] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1755] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1756] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1757] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1758] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1759] The following is further disclosed regarding the above embodiment.

[1760] (Claim 1)

[1761] a means for accepting voice or written input;

[1762] A means for converting audio data into text data;

[1763] A means for performing sentiment analysis on input text data;

[1764] A means for storing the sentiment-analyzed text data in a database;

[1765] A means for retrieving stored data based on specific conditions;

[1766] a means for reproducing the retrieved data in text or audio;

[1767] A system including:

[1768] (Claim 2)

[1769] The system of claim 1, wherein the user's voice input is sent to a voice recognition engine and converted into text data.

[1770] (Claim 3)

[1771] The system of claim 1, wherein the text data is sent to a speech synthesis engine and converted into speech data.

[1772] "Example 1"

[1773] (Claim 1)

[1774] a means for accepting voice or written input;

[1775] A means for converting audio data into text data;

[1776] A means for performing sentiment analysis on input text data;

[1777] A means for storing the sentiment-analyzed text data in a database;

[1778] A means for retrieving stored data based on specific conditions;

[1779] a means for reproducing the retrieved data in text or audio;

[1780] A means for a user to use a terminal to send a request to search for data based on criteria;

[1781] A means for the terminal to retrieve data based on a request sent to the server and present it to the user;

[1782] A system including:

[1783] (Claim 2)

[1784] The system of claim 1, wherein the system transmits a user's voice input to a speech recognition engine and converts the voice input into text data.

[1785] (Claim 3)

[1786] The system of claim 1, wherein the text data is sent to a speech synthesis engine and converted into speech data.

[1787] "Application Example 1"

[1788] (Claim 1)

[1789] a means for accepting voice or written input;

[1790] A means for converting audio data into text data;

[1791] A means for performing sentiment analysis on input text data;

[1792] A means for storing the sentiment-analyzed text data in a database;

[1793] A means for retrieving stored data based on specific conditions;

[1794] a means for reproducing the retrieved data in text or audio;

[1795] A means to record customer voice or text input during a store visit and retrieve and play back the data to provide special service the next time the customer visits; and

[1796] A system including:

[1797] (Claim 2)

[1798] The system of claim 1, wherein the user's voice input is sent to a voice recognition engine and converted into text data.

[1799] (Claim 3)

[1800] The system of claim 1, wherein the text data is sent to a speech synthesis engine and converted into speech data.

[1801] "Example 2: Combining Emotion Engines"

[1802] (Claim 1)

[1803] a means for accepting voice or written input;

[1804] A means for converting the received voice data into text data using a voice recognition engine;

[1805] [Using an emotion recognition engine that generates real-time emotion information based on text data and the user's voice tone];

[1806] A means for performing emotion analysis on the generated emotion information together with text data and assigning emotion tags;

[1807] A means for storing the emotion-analyzed text data and emotion information in a database together with date information;

[1808] A means to search and retrieve stored data based on specific criteria, based on user requests;

[1809] a means for converting the acquired data into voice using a text or voice synthesis engine, and

[1810] A system including:

[1811] (Claim 2)

[1812] The system of claim 1, wherein the system receives a user's voice input and converts it into text data using a speech recognition engine.

[1813] (Claim 3)

[1814] The system of claim 1 [generates user emotion information in real time using an emotion recognition engine and transmits it together with text data].

[1815] "Application example 2 when combining emotion engines"

[1816] (Claim 1)

[1817] a means for accepting voice or written input;

[1818] A means for converting audio data into text data;

[1819] A means for performing sentiment analysis on input text data;

[1820] A means for storing the sentiment-analyzed text data in a database;

[1821] A means for retrieving stored data based on specific conditions;

[1822] a means for reproducing the retrieved data in text or audio;

[1823] a means for recommending relevant information based on the stored emotion data;

[1824] A system including:

[1825] (Claim 2)

[1826] The system of claim 1, wherein the user's voice input is sent to a voice recognition engine and converted into text data.

[1827] (Claim 3)

[1828] The system of claim 1, wherein the text data is sent to a speech synthesis engine and converted into speech data. [Explanation of symbols]

[1829] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for accepting voice or text input; means for converting voice data into text data; A means for performing sentiment analysis on input text data; A means for storing the sentiment-analyzed text data in a database; A means for retrieving the stored data based on specific conditions; means for reproducing the retrieved data in text or audio; A system including:

2. 10. The system of claim 1, wherein the user's voice input is sent to a voice recognition engine and converted into text data.

3. 10. The system of claim 1, wherein the text data is transmitted to a speech synthesis engine and converted into speech data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A