System
The system addresses inefficient and potentially harmful information intake by allowing users to select and receive personalized audio information, filtering and generating voice scripts, thereby enhancing user experience and health safety.
Patent Information
- Application Number
- JP2024118207
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
General radio programs contain unnecessary information for individual users, posing a challenge in efficient information intake and risking health hazards due to blue light exposure from smartphones during sleep.
A system that allows users to select desired information categories, collects relevant data, filters it, generates an audio script, and plays it, incorporating personalized greetings to facilitate efficient and healthy information acquisition without screen use.
Enables users to obtain necessary information efficiently and healthily by voice, reducing the risk of using smartphones while sleeping and enhancing user experience.
Smart Images

Figure 2026017425000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] General radio programs contain a lot of unnecessary information for individual users, making it difficult to efficiently input information. There is also the risk of "using your smartphone while sleeping," where the blue light from smartphones has a negative impact on sleep. By solving these problems and providing a means for users to efficiently obtain the information they want to receive, it is necessary to reduce wasted time and prevent health hazards. [Means for solving the problem]
[0005] This invention provides a system including: means for selecting an information category a user wants to receive; means for collecting related information from the Internet based on the selected information category; means for filtering the collected information and extracting only information relevant to the user; means for generating an original voice script based on the extracted information; means for converting the generated script into an audio file using a voice synthesis engine; means for transmitting the audio file to the user's terminal; and means for playing the transmitted audio file. The voice script can also include a user's name or a friendly greeting, enhancing a sense of familiarity with the user. This allows users to efficiently obtain necessary information without looking at the screen, thereby reducing the risk of "using a smartphone while sleeping."
[0006] "User" refers to any person or entity that uses this system to receive information.
[0007] The "category of information desired to be received" refers to a specific type of information that the user wishes to obtain, such as news, weather, schedule, etc.
[0008] The "Internet" refers to a global network that connects computers and networks around the world to each other and allows information to be exchanged.
[0009] "Related information" refers to information collected based on the information category selected by the user.
[0010] "Means of collection" refers to the technology or method used to obtain the required data from a particular source (e.g., news API, weather service, calendar service).
[0011] "Filtering means" refers to a technique or method for filtering out unnecessary information from collected information based on specific criteria and selecting only necessary information.
[0012] "Means for extracting" refers to a technique or method for extracting specific data from the filtered information.
[0013] "Audio script" refers to text to be read aloud that is created based on collected and filtered information.
[0014] "Speech synthesis engine" refers to software or technology for converting text into speech.
[0015] "Audio file" refers to an acoustic data file generated by a speech synthesis engine.
[0016] "Device" refers to the device (e.g., smartphone, tablet, computer) used by a User to receive information.
[0017] "Means for playing" refers to the technology or method for playing an audio file on a terminal and allowing a user to listen to the audio.
[0018] "Username and Friendly Callout" refers to a personalized greeting or callout for an individual user that is included in the voice script. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] This invention relates to a system that allows a user to select the information category they wish to receive, collects related information from the Internet based on the selected information category, filters it, generates an audio script, and plays an audio file. The program processing of this system is explained below in natural language.
[0041] 1. User configuration
[0042] Users install the app and select the categories of information they want to receive (news, weather, schedule, etc.) when they first launch it. They can also set a username and a friendly greeting, allowing them to receive information that is individually customized.
[0043] 2. Collection of information
[0044] The device sends the user's configuration information to the server, which then collects the latest information from data providers such as news APIs, weather information services, and calendar services based on the configured information categories.
[0045] 3. Filtering and organizing information
[0046] The server analyzes the collected information and uses natural language processing technology to filter out unnecessary information. For example, in the case of news, it extracts only articles related to topics that interest the user, and in the case of weather information, it extracts the forecast for the user's location. In addition, in the case of schedule information, it lists important events and organizes them based on the user's interests.
[0047] 4. Generate voice script
[0048] Based on the filtered information, the server generates an original voice script containing friendly speech, such as "Good morning, Mr. / Ms. X. Today is X / X. First, I'll tell you today's news..."
[0049] 5. Speech synthesis and transmission
[0050] The server sends the generated script to a speech synthesis engine, which converts it into an audio file, which is then sent to the user's device.
[0051] 6. Audio playback
[0052] Users can play the audio file sent through their device and receive the information they have set by voice. This allows them to receive the necessary information without looking at the smartphone screen, reducing the risks of using their smartphone while sleeping and providing users with a convenient and healthy way to obtain information.
[0053] Specific examples
[0054] For example, if a user named "Tanaka-san" specifies that they want to receive news, weather, and today's schedule, the server retrieves the latest news from the news API and uses natural language processing to extract only articles that are highly relevant to Tanaka-san. It also retrieves today's weather forecast from the weather information service and organizes weather information related to Tanaka-san's location. Finally, it retrieves today's schedule from the calendar service and lists important events.
[0055] Finally, a voice script is generated that says, "Good morning, Tanaka-san. Today is October 1, 2023. Let me start with today's news..." and converted into an audio file using a speech synthesis engine. This file is then sent to Tanaka's device, allowing him to receive the information by voice without having to look at his smartphone.
[0056] The above is a specific embodiment for carrying out the present invention. By configuring the system in this way, users can efficiently obtain the necessary information and reduce the risk of using their smartphone while sleeping.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] User-defined
[0060] Users install the app and upon first launching it, select and enter the information categories they want to receive (news, weather, schedule, etc.), as well as their name and a familiar greeting. This allows for individually customized information acquisition.
[0061] Step 2:
[0062] Sending setting information from the device to the server
[0063] The device sends the user-entered configuration information, including selected information categories and user name, to the server, which stores this information in a database as a user profile.
[0064] Step 3:
[0065] Information gathering
[0066] Based on the user profile, the server collects the latest data from sources such as news APIs, weather information services, and calendar services. For example, it gets the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0067] Step 4:
[0068] Information Analysis and Filtering
[0069] The server analyzes the collected information and uses natural language processing technology to extract only information relevant to the user. For example, news information might extract articles that match the user's interest categories, weather information might extract weather forecasts for the user's location, and schedule information might extract important events.
[0070] Step 5:
[0071] Generate voice scripts
[0072] The server generates an original voice script based on the filtered information. This script includes a friendly greeting (e.g., "Good morning, Tanaka-san") and selected information categories. For example, "Today is the ____ day of the month. Here's today's news..."
[0073] Step 6:
[0074] Speech synthesis
[0075] The server sends the generated script to a speech synthesis engine (e.g., a Text-to-Speech engine) and converts it into an audio file, which is generated in a format that is easy for the user to listen to (e.g., MP3 format).
[0076] Step 7:
[0077] Sending an audio file
[0078] The server sends the generated audio file to the user's device, which then stores the received audio file in the appropriate folder.
[0079] Step 8:
[0080] Playing audio
[0081] The user plays the audio file sent through the device. By pressing the play button, the audio begins, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This provides a healthy way to obtain information and prevents excessive smartphone use.
[0082] Example 1
[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0084] In modern society, users must efficiently obtain a large amount of information, but visual information acquisition methods involve health risks and hassle, such as using a smartphone while sleeping. Furthermore, because the information users want to receive is diverse, there is a need for a method that can centrally extract only the necessary information and provide it in a user-friendly format. The purpose of this invention is to solve these problems and provide a means for users to acquire information efficiently and healthily.
[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0086] In this invention, the server includes: means for allowing a user to select an information category they wish to receive; means for collecting related information from a network based on the selected information category; and means for filtering the collected information using natural language processing technology to extract only information relevant to the user. This allows users to efficiently obtain information that is appropriate for them. The server also includes means for generating a friendly voice script using a generative AI model; means for converting the generated script into an audio file using a speech synthesis engine; means for transmitting the audio file to the user's device; and means for playing the transmitted audio file. This allows users to obtain information in a healthy way without relying on visual information acquisition, enabling them to use information efficiently and conveniently.
[0087] "User" refers to a person who uses the system to receive information.
[0088] "Information category" refers to the type of information that a user wants to receive, such as news, weather forecast, schedule information, etc.
[0089] "Network" refers to various communication infrastructures, including the Internet, and is used as a means for collecting and transmitting information.
[0090] "Related information" refers to data that is relevant to the user and is obtained based on the information category set by the user.
[0091] "Natural language processing technology" refers to the technology of analyzing text data to understand meaning and extract information.
[0092] "Generative AI model" refers to an artificial intelligence model that generates natural language based on given prompts.
[0093] "Voice script" refers to a sentence that expresses information provided to a user in a voice form.
[0094] A "speech synthesis engine" refers to a technology or system that converts text data into voice data.
[0095] "Audio file" refers to a file that stores audio data generated by speech synthesis.
[0096] "Device" refers to an electronic device that a user possesses and that receives and reproduces information.
[0097] This invention relates to a system that allows a user to select the information category they wish to receive, collects related information from a network based on the selected information category, filters it, generates a voice script, and plays back a voice file. The program processing of this system is explained below in natural language.
[0098] First, the user installs and launches a dedicated application on a device such as a smartphone or tablet. When launching the application for the first time, the user selects the categories of information they wish to receive (news, weather forecasts, schedule information, etc.) on the account settings screen, and also sets their name and how they want to be addressed. This information is sent from the user's device to the server via the network.
[0099] The server collects related information from various information providers based on the received user setting information. For example, news information is obtained using a news API (e.g., NewsAPI or Google News API), weather forecast information is obtained from a weather service (e.g., OpenWeatherMap), and schedule information is obtained from a calendar service (e.g., Google Calendar).
[0100] The server analyzes the large amount of collected data using natural language processing technology (e.g., SpaCy or IBM Watson NLP) and filters out unnecessary information. During this process, it extracts only news articles related to topics that interest the user, extracts information about the user's location from weather forecast information, and lists important schedule information.
[0101] Next, based on the filtered information, the server uses a generative AI model (e.g., OpenAI's ChatGPT) to generate a friendly voice script, using prompts such as "Generate a friendly script that starts with a morning greeting and covers today's news, weather, and schedule."
[0102] The generated script is passed to a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech) and converted into an audio file, which is then sent from the server to the user's device.
[0103] Finally, users can play audio files sent through the device and receive the information they have set by voice, which allows them to obtain information efficiently and healthily without relying on visual information acquisition.
[0104] For example, if a user named "Tanaka-san" requests to receive news, weather forecasts, and schedule information, the server retrieves the latest news from the news API and uses natural language processing technology to extract only articles that are relevant to Tanaka-san. It then retrieves today's weather forecast from a weather service and organizes information about Tanaka-san's location. It then retrieves today's schedule from a calendar service and lists important events. Finally, it generates a voice script that says, "Good morning, Tanaka-san. Today is October 1, 2023. Let's start with today's news..." and converts it into an audio file using a speech synthesis engine. This file is then sent to Tanaka-san's device, allowing her to receive the information by voice without looking at her smartphone.
[0105] The above is a specific embodiment for carrying out the present invention. By configuring such a system, users can efficiently obtain necessary information, reduce the risk of using a smartphone while sleeping, and provide a convenient and healthy means of obtaining information.
[0106] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0107] Step 1:
[0108] The user installs and launches the dedicated application. When the application is launched for the first time, the user selects the information categories they wish to receive (news, weather forecasts, schedule information, etc.) on the account settings screen, and sets their name and how they want to be addressed. This input information includes data such as the user name, information category, and how they want to be addressed. The device receives this setting information as input, generates an API request to send it to the server, and then sends it.
[0109] Step 2:
[0110] The server receives the setting information sent from the device and stores it in a database. Next, based on the user's setting information, it retrieves news data from a news API (e.g., NewsAPI or Google News API). Specifically, it creates an API request to collect news data. This collected news data becomes the input.
[0111] Step 3:
[0112] Similarly, the server retrieves weather forecast data from OpenWeatherMap and schedule information from Google Calendar. It sends an API request to each service and retrieves the respective data (weather forecast data, schedule information data). This becomes the input for each service.
[0113] Step 4:
[0114] The server analyzes the collected news data, weather forecast data, and schedule information data, and filters out unnecessary information using natural language processing technology (e.g., SpaCy or IBM Watson NLP). It receives the collected data as input and extracts only the information that is highly relevant to the user through a filtering process. This is the filtered data that is output.
[0115] Step 5:
[0116] The server uses a generative AI model (e.g., OpenAI's ChatGPT) to generate a friendly voice script based on the filtered data. The input to this process is the filtered data, and the prompt "Generate a friendly script that starts with a morning greeting and covers today's news, weather, and schedule." The output is the generated voice script.
[0117] Step 6:
[0118] The server passes the generated voice script to a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech) and converts it into an audio file. The input for this process is the voice script, which the speech synthesis engine analyzes and outputs an audio file. Specifically, the speech synthesis engine generates voice data based on text data and saves it in an audio file.
[0119] Step 7:
[0120] The server creates and sends an API request to send the generated audio file to the user's device. The input of this process is the audio file, and the output is the audio file sent to the user's device.
[0121] Step 8:
[0122] The user's device receives and saves the audio file sent from the server. When the user operates the application to instruct audio playback, the device plays the saved audio file. The input for this process is the audio file, and the output is audio information provided to the user. The user can receive the necessary information by audio without looking at their smartphone.
[0123] The above are the specific processing steps of this system. Users can acquire information efficiently and are provided with a healthy information acquisition method that does not rely on visual information acquisition.
[0124] (Application example 1)
[0125] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0126] Modern users are seeking ways to efficiently and comfortably gather information amid their busy daily lives. However, conventional information gathering systems have limitations in providing personalized information based on users' preferences and interests, and require users to manually search for information. Furthermore, obtaining product information in a virtual store primarily requires visual confirmation, which requires significant time and effort. Therefore, the present invention aims to provide a system that automates the provision of information based on users' preferences and provides product information in a virtual store via voice, allowing users to efficiently and comfortably obtain the information they need.
[0127] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0128] In this invention, the server includes means for selecting an information category that the user wishes to receive, means for collecting related information from the Internet based on the selected information category, means for filtering the collected information and extracting only information relevant to the user, means for generating an original voice script based on the extracted information, means for converting the generated script into a voice file using a voice synthesis engine, means for transmitting the voice file to the user's terminal, means for playing the transmitted voice file, means for selecting a product category that the user wishes to receive and for providing voice information about products in the virtual store, and means for automatically playing voice information about the products when the user approaches a product in the virtual store. This allows the user to efficiently obtain information regardless of time or place, and to obtain voice information about products in the virtual store without relying on vision.
[0129] A "user" is an entity that selects an information category, collects related information from the Internet, and receives the information in audio format.
[0130] "Information category" indicates the type of information that the user wants to receive, and includes news, weather information, schedule information, product information, and the like.
[0131] The term "means" refers to components or methods required to achieve a specific function, and in the present invention includes functions such as information collection, filtering, voice generation, and voice playback.
[0132] The "Internet" is a global network that connects computers around the world and enables the exchange of information.
[0133] "Related information" is data collected based on an information category selected by the user, and is information that contains content that is useful to the user.
[0134] "Filtering" is the process of selecting only information relevant to the user from the collected information and removing unnecessary information.
[0135] A "voice script" is a text for communicating information by voice that is generated based on collected and filtered information.
[0136] A "speech synthesis engine" is a software and hardware component for converting text information into speech data.
[0137] "Audio file" means a digital file generated by a speech synthesis engine to provide audio information to a user.
[0138] A "terminal" is a device that allows a user to receive audio files and play the information, including smartphones, smart glasses, and head-mounted displays.
[0139] A "virtual store" is a commercial facility built in a virtual space where users can browse and purchase products online.
[0140] "Product category" indicates the type of product in which the user is interested, and includes home appliances, fashion, accessories, and the like.
[0141] "Detailed information" refers to information that a user needs to make a decision about purchasing a product, such as product features, price, specifications, and usage.
[0142] The system for implementing this invention allows a user to select a category of information they wish to receive, collects related information from the Internet based on that category, generates a voice script, and transmits it to the user's terminal for playback. This system also has the function of providing voice information about products in a virtual store. The specific configuration and operation of the system are described below.
[0143] System configuration
[0144] This system mainly uses the following hardware and software:
[0145] Hardware: Smartphones, smart glasses, head-mounted displays (HMDs)
[0146] Software: Natural language processing (NLP) engine, speech synthesis engine (e.g., Google Text-to-Speech API), user data management server, product information database
[0147] Program processing
[0148] 1. User configuration
[0149] Users install the app and select the desired information category (news, weather, product information, etc.) when they first start it. They can also set a name and how they want to be called.
[0150] 2. Information collection by the server
[0151] The device sends the user's setting information to the server, which then collects the latest information from related services (news API, weather information service, product information database) based on the set information category.
[0152] 3. Server filtering and organization of information
[0153] The server analyzes the collected information and uses NLP technology to filter out unnecessary information. For example, in the case of news, it extracts only articles related to topics that the user is interested in, and in the case of product information, it retrieves detailed information about products that interest the user.
[0154] 4. Server-generated voice script
[0155] Based on the filtered information, the server generates an original voice script containing friendly speech, such as "Good morning, Mr. / Ms. X. Today is X / X. First, I'll tell you today's news..." including the user's name.
[0156] 5. Server-based speech synthesis and transmission
[0157] The server sends the generated script to a speech synthesis engine to generate an audio file, which is then sent to the user's device.
[0158] 6. Playing Audio on the Device
[0159] Users can play the audio file sent through the device and receive the set information by voice, allowing them to receive information without looking at the smartphone screen.
[0160] 7. Audio guide function in virtual stores
[0161] In particular, within the virtual store, product information is provided based on product categories set by the user. When the user approaches a product, detailed information about that product is automatically played back in audio.
[0162] Examples and prompts
[0163] Specific examples
[0164] For example, if a user named "Tanaka-san" requests news, weather, and information about home appliances in a virtual store, the server retrieves the latest news from the news API and uses natural language processing to extract only articles most relevant to Tanaka-san. It then retrieves today's weather forecast from a weather information service and organizes weather information relevant to Tanaka-san's location. For home appliance information, the server collects the latest product information within the virtual store and lists important specifications and pricing information. Finally, it generates a voice script that reads, "Good morning, Tanaka-san. Today is October 1, 2023. Let me start with today's news..." and converts it into an audio file using a speech synthesis engine (Google Text-to-Speech API). This file is then sent to Tanaka-san's device, allowing him to easily access the information by voice via his smartphone or HMD.
[0165] Prompt Sentence Examples
[0166] Username: Tanaka-san
[0167] Favorite product category: Home appliances
[0168] Provides the latest information on home appliances via voice
[0169] Example of generated voice script: "Hey Tanaka, here are some recommended new home appliances..."
[0170] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0171] Step 1:
[0172] User category selection and individual settings
[0173] (Input) The user installs the app and selects the desired information category (news, weather, product information, etc.) when launching it for the first time. They also set a name and how they want to be called.
[0174] (Processing) The app collects user selections and settings.
[0175] (Output) The collected setting information is stored in the terminal and also sent to the server.
[0176] Step 2:
[0177] Server collection of information
[0178] (Input) The server collects related information based on the selected information category, based on the user setting information received from the terminal.
[0179] The (processing) server obtains the necessary data from news APIs, weather information services, product information databases, etc.
[0180] (Output) The acquired data is temporarily stored on the server.
[0181] Step 3:
[0182] Filtering and organizing information
[0183] (Input) Collected raw data (news articles, weather information, product information, etc.).
[0184] The (processing) server uses a natural language processing (NLP) engine to extract only the information relevant to the user and filter out unnecessary information. For example, in the case of a news category, it selects only important and reliable articles. In the case of product information, it organizes detailed information related to items that the user is interested in.
[0185] (Output) The filtered and organized information is stored on the server.
[0186] Step 4:
[0187] Generate voice scripts
[0188] (Input) Filtered information.
[0189] The (processing) server generates a friendly, customized voice script based on the received information. For example, it creates text in the form of "Good morning, Mr. / Ms. XX. Today is the XXth month. First, I'll tell you today's news..."
[0190] (Output) The generated voice script is saved on the server.
[0191] Step 5:
[0192] Speech synthesis and transmission
[0193] (Input) The generated voice script.
[0194] The (processing) server uses a speech synthesis engine (e.g., Google Text-to-Speech API) to convert the text-based voice script into an audio file.
[0195] (Output) The converted audio file is generated and sent to the user's device.
[0196] Step 6:
[0197] Device audio playback
[0198] (Input) The audio file sent from the server.
[0199] (Processing) The user terminal receives the audio file and plays the audio using the built-in media player.
[0200] (Output) The user can receive the information set through the terminal by voice.
[0201] Step 7:
[0202] Audio guide function in virtual stores
[0203] (Input) Product category and its location information set by the user.
[0204] (Processing) When a user walks around the virtual store and comes close to a certain product, detailed information about that product is automatically played back using the device's sensors and location information. Specifically, the device uses distance sensors and GPS data to determine the user's current location, and plays back related product information as an audio file.
[0205] (Output) The user can receive voice information about products that interest them in the virtual store.
[0206] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0207] This invention combines a system in which a user selects the information category they wish to receive, collects related information from the Internet based on the selected information category, filters it, generates voice scripts, and plays voice files, with an emotion engine that recognizes the user's emotions and adjusts the information and voice based on those emotions. Below, the program processing of this system is explained in natural language.
[0208] 1. User configuration
[0209] When users install the app and start it for the first time, they select and enter the categories of information they want to receive (news, weather, schedule, etc.), as well as their own name and a familiar greeting. This allows them to receive information that is individually customized.
[0210] 2. Emotion Recognition by Emotion Engine
[0211] To recognize the user's emotions, the device analyzes data acquired from the camera and microphone in real time, thereby detecting the user's current emotions (e.g., joy, sadness, anger, surprise, etc.). The emotion engine analyzes this data using an emotion analysis model to determine the user's emotional state.
[0212] 3. Sending setting information from the device to the server
[0213] The device sends the user-entered setting information and the emotion information recognized by the emotion engine to the server, which then stores this information in a database as a user profile.
[0214] 4. Information Collection
[0215] Based on the user profile, the server collects the latest data from sources such as news APIs, weather information services, and calendar services. For example, it gets the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0216] 5. Information Analysis and Filtering
[0217] The server analyzes the collected information and uses natural language processing technology to extract only information relevant to the user. For example, news information will extract articles that match the user's interest categories, weather information will extract weather forecasts for the user's location, and schedule information will extract important appointments. Additionally, based on the emotional information provided by the emotion engine, the server prioritizes and selects information that best suits the user's current emotional state.
[0218] 6. Generate voice script
[0219] The server generates an original voice script based on the filtered information. This script includes a friendly greeting (e.g., "Good morning, Tanaka-san") and selected information categories. For example, "Today is the ____ day of the month. Here's today's news..." Furthermore, the script incorporates phrases with adjusted tone and speed according to the emotions recognized by the emotion engine.
[0220] 7. Speech Synthesis
[0221] The server sends the generated script to a speech synthesis engine (e.g., a Text-to-Speech engine) and converts it into an audio file, which is generated in a format that is easy for the user to listen to (e.g., MP3 format).
[0222] 8. Sending audio files
[0223] The server sends the generated audio file to the user's device, which then stores the received audio file in the appropriate folder.
[0224] 9. Audio playback
[0225] The user plays the audio file sent through the device. By pressing the play button, the audio begins, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This provides a healthy way to obtain information and prevents excessive smartphone use.
[0226] Specific examples
[0227] For example, if a user has the name "Tanaka-san" and has set that they want to receive news, weather, and today's schedule, and the emotion engine recognizes that the current emotion is "happy," the server will retrieve the latest news from the news API and use natural language processing to extract only articles that are highly relevant to Tanaka-san. It will also retrieve today's weather forecast from the weather information service and organize the weather information for Tanaka-san's location. It will then retrieve today's schedule from the calendar service and list important events. The emotion engine will then synthesize speech in a positive tone to match the user's "happy" state.
[0228] Finally, a voice script is generated that says, "Good morning, Tanaka. Today is October 1, 2023. Today's news is very interesting..." and converted into an audio file using a speech synthesis engine. This file is then sent to Tanaka's device, allowing him to receive the information by voice without having to look at his smartphone.
[0229] The above is a specific embodiment for carrying out the present invention. By configuring the system in this way, it is possible to efficiently obtain necessary information in a manner that is in line with the user's emotions, and reduce the risk of "using a smartphone while sleeping."
[0230] The processing flow will be explained below.
[0231] Step 1:
[0232] User-defined
[0233] When users install the app and launch it for the first time, they select and input the categories of information they want to receive (news, weather, schedule, etc.), as well as their own name and a familiar greeting. This allows them to receive information that is individually customized.
[0234] Step 2:
[0235] Emotion recognition by emotion engine
[0236] The device captures and analyzes data from the camera and microphone in real time to recognize the user's emotions, such as whether the user is happy, sad, angry, or surprised, by analyzing facial expressions and voice tone.
[0237] Step 3:
[0238] Sending setting information from the device to the server
[0239] The device transmits the user's setting information and the emotion information recognized by the emotion engine to the server, where the transmitted data is stored in the user profile and used for subsequent processing.
[0240] Step 4:
[0241] Information gathering
[0242] Based on the user profile, the server collects the latest data from information sources such as a news API, weather information service, and calendar service. For example, it obtains the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0243] Step 5:
[0244] Information Analysis and Filtering
[0245] The server analyzes the collected information and uses natural language processing technology to extract only information that is highly relevant to the user. Furthermore, based on the emotional data provided by the emotion engine, it prioritizes information that matches the user's current emotional state. For example, if the user is "happy," it prioritizes positive news and information that enhances emotions.
[0246] Step 6:
[0247] Generate voice scripts
[0248] The server generates an original voice script based on the filtered information. This script includes friendly greetings (e.g., "Good morning, Tanaka-san") and phrases that correspond to emotions. For example, if the user is in a "happy" state, the script might say, "I have some very exciting news for you today..."
[0249] Step 7:
[0250] Speech synthesis
[0251] The server then sends the generated script to a speech synthesis engine (e.g., a text-to-speech engine) and converts it into an audio file. This audio file is generated in a smooth listening format (e.g., MP3 format), and the tone and speed of the audio are adjusted based on the user's emotions.
[0252] Step 8:
[0253] Sending an audio file
[0254] The server sends the generated audio file to the user's device, which stores the received audio file in an appropriate folder and prepares it for playback.
[0255] Step 9:
[0256] Playing audio
[0257] The user plays the audio file sent through the device. By pressing the play button, a voice will play saying, "Good morning, Tanaka-san..." and the user can listen to the necessary information without looking at the screen. This prevents excessive smartphone use and provides a healthy way to obtain information.
[0258] In this way, the system of the present invention can adjust the content of information and the tone of voice according to the user's emotional state, providing an optimal information acquisition experience.
[0259] Example 2
[0260] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0261] In modern society, many users use the Internet to obtain information, but the overwhelming amount of information available makes it difficult to quickly find relevant information. Users who want to easily obtain information need a means to utilize not only their eyes but also their ears, but there is a lack of systems that can accommodate individual needs and emotional states. Furthermore, because excessive smartphone use can cause health problems, there is a need for a means to obtain information without looking at the screen.
[0262] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0263] In this invention, the server includes means for allowing a user to select an information category they wish to receive, means for collecting related information from a data network based on the selected information category, and means for analyzing the collected information using natural language processing technology and extracting only information relevant to the user. This allows the user to easily obtain only the information that suits them from an excess of information and to receive that information by voice.
[0264] The "means for selecting the information category that the user wishes to receive" refers to an input device or interface that allows the user to specify the type of information in which the user is interested (for example, news, weather, schedule, etc.).
[0265] A "data network" is a communication path through which multiple computing resources exchange information with each other, including the Internet and other information and communication networks.
[0266] "Natural language processing technology" refers to computer technology for understanding, analyzing, and generating human language (e.g., keyword extraction, document classification, sentiment analysis, etc.).
[0267] "Means for recognizing an emotional state and adjusting information based on that emotional state" refers to technology or a device that reads emotions from a user's facial expressions, voice, etc., and changes the content and expression of the information provided according to those emotions.
[0268] "Means for generating voice scripts" refers to technology or devices that automatically create text to be read aloud based on collected and analyzed information.
[0269] A "speech synthesis engine" is a computer technology (for example, text-to-speech technology) for converting text in text format into voice data.
[0270] A "terminal" is an information processing device (for example, a smartphone, tablet, or PC) that can be directly operated by a user.
[0271] "News information" is information that reports the latest facts and events related to society, economy, sports, culture, etc.
[0272] "Weather information" refers to data related to the weather, such as weather forecasts, temperature, precipitation, and wind speed.
[0273] "Schedule information" is information about schedules and events listed in a user's schedule or calendar.
[0274] "Means for including a user name or friendly nickname" refers to technology or a device for incorporating a name set by the user or a nickname that gives a sense of familiarity to the user (for example, "san" or "kun") into the voice script.
[0275] This invention combines a system that allows a user to select an information category of interest, collects related information from the Internet based on that information, filters it, generates a voice script, and plays back a voice file, with an emotion engine that recognizes the user's emotions and adjusts the information and voice based on those emotions. A specific method for implementing this system will be described below.
[0276] User-defined
[0277] Users install a dedicated application on their device and, when they first start it up, set the categories of information they want to receive (e.g., news, weather, schedule), as well as their own name and a friendly greeting. This allows for individually customized information to be provided.
[0278] Emotion recognition by emotion engine
[0279] The device uses a camera and microphone to capture the user's facial expressions and voice data, and analyzes it in real time using an emotion analysis model (e.g., OpenFace or DeepFace). This identifies the user's emotional state and sends that information to the server as data. This data is used for information filtering and voice script generation, which will be described later.
[0280] Sending configuration information to the server
[0281] The device sends the information entered and set by the user, as well as the emotional data recognized by the emotion engine, to the server, which then stores this information in a database as a user profile and uses it to provide the most appropriate information to each individual user.
[0282] Information gathering
[0283] The server collects the necessary data from external sources such as news APIs, weather information services, and calendar services. For example, it obtains the latest news from a news API (e.g., Google News API), local weather forecasts from a weather information service (e.g., OpenWeatherMap API), and user schedule information from a calendar service (e.g., Google Calendar API).
[0284] Information Analysis and Filtering
[0285] The server analyzes the collected data using natural language processing technology and scrutinizes information relevant to the user. For example, for news information, keywords are extracted to extract articles that match the user's categories of interest, and for weather information, only information related to the user's location is extracted. Based on the recognition results of the emotion engine, information that matches the user's current emotions is selected.
[0286] Generate voice scripts
[0287] The server generates a voice script based on the filtered information. This voice script includes the user's name and a friendly greeting, and incorporates tones and expressions that correspond to the user's emotional state. For example, it includes gentle expressions such as, "It's a beautiful day today. Let's have a good day."
[0288] Speech synthesis
[0289] The server sends the generated script to a speech synthesis engine (e.g., Google Text-to-Speech API) and converts it into an audio file, which is generated in, for example, MP3 format.
[0290] Sending an audio file
[0291] The server sends the generated audio file to the user's device, which stores the received audio file in an appropriate folder and prepares it for playback.
[0292] Playing audio
[0293] Users can play the audio files sent through the device. By pressing the play button, they can hear a voice such as, "Good morning, Tanaka-san. Here's today's news..." and can listen to the necessary information without looking at the screen.
[0294] Specific examples
[0295] For example, if a user named "Tanaka-san" selects to receive news, weather, and today's schedule, and the emotion engine recognizes that the current emotion is "happy," the server collects and filters data from each source. Based on the emotion, a voice script with a positive tone is generated. Finally, the voice script, "Good morning, Tanaka-san. Today is October 1, 2023. Today's news is very interesting...," is converted into an audio file and sent to Tanaka-san's device.
[0296] Prompt Sentence Examples
[0297] For example, the following prompt sentence is fed into the generative AI model:
[0298] "Includes news, weather, and schedule categories. Current emotion is joy. Name is Tanaka. Uses familiar greeting. Send generated audio file to device."
[0299] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0300] Step 1: User Setup
[0301] A user launches the application on their device and selects the information categories they want to receive (e.g., news, weather, schedule), and also sets their name and a friendly greeting (e.g., "Mr. Tanaka"). This set information is stored in the device's local database.
[0302] Input: Category of information you want to receive, username, call
[0303] Output: Configuration information stored in the local database
[0304] Step 2: Emotion recognition by the emotion engine
[0305] The device uses a camera and microphone to collect the user's facial and voice data in real time. The collected data is analyzed using an emotion analysis model (e.g., OpenFace or DeepFace) to identify the user's emotional state (e.g., joy, sadness, anger). The analysis results are temporarily stored on the device for later use.
[0306] Input: User's facial expression data, voice data
[0307] Output: Parsed emotional state data
[0308] Step 3: Sending configuration information to the server
[0309] The terminal transmits the user's selected information categories and analyzed emotional state data to the server, which receives this information and stores it in a database as a user profile.
[0310] Input: Setting information stored in a local database, analyzed emotional state data
[0311] Output: User profile data sent to the server
[0312] Step 4: Gather information
[0313] Based on the collected user profile data, the server collects relevant information from external sources such as news APIs, weather information services, calendar services, etc. For example, the server retrieves the latest news articles from the news API and local weather forecasts from the weather information service.
[0314] Input: User profile data
[0315] Output: Raw data collected from external sources (news, weather, schedule information)
[0316] Step 5: Information analysis and filtering
[0317] The server analyzes the collected raw data using natural language processing technology. For example, it can extract articles that match the user's interest categories from news data, extract information related to the user's location from weather data, and adjust the priority and content of information based on the user's emotional state.
[0318] Input: Raw data collected from external sources, user emotional state data
[0319] Output: filtered and adjusted information data
[0320] Step 6: Generate the voice script
[0321] The server generates a voice script based on the filtered and adjusted data, which includes the user's name and a friendly greeting, and adjusts the tone and expression depending on the user's emotional state.
[0322] Input: Filtered and conditioned information data
[0323] Output: The generated voice script
[0324] Step 7: Text-to-Speech
[0325] The server sends the generated voice script to a text-to-speech engine (e.g., Google Text-to-Speech API) and converts it into an audio file (e.g., MP3 format). The converted audio file is temporarily stored on the server.
[0326] Input: Generated voice script
[0327] Output: Converted audio file
[0328] Step 8: Send the audio file
[0329] The server sends the generated voice file to the user's device, which saves the received voice file in the appropriate folder (e.g., " / user / voices / ").
[0330] Input: Converted audio file
[0331] Output: Audio file saved on your device
[0332] Step 9: Playing Audio
[0333] The user plays the audio file through the device. By pressing the play button, the audio will play, "Good morning, Tanaka-san. Here's today's news..." and the user can hear the necessary information without looking at the screen.
[0334] Input: Audio files stored on the device
[0335] Output: Audio information played to the user
[0336] (Application example 2)
[0337] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0338] In order to provide optimal services based on customer emotions, commercial facilities and brick-and-mortar stores require a system that recognizes customer emotions in real time and collects and provides information based on those emotions. However, current systems lack the functionality to recognize and appropriately reflect customer emotions, making generalization and automation difficult. Furthermore, the collected information is not optimized for customer emotions, limiting the improvement of customer satisfaction.
[0339] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0340] In this invention, the server includes means for selecting an information category that the user wishes to receive, means for collecting related information from the Internet based on the selected information category, means for filtering the collected information and extracting only information relevant to the user, means for generating an original voice script based on the extracted information, means for converting the generated script into a voice file using a voice synthesis engine, means for transmitting the voice file to the user's terminal, means for playing the transmitted voice file, means for analyzing data acquired from a camera or microphone and recognizing the user's emotions, and means for adjusting information and voice based on the recognized emotional information. This makes it possible to efficiently provide appropriate information in accordance with the customer's emotions and improve customer satisfaction.
[0341] The "means for selecting the information category that the user wants to receive" is a function that allows the user to select the type of information that the user wants to receive based on their own interests and concerns.
[0342] The "means for collecting related information from the Internet" is a function for obtaining data related to the information category selected by the user from various information sources on the Internet.
[0343] "Means of filtering collected information and extracting only information relevant to the user" refers to a function that selects the most relevant information from the acquired information based on the user's interests and preferences.
[0344] The "means for generating an original voice script based on extracted information" is a function for generating a script based on filtered information to be read aloud in a form that is familiar to the user.
[0345] "Means for converting the generated script into an audio file using a voice synthesis engine" is a function for converting the generated voice script into actual voice data using voice synthesis technology.
[0346] The "means for transmitting the audio file to the user's terminal" is a function for transmitting the generated audio file to the terminal used by the user via the Internet or other communication means.
[0347] The "means for playing transmitted audio files" is a function for playing audio files stored in the user's terminal, allowing the user to listen to audio information.
[0348] "Means for analyzing data obtained from a camera or microphone and recognizing the user's emotions" refers to a function that recognizes the user's emotional state by capturing and analyzing the user's facial expressions and voice using a camera or microphone installed on the device.
[0349] The "means for adjusting information or voice based on recognized emotional information" is a function for presenting information in an optimal form or adjusting the tone or content of voice based on the acquired emotional information of the user.
[0350] A system for implementing this invention allows a user to select the information categories they wish to receive, collects and filters relevant information based on those categories, generates audio scripts and plays audio files, and recognizes the user's emotions and adjusts the information and audio based on those emotions.
[0351] 1. User Settings
[0352] When the user starts the system for the first time, they set the category of information they want to receive (news, weather, schedule, etc.), their name, and a friendly greeting. These settings allow the system to retrieve and provide information tailored to the user.
[0353] 2. Emotion recognition
[0354] The device uses a camera and microphone to recognize emotions in real time from the user's facial expressions and voice. Specifically, it uses an emotion engine to analyze facial expressions using images captured by the camera and emotions from the voice. A library called DeepFace is used for emotion analysis, and a voice recognition API is used for voice analysis.
[0355] 3. Information gathering
[0356] The server collects the latest data from various information sources, such as a news API, weather information service, and calendar service, based on the user's settings and emotion information. For example, it obtains the latest articles from the news API, the local weather forecast from the weather information service, and schedule information from the calendar service.
[0357] 4. Information analysis and filtering
[0358] The server analyzes the collected information using natural language processing technology to extract only the information relevant to the user. At the same time, based on the emotional information recognized by the emotion engine, it prioritizes and selects the information that best suits the user's current emotional state.
[0359] 5. Voice script generation
[0360] The server generates an original voice script based on the filtered information, which includes the user's name, a friendly greeting, the selected information category, and phrases with adjusted tone and speed depending on the emotional information.
[0361] 6. Speech synthesis
[0362] The generated script is sent to a speech synthesis engine (e.g., a text-to-speech engine) and converted into an audio file, which generates an audio file in a format that is easy for users to listen to.
[0363] 7. Sending and playing audio files
[0364] The generated audio file is sent to the user's device, which receives it and saves it in the appropriate folder. The user can immediately listen to the audio information by pressing the play button.
[0365] Specific examples
[0366] For example, if a user named "Sato" specifies that they would like to receive news, weather, and today's schedule information, and the emotion engine recognizes that the current emotion is "happy," the server retrieves the latest news articles from the news API and uses natural language processing technology to extract only those articles that are highly relevant to Mr. Sato. It also retrieves weather information related to Mr. Sato's location from the weather information service and lists today's schedule from the calendar service. The emotion engine synthesizes voice in a positive tone that matches the user's "happy" state.
[0367] Example prompts for generative AI models
[0368] "The user's name is Sato. I'm happy to hear from Sato. Please let me know about new products that might interest Sato."
[0369] This makes it possible to efficiently provide appropriate information in a manner that is in line with the customer's feelings, thereby improving customer satisfaction.
[0370] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0371] Step 1: User Setup
[0372] Users launch the application installed on their device, select the category of information they want to receive (news, weather, schedule, etc.), and enter their name or a friendly greeting.
[0373] Input (user): Information category, name, call
[0374] Output (terminal): Save as configuration information
[0375] Step 2: Emotion Recognition
[0376] The device uses a camera and microphone to capture the user's facial expressions and voice in real time, and then analyzes the captured data using an emotion engine (e.g., the DeepFace library or a voice recognition API) to recognize the user's emotions.
[0377] Input (terminal): Camera video, audio data
[0378] Output (terminal): Emotion recognition results (e.g., joy, sadness, anger)
[0379] Step 3: Sending preferences and emotions
[0380] The device sends the setting information entered by the user and the recognized emotion information to the server, where they are stored in a database as a user profile.
[0381] Input (device): setting information, emotion recognition results
[0382] Output (Server): Save as user profile
[0383] Step 4: Gather information
[0384] Based on the user profile, the server collects the latest data from news APIs, weather information services, calendar services, etc. For example, it obtains the latest articles from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0385] Input (server): User profile
[0386] Output (server): Collected information (news articles, weather forecasts, schedules)
[0387] Step 5: Information analysis and filtering
[0388] The server analyzes the collected information using natural language processing technology to extract only the information relevant to the user. At the same time, it prioritizes and selects the information that best suits the user's current emotional state based on the emotional information provided by the emotion engine.
[0389] Input (server): Collected information, emotional information
[0390] Output (server): Filtered information
[0391] Step 6: Generate voice script
[0392] The server generates an original voice script based on the filtered information, including a user-friendly call and selected information categories, and adjusts the tone and speed of the voice based on the emotional information.
[0393] Input (server): filtered information, emotional information
[0394] Output (server): Voice script
[0395] Step 7: Text-to-Speech
[0396] The server sends the generated script to a speech synthesis engine and converts it into an audio file (e.g., MP3 format), which generates an audio file in a format that is easy for users to listen to.
[0397] Input (server): Voice script
[0398] Output (server): Audio file
[0399] Step 8: Send and save the audio file
[0400] The server sends the generated audio file to the user's terminal, which receives it and saves it in an appropriate folder.
[0401] Input (server): Audio file
[0402] Output (Device): Saved audio file
[0403] Step 9: Play the audio file
[0404] The user plays the audio file stored on the device and can listen to the information by pressing the play button.
[0405] Input (user): Playback instructions
[0406] Output (terminal): Playback of audio information
[0407] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0408] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0409] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0410] [Second embodiment]
[0411] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0412] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0413] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0414] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0415] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0416] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0417] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0418] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0419] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0420] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0421] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0422] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0423] This invention relates to a system that allows a user to select the information category they wish to receive, collects related information from the Internet based on the selected information category, filters it, generates an audio script, and plays an audio file. The program processing of this system is explained below in natural language.
[0424] 1. User configuration
[0425] Users install the app and select the categories of information they want to receive (news, weather, schedule, etc.) when they first launch it. They can also set a username and a friendly greeting, allowing them to receive information that is individually customized.
[0426] 2. Collection of information
[0427] The device sends the user's configuration information to the server, which then collects the latest information from data providers such as news APIs, weather information services, and calendar services based on the configured information categories.
[0428] 3. Filtering and organizing information
[0429] The server analyzes the collected information and uses natural language processing technology to filter out unnecessary information. For example, in the case of news, it extracts only articles related to topics that interest the user, and in the case of weather information, it extracts the forecast for the user's location. In addition, in the case of schedule information, it lists important events and organizes them based on the user's interests.
[0430] 4. Generate voice script
[0431] Based on the filtered information, the server generates an original voice script containing friendly speech, such as "Good morning, Mr. / Ms. X. Today is X / X. First, I'll tell you today's news..."
[0432] 5. Speech synthesis and transmission
[0433] The server sends the generated script to a speech synthesis engine, which converts it into an audio file, which is then sent to the user's device.
[0434] 6. Audio playback
[0435] Users can play the audio file sent through their device and receive the information they have set by voice. This allows them to receive the necessary information without looking at the smartphone screen, reducing the risks of using their smartphone while sleeping and providing users with a convenient and healthy way to obtain information.
[0436] Specific examples
[0437] For example, if a user named "Tanaka-san" specifies that they want to receive news, weather, and today's schedule, the server retrieves the latest news from the news API and uses natural language processing to extract only articles that are highly relevant to Tanaka-san. It also retrieves today's weather forecast from the weather information service and organizes weather information related to Tanaka-san's location. Finally, it retrieves today's schedule from the calendar service and lists important events.
[0438] Finally, a voice script is generated that says, "Good morning, Tanaka-san. Today is October 1, 2023. Let me start with today's news..." and converted into an audio file using a speech synthesis engine. This file is then sent to Tanaka's device, allowing him to receive the information by voice without having to look at his smartphone.
[0439] The above is a specific embodiment for carrying out the present invention. By configuring the system in this way, users can efficiently obtain the necessary information and reduce the risk of using their smartphone while sleeping.
[0440] The processing flow will be explained below.
[0441] Step 1:
[0442] User-defined
[0443] Users install the app and upon first launching it, select and enter the information categories they want to receive (news, weather, schedule, etc.), as well as their name and a familiar greeting. This allows for individually customized information acquisition.
[0444] Step 2:
[0445] Sending setting information from the device to the server
[0446] The device sends the user-entered configuration information, including selected information categories and user name, to the server, which stores this information in a database as a user profile.
[0447] Step 3:
[0448] Information gathering
[0449] Based on the user profile, the server collects the latest data from sources such as news APIs, weather information services, and calendar services. For example, it gets the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0450] Step 4:
[0451] Information Analysis and Filtering
[0452] The server analyzes the collected information and uses natural language processing technology to extract only information relevant to the user. For example, news information might extract articles that match the user's interest categories, weather information might extract weather forecasts for the user's location, and schedule information might extract important events.
[0453] Step 5:
[0454] Generate voice scripts
[0455] The server generates an original voice script based on the filtered information. This script includes a friendly greeting (e.g., "Good morning, Tanaka-san") and selected information categories. For example, "Today is the ____ day of the month. Here's today's news..."
[0456] Step 6:
[0457] Speech synthesis
[0458] The server sends the generated script to a speech synthesis engine (e.g., a Text-to-Speech engine) and converts it into an audio file, which is generated in a format that is easy for the user to listen to (e.g., MP3 format).
[0459] Step 7:
[0460] Sending an audio file
[0461] The server sends the generated audio file to the user's device, which then stores the received audio file in the appropriate folder.
[0462] Step 8:
[0463] Playing audio
[0464] The user plays the audio file sent through the device. By pressing the play button, the audio begins, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This provides a healthy way to obtain information and prevents excessive smartphone use.
[0465] Example 1
[0466] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0467] In modern society, users must efficiently obtain a large amount of information, but visual information acquisition methods involve health risks and hassle, such as using a smartphone while sleeping. Furthermore, because the information users want to receive is diverse, there is a need for a method that can centrally extract only the necessary information and provide it in a user-friendly format. The purpose of this invention is to solve these problems and provide a means for users to acquire information efficiently and healthily.
[0468] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0469] In this invention, the server includes: means for allowing a user to select an information category they wish to receive; means for collecting related information from a network based on the selected information category; and means for filtering the collected information using natural language processing technology to extract only information relevant to the user. This allows users to efficiently obtain information that is appropriate for them. The server also includes means for generating a friendly voice script using a generative AI model; means for converting the generated script into an audio file using a speech synthesis engine; means for transmitting the audio file to the user's device; and means for playing the transmitted audio file. This allows users to obtain information in a healthy way without relying on visual information acquisition, enabling them to use information efficiently and conveniently.
[0470] "User" refers to a person who uses the system to receive information.
[0471] "Information category" refers to the type of information that a user wants to receive, such as news, weather forecast, schedule information, etc.
[0472] "Network" refers to various communication infrastructures, including the Internet, and is used as a means for collecting and transmitting information.
[0473] "Related information" refers to data that is relevant to the user and is obtained based on the information category set by the user.
[0474] "Natural language processing technology" refers to the technology of analyzing text data to understand meaning and extract information.
[0475] "Generative AI model" refers to an artificial intelligence model that generates natural language based on given prompts.
[0476] "Voice script" refers to a sentence that expresses information provided to a user in a voice form.
[0477] A "speech synthesis engine" refers to a technology or system that converts text data into voice data.
[0478] "Audio file" refers to a file that stores audio data generated by speech synthesis.
[0479] "Device" refers to an electronic device that a user possesses and that receives and reproduces information.
[0480] This invention relates to a system that allows a user to select the information category they wish to receive, collects related information from a network based on the selected information category, filters it, generates a voice script, and plays back a voice file. The program processing of this system is explained below in natural language.
[0481] First, the user installs and launches a dedicated application on a device such as a smartphone or tablet. When launching the application for the first time, the user selects the categories of information they wish to receive (news, weather forecasts, schedule information, etc.) on the account settings screen, and also sets their name and how they want to be addressed. This information is sent from the user's device to the server via the network.
[0482] The server collects related information from various information providers based on the received user setting information. For example, news information is obtained using a news API (e.g., NewsAPI or Google News API), weather forecast information is obtained from a weather service (e.g., OpenWeatherMap), and schedule information is obtained from a calendar service (e.g., Google Calendar).
[0483] The server analyzes the large amount of collected data using natural language processing technology (e.g., SpaCy or IBM Watson NLP) and filters out unnecessary information. During this process, it extracts only news articles related to topics that interest the user, extracts information about the user's location from weather forecast information, and lists important schedule information.
[0484] Next, based on the filtered information, the server uses a generative AI model (e.g., OpenAI's ChatGPT) to generate a friendly voice script, using prompts such as "Generate a friendly script that starts with a morning greeting and covers today's news, weather, and schedule."
[0485] The generated script is passed to a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech) and converted into an audio file, which is then sent from the server to the user's device.
[0486] Finally, users can play audio files sent through the device and receive the information they have set by voice, which allows them to obtain information efficiently and healthily without relying on visual information acquisition.
[0487] For example, if a user named "Tanaka-san" requests to receive news, weather forecasts, and schedule information, the server retrieves the latest news from the news API and uses natural language processing technology to extract only articles that are relevant to Tanaka-san. It then retrieves today's weather forecast from a weather service and organizes information about Tanaka-san's location. It then retrieves today's schedule from a calendar service and lists important events. Finally, it generates a voice script that says, "Good morning, Tanaka-san. Today is October 1, 2023. Let's start with today's news..." and converts it into an audio file using a speech synthesis engine. This file is then sent to Tanaka-san's device, allowing her to receive the information by voice without looking at her smartphone.
[0488] The above is a specific embodiment for carrying out the present invention. By configuring such a system, users can efficiently obtain necessary information, reduce the risk of using a smartphone while sleeping, and provide a convenient and healthy means of obtaining information.
[0489] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0490] Step 1:
[0491] The user installs and launches the dedicated application. When the application is launched for the first time, the user selects the information categories they wish to receive (news, weather forecasts, schedule information, etc.) on the account settings screen, and sets their name and how they want to be addressed. This input information includes data such as the user name, information category, and how they want to be addressed. The device receives this setting information as input, generates an API request to send it to the server, and then sends it.
[0492] Step 2:
[0493] The server receives the setting information sent from the device and stores it in a database. Next, based on the user's setting information, it retrieves news data from a news API (e.g., NewsAPI or Google News API). Specifically, it creates an API request to collect news data. This collected news data becomes the input.
[0494] Step 3:
[0495] Similarly, the server retrieves weather forecast data from OpenWeatherMap and schedule information from Google Calendar. It sends an API request to each service and retrieves the respective data (weather forecast data, schedule information data). This becomes the input for each service.
[0496] Step 4:
[0497] The server analyzes the collected news data, weather forecast data, and schedule information data, and filters out unnecessary information using natural language processing technology (e.g., SpaCy or IBM Watson NLP). It receives the collected data as input and extracts only the information that is highly relevant to the user through a filtering process. This is the filtered data that is output.
[0498] Step 5:
[0499] The server uses a generative AI model (e.g., OpenAI's ChatGPT) to generate a friendly voice script based on the filtered data. The input to this process is the filtered data, and the prompt "Generate a friendly script that starts with a morning greeting and covers today's news, weather, and schedule." The output is the generated voice script.
[0500] Step 6:
[0501] The server passes the generated voice script to a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech) and converts it into an audio file. The input for this process is the voice script, which the speech synthesis engine analyzes and outputs an audio file. Specifically, the speech synthesis engine generates voice data based on text data and saves it in an audio file.
[0502] Step 7:
[0503] The server creates and sends an API request to send the generated audio file to the user's device. The input of this process is the audio file, and the output is the audio file sent to the user's device.
[0504] Step 8:
[0505] The user's device receives and saves the audio file sent from the server. When the user operates the application to instruct audio playback, the device plays the saved audio file. The input for this process is the audio file, and the output is audio information provided to the user. The user can receive the necessary information by audio without looking at their smartphone.
[0506] The above are the specific processing steps of this system. Users can acquire information efficiently and are provided with a healthy information acquisition method that does not rely on visual information acquisition.
[0507] (Application example 1)
[0508] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0509] Modern users are seeking ways to efficiently and comfortably gather information amid their busy daily lives. However, conventional information gathering systems have limitations in providing personalized information based on users' preferences and interests, and require users to manually search for information. Furthermore, obtaining product information in a virtual store primarily requires visual confirmation, which requires significant time and effort. Therefore, the present invention aims to provide a system that automates the provision of information based on users' preferences and provides product information in a virtual store via voice, allowing users to efficiently and comfortably obtain the information they need.
[0510] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0511] In this invention, the server includes means for selecting an information category that the user wishes to receive, means for collecting related information from the Internet based on the selected information category, means for filtering the collected information and extracting only information relevant to the user, means for generating an original voice script based on the extracted information, means for converting the generated script into a voice file using a voice synthesis engine, means for transmitting the voice file to the user's terminal, means for playing the transmitted voice file, means for selecting a product category that the user wishes to receive and for providing voice information about products in the virtual store, and means for automatically playing voice information about the products when the user approaches a product in the virtual store. This allows the user to efficiently obtain information regardless of time or place, and to obtain voice information about products in the virtual store without relying on vision.
[0512] A "user" is an entity that selects an information category, collects related information from the Internet, and receives the information in audio format.
[0513] "Information category" indicates the type of information that the user wants to receive, and includes news, weather information, schedule information, product information, and the like.
[0514] The term "means" refers to components or methods required to achieve a specific function, and in the present invention includes functions such as information collection, filtering, voice generation, and voice playback.
[0515] The "Internet" is a global network that connects computers around the world and enables the exchange of information.
[0516] "Related information" is data collected based on an information category selected by the user, and is information that contains content that is useful to the user.
[0517] "Filtering" is the process of selecting only information relevant to the user from the collected information and removing unnecessary information.
[0518] A "voice script" is a text for communicating information by voice that is generated based on collected and filtered information.
[0519] A "speech synthesis engine" is a software and hardware component for converting text information into speech data.
[0520] "Audio file" means a digital file generated by a speech synthesis engine to provide audio information to a user.
[0521] A "terminal" is a device that allows a user to receive audio files and play the information, including smartphones, smart glasses, and head-mounted displays.
[0522] A "virtual store" is a commercial facility built in a virtual space where users can browse and purchase products online.
[0523] "Product category" indicates the type of product in which the user is interested, and includes home appliances, fashion, accessories, and the like.
[0524] "Detailed information" refers to information that a user needs to make a decision about purchasing a product, such as product features, price, specifications, and usage.
[0525] The system for implementing this invention allows a user to select a category of information they wish to receive, collects related information from the Internet based on that category, generates a voice script, and transmits it to the user's terminal for playback. This system also has the function of providing voice information about products in a virtual store. The specific configuration and operation of the system are described below.
[0526] System configuration
[0527] This system mainly uses the following hardware and software:
[0528] Hardware: Smartphones, smart glasses, head-mounted displays (HMDs)
[0529] Software: Natural language processing (NLP) engine, speech synthesis engine (e.g., Google Text-to-Speech API), user data management server, product information database
[0530] Program processing
[0531] 1. User configuration
[0532] Users install the app and select the desired information category (news, weather, product information, etc.) when they first start it. They can also set a name and how they want to be called.
[0533] 2. Information collection by the server
[0534] The device sends the user's setting information to the server, which then collects the latest information from related services (news API, weather information service, product information database) based on the set information category.
[0535] 3. Server filtering and organization of information
[0536] The server analyzes the collected information and uses NLP technology to filter out unnecessary information. For example, in the case of news, it extracts only articles related to topics that the user is interested in, and in the case of product information, it retrieves detailed information about products that interest the user.
[0537] 4. Server-generated voice script
[0538] Based on the filtered information, the server generates an original voice script containing friendly speech, such as "Good morning, Mr. / Ms. X. Today is X / X. First, I'll tell you today's news..." including the user's name.
[0539] 5. Server-based speech synthesis and transmission
[0540] The server sends the generated script to a speech synthesis engine to generate an audio file, which is then sent to the user's device.
[0541] 6. Playing Audio on the Device
[0542] Users can play the audio file sent through the device and receive the set information by voice, allowing them to receive information without looking at the smartphone screen.
[0543] 7. Audio guide function in virtual stores
[0544] In particular, within the virtual store, product information is provided based on product categories set by the user. When the user approaches a product, detailed information about that product is automatically played back in audio.
[0545] Examples and prompts
[0546] Specific examples
[0547] For example, if a user named "Tanaka-san" requests news, weather, and information about home appliances in a virtual store, the server retrieves the latest news from the news API and uses natural language processing to extract only articles most relevant to Tanaka-san. It then retrieves today's weather forecast from a weather information service and organizes weather information relevant to Tanaka-san's location. For home appliance information, the server collects the latest product information within the virtual store and lists important specifications and pricing information. Finally, it generates a voice script that reads, "Good morning, Tanaka-san. Today is October 1, 2023. Let me start with today's news..." and converts it into an audio file using a speech synthesis engine (Google Text-to-Speech API). This file is then sent to Tanaka-san's device, allowing him to easily access the information by voice via his smartphone or HMD.
[0548] Prompt Sentence Examples
[0549] Username: Tanaka-san
[0550] Favorite product category: Home appliances
[0551] Provides the latest information on home appliances via voice
[0552] Example of generated voice script: "Hey Tanaka, here are some recommended new home appliances..."
[0553] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0554] Step 1:
[0555] User category selection and individual settings
[0556] (Input) The user installs the app and selects the desired information category (news, weather, product information, etc.) when launching it for the first time. They also set a name and how they want to be called.
[0557] (Processing) The app collects user selections and settings.
[0558] (Output) The collected setting information is stored in the terminal and also sent to the server.
[0559] Step 2:
[0560] Server collection of information
[0561] (Input) The server collects related information based on the selected information category, based on the user setting information received from the terminal.
[0562] The (processing) server obtains the necessary data from news APIs, weather information services, product information databases, etc.
[0563] (Output) The acquired data is temporarily stored on the server.
[0564] Step 3:
[0565] Filtering and organizing information
[0566] (Input) Collected raw data (news articles, weather information, product information, etc.).
[0567] The (processing) server uses a natural language processing (NLP) engine to extract only the information relevant to the user and filter out unnecessary information. For example, in the case of a news category, it selects only important and reliable articles. In the case of product information, it organizes detailed information related to items that the user is interested in.
[0568] (Output) The filtered and organized information is stored on the server.
[0569] Step 4:
[0570] Generate voice scripts
[0571] (Input) Filtered information.
[0572] The (processing) server generates a friendly, customized voice script based on the received information. For example, it creates text in the form of "Good morning, Mr. / Ms. XX. Today is the XXth month. First, I'll tell you today's news..."
[0573] (Output) The generated voice script is saved on the server.
[0574] Step 5:
[0575] Speech synthesis and transmission
[0576] (Input) The generated voice script.
[0577] The (processing) server uses a speech synthesis engine (e.g., Google Text-to-Speech API) to convert the text-based voice script into an audio file.
[0578] (Output) The converted audio file is generated and sent to the user's device.
[0579] Step 6:
[0580] Playing audio through the device
[0581] (Input) The audio file sent from the server.
[0582] (Processing) The user terminal receives the audio file and plays the audio using the built-in media player.
[0583] (Output) The user can receive the information set through the terminal by voice.
[0584] Step 7:
[0585] Audio guide function in virtual stores
[0586] (Input) Product category and its location information set by the user.
[0587] (Processing) When a user walks around the virtual store and comes close to a certain product, detailed information about that product is automatically played back using the device's sensors and location information. Specifically, the device uses distance sensors and GPS data to determine the user's current location, and plays back related product information as an audio file.
[0588] (Output) The user can receive voice information about products that interest them in the virtual store.
[0589] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0590] This invention combines a system in which a user selects the information category they wish to receive, collects related information from the Internet based on the selected information category, filters it, generates voice scripts, and plays voice files, with an emotion engine that recognizes the user's emotions and adjusts the information and voice based on those emotions. Below, the program processing of this system is explained in natural language.
[0591] 1. User configuration
[0592] When users install the app and launch it for the first time, they select and enter the categories of information they want to receive (news, weather, schedule, etc.), as well as their own name and a familiar greeting. This allows them to receive information that is individually customized.
[0593] 2. Emotion Recognition by Emotion Engine
[0594] To recognize the user's emotions, the device analyzes data acquired from the camera and microphone in real time, thereby detecting the user's current emotions (e.g., joy, sadness, anger, surprise, etc.). The emotion engine analyzes this data using an emotion analysis model to determine the user's emotional state.
[0595] 3. Sending setting information from the device to the server
[0596] The device sends the user-entered setting information and the emotion information recognized by the emotion engine to the server, which then stores this information in a database as a user profile.
[0597] 4. Information Collection
[0598] Based on the user profile, the server collects the latest data from sources such as news APIs, weather information services, and calendar services. For example, it gets the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0599] 5. Information Analysis and Filtering
[0600] The server analyzes the collected information and uses natural language processing technology to extract only information relevant to the user. For example, news information will extract articles that match the user's interest categories, weather information will extract weather forecasts for the user's location, and schedule information will extract important appointments. Additionally, based on the emotional information provided by the emotion engine, the server prioritizes and selects information that best suits the user's current emotional state.
[0601] 6. Generate voice script
[0602] The server generates an original voice script based on the filtered information. This script includes a friendly greeting (e.g., "Good morning, Tanaka-san") and selected information categories. For example, "Today is the ____ day of the month. Here's today's news..." Furthermore, the script incorporates phrases with adjusted tone and speed according to the emotions recognized by the emotion engine.
[0603] 7. Speech Synthesis
[0604] The server sends the generated script to a speech synthesis engine (e.g., a Text-to-Speech engine) and converts it into an audio file, which is generated in a format that is easy for the user to listen to (e.g., MP3 format).
[0605] 8. Sending audio files
[0606] The server sends the generated audio file to the user's device, which then stores the received audio file in the appropriate folder.
[0607] 9. Audio playback
[0608] The user plays the audio file sent through the device. By pressing the play button, the audio begins, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This provides a healthy way to obtain information and prevents excessive smartphone use.
[0609] Specific examples
[0610] For example, if a user has the name "Tanaka-san" and has set that they want to receive news, weather, and today's schedule, and the emotion engine recognizes that the current emotion is "happy," the server will retrieve the latest news from the news API and use natural language processing to extract only articles that are highly relevant to Tanaka-san. It will also retrieve today's weather forecast from the weather information service and organize the weather information for Tanaka-san's location. It will then retrieve today's schedule from the calendar service and list important events. The emotion engine will then synthesize speech in a positive tone to match the user's "happy" state.
[0611] Finally, a voice script is generated that says, "Good morning, Tanaka. Today is October 1, 2023. Today's news is very interesting..." and converted into an audio file using a speech synthesis engine. This file is then sent to Tanaka's device, allowing him to receive the information by voice without having to look at his smartphone.
[0612] The above is a specific embodiment for carrying out the present invention. By configuring the system in this way, it is possible to efficiently obtain necessary information in a manner that is in line with the user's emotions, and reduce the risk of "using a smartphone while sleeping."
[0613] The processing flow will be explained below.
[0614] Step 1:
[0615] User-defined
[0616] When users install the app and launch it for the first time, they select and input the categories of information they want to receive (news, weather, schedule, etc.), as well as their own name and a familiar greeting. This allows them to receive information that is individually customized.
[0617] Step 2:
[0618] Emotion recognition by emotion engine
[0619] The device captures and analyzes data from the camera and microphone in real time to recognize the user's emotions, such as whether the user is happy, sad, angry, or surprised, by analyzing facial expressions and voice tone.
[0620] Step 3:
[0621] Sending setting information from the device to the server
[0622] The device transmits the user's setting information and the emotion information recognized by the emotion engine to the server, where the transmitted data is stored in the user profile and used for subsequent processing.
[0623] Step 4:
[0624] Information gathering
[0625] Based on the user profile, the server collects the latest data from information sources such as a news API, weather information service, and calendar service. For example, it obtains the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0626] Step 5:
[0627] Information Analysis and Filtering
[0628] The server analyzes the collected information and uses natural language processing technology to extract only information that is highly relevant to the user. Furthermore, based on the emotional data provided by the emotion engine, it prioritizes information that matches the user's current emotional state. For example, if the user is "happy," it prioritizes positive news and information that enhances emotions.
[0629] Step 6:
[0630] Generate voice scripts
[0631] The server generates an original voice script based on the filtered information. This script includes friendly greetings (e.g., "Good morning, Tanaka-san") and phrases that correspond to emotions. For example, if the user is in a "happy" state, the script might say, "I have some very exciting news for you today..."
[0632] Step 7:
[0633] Speech synthesis
[0634] The server then sends the generated script to a speech synthesis engine (e.g., a text-to-speech engine) and converts it into an audio file. This audio file is generated in a smooth listening format (e.g., MP3 format), and the tone and speed of the audio are adjusted based on the user's emotions.
[0635] Step 8:
[0636] Sending an audio file
[0637] The server sends the generated audio file to the user's device, which stores the received audio file in an appropriate folder and prepares it for playback.
[0638] Step 9:
[0639] Playing audio
[0640] The user plays the audio file sent through the device. By pressing the play button, a voice will play saying, "Good morning, Tanaka-san..." and the user can listen to the necessary information without looking at the screen. This prevents excessive smartphone use and provides a healthy way to obtain information.
[0641] In this way, the system of the present invention can adjust the content of information and the tone of voice according to the user's emotional state, providing an optimal information acquisition experience.
[0642] Example 2
[0643] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0644] In modern society, many users use the Internet to obtain information, but the overwhelming amount of information available makes it difficult to quickly find relevant information. Users who want to easily obtain information need a means to utilize not only their eyes but also their ears, but there is a lack of systems that can accommodate individual needs and emotional states. Furthermore, because excessive smartphone use can cause health problems, there is a need for a means to obtain information without looking at the screen.
[0645] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0646] In this invention, the server includes means for allowing a user to select an information category they wish to receive, means for collecting related information from a data network based on the selected information category, and means for analyzing the collected information using natural language processing technology and extracting only information relevant to the user. This allows the user to easily obtain only the information that suits them from an excess of information and to receive that information by voice.
[0647] The "means for selecting the information category that the user wishes to receive" refers to an input device or interface that allows the user to specify the type of information in which the user is interested (for example, news, weather, schedule, etc.).
[0648] A "data network" is a communication path through which multiple computing resources exchange information with each other, including the Internet and other information and communication networks.
[0649] "Natural language processing technology" refers to computer technology for understanding, analyzing, and generating human language (e.g., keyword extraction, document classification, sentiment analysis, etc.).
[0650] "Means for recognizing an emotional state and adjusting information based on that emotional state" refers to technology or a device that reads emotions from a user's facial expressions, voice, etc., and changes the content and expression of the information provided according to those emotions.
[0651] "Means for generating voice scripts" refers to technology or devices that automatically create text to be read aloud based on collected and analyzed information.
[0652] A "speech synthesis engine" is a computer technology (for example, text-to-speech technology) for converting text in text format into voice data.
[0653] A "terminal" is an information processing device (for example, a smartphone, tablet, or PC) that can be directly operated by a user.
[0654] "News information" is information that reports the latest facts and events related to society, economy, sports, culture, etc.
[0655] "Weather information" refers to data related to the weather, such as weather forecasts, temperature, precipitation, and wind speed.
[0656] "Schedule information" is information about schedules and events listed in a user's schedule or calendar.
[0657] "Means for including a user name or friendly nickname" refers to technology or a device for incorporating a name set by the user or a nickname that gives a sense of familiarity to the user (for example, "san" or "kun") into the voice script.
[0658] This invention combines a system that allows a user to select an information category of interest, collects related information from the Internet based on that information, filters it, generates a voice script, and plays back a voice file, with an emotion engine that recognizes the user's emotions and adjusts the information and voice based on those emotions. A specific method for implementing this system will be described below.
[0659] User-defined
[0660] Users install a dedicated application on their device and, when they first start it up, set the categories of information they want to receive (e.g., news, weather, schedule), as well as their own name and a friendly greeting. This allows for individually customized information to be provided.
[0661] Emotion recognition by emotion engine
[0662] The device uses a camera and microphone to capture the user's facial expressions and voice data, and analyzes it in real time using an emotion analysis model (e.g., OpenFace or DeepFace). This identifies the user's emotional state and sends that information to the server as data. This data is used for information filtering and voice script generation, which will be described later.
[0663] Sending configuration information to the server
[0664] The device sends the information entered and set by the user, as well as the emotional data recognized by the emotion engine, to the server, which then stores this information in a database as a user profile and uses it to provide the most appropriate information to each individual user.
[0665] Information gathering
[0666] The server collects the necessary data from external sources such as news APIs, weather information services, and calendar services. For example, it obtains the latest news from a news API (e.g., Google News API), local weather forecasts from a weather information service (e.g., OpenWeatherMap API), and user schedule information from a calendar service (e.g., Google Calendar API).
[0667] Information Analysis and Filtering
[0668] The server analyzes the collected data using natural language processing technology and scrutinizes information relevant to the user. For example, for news information, keywords are extracted to extract articles that match the user's categories of interest, and for weather information, only information related to the user's location is extracted. Based on the recognition results of the emotion engine, information that matches the user's current emotions is selected.
[0669] Generate voice scripts
[0670] The server generates a voice script based on the filtered information. This voice script includes the user's name and a friendly greeting, and incorporates tones and expressions that correspond to the user's emotional state. For example, it includes gentle expressions such as, "It's a beautiful day today. Let's have a good day."
[0671] Speech synthesis
[0672] The server sends the generated script to a speech synthesis engine (e.g., Google Text-to-Speech API) and converts it into an audio file, which is generated in, for example, MP3 format.
[0673] Sending an audio file
[0674] The server sends the generated audio file to the user's device, which stores the received audio file in an appropriate folder and prepares it for playback.
[0675] Playing audio
[0676] Users can play the audio files sent through the device. By pressing the play button, they can hear a voice such as, "Good morning, Tanaka-san. Here's today's news..." and can listen to the necessary information without looking at the screen.
[0677] Specific examples
[0678] For example, if a user named "Tanaka-san" selects to receive news, weather, and today's schedule, and the emotion engine recognizes that the current emotion is "happy," the server collects and filters data from each source. Based on the emotion, a voice script with a positive tone is generated. Finally, the voice script, "Good morning, Tanaka-san. Today is October 1, 2023. Today's news is very interesting...," is converted into an audio file and sent to Tanaka-san's device.
[0679] Prompt Sentence Examples
[0680] For example, the following prompt sentence is fed into the generative AI model:
[0681] "Includes news, weather, and schedule categories. Current emotion is joy. Name is Tanaka. Uses familiar greeting. Send generated audio file to device."
[0682] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0683] Step 1: User Setup
[0684] A user launches the application on their device and selects the information categories they want to receive (e.g., news, weather, schedule), and also sets their name and a friendly greeting (e.g., "Mr. Tanaka"). This set information is stored in the device's local database.
[0685] Input: Category of information you want to receive, username, call
[0686] Output: Configuration information stored in the local database
[0687] Step 2: Emotion recognition by the emotion engine
[0688] The device uses a camera and microphone to collect the user's facial and voice data in real time. The collected data is analyzed using an emotion analysis model (e.g., OpenFace or DeepFace) to identify the user's emotional state (e.g., joy, sadness, anger). The analysis results are temporarily stored on the device for later use.
[0689] Input: User's facial expression data, voice data
[0690] Output: Parsed emotional state data
[0691] Step 3: Sending configuration information to the server
[0692] The terminal transmits the user's selected information categories and analyzed emotional state data to the server, which receives this information and stores it in a database as a user profile.
[0693] Input: Setting information stored in a local database, analyzed emotional state data
[0694] Output: User profile data sent to the server
[0695] Step 4: Gather information
[0696] Based on the collected user profile data, the server collects relevant information from external sources such as news APIs, weather information services, calendar services, etc. For example, the server retrieves the latest news articles from the news API and local weather forecasts from the weather information service.
[0697] Input: User profile data
[0698] Output: Raw data collected from external sources (news, weather, schedule information)
[0699] Step 5: Information analysis and filtering
[0700] The server analyzes the collected raw data using natural language processing technology. For example, it can extract articles that match the user's interest categories from news data, extract information related to the user's location from weather data, and adjust the priority and content of information based on the user's emotional state.
[0701] Input: Raw data collected from external sources, user emotional state data
[0702] Output: filtered and adjusted information data
[0703] Step 6: Generate the voice script
[0704] The server generates a voice script based on the filtered and adjusted data, which includes the user's name and a friendly greeting, and adjusts the tone and expression depending on the user's emotional state.
[0705] Input: Filtered and conditioned information data
[0706] Output: The generated voice script
[0707] Step 7: Text-to-Speech
[0708] The server sends the generated voice script to a text-to-speech engine (e.g., Google Text-to-Speech API) and converts it into an audio file (e.g., MP3 format). The converted audio file is temporarily stored on the server.
[0709] Input: Generated voice script
[0710] Output: Converted audio file
[0711] Step 8: Send the audio file
[0712] The server sends the generated voice file to the user's device, which saves the received voice file in the appropriate folder (e.g., " / user / voices / ").
[0713] Input: Converted audio file
[0714] Output: Audio file saved on your device
[0715] Step 9: Playing Audio
[0716] The user plays the audio file through the device. By pressing the play button, the audio will play, "Good morning, Tanaka-san. Here's today's news..." and the user can hear the necessary information without looking at the screen.
[0717] Input: Audio files stored on the device
[0718] Output: Audio information played to the user
[0719] (Application example 2)
[0720] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0721] In order to provide optimal services based on customer emotions, commercial facilities and brick-and-mortar stores require a system that recognizes customer emotions in real time and collects and provides information based on those emotions. However, current systems lack the functionality to recognize and appropriately reflect customer emotions, making generalization and automation difficult. Furthermore, the collected information is not optimized for customer emotions, limiting the improvement of customer satisfaction.
[0722] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0723] In this invention, the server includes means for selecting an information category that the user wishes to receive, means for collecting related information from the Internet based on the selected information category, means for filtering the collected information and extracting only information relevant to the user, means for generating an original voice script based on the extracted information, means for converting the generated script into a voice file using a voice synthesis engine, means for transmitting the voice file to the user's terminal, means for playing the transmitted voice file, means for analyzing data acquired from a camera or microphone and recognizing the user's emotions, and means for adjusting information and voice based on the recognized emotional information. This makes it possible to efficiently provide appropriate information in accordance with the customer's emotions and improve customer satisfaction.
[0724] The "means for selecting the information category that the user wants to receive" is a function that allows the user to select the type of information that the user wants to receive based on their own interests and concerns.
[0725] The "means for collecting related information from the Internet" is a function for obtaining data related to the information category selected by the user from various information sources on the Internet.
[0726] "Means of filtering collected information and extracting only information relevant to the user" refers to a function that selects the most relevant information from the acquired information based on the user's interests and preferences.
[0727] The "means for generating an original voice script based on extracted information" is a function for generating a script based on filtered information to be read aloud in a form that is familiar to the user.
[0728] "Means for converting the generated script into an audio file using a voice synthesis engine" is a function for converting the generated voice script into actual voice data using voice synthesis technology.
[0729] The "means for transmitting the audio file to the user's terminal" is a function for transmitting the generated audio file to the terminal used by the user via the Internet or other communication means.
[0730] The "means for playing transmitted audio files" is a function for playing audio files stored in the user's terminal, allowing the user to listen to audio information.
[0731] "Means for analyzing data obtained from a camera or microphone and recognizing the user's emotions" refers to a function that recognizes the user's emotional state by capturing and analyzing the user's facial expressions and voice using a camera or microphone installed on the device.
[0732] The "means for adjusting information or voice based on recognized emotional information" is a function for presenting information in an optimal form or adjusting the tone or content of voice based on the acquired emotional information of the user.
[0733] A system for implementing this invention allows a user to select the information categories they wish to receive, collects and filters relevant information based on those categories, generates audio scripts and plays audio files, and recognizes the user's emotions and adjusts the information and audio based on those emotions.
[0734] 1. User Settings
[0735] When the user starts the system for the first time, they set the category of information they want to receive (news, weather, schedule, etc.), their name, and a friendly greeting. These settings allow the system to retrieve and provide information tailored to the user.
[0736] 2. Emotion recognition
[0737] The device uses a camera and microphone to recognize emotions in real time from the user's facial expressions and voice. Specifically, it uses an emotion engine to analyze facial expressions using images captured by the camera and emotions from the voice. A library called DeepFace is used for emotion analysis, and a voice recognition API is used for voice analysis.
[0738] 3. Information gathering
[0739] The server collects the latest data from various information sources, such as a news API, weather information service, and calendar service, based on the user's settings and emotion information. For example, it obtains the latest articles from the news API, the local weather forecast from the weather information service, and schedule information from the calendar service.
[0740] 4. Information analysis and filtering
[0741] The server analyzes the collected information using natural language processing technology to extract only the information relevant to the user. At the same time, based on the emotional information recognized by the emotion engine, it prioritizes and selects the information that best suits the user's current emotional state.
[0742] 5. Voice script generation
[0743] The server generates an original voice script based on the filtered information, which includes the user's name, a friendly greeting, the selected information category, and phrases with adjusted tone and speed depending on the emotional information.
[0744] 6. Speech synthesis
[0745] The generated script is sent to a speech synthesis engine (e.g., a text-to-speech engine) and converted into an audio file, which generates an audio file in a format that is easy for users to listen to.
[0746] 7. Sending and playing audio files
[0747] The generated audio file is sent to the user's device, which receives it and saves it in the appropriate folder. The user can immediately listen to the audio information by pressing the play button.
[0748] Specific examples
[0749] For example, if a user named "Sato" specifies that they would like to receive news, weather, and today's schedule information, and the emotion engine recognizes that the current emotion is "happy," the server retrieves the latest news articles from the news API and uses natural language processing technology to extract only those articles that are highly relevant to Mr. Sato. It also retrieves weather information related to Mr. Sato's location from the weather information service and lists today's schedule from the calendar service. The emotion engine synthesizes voice in a positive tone that matches the user's "happy" state.
[0750] Example prompts for generative AI models
[0751] "The user's name is Sato. I'm happy to hear from Sato. Please let me know about new products that might interest Sato."
[0752] This makes it possible to efficiently provide appropriate information in a manner that is in line with the customer's feelings, thereby improving customer satisfaction.
[0753] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0754] Step 1: User Setup
[0755] Users launch the application installed on their device, select the category of information they want to receive (news, weather, schedule, etc.), and enter their name or a friendly greeting.
[0756] Input (user): Information category, name, call
[0757] Output (terminal): Save as configuration information
[0758] Step 2: Emotion Recognition
[0759] The device uses a camera and microphone to capture the user's facial expressions and voice in real time, and then analyzes the captured data using an emotion engine (e.g., the DeepFace library or a voice recognition API) to recognize the user's emotions.
[0760] Input (terminal): Camera video, audio data
[0761] Output (terminal): Emotion recognition results (e.g., joy, sadness, anger)
[0762] Step 3: Sending preferences and emotions
[0763] The device sends the setting information entered by the user and the recognized emotion information to the server, where they are stored in a database as a user profile.
[0764] Input (device): setting information, emotion recognition results
[0765] Output (Server): Save as user profile
[0766] Step 4: Gather information
[0767] Based on the user profile, the server collects the latest data from news APIs, weather information services, calendar services, etc. For example, it obtains the latest articles from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0768] Input (server): User profile
[0769] Output (server): Collected information (news articles, weather forecasts, schedules)
[0770] Step 5: Information analysis and filtering
[0771] The server analyzes the collected information using natural language processing technology to extract only the information relevant to the user. At the same time, it prioritizes and selects the information that best suits the user's current emotional state based on the emotional information provided by the emotion engine.
[0772] Input (server): Collected information, emotional information
[0773] Output (server): Filtered information
[0774] Step 6: Generate voice script
[0775] The server generates an original voice script based on the filtered information, including a user-friendly call and selected information categories, and adjusts the tone and speed of the voice based on the emotional information.
[0776] Input (server): filtered information, emotional information
[0777] Output (server): Voice script
[0778] Step 7: Text-to-Speech
[0779] The server sends the generated script to a speech synthesis engine and converts it into an audio file (e.g., MP3 format), which generates an audio file in a format that is easy for users to listen to.
[0780] Input (server): Voice script
[0781] Output (server): Audio file
[0782] Step 8: Send and save the audio file
[0783] The server sends the generated audio file to the user's terminal, which receives it and saves it in an appropriate folder.
[0784] Input (server): Audio file
[0785] Output (Device): Saved audio file
[0786] Step 9: Play the audio file
[0787] The user plays the audio file stored on the device and can listen to the information by pressing the play button.
[0788] Input (user): Playback instructions
[0789] Output (terminal): Playback of audio information
[0790] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0791] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0792] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0793] [Third embodiment]
[0794] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0795] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0796] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0797] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0798] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0799] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0800] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0801] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0802] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0803] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0804] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0805] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0806] This invention relates to a system that allows a user to select the information category they wish to receive, collects related information from the Internet based on the selected information category, filters it, generates an audio script, and plays an audio file. The program processing of this system is explained below in natural language.
[0807] 1. User configuration
[0808] Users install the app and select the categories of information they want to receive (news, weather, schedule, etc.) when they first launch it. They can also set a username and a friendly greeting, allowing them to receive information that is individually customized.
[0809] 2. Collection of information
[0810] The device sends the user's configuration information to the server, which then collects the latest information from data providers such as news APIs, weather information services, and calendar services based on the configured information categories.
[0811] 3. Filtering and organizing information
[0812] The server analyzes the collected information and uses natural language processing technology to filter out unnecessary information. For example, in the case of news, it extracts only articles related to topics that interest the user, and in the case of weather information, it extracts the forecast for the user's location. In addition, in the case of schedule information, it lists important events and organizes them based on the user's interests.
[0813] 4. Generate voice script
[0814] Based on the filtered information, the server generates an original voice script containing friendly speech, such as "Good morning, Mr. / Ms. X. Today is X / X. First, I'll tell you today's news..."
[0815] 5. Speech synthesis and transmission
[0816] The server sends the generated script to a speech synthesis engine, which converts it into an audio file, which is then sent to the user's device.
[0817] 6. Audio playback
[0818] Users can play the audio file sent through their device and receive the information they have set by voice. This allows them to receive the necessary information without looking at the smartphone screen, reducing the risks of using their smartphone while sleeping and providing users with a convenient and healthy way to obtain information.
[0819] Specific examples
[0820] For example, if a user named "Tanaka-san" specifies that they want to receive news, weather, and today's schedule, the server retrieves the latest news from the news API and uses natural language processing to extract only articles that are highly relevant to Tanaka-san. It also retrieves today's weather forecast from the weather information service and organizes weather information related to Tanaka-san's location. Finally, it retrieves today's schedule from the calendar service and lists important events.
[0821] Finally, a voice script is generated that says, "Good morning, Tanaka-san. Today is October 1, 2023. Let me start with today's news..." and converted into an audio file using a speech synthesis engine. This file is then sent to Tanaka's device, allowing him to receive the information by voice without having to look at his smartphone.
[0822] The above is a specific embodiment for carrying out the present invention. By configuring the system in this way, users can efficiently obtain the necessary information and reduce the risk of using their smartphone while sleeping.
[0823] The processing flow will be explained below.
[0824] Step 1:
[0825] User-defined
[0826] Users install the app and upon first launching it, select and enter the information categories they want to receive (news, weather, schedule, etc.), as well as their name and a familiar greeting. This allows for individually customized information acquisition.
[0827] Step 2:
[0828] Sending setting information from the device to the server
[0829] The device sends the user-entered configuration information, including selected information categories and user name, to the server, which stores this information in a database as a user profile.
[0830] Step 3:
[0831] Information gathering
[0832] Based on the user profile, the server collects the latest data from sources such as news APIs, weather information services, and calendar services. For example, it gets the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0833] Step 4:
[0834] Information Analysis and Filtering
[0835] The server analyzes the collected information and uses natural language processing technology to extract only information relevant to the user. For example, news information might extract articles that match the user's interest categories, weather information might extract weather forecasts for the user's location, and schedule information might extract important events.
[0836] Step 5:
[0837] Generate voice scripts
[0838] The server generates an original voice script based on the filtered information. This script includes a friendly greeting (e.g., "Good morning, Tanaka-san") and selected information categories. For example, "Today is the ____ day of the month. Here's today's news..."
[0839] Step 6:
[0840] Speech synthesis
[0841] The server sends the generated script to a speech synthesis engine (e.g., a Text-to-Speech engine) and converts it into an audio file, which is generated in a format that is easy for the user to listen to (e.g., MP3 format).
[0842] Step 7:
[0843] Sending an audio file
[0844] The server sends the generated audio file to the user's device, which then stores the received audio file in the appropriate folder.
[0845] Step 8:
[0846] Playing audio
[0847] The user plays the audio file sent through the device. By pressing the play button, the audio begins, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This provides a healthy way to obtain information and prevents excessive smartphone use.
[0848] Example 1
[0849] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0850] In modern society, users must efficiently obtain a large amount of information, but visual information acquisition methods involve health risks and hassle, such as using a smartphone while sleeping. Furthermore, because the information users want to receive is diverse, there is a need for a method that can centrally extract only the necessary information and provide it in a user-friendly format. The purpose of this invention is to solve these problems and provide a means for users to acquire information efficiently and healthily.
[0851] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0852] In this invention, the server includes: means for allowing a user to select an information category they wish to receive; means for collecting related information from a network based on the selected information category; and means for filtering the collected information using natural language processing technology to extract only information relevant to the user. This allows users to efficiently obtain information that is appropriate for them. The server also includes means for generating a friendly voice script using a generative AI model; means for converting the generated script into an audio file using a speech synthesis engine; means for transmitting the audio file to the user's device; and means for playing the transmitted audio file. This allows users to obtain information in a healthy way without relying on visual information acquisition, enabling them to use information efficiently and conveniently.
[0853] "User" refers to a person who uses the system to receive information.
[0854] "Information category" refers to the type of information that a user wants to receive, such as news, weather forecast, schedule information, etc.
[0855] "Network" refers to various communication infrastructures, including the Internet, and is used as a means for collecting and transmitting information.
[0856] "Related information" refers to data that is relevant to the user and is obtained based on the information category set by the user.
[0857] "Natural language processing technology" refers to the technology of analyzing text data to understand meaning and extract information.
[0858] "Generative AI model" refers to an artificial intelligence model that generates natural language based on given prompts.
[0859] "Voice script" refers to a sentence that expresses information provided to a user in a voice form.
[0860] A "speech synthesis engine" refers to a technology or system that converts text data into voice data.
[0861] "Audio file" refers to a file that stores audio data generated by speech synthesis.
[0862] "Device" refers to an electronic device that a user possesses and that receives and reproduces information.
[0863] This invention relates to a system that allows a user to select the information category they wish to receive, collects related information from a network based on the selected information category, filters it, generates a voice script, and plays back a voice file. The program processing of this system is explained below in natural language.
[0864] First, the user installs and launches a dedicated application on a device such as a smartphone or tablet. When launching the application for the first time, the user selects the categories of information they wish to receive (news, weather forecasts, schedule information, etc.) on the account settings screen, and also sets their name and how they want to be addressed. This information is sent from the user's device to the server via the network.
[0865] The server collects related information from various information providers based on the received user setting information. For example, news information is obtained using a news API (e.g., NewsAPI or Google News API), weather forecast information is obtained from a weather service (e.g., OpenWeatherMap), and schedule information is obtained from a calendar service (e.g., Google Calendar).
[0866] The server analyzes the large amount of collected data using natural language processing technology (e.g., SpaCy or IBM Watson NLP) and filters out unnecessary information. During this process, it extracts only news articles related to topics that interest the user, extracts information about the user's location from weather forecast information, and lists important schedule information.
[0867] Next, based on the filtered information, the server uses a generative AI model (e.g., OpenAI's ChatGPT) to generate a friendly voice script, using prompts such as "Generate a friendly script that starts with a morning greeting and covers today's news, weather, and schedule."
[0868] The generated script is passed to a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech) and converted into an audio file, which is then sent from the server to the user's device.
[0869] Finally, users can play audio files sent through the device and receive the information they have set by voice, which allows them to obtain information efficiently and healthily without relying on visual information acquisition.
[0870] For example, if a user named "Tanaka-san" requests to receive news, weather forecasts, and schedule information, the server retrieves the latest news from the news API and uses natural language processing technology to extract only articles that are relevant to Tanaka-san. It then retrieves today's weather forecast from a weather service and organizes information about Tanaka-san's location. It then retrieves today's schedule from a calendar service and lists important events. Finally, it generates a voice script that says, "Good morning, Tanaka-san. Today is October 1, 2023. Let's start with today's news..." and converts it into an audio file using a speech synthesis engine. This file is then sent to Tanaka-san's device, allowing her to receive the information by voice without looking at her smartphone.
[0871] The above is a specific embodiment for carrying out the present invention. By configuring such a system, users can efficiently obtain necessary information, reduce the risk of using a smartphone while sleeping, and provide a convenient and healthy means of obtaining information.
[0872] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0873] Step 1:
[0874] The user installs and launches the dedicated application. When the application is launched for the first time, the user selects the information categories they wish to receive (news, weather forecasts, schedule information, etc.) on the account settings screen, and sets their name and how they want to be addressed. This input information includes data such as the user name, information category, and how they want to be addressed. The device receives this setting information as input, generates an API request to send it to the server, and then sends it.
[0875] Step 2:
[0876] The server receives the setting information sent from the device and stores it in a database. Next, based on the user's setting information, it retrieves news data from a news API (e.g., NewsAPI or Google News API). Specifically, it creates an API request to collect news data. This collected news data becomes the input.
[0877] Step 3:
[0878] Similarly, the server retrieves weather forecast data from OpenWeatherMap and schedule information from Google Calendar. It sends an API request to each service and retrieves the respective data (weather forecast data, schedule information data). This becomes the input for each service.
[0879] Step 4:
[0880] The server analyzes the collected news data, weather forecast data, and schedule information data, and filters out unnecessary information using natural language processing technology (e.g., SpaCy or IBM Watson NLP). It receives the collected data as input and extracts only the information that is highly relevant to the user through a filtering process. This is the filtered data that is output.
[0881] Step 5:
[0882] The server uses a generative AI model (e.g., OpenAI's ChatGPT) to generate a friendly voice script based on the filtered data. The input to this process is the filtered data, and the prompt "Generate a friendly script that starts with a morning greeting and covers today's news, weather, and schedule." The output is the generated voice script.
[0883] Step 6:
[0884] The server passes the generated voice script to a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech) and converts it into an audio file. The input for this process is the voice script, which the speech synthesis engine analyzes and outputs an audio file. Specifically, the speech synthesis engine generates voice data based on text data and saves it in an audio file.
[0885] Step 7:
[0886] The server creates and sends an API request to send the generated audio file to the user's device. The input of this process is the audio file, and the output is the audio file sent to the user's device.
[0887] Step 8:
[0888] The user's device receives and saves the audio file sent from the server. When the user operates the application to instruct audio playback, the device plays the saved audio file. The input for this process is the audio file, and the output is audio information provided to the user. The user can receive the necessary information by audio without looking at their smartphone.
[0889] The above are the specific processing steps of this system. Users can acquire information efficiently and are provided with a healthy information acquisition method that does not rely on visual information acquisition.
[0890] (Application example 1)
[0891] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0892] Modern users are seeking ways to efficiently and comfortably gather information amid their busy daily lives. However, conventional information gathering systems have limitations in providing personalized information based on users' preferences and interests, and require users to manually search for information. Furthermore, obtaining product information in a virtual store primarily requires visual confirmation, which requires significant time and effort. Therefore, the present invention aims to provide a system that automates the provision of information based on users' preferences and provides product information in a virtual store via voice, allowing users to efficiently and comfortably obtain the information they need.
[0893] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0894] In this invention, the server includes means for selecting an information category that the user wishes to receive, means for collecting related information from the Internet based on the selected information category, means for filtering the collected information and extracting only information relevant to the user, means for generating an original voice script based on the extracted information, means for converting the generated script into a voice file using a voice synthesis engine, means for transmitting the voice file to the user's terminal, means for playing the transmitted voice file, means for selecting a product category that the user wishes to receive and for providing voice information about products in the virtual store, and means for automatically playing voice information about the products when the user approaches a product in the virtual store. This allows the user to efficiently obtain information regardless of time or place, and to obtain voice information about products in the virtual store without relying on vision.
[0895] A "user" is an entity that selects an information category, collects related information from the Internet, and receives the information in audio format.
[0896] "Information category" indicates the type of information that the user wants to receive, and includes news, weather information, schedule information, product information, and the like.
[0897] The term "means" refers to components or methods required to achieve a specific function, and in the present invention includes functions such as information collection, filtering, voice generation, and voice playback.
[0898] The "Internet" is a global network that connects computers around the world and enables the exchange of information.
[0899] "Related information" is data collected based on an information category selected by the user, and is information that contains content that is useful to the user.
[0900] "Filtering" is the process of selecting only information relevant to the user from the collected information and removing unnecessary information.
[0901] A "voice script" is a text for communicating information by voice that is generated based on collected and filtered information.
[0902] A "speech synthesis engine" is a software and hardware component for converting text information into speech data.
[0903] "Audio file" means a digital file generated by a speech synthesis engine to provide audio information to a user.
[0904] A "terminal" is a device that allows a user to receive audio files and play the information, including smartphones, smart glasses, and head-mounted displays.
[0905] A "virtual store" is a commercial facility built in a virtual space where users can browse and purchase products online.
[0906] "Product category" indicates the type of product in which the user is interested, and includes home appliances, fashion, accessories, and the like.
[0907] "Detailed information" refers to information that a user needs to make a decision about purchasing a product, such as product features, price, specifications, and usage.
[0908] The system for implementing this invention allows a user to select a category of information they wish to receive, collects related information from the Internet based on that category, generates a voice script, and transmits it to the user's terminal for playback. This system also has the function of providing voice information about products in a virtual store. The specific configuration and operation of the system are described below.
[0909] System configuration
[0910] This system mainly uses the following hardware and software:
[0911] Hardware: Smartphones, smart glasses, head-mounted displays (HMDs)
[0912] Software: Natural language processing (NLP) engine, speech synthesis engine (e.g., Google Text-to-Speech API), user data management server, product information database
[0913] Program processing
[0914] 1. User configuration
[0915] Users install the app and select the desired information category (news, weather, product information, etc.) when they first start it. They can also set a name and how they want to be called.
[0916] 2. Information collection by the server
[0917] The device sends the user's setting information to the server, which then collects the latest information from related services (news API, weather information service, product information database) based on the set information category.
[0918] 3. Server filtering and organization of information
[0919] The server analyzes the collected information and uses NLP technology to filter out unnecessary information. For example, in the case of news, it extracts only articles related to topics that the user is interested in, and in the case of product information, it retrieves detailed information about products that interest the user.
[0920] 4. Server-generated voice script
[0921] Based on the filtered information, the server generates an original voice script containing friendly speech, such as "Good morning, Mr. / Ms. X. Today is X / X. First, I'll tell you today's news..." including the user's name.
[0922] 5. Server-based speech synthesis and transmission
[0923] The server sends the generated script to a speech synthesis engine to generate an audio file, which is then sent to the user's device.
[0924] 6. Playing Audio on the Device
[0925] Users can play the audio file sent through the device and receive the set information by voice, allowing them to receive information without looking at the smartphone screen.
[0926] 7. Audio guide function in virtual stores
[0927] In particular, within the virtual store, product information is provided based on product categories set by the user. When the user approaches a product, detailed information about that product is automatically played back in audio.
[0928] Examples and prompts
[0929] Specific examples
[0930] For example, if a user named "Tanaka-san" requests news, weather, and information about home appliances in a virtual store, the server retrieves the latest news from the news API and uses natural language processing to extract only articles most relevant to Tanaka-san. It then retrieves today's weather forecast from a weather information service and organizes weather information relevant to Tanaka-san's location. For home appliance information, the server collects the latest product information within the virtual store and lists important specifications and pricing information. Finally, it generates a voice script that reads, "Good morning, Tanaka-san. Today is October 1, 2023. Let me start with today's news..." and converts it into an audio file using a speech synthesis engine (Google Text-to-Speech API). This file is then sent to Tanaka-san's device, allowing him to easily access the information by voice via his smartphone or HMD.
[0931] Prompt Sentence Examples
[0932] Username: Tanaka-san
[0933] Favorite product category: Home appliances
[0934] Provides the latest information on home appliances via voice
[0935] Example of generated voice script: "Hey Tanaka, here are some recommended new home appliances..."
[0936] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0937] Step 1:
[0938] User category selection and individual settings
[0939] (Input) The user installs the app and selects the desired information category (news, weather, product information, etc.) when launching it for the first time. They also set a name and how they want to be called.
[0940] (Processing) The app collects user selections and settings.
[0941] (Output) The collected setting information is stored in the terminal and also sent to the server.
[0942] Step 2:
[0943] Server collection of information
[0944] (Input) The server collects related information based on the selected information category, based on the user setting information received from the terminal.
[0945] The (processing) server obtains the necessary data from news APIs, weather information services, product information databases, etc.
[0946] (Output) The acquired data is temporarily stored on the server.
[0947] Step 3:
[0948] Filtering and organizing information
[0949] (Input) Collected raw data (news articles, weather information, product information, etc.).
[0950] The (processing) server uses a natural language processing (NLP) engine to extract only the information relevant to the user and filter out unnecessary information. For example, in the case of a news category, it selects only important and reliable articles. In the case of product information, it organizes detailed information related to items that the user is interested in.
[0951] (Output) The filtered and organized information is stored on the server.
[0952] Step 4:
[0953] Generate voice scripts
[0954] (Input) Filtered information.
[0955] The (processing) server generates a friendly, customized voice script based on the received information. For example, it creates text in the form of "Good morning, Mr. / Ms. XX. Today is the XXth month. First, I'll tell you today's news..."
[0956] (Output) The generated voice script is saved on the server.
[0957] Step 5:
[0958] Speech synthesis and transmission
[0959] (Input) The generated voice script.
[0960] The (processing) server uses a speech synthesis engine (e.g., Google Text-to-Speech API) to convert the text-based voice script into an audio file.
[0961] (Output) The converted audio file is generated and sent to the user's device.
[0962] Step 6:
[0963] Device audio playback
[0964] (Input) The audio file sent from the server.
[0965] (Processing) The user terminal receives the audio file and plays the audio using the built-in media player.
[0966] (Output) The user can receive the information set through the terminal by voice.
[0967] Step 7:
[0968] Audio guide function in virtual stores
[0969] (Input) Product category and its location information set by the user.
[0970] (Processing) When a user walks around the virtual store and comes close to a certain product, detailed information about that product is automatically played back using the device's sensors and location information. Specifically, the device uses distance sensors and GPS data to determine the user's current location, and plays back related product information as an audio file.
[0971] (Output) The user can receive voice information about products that interest them in the virtual store.
[0972] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0973] This invention combines a system in which a user selects the information category they wish to receive, collects related information from the Internet based on the selected information category, filters it, generates voice scripts, and plays voice files, with an emotion engine that recognizes the user's emotions and adjusts the information and voice based on those emotions. Below, the program processing of this system is explained in natural language.
[0974] 1. User configuration
[0975] When users install the app and start it for the first time, they select and enter the categories of information they want to receive (news, weather, schedule, etc.), as well as their own name and a familiar greeting. This allows them to receive information that is individually customized.
[0976] 2. Emotion Recognition by Emotion Engine
[0977] To recognize the user's emotions, the device analyzes data acquired from the camera and microphone in real time, thereby detecting the user's current emotions (e.g., joy, sadness, anger, surprise, etc.). The emotion engine analyzes this data using an emotion analysis model to determine the user's emotional state.
[0978] 3. Sending setting information from the device to the server
[0979] The device sends the user-entered setting information and the emotion information recognized by the emotion engine to the server, which then stores this information in a database as a user profile.
[0980] 4. Information Collection
[0981] Based on the user profile, the server collects the latest data from sources such as news APIs, weather information services, and calendar services. For example, it gets the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[0982] 5. Information Analysis and Filtering
[0983] The server analyzes the collected information and uses natural language processing technology to extract only information relevant to the user. For example, news information will extract articles that match the user's interest categories, weather information will extract weather forecasts for the user's location, and schedule information will extract important appointments. Additionally, based on the emotional information provided by the emotion engine, the server prioritizes and selects information that best suits the user's current emotional state.
[0984] 6. Generate voice script
[0985] The server generates an original voice script based on the filtered information. This script includes a friendly greeting (e.g., "Good morning, Tanaka-san") and selected information categories. For example, "Today is the ____ day of the month. Here's today's news..." Furthermore, the script incorporates phrases with adjusted tone and speed according to the emotions recognized by the emotion engine.
[0986] 7. Speech Synthesis
[0987] The server sends the generated script to a speech synthesis engine (e.g., a Text-to-Speech engine) and converts it into an audio file, which is generated in a format that is easy for the user to listen to (e.g., MP3 format).
[0988] 8. Sending audio files
[0989] The server sends the generated audio file to the user's device, which then stores the received audio file in the appropriate folder.
[0990] 9. Audio playback
[0991] The user plays the audio file sent through the device. By pressing the play button, the audio begins, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This provides a healthy way to obtain information and prevents excessive smartphone use.
[0992] Specific examples
[0993] For example, if a user has the name "Tanaka-san" and has set that they want to receive news, weather, and today's schedule, and the emotion engine recognizes that the current emotion is "happy," the server will retrieve the latest news from the news API and use natural language processing to extract only articles that are highly relevant to Tanaka-san. It will also retrieve today's weather forecast from the weather information service and organize the weather information for Tanaka-san's location. It will then retrieve today's schedule from the calendar service and list important events. The emotion engine will then synthesize speech in a positive tone to match the user's "happy" state.
[0994] Finally, a voice script is generated that says, "Good morning, Tanaka. Today is October 1, 2023. Today's news is very interesting..." and converted into an audio file using a speech synthesis engine. This file is then sent to Tanaka's device, allowing him to receive the information by voice without having to look at his smartphone.
[0995] The above is a specific embodiment for carrying out the present invention. By configuring the system in this way, it is possible to efficiently obtain necessary information in a manner that is in line with the user's emotions, and reduce the risk of "using a smartphone while sleeping."
[0996] The processing flow will be explained below.
[0997] Step 1:
[0998] User-defined
[0999] When users install the app and launch it for the first time, they select and input the categories of information they want to receive (news, weather, schedule, etc.), as well as their own name and a familiar greeting. This allows them to receive information that is individually customized.
[1000] Step 2:
[1001] Emotion recognition by emotion engine
[1002] The device captures and analyzes data from the camera and microphone in real time to recognize the user's emotions, such as whether the user is happy, sad, angry, or surprised, by analyzing facial expressions and voice tone.
[1003] Step 3:
[1004] Sending setting information from the device to the server
[1005] The device transmits the user's setting information and the emotion information recognized by the emotion engine to the server, where the transmitted data is stored in the user profile and used for subsequent processing.
[1006] Step 4:
[1007] Information gathering
[1008] Based on the user profile, the server collects the latest data from information sources such as a news API, weather information service, and calendar service. For example, it obtains the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[1009] Step 5:
[1010] Information Analysis and Filtering
[1011] The server analyzes the collected information and uses natural language processing technology to extract only information that is highly relevant to the user. Furthermore, based on the emotional data provided by the emotion engine, it prioritizes information that matches the user's current emotional state. For example, if the user is "happy," it prioritizes positive news and information that enhances emotions.
[1012] Step 6:
[1013] Generate voice scripts
[1014] The server generates an original voice script based on the filtered information. This script includes friendly greetings (e.g., "Good morning, Tanaka-san") and phrases that correspond to emotions. For example, if the user is in a "happy" state, the script might say, "I have some very exciting news for you today..."
[1015] Step 7:
[1016] Speech synthesis
[1017] The server then sends the generated script to a speech synthesis engine (e.g., a text-to-speech engine) and converts it into an audio file. This audio file is generated in a smooth listening format (e.g., MP3 format), and the tone and speed of the audio are adjusted based on the user's emotions.
[1018] Step 8:
[1019] Sending an audio file
[1020] The server sends the generated audio file to the user's device, which stores the received audio file in an appropriate folder and prepares it for playback.
[1021] Step 9:
[1022] Playing audio
[1023] The user plays the audio file sent through the device. By pressing the play button, a voice will play saying, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This prevents excessive smartphone use and provides a healthy way to obtain information.
[1024] In this way, the system of the present invention can adjust the content of information and the tone of voice according to the user's emotional state, providing an optimal information acquisition experience.
[1025] Example 2
[1026] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1027] In modern society, many users use the Internet to obtain information, but the overwhelming amount of information available makes it difficult to quickly find relevant information. Users who want to easily obtain information need a means to utilize not only their eyes but also their ears, but there is a lack of systems that can accommodate individual needs and emotional states. Furthermore, because excessive smartphone use can cause health problems, there is a need for a means to obtain information without looking at the screen.
[1028] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1029] In this invention, the server includes means for allowing a user to select an information category they wish to receive, means for collecting related information from a data network based on the selected information category, and means for analyzing the collected information using natural language processing technology and extracting only information relevant to the user. This allows the user to easily obtain only the information that suits them from an excess of information and to receive that information by voice.
[1030] The "means for selecting the information category that the user wishes to receive" refers to an input device or interface that allows the user to specify the type of information in which the user is interested (for example, news, weather, schedule, etc.).
[1031] A "data network" is a communication path through which multiple computing resources exchange information with each other, including the Internet and other information and communication networks.
[1032] "Natural language processing technology" refers to computer technology for understanding, analyzing, and generating human language (e.g., keyword extraction, document classification, sentiment analysis, etc.).
[1033] "Means for recognizing an emotional state and adjusting information based on that emotional state" refers to technology or a device that reads emotions from a user's facial expressions, voice, etc., and changes the content and expression of the information provided according to those emotions.
[1034] "Means for generating voice scripts" refers to technology or devices that automatically create text to be read aloud based on collected and analyzed information.
[1035] A "speech synthesis engine" is a computer technology (for example, text-to-speech technology) for converting text in text format into voice data.
[1036] A "terminal" is an information processing device (for example, a smartphone, tablet, or PC) that can be directly operated by a user.
[1037] "News information" is information that reports the latest facts and events related to society, economy, sports, culture, etc.
[1038] "Weather information" refers to data related to the weather, such as weather forecasts, temperature, precipitation, and wind speed.
[1039] "Schedule information" is information about schedules and events listed in a user's schedule or calendar.
[1040] "Means for including a user name or friendly nickname" refers to technology or a device for incorporating a name set by the user or a nickname that gives a sense of familiarity to the user (for example, "san" or "kun") into the voice script.
[1041] This invention combines a system that allows a user to select an information category of interest, collects related information from the Internet based on that information, filters it, generates a voice script, and plays back a voice file, with an emotion engine that recognizes the user's emotions and adjusts the information and voice based on those emotions. A specific method for implementing this system will be described below.
[1042] User-defined
[1043] Users install a dedicated application on their device and, when they first start it up, set the categories of information they want to receive (e.g., news, weather, schedule), as well as their own name and a friendly greeting. This allows for individually customized information to be provided.
[1044] Emotion recognition by emotion engine
[1045] The device uses a camera and microphone to capture the user's facial expressions and voice data, and analyzes it in real time using an emotion analysis model (e.g., OpenFace or DeepFace). This identifies the user's emotional state and sends that information to the server as data. This data is used for information filtering and voice script generation, which will be described later.
[1046] Sending configuration information to the server
[1047] The device sends the information entered and set by the user, as well as the emotional data recognized by the emotion engine, to the server, which then stores this information in a database as a user profile and uses it to provide the most appropriate information to each individual user.
[1048] Information gathering
[1049] The server collects the necessary data from external sources such as news APIs, weather information services, and calendar services. For example, it obtains the latest news from a news API (e.g., Google News API), local weather forecasts from a weather information service (e.g., OpenWeatherMap API), and user schedule information from a calendar service (e.g., Google Calendar API).
[1050] Information Analysis and Filtering
[1051] The server analyzes the collected data using natural language processing technology and scrutinizes information relevant to the user. For example, for news information, keywords are extracted to extract articles that match the user's categories of interest, and for weather information, only information related to the user's location is extracted. Based on the recognition results of the emotion engine, information that matches the user's current emotions is selected.
[1052] Generate voice scripts
[1053] The server generates a voice script based on the filtered information. This voice script includes the user's name and a friendly greeting, and incorporates tones and expressions that correspond to the user's emotional state. For example, it includes gentle expressions such as, "It's a beautiful day today. Let's have a good day."
[1054] Speech synthesis
[1055] The server sends the generated script to a speech synthesis engine (e.g., Google Text-to-Speech API) and converts it into an audio file, which is generated in, for example, MP3 format.
[1056] Sending an audio file
[1057] The server sends the generated audio file to the user's device, which stores the received audio file in an appropriate folder and prepares it for playback.
[1058] Playing audio
[1059] Users can play the audio files sent through the device. By pressing the play button, they can hear a voice such as, "Good morning, Tanaka-san. Here's today's news..." and can listen to the necessary information without looking at the screen.
[1060] Specific examples
[1061] For example, if a user named "Tanaka-san" selects to receive news, weather, and today's schedule, and the emotion engine recognizes that the current emotion is "happy," the server collects and filters data from each source. Based on the emotion, a voice script with a positive tone is generated. Finally, the voice script, "Good morning, Tanaka-san. Today is October 1, 2023. Today's news is very interesting...," is converted into an audio file and sent to Tanaka-san's device.
[1062] Prompt Sentence Examples
[1063] For example, the following prompt sentence is fed into the generative AI model:
[1064] "Includes news, weather, and schedule categories. Current emotion is joy. Name is Tanaka. Uses familiar greeting. Send generated audio file to device."
[1065] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1066] Step 1: User Setup
[1067] A user launches the application on their device and selects the information categories they want to receive (e.g., news, weather, schedule), and also sets their name and a friendly greeting (e.g., "Mr. Tanaka"). This set information is stored in the device's local database.
[1068] Input: Category of information you want to receive, username, call
[1069] Output: Configuration information stored in the local database
[1070] Step 2: Emotion recognition by the emotion engine
[1071] The device uses a camera and microphone to collect the user's facial and voice data in real time. The collected data is analyzed using an emotion analysis model (e.g., OpenFace or DeepFace) to identify the user's emotional state (e.g., joy, sadness, anger). The analysis results are temporarily stored on the device for later use.
[1072] Input: User's facial expression data, voice data
[1073] Output: Parsed emotional state data
[1074] Step 3: Sending configuration information to the server
[1075] The terminal transmits the user's selected information categories and analyzed emotional state data to the server, which receives this information and stores it in a database as a user profile.
[1076] Input: Setting information stored in a local database, analyzed emotional state data
[1077] Output: User profile data sent to the server
[1078] Step 4: Gather information
[1079] Based on the collected user profile data, the server collects relevant information from external sources such as news APIs, weather information services, calendar services, etc. For example, the server retrieves the latest news articles from the news API and local weather forecasts from the weather information service.
[1080] Input: User profile data
[1081] Output: Raw data collected from external sources (news, weather, schedule information)
[1082] Step 5: Information analysis and filtering
[1083] The server analyzes the collected raw data using natural language processing technology. For example, it can extract articles that match the user's interest categories from news data, extract information related to the user's location from weather data, and adjust the priority and content of information based on the user's emotional state.
[1084] Input: Raw data collected from external sources, user emotional state data
[1085] Output: filtered and adjusted information data
[1086] Step 6: Generate the voice script
[1087] The server generates a voice script based on the filtered and adjusted data, which includes the user's name and a friendly greeting, and adjusts the tone and expression depending on the user's emotional state.
[1088] Input: Filtered and conditioned information data
[1089] Output: The generated voice script
[1090] Step 7: Text-to-Speech
[1091] The server sends the generated voice script to a text-to-speech engine (e.g., Google Text-to-Speech API) and converts it into an audio file (e.g., MP3 format). The converted audio file is temporarily stored on the server.
[1092] Input: Generated voice script
[1093] Output: Converted audio file
[1094] Step 8: Send the audio file
[1095] The server sends the generated voice file to the user's device, which saves the received voice file in the appropriate folder (e.g., " / user / voices / ").
[1096] Input: Converted audio file
[1097] Output: Audio file saved on your device
[1098] Step 9: Playing Audio
[1099] The user plays the audio file through the device. By pressing the play button, the audio will play, "Good morning, Tanaka-san. Here's today's news..." and the user can hear the necessary information without looking at the screen.
[1100] Input: Audio files stored on the device
[1101] Output: Audio information played to the user
[1102] (Application example 2)
[1103] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1104] In order to provide optimal services based on customer emotions, commercial facilities and brick-and-mortar stores require a system that recognizes customer emotions in real time and collects and provides information based on those emotions. However, current systems lack the functionality to recognize and appropriately reflect customer emotions, making generalization and automation difficult. Furthermore, the collected information is not optimized for customer emotions, limiting the improvement of customer satisfaction.
[1105] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1106] In this invention, the server includes means for selecting an information category that the user wishes to receive, means for collecting related information from the Internet based on the selected information category, means for filtering the collected information and extracting only information relevant to the user, means for generating an original voice script based on the extracted information, means for converting the generated script into a voice file using a voice synthesis engine, means for transmitting the voice file to the user's terminal, means for playing the transmitted voice file, means for analyzing data acquired from a camera or microphone and recognizing the user's emotions, and means for adjusting information and voice based on the recognized emotional information. This makes it possible to efficiently provide appropriate information in accordance with the customer's emotions and improve customer satisfaction.
[1107] The "means for selecting the information category that the user wants to receive" is a function that allows the user to select the type of information that the user wants to receive based on their own interests and concerns.
[1108] The "means for collecting related information from the Internet" is a function for obtaining data related to the information category selected by the user from various information sources on the Internet.
[1109] "Means of filtering collected information and extracting only information relevant to the user" refers to a function that selects the most relevant information from the acquired information based on the user's interests and preferences.
[1110] The "means for generating an original voice script based on extracted information" is a function for generating a script based on filtered information to be read aloud in a form that is familiar to the user.
[1111] "Means for converting the generated script into an audio file using a voice synthesis engine" is a function for converting the generated voice script into actual voice data using voice synthesis technology.
[1112] The "means for transmitting the audio file to the user's terminal" is a function for transmitting the generated audio file to the terminal used by the user via the Internet or other communication means.
[1113] The "means for playing transmitted audio files" is a function for playing audio files stored in the user's terminal, allowing the user to listen to audio information.
[1114] "Means for analyzing data obtained from a camera or microphone and recognizing the user's emotions" refers to a function that recognizes the user's emotional state by capturing and analyzing the user's facial expressions and voice using a camera or microphone installed on the device.
[1115] The "means for adjusting information or voice based on recognized emotional information" is a function for presenting information in an optimal form or adjusting the tone or content of voice based on the acquired emotional information of the user.
[1116] A system for implementing this invention allows a user to select the information categories they wish to receive, collects and filters relevant information based on those categories, generates audio scripts and plays audio files, and recognizes the user's emotions and adjusts the information and audio based on those emotions.
[1117] 1. User Settings
[1118] When the user starts the system for the first time, they set the category of information they want to receive (news, weather, schedule, etc.), their name, and a friendly greeting. These settings allow the system to retrieve and provide information tailored to the user.
[1119] 2. Emotion recognition
[1120] The device uses a camera and microphone to recognize emotions in real time from the user's facial expressions and voice. Specifically, it uses an emotion engine to analyze facial expressions using images captured by the camera and emotions from the voice. A library called DeepFace is used for emotion analysis, and a voice recognition API is used for voice analysis.
[1121] 3. Information gathering
[1122] The server collects the latest data from various information sources, such as a news API, weather information service, and calendar service, based on the user's settings and emotion information. For example, it obtains the latest articles from the news API, the local weather forecast from the weather information service, and schedule information from the calendar service.
[1123] 4. Information analysis and filtering
[1124] The server analyzes the collected information using natural language processing technology to extract only the information relevant to the user. At the same time, based on the emotional information recognized by the emotion engine, it prioritizes and selects the information that best suits the user's current emotional state.
[1125] 5. Voice script generation
[1126] The server generates an original voice script based on the filtered information, which includes the user's name, a friendly greeting, the selected information category, and phrases with adjusted tone and speed depending on the emotional information.
[1127] 6. Speech synthesis
[1128] The generated script is sent to a speech synthesis engine (e.g., a text-to-speech engine) and converted into an audio file, which generates an audio file in a format that is easy for users to listen to.
[1129] 7. Sending and playing audio files
[1130] The generated audio file is sent to the user's device, which receives it and saves it in the appropriate folder. The user can immediately listen to the audio information by pressing the play button.
[1131] Specific examples
[1132] For example, if a user named "Sato" specifies that they would like to receive news, weather, and today's schedule information, and the emotion engine recognizes that the current emotion is "happy," the server retrieves the latest news articles from the news API and uses natural language processing technology to extract only those articles that are highly relevant to Mr. Sato. It also retrieves weather information related to Mr. Sato's location from the weather information service and lists today's schedule from the calendar service. The emotion engine synthesizes voice in a positive tone that matches the user's "happy" state.
[1133] Example prompts for generative AI models
[1134] "The user's name is Sato. I'm happy to hear from Sato. Please let me know about new products that might interest Sato."
[1135] This makes it possible to efficiently provide appropriate information in a manner that is in line with the customer's feelings, thereby improving customer satisfaction.
[1136] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1137] Step 1: User Setup
[1138] Users launch the application installed on their device, select the category of information they want to receive (news, weather, schedule, etc.), and enter their name or a friendly greeting.
[1139] Input (user): Information category, name, call
[1140] Output (terminal): Save as configuration information
[1141] Step 2: Emotion Recognition
[1142] The device uses a camera and microphone to capture the user's facial expressions and voice in real time, and then analyzes the captured data using an emotion engine (e.g., the DeepFace library or a voice recognition API) to recognize the user's emotions.
[1143] Input (terminal): Camera video, audio data
[1144] Output (terminal): Emotion recognition results (e.g., joy, sadness, anger)
[1145] Step 3: Sending preferences and emotions
[1146] The device sends the setting information entered by the user and the recognized emotion information to the server, where they are stored in a database as a user profile.
[1147] Input (device): setting information, emotion recognition results
[1148] Output (Server): Save as user profile
[1149] Step 4: Gather information
[1150] Based on the user profile, the server collects the latest data from news APIs, weather information services, calendar services, etc. For example, it obtains the latest articles from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[1151] Input (server): User profile
[1152] Output (server): Collected information (news articles, weather forecasts, schedules)
[1153] Step 5: Information analysis and filtering
[1154] The server analyzes the collected information using natural language processing technology to extract only the information relevant to the user. At the same time, it prioritizes and selects the information that best suits the user's current emotional state based on the emotional information provided by the emotion engine.
[1155] Input (server): Collected information, emotional information
[1156] Output (server): Filtered information
[1157] Step 6: Generate voice script
[1158] The server generates an original voice script based on the filtered information, including a user-friendly call and selected information categories, and adjusts the tone and speed of the voice based on the emotional information.
[1159] Input (server): filtered information, emotional information
[1160] Output (server): Voice script
[1161] Step 7: Text-to-Speech
[1162] The server sends the generated script to a speech synthesis engine and converts it into an audio file (e.g., MP3 format), which generates an audio file in a format that is easy for users to listen to.
[1163] Input (server): Voice script
[1164] Output (server): Audio file
[1165] Step 8: Send and save the audio file
[1166] The server sends the generated audio file to the user's terminal, which receives it and saves it in an appropriate folder.
[1167] Input (server): Audio file
[1168] Output (Device): Saved audio file
[1169] Step 9: Play the audio file
[1170] The user plays the audio file stored on the device and can listen to the information by pressing the play button.
[1171] Input (user): Playback instructions
[1172] Output (terminal): Playback of audio information
[1173] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1174] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1175] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1176] [Fourth embodiment]
[1177] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1178] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1179] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1180] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1181] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1182] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1183] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1184] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1185] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1186] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1187] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1188] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1189] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1190] This invention relates to a system that allows a user to select the information category they wish to receive, collects related information from the Internet based on the selected information category, filters it, generates an audio script, and plays an audio file. The program processing of this system is explained below in natural language.
[1191] 1. User configuration
[1192] Users install the app and select the categories of information they want to receive (news, weather, schedule, etc.) when they first launch it. They can also set a username and a friendly greeting, allowing them to receive information that is individually customized.
[1193] 2. Collection of information
[1194] The device sends the user's configuration information to the server, which then collects the latest information from data providers such as news APIs, weather information services, and calendar services based on the configured information categories.
[1195] 3. Filtering and organizing information
[1196] The server analyzes the collected information and uses natural language processing technology to filter out unnecessary information. For example, in the case of news, it extracts only articles related to topics that interest the user, and in the case of weather information, it extracts the forecast for the user's location. In addition, in the case of schedule information, it lists important events and organizes them based on the user's interests.
[1197] 4. Generate voice script
[1198] Based on the filtered information, the server generates an original voice script containing friendly speech, such as "Good morning, Mr. / Ms. X. Today is X / X. First, I'll tell you today's news..."
[1199] 5. Speech synthesis and transmission
[1200] The server sends the generated script to a speech synthesis engine, which converts it into an audio file, which is then sent to the user's device.
[1201] 6. Audio playback
[1202] Users can play the audio file sent through their device and receive the information they have set by voice. This allows them to receive the necessary information without looking at the smartphone screen, reducing the risks of using their smartphone while sleeping and providing users with a convenient and healthy way to obtain information.
[1203] Specific examples
[1204] For example, if a user named "Tanaka-san" specifies that they want to receive news, weather, and today's schedule, the server retrieves the latest news from the news API and uses natural language processing to extract only articles that are highly relevant to Tanaka-san. It also retrieves today's weather forecast from the weather information service and organizes weather information related to Tanaka-san's location. Finally, it retrieves today's schedule from the calendar service and lists important events.
[1205] Finally, a voice script is generated that says, "Good morning, Tanaka-san. Today is October 1, 2023. Let me start with today's news..." and converted into an audio file using a speech synthesis engine. This file is then sent to Tanaka's device, allowing him to receive the information by voice without having to look at his smartphone.
[1206] The above is a specific embodiment for carrying out the present invention. By configuring the system in this way, users can efficiently obtain the necessary information and reduce the risk of using their smartphone while sleeping.
[1207] The processing flow will be explained below.
[1208] Step 1:
[1209] User-defined
[1210] Users install the app and upon first launching it, select and enter the information categories they want to receive (news, weather, schedule, etc.), as well as their name and a familiar greeting. This allows for individually customized information acquisition.
[1211] Step 2:
[1212] Sending setting information from the device to the server
[1213] The device sends the user-entered configuration information, including selected information categories and user name, to the server, which stores this information in a database as a user profile.
[1214] Step 3:
[1215] Information gathering
[1216] Based on the user profile, the server collects the latest data from sources such as news APIs, weather information services, and calendar services. For example, it gets the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[1217] Step 4:
[1218] Information Analysis and Filtering
[1219] The server analyzes the collected information and uses natural language processing technology to extract only information relevant to the user. For example, news information might extract articles that match the user's interest categories, weather information might extract weather forecasts for the user's location, and schedule information might extract important events.
[1220] Step 5:
[1221] Generate voice scripts
[1222] The server generates an original voice script based on the filtered information. This script includes a friendly greeting (e.g., "Good morning, Tanaka-san") and selected information categories. For example, "Today is the ____ day of the month. Here's today's news..."
[1223] Step 6:
[1224] Speech synthesis
[1225] The server sends the generated script to a speech synthesis engine (e.g., a Text-to-Speech engine) and converts it into an audio file, which is generated in a format that is easy for the user to listen to (e.g., MP3 format).
[1226] Step 7:
[1227] Sending an audio file
[1228] The server sends the generated audio file to the user's device, which then stores the received audio file in the appropriate folder.
[1229] Step 8:
[1230] Playing audio
[1231] The user plays the audio file sent through the device. By pressing the play button, the audio begins, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This provides a healthy way to obtain information and prevents excessive smartphone use.
[1232] Example 1
[1233] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1234] In modern society, users must efficiently obtain a large amount of information, but visual information acquisition methods involve health risks and hassle, such as using a smartphone while sleeping. Furthermore, because the information users want to receive is diverse, there is a need for a method that can centrally extract only the necessary information and provide it in a user-friendly format. The purpose of this invention is to solve these problems and provide a means for users to acquire information efficiently and healthily.
[1235] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1236] In this invention, the server includes: means for allowing a user to select an information category they wish to receive; means for collecting related information from a network based on the selected information category; and means for filtering the collected information using natural language processing technology to extract only information relevant to the user. This allows users to efficiently obtain information that is appropriate for them. The server also includes means for generating a friendly voice script using a generative AI model; means for converting the generated script into an audio file using a speech synthesis engine; means for transmitting the audio file to the user's device; and means for playing the transmitted audio file. This allows users to obtain information in a healthy way without relying on visual information acquisition, enabling them to use information efficiently and conveniently.
[1237] "User" refers to a person who uses the system to receive information.
[1238] "Information category" refers to the type of information that a user wants to receive, such as news, weather forecast, schedule information, etc.
[1239] "Network" refers to various communication infrastructures, including the Internet, and is used as a means for collecting and transmitting information.
[1240] "Related information" refers to data that is relevant to the user and is obtained based on the information category set by the user.
[1241] "Natural language processing technology" refers to the technology of analyzing text data to understand meaning and extract information.
[1242] "Generative AI model" refers to an artificial intelligence model that generates natural language based on given prompts.
[1243] "Voice script" refers to a sentence that expresses information provided to a user in a voice form.
[1244] A "speech synthesis engine" refers to a technology or system that converts text data into voice data.
[1245] "Audio file" refers to a file that stores audio data generated by speech synthesis.
[1246] "Device" refers to an electronic device that a user possesses and that receives and reproduces information.
[1247] This invention relates to a system that allows a user to select the information category they wish to receive, collects related information from a network based on the selected information category, filters it, generates a voice script, and plays back a voice file. The program processing of this system is explained below in natural language.
[1248] First, the user installs and launches a dedicated application on a device such as a smartphone or tablet. When launching the application for the first time, the user selects the categories of information they wish to receive (news, weather forecasts, schedule information, etc.) on the account settings screen, and also sets their name and how they want to be addressed. This information is sent from the user's device to the server via the network.
[1249] The server collects related information from various information providers based on the received user setting information. For example, news information is obtained using a news API (e.g., NewsAPI or Google News API), weather forecast information is obtained from a weather service (e.g., OpenWeatherMap), and schedule information is obtained from a calendar service (e.g., Google Calendar).
[1250] The server analyzes the large amount of collected data using natural language processing technology (e.g., SpaCy or IBM Watson NLP) and filters out unnecessary information. During this process, it extracts only news articles related to topics that interest the user, extracts information about the user's location from weather forecast information, and lists important schedule information.
[1251] Next, based on the filtered information, the server uses a generative AI model (e.g., OpenAI's ChatGPT) to generate a friendly voice script, using prompts such as "Generate a friendly script that starts with a morning greeting and covers today's news, weather, and schedule."
[1252] The generated script is passed to a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech) and converted into an audio file, which is then sent from the server to the user's device.
[1253] Finally, users can play audio files sent through the device and receive the information they have set by voice, which allows them to obtain information efficiently and healthily without relying on visual information acquisition.
[1254] For example, if a user named "Tanaka-san" requests to receive news, weather forecasts, and schedule information, the server retrieves the latest news from the news API and uses natural language processing technology to extract only articles that are relevant to Tanaka-san. It then retrieves today's weather forecast from a weather service and organizes information about Tanaka-san's location. It then retrieves today's schedule from a calendar service and lists important events. Finally, it generates a voice script that says, "Good morning, Tanaka-san. Today is October 1, 2023. Let's start with today's news..." and converts it into an audio file using a speech synthesis engine. This file is then sent to Tanaka-san's device, allowing her to receive the information by voice without looking at her smartphone.
[1255] The above is a specific embodiment for carrying out the present invention. By configuring such a system, users can efficiently obtain necessary information, reduce the risk of using a smartphone while sleeping, and provide a convenient and healthy means of obtaining information.
[1256] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1257] Step 1:
[1258] The user installs and launches the dedicated application. When the application is launched for the first time, the user selects the information categories they wish to receive (news, weather forecasts, schedule information, etc.) on the account settings screen, and sets their name and how they want to be addressed. This input information includes data such as the user name, information category, and how they want to be addressed. The device receives this setting information as input, generates an API request to send it to the server, and then sends it.
[1259] Step 2:
[1260] The server receives the setting information sent from the device and stores it in a database. Next, based on the user's setting information, it retrieves news data from a news API (e.g., NewsAPI or Google News API). Specifically, it creates an API request to collect news data. This collected news data becomes the input.
[1261] Step 3:
[1262] Similarly, the server retrieves weather forecast data from OpenWeatherMap and schedule information from Google Calendar. It sends an API request to each service and retrieves the respective data (weather forecast data, schedule information data). This becomes the input for each service.
[1263] Step 4:
[1264] The server analyzes the collected news data, weather forecast data, and schedule information data, and filters out unnecessary information using natural language processing technology (e.g., SpaCy or IBM Watson NLP). It receives the collected data as input and extracts only the information that is highly relevant to the user through a filtering process. This is the filtered data that is output.
[1265] Step 5:
[1266] The server uses a generative AI model (e.g., OpenAI's ChatGPT) to generate a friendly voice script based on the filtered data. The input to this process is the filtered data, and the prompt "Generate a friendly script that starts with a morning greeting and covers today's news, weather, and schedule." The output is the generated voice script.
[1267] Step 6:
[1268] The server passes the generated voice script to a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech) and converts it into an audio file. The input for this process is the voice script, which the speech synthesis engine analyzes and outputs an audio file. Specifically, the speech synthesis engine generates voice data based on text data and saves it in an audio file.
[1269] Step 7:
[1270] The server creates and sends an API request to send the generated audio file to the user's device. The input of this process is the audio file, and the output is the audio file sent to the user's device.
[1271] Step 8:
[1272] The user's device receives and saves the audio file sent from the server. When the user operates the application to instruct audio playback, the device plays the saved audio file. The input for this process is the audio file, and the output is audio information provided to the user. The user can receive the necessary information by audio without looking at their smartphone.
[1273] The above are the specific processing steps of this system. Users can acquire information efficiently and are provided with a healthy information acquisition method that does not rely on visual information acquisition.
[1274] (Application example 1)
[1275] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1276] Modern users are seeking ways to efficiently and comfortably gather information amid their busy daily lives. However, conventional information gathering systems have limitations in providing personalized information based on users' preferences and interests, and require users to manually search for information. Furthermore, obtaining product information in a virtual store primarily requires visual confirmation, which requires significant time and effort. Therefore, the present invention aims to provide a system that automates the provision of information based on users' preferences and provides product information in a virtual store via voice, allowing users to efficiently and comfortably obtain the information they need.
[1277] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1278] In this invention, the server includes means for selecting an information category that the user wishes to receive, means for collecting related information from the Internet based on the selected information category, means for filtering the collected information and extracting only information relevant to the user, means for generating an original voice script based on the extracted information, means for converting the generated script into a voice file using a voice synthesis engine, means for transmitting the voice file to the user's terminal, means for playing the transmitted voice file, means for selecting a product category that the user wishes to receive and for providing voice information about products in the virtual store, and means for automatically playing voice information about the products when the user approaches a product in the virtual store. This allows the user to efficiently obtain information regardless of time or place, and to obtain voice information about products in the virtual store without relying on vision.
[1279] A "user" is an entity that selects an information category, collects related information from the Internet, and receives the information in audio format.
[1280] "Information category" indicates the type of information that the user wants to receive, and includes news, weather information, schedule information, product information, and the like.
[1281] The term "means" refers to components or methods required to achieve a specific function, and in the present invention includes functions such as information collection, filtering, voice generation, and voice playback.
[1282] The "Internet" is a global network that connects computers around the world and enables the exchange of information.
[1283] "Related information" is data collected based on an information category selected by the user, and is information that contains content that is useful to the user.
[1284] "Filtering" is the process of selecting only information relevant to the user from the collected information and removing unnecessary information.
[1285] A "voice script" is a text for communicating information by voice that is generated based on collected and filtered information.
[1286] A "speech synthesis engine" is a software and hardware component for converting text information into speech data.
[1287] "Audio file" means a digital file generated by a speech synthesis engine to provide audio information to a user.
[1288] A "terminal" is a device that allows a user to receive audio files and play the information, including smartphones, smart glasses, and head-mounted displays.
[1289] A "virtual store" is a commercial facility built in a virtual space where users can browse and purchase products online.
[1290] "Product category" indicates the type of product in which the user is interested, and includes home appliances, fashion, accessories, and the like.
[1291] "Detailed information" refers to information that a user needs to make a decision about purchasing a product, such as product features, price, specifications, and usage.
[1292] The system for implementing this invention allows a user to select a category of information they wish to receive, collects related information from the Internet based on that category, generates a voice script, and transmits it to the user's terminal for playback. This system also has the function of providing voice information about products in a virtual store. The specific configuration and operation of the system are described below.
[1293] System configuration
[1294] This system mainly uses the following hardware and software:
[1295] Hardware: Smartphones, smart glasses, head-mounted displays (HMDs)
[1296] Software: Natural language processing (NLP) engine, speech synthesis engine (e.g., Google Text-to-Speech API), user data management server, product information database
[1297] Program processing
[1298] 1. User configuration
[1299] Users install the app and select the desired information category (news, weather, product information, etc.) when they first start it. They can also set a name and how they want to be called.
[1300] 2. Information collection by the server
[1301] The device sends the user's setting information to the server, which then collects the latest information from related services (news API, weather information service, product information database) based on the set information category.
[1302] 3. Server filtering and organization of information
[1303] The server analyzes the collected information and uses NLP technology to filter out unnecessary information. For example, in the case of news, it extracts only articles related to topics that the user is interested in, and in the case of product information, it retrieves detailed information about products that interest the user.
[1304] 4. Server-generated voice script
[1305] Based on the filtered information, the server generates an original voice script containing friendly speech, such as "Good morning, Mr. / Ms. X. Today is X / X. First, I'll tell you today's news..." including the user's name.
[1306] 5. Server-based speech synthesis and transmission
[1307] The server sends the generated script to a speech synthesis engine to generate an audio file, which is then sent to the user's device.
[1308] 6. Playing Audio on the Device
[1309] Users can play the audio file sent through the device and receive the set information by voice, allowing them to receive information without looking at the smartphone screen.
[1310] 7. Audio guide function in virtual stores
[1311] In particular, within the virtual store, product information is provided based on product categories set by the user. When the user approaches a product, detailed information about that product is automatically played back in audio.
[1312] Examples and prompts
[1313] Specific examples
[1314] For example, if a user named "Tanaka-san" requests news, weather, and information about home appliances in a virtual store, the server retrieves the latest news from the news API and uses natural language processing to extract only articles most relevant to Tanaka-san. It then retrieves today's weather forecast from a weather information service and organizes weather information relevant to Tanaka-san's location. For home appliance information, the server collects the latest product information within the virtual store and lists important specifications and pricing information. Finally, it generates a voice script that reads, "Good morning, Tanaka-san. Today is October 1, 2023. Let me start with today's news..." and converts it into an audio file using a speech synthesis engine (Google Text-to-Speech API). This file is then sent to Tanaka-san's device, allowing him to easily access the information by voice via his smartphone or HMD.
[1315] Prompt Sentence Examples
[1316] Username: Tanaka-san
[1317] Favorite product category: Home appliances
[1318] Provides the latest information on home appliances via voice
[1319] Example of generated voice script: "Hey Tanaka, here are some recommended new home appliances..."
[1320] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1321] Step 1:
[1322] User category selection and individual settings
[1323] (Input) The user installs the app and selects the desired information category (news, weather, product information, etc.) when launching it for the first time. They also set a name and how they want to be called.
[1324] (Processing) The app collects user selections and settings.
[1325] (Output) The collected setting information is stored in the terminal and also sent to the server.
[1326] Step 2:
[1327] Server collection of information
[1328] (Input) The server collects related information based on the selected information category, based on the user setting information received from the terminal.
[1329] The (processing) server obtains the necessary data from news APIs, weather information services, product information databases, etc.
[1330] (Output) The acquired data is temporarily stored on the server.
[1331] Step 3:
[1332] Filtering and organizing information
[1333] (Input) Collected raw data (news articles, weather information, product information, etc.).
[1334] The (processing) server uses a natural language processing (NLP) engine to extract only the information relevant to the user and filter out unnecessary information. For example, in the case of a news category, it selects only important and reliable articles. In the case of product information, it organizes detailed information related to items that the user is interested in.
[1335] (Output) The filtered and organized information is stored on the server.
[1336] Step 4:
[1337] Generate voice scripts
[1338] (Input) Filtered information.
[1339] The (processing) server generates a friendly, customized voice script based on the received information. For example, it creates text in the form of "Good morning, Mr. / Ms. XX. Today is the XXth month. First, I'll tell you today's news..."
[1340] (Output) The generated voice script is saved on the server.
[1341] Step 5:
[1342] Speech synthesis and transmission
[1343] (Input) The generated voice script.
[1344] The (processing) server uses a speech synthesis engine (e.g., Google Text-to-Speech API) to convert the text-based voice script into an audio file.
[1345] (Output) The converted audio file is generated and sent to the user's device.
[1346] Step 6:
[1347] Device audio playback
[1348] (Input) The audio file sent from the server.
[1349] (Processing) The user terminal receives the audio file and plays the audio using the built-in media player.
[1350] (Output) The user can receive the information set through the terminal by voice.
[1351] Step 7:
[1352] Audio guide function in virtual stores
[1353] (Input) Product category and its location information set by the user.
[1354] (Processing) When a user walks around the virtual store and comes close to a certain product, detailed information about that product is automatically played back using the device's sensors and location information. Specifically, the device uses distance sensors and GPS data to determine the user's current location, and plays back related product information as an audio file.
[1355] (Output) The user can receive voice information about products that interest them in the virtual store.
[1356] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1357] This invention combines a system in which a user selects the information category they wish to receive, collects related information from the Internet based on the selected information category, filters it, generates voice scripts, and plays voice files, with an emotion engine that recognizes the user's emotions and adjusts the information and voice based on those emotions. Below, the program processing of this system is explained in natural language.
[1358] 1. User configuration
[1359] When users install the app and launch it for the first time, they select and enter the categories of information they want to receive (news, weather, schedule, etc.), as well as their own name and a familiar greeting. This allows them to receive information that is individually customized.
[1360] 2. Emotion Recognition by Emotion Engine
[1361] To recognize the user's emotions, the device analyzes data acquired from the camera and microphone in real time, thereby detecting the user's current emotions (e.g., joy, sadness, anger, surprise, etc.). The emotion engine analyzes this data using an emotion analysis model to determine the user's emotional state.
[1362] 3. Sending setting information from the device to the server
[1363] The device sends the user-entered setting information and the emotion information recognized by the emotion engine to the server, which then stores this information in a database as a user profile.
[1364] 4. Information Collection
[1365] Based on the user profile, the server collects the latest data from sources such as news APIs, weather information services, and calendar services. For example, it gets the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[1366] 5. Information Analysis and Filtering
[1367] The server analyzes the collected information and uses natural language processing technology to extract only information relevant to the user. For example, news information will extract articles that match the user's interest categories, weather information will extract weather forecasts for the user's location, and schedule information will extract important appointments. Additionally, based on the emotional information provided by the emotion engine, the server prioritizes and selects information that best suits the user's current emotional state.
[1368] 6. Generate voice script
[1369] The server generates an original voice script based on the filtered information. This script includes a friendly greeting (e.g., "Good morning, Tanaka-san") and selected information categories. For example, "Today is the ____ day of the month. Here's today's news..." Furthermore, the script incorporates phrases with adjusted tone and speed according to the emotions recognized by the emotion engine.
[1370] 7. Speech Synthesis
[1371] The server sends the generated script to a speech synthesis engine (e.g., a Text-to-Speech engine) and converts it into an audio file, which is generated in a format that is easy for the user to listen to (e.g., MP3 format).
[1372] 8. Sending audio files
[1373] The server sends the generated audio file to the user's device, which then stores the received audio file in the appropriate folder.
[1374] 9. Audio playback
[1375] The user plays the audio file sent through the device. By pressing the play button, the audio begins, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This provides a healthy way to obtain information and prevents excessive smartphone use.
[1376] Specific examples
[1377] For example, if a user has the name "Tanaka-san" and has set that they want to receive news, weather, and today's schedule, and the emotion engine recognizes that the current emotion is "happy," the server will retrieve the latest news from the news API and use natural language processing to extract only articles that are highly relevant to Tanaka-san. It will also retrieve today's weather forecast from the weather information service and organize the weather information for Tanaka-san's location. It will then retrieve today's schedule from the calendar service and list important events. The emotion engine will then synthesize speech in a positive tone to match the user's "happy" state.
[1378] Finally, a voice script is generated that says, "Good morning, Tanaka. Today is October 1, 2023. Today's news is very interesting..." and converted into an audio file using a speech synthesis engine. This file is then sent to Tanaka's device, allowing him to receive the information by voice without having to look at his smartphone.
[1379] The above is a specific embodiment for carrying out the present invention. By configuring the system in this way, it is possible to efficiently obtain necessary information in a manner that is in line with the user's emotions, and reduce the risk of "using a smartphone while sleeping."
[1380] The processing flow will be explained below.
[1381] Step 1:
[1382] User-defined
[1383] When users install the app and launch it for the first time, they select and input the categories of information they want to receive (news, weather, schedule, etc.), as well as their own name and a familiar greeting. This allows them to receive information that is individually customized.
[1384] Step 2:
[1385] Emotion recognition by emotion engine
[1386] The device captures and analyzes data from the camera and microphone in real time to recognize the user's emotions, such as whether the user is happy, sad, angry, or surprised, by analyzing facial expressions and voice tone.
[1387] Step 3:
[1388] Sending setting information from the device to the server
[1389] The device transmits the user's setting information and the emotion information recognized by the emotion engine to the server, where the transmitted data is stored in the user profile and used for subsequent processing.
[1390] Step 4:
[1391] Information gathering
[1392] Based on the user profile, the server collects the latest data from information sources such as a news API, weather information service, and calendar service. For example, it obtains the latest news from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[1393] Step 5:
[1394] Information Analysis and Filtering
[1395] The server analyzes the collected information and uses natural language processing technology to extract only information that is highly relevant to the user. Furthermore, based on the emotional data provided by the emotion engine, it prioritizes information that matches the user's current emotional state. For example, if the user is "happy," it prioritizes positive news and information that enhances emotions.
[1396] Step 6:
[1397] Generate voice scripts
[1398] The server generates an original voice script based on the filtered information. This script includes friendly greetings (e.g., "Good morning, Tanaka-san") and phrases that correspond to emotions. For example, if the user is in a "happy" state, the script might say, "I have some very exciting news for you today..."
[1399] Step 7:
[1400] Speech synthesis
[1401] The server then sends the generated script to a speech synthesis engine (e.g., a text-to-speech engine) and converts it into an audio file. This audio file is generated in a smooth listening format (e.g., MP3 format), and the tone and speed of the audio are adjusted based on the user's emotions.
[1402] Step 8:
[1403] Sending an audio file
[1404] The server sends the generated audio file to the user's device, which stores the received audio file in an appropriate folder and prepares it for playback.
[1405] Step 9:
[1406] Playing audio
[1407] The user plays the audio file sent through the device. By pressing the play button, a voice will play saying, "Good morning, Tanaka-san..." and the user can hear the necessary information without looking at the screen. This prevents excessive smartphone use and provides a healthy way to obtain information.
[1408] In this way, the system of the present invention can adjust the content of information and the tone of voice according to the user's emotional state, providing an optimal information acquisition experience.
[1409] Example 2
[1410] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1411] In modern society, many users use the Internet to obtain information, but the overwhelming amount of information available makes it difficult to quickly find relevant information. Users who want to easily obtain information need a means to utilize not only their eyes but also their ears, but there is a lack of systems that can accommodate individual needs and emotional states. Furthermore, because excessive smartphone use can cause health problems, there is a need for a means to obtain information without looking at the screen.
[1412] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1413] In this invention, the server includes means for allowing a user to select an information category they wish to receive, means for collecting related information from a data network based on the selected information category, and means for analyzing the collected information using natural language processing technology and extracting only information relevant to the user. This allows the user to easily obtain only the information that suits them from an excess of information and to receive that information by voice.
[1414] The "means for selecting the information category that the user wishes to receive" refers to an input device or interface that allows the user to specify the type of information in which the user is interested (for example, news, weather, schedule, etc.).
[1415] A "data network" is a communication path through which multiple computing resources exchange information with each other, including the Internet and other information and communication networks.
[1416] "Natural language processing technology" refers to computer technology for understanding, analyzing, and generating human language (e.g., keyword extraction, document classification, sentiment analysis, etc.).
[1417] "Means for recognizing an emotional state and adjusting information based on that emotional state" refers to technology or a device that reads emotions from a user's facial expressions, voice, etc., and changes the content and expression of the information provided according to those emotions.
[1418] "Means for generating voice scripts" refers to technology or devices that automatically create text to be read aloud based on collected and analyzed information.
[1419] A "speech synthesis engine" is a computer technology (for example, text-to-speech technology) for converting text in text format into voice data.
[1420] A "terminal" is an information processing device (for example, a smartphone, tablet, or PC) that can be directly operated by a user.
[1421] "News information" is information that reports the latest facts and events related to society, economy, sports, culture, etc.
[1422] "Weather information" refers to data related to the weather, such as weather forecasts, temperature, precipitation, and wind speed.
[1423] "Schedule information" is information about schedules and events listed in a user's schedule or calendar.
[1424] "Means for including a user name or friendly nickname" refers to technology or a device for incorporating a name set by the user or a nickname that gives a sense of familiarity to the user (for example, "san" or "kun") into the voice script.
[1425] This invention combines a system that allows a user to select an information category of interest, collects related information from the Internet based on that information, filters it, generates a voice script, and plays back a voice file, with an emotion engine that recognizes the user's emotions and adjusts the information and voice based on those emotions. A specific method for implementing this system will be described below.
[1426] User-defined
[1427] Users install a dedicated application on their device and, when they first start it up, set the categories of information they want to receive (e.g., news, weather, schedule), as well as their own name and a friendly greeting. This allows for individually customized information to be provided.
[1428] Emotion recognition by emotion engine
[1429] The device uses a camera and microphone to capture the user's facial expressions and voice data, and analyzes it in real time using an emotion analysis model (e.g., OpenFace or DeepFace). This identifies the user's emotional state and sends that information to the server as data. This data is used for information filtering and voice script generation, which will be described later.
[1430] Sending configuration information to the server
[1431] The device sends the information entered and set by the user, as well as the emotional data recognized by the emotion engine, to the server, which then stores this information in a database as a user profile and uses it to provide the most appropriate information to each individual user.
[1432] Information gathering
[1433] The server collects the necessary data from external sources such as news APIs, weather information services, and calendar services. For example, it obtains the latest news from a news API (e.g., Google News API), local weather forecasts from a weather information service (e.g., OpenWeatherMap API), and user schedule information from a calendar service (e.g., Google Calendar API).
[1434] Information Analysis and Filtering
[1435] The server analyzes the collected data using natural language processing technology and scrutinizes information relevant to the user. For example, for news information, keywords are extracted to extract articles that match the user's categories of interest, and for weather information, only information related to the user's location is extracted. Based on the recognition results of the emotion engine, information that matches the user's current emotions is selected.
[1436] Generate voice scripts
[1437] The server generates a voice script based on the filtered information. This voice script includes the user's name and a friendly greeting, and incorporates tones and expressions that correspond to the user's emotional state. For example, it includes gentle expressions such as, "It's a beautiful day today. Let's have a good day."
[1438] Speech synthesis
[1439] The server sends the generated script to a speech synthesis engine (e.g., Google Text-to-Speech API) and converts it into an audio file, which is generated in, for example, MP3 format.
[1440] Sending an audio file
[1441] The server sends the generated audio file to the user's device, which stores the received audio file in an appropriate folder and prepares it for playback.
[1442] Playing audio
[1443] Users can play the audio files sent through the device. By pressing the play button, they can hear a voice such as, "Good morning, Tanaka-san. Here's today's news..." and can listen to the necessary information without looking at the screen.
[1444] Specific examples
[1445] For example, if a user named "Tanaka-san" selects to receive news, weather, and today's schedule, and the emotion engine recognizes that the current emotion is "happy," the server collects and filters data from each source. Based on the emotion, a voice script with a positive tone is generated. Finally, the voice script, "Good morning, Tanaka-san. Today is October 1, 2023. Today's news is very interesting...," is converted into an audio file and sent to Tanaka-san's device.
[1446] Prompt Sentence Examples
[1447] For example, the following prompt sentence is fed into the generative AI model:
[1448] "Includes news, weather, and schedule categories. Current emotion is joy. Name is Tanaka. Uses familiar greeting. Send generated audio file to device."
[1449] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1450] Step 1: User Setup
[1451] A user launches the application on their device and selects the information categories they want to receive (e.g., news, weather, schedule), and also sets their name and a friendly greeting (e.g., "Mr. Tanaka"). This set information is stored in the device's local database.
[1452] Input: Category of information you want to receive, username, call
[1453] Output: Configuration information stored in the local database
[1454] Step 2: Emotion recognition by the emotion engine
[1455] The device uses a camera and microphone to collect the user's facial and voice data in real time. The collected data is analyzed using an emotion analysis model (e.g., OpenFace or DeepFace) to identify the user's emotional state (e.g., joy, sadness, anger). The analysis results are temporarily stored on the device for later use.
[1456] Input: User's facial expression data, voice data
[1457] Output: Parsed emotional state data
[1458] Step 3: Sending configuration information to the server
[1459] The terminal transmits the user's selected information categories and analyzed emotional state data to the server, which receives this information and stores it in a database as a user profile.
[1460] Input: Setting information stored in a local database, analyzed emotional state data
[1461] Output: User profile data sent to the server
[1462] Step 4: Gather information
[1463] Based on the collected user profile data, the server collects relevant information from external sources such as news APIs, weather information services, calendar services, etc. For example, the server retrieves the latest news articles from the news API and local weather forecasts from the weather information service.
[1464] Input: User profile data
[1465] Output: Raw data collected from external sources (news, weather, schedule information)
[1466] Step 5: Information analysis and filtering
[1467] The server analyzes the collected raw data using natural language processing technology. For example, it can extract articles that match the user's interest categories from news data, extract information related to the user's location from weather data, and adjust the priority and content of information based on the user's emotional state.
[1468] Input: Raw data collected from external sources, user emotional state data
[1469] Output: filtered and adjusted information data
[1470] Step 6: Generate the voice script
[1471] The server generates a voice script based on the filtered and adjusted data, which includes the user's name and a friendly greeting, and adjusts the tone and expression depending on the user's emotional state.
[1472] Input: Filtered and conditioned information data
[1473] Output: The generated voice script
[1474] Step 7: Text-to-Speech
[1475] The server sends the generated voice script to a text-to-speech engine (e.g., Google Text-to-Speech API) and converts it into an audio file (e.g., MP3 format). The converted audio file is temporarily stored on the server.
[1476] Input: Generated voice script
[1477] Output: Converted audio file
[1478] Step 8: Send the audio file
[1479] The server sends the generated voice file to the user's device, which saves the received voice file in the appropriate folder (e.g., " / user / voices / ").
[1480] Input: Converted audio file
[1481] Output: Audio file saved on your device
[1482] Step 9: Playing Audio
[1483] The user plays the audio file through the device. By pressing the play button, the audio will play, "Good morning, Tanaka-san. Here's today's news..." and the user can hear the necessary information without looking at the screen.
[1484] Input: Audio files stored on the device
[1485] Output: Audio information played to the user
[1486] (Application example 2)
[1487] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1488] In order to provide optimal services based on customer emotions, commercial facilities and brick-and-mortar stores require a system that recognizes customer emotions in real time and collects and provides information based on those emotions. However, current systems lack the functionality to recognize and appropriately reflect customer emotions, making generalization and automation difficult. Furthermore, the collected information is not optimized for customer emotions, limiting the improvement of customer satisfaction.
[1489] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1490] In this invention, the server includes means for selecting an information category that the user wishes to receive, means for collecting related information from the Internet based on the selected information category, means for filtering the collected information and extracting only information relevant to the user, means for generating an original voice script based on the extracted information, means for converting the generated script into a voice file using a voice synthesis engine, means for transmitting the voice file to the user's terminal, means for playing the transmitted voice file, means for analyzing data acquired from a camera or microphone and recognizing the user's emotions, and means for adjusting information and voice based on the recognized emotional information. This makes it possible to efficiently provide appropriate information in accordance with the customer's emotions and improve customer satisfaction.
[1491] The "means for selecting the information category that the user wants to receive" is a function that allows the user to select the type of information that the user wants to receive based on their own interests and concerns.
[1492] The "means for collecting related information from the Internet" is a function for obtaining data related to the information category selected by the user from various information sources on the Internet.
[1493] "Means of filtering collected information and extracting only information relevant to the user" refers to a function that selects the most relevant information from the acquired information based on the user's interests and preferences.
[1494] The "means for generating an original voice script based on extracted information" is a function for generating a script based on filtered information to be read aloud in a form that is familiar to the user.
[1495] "Means for converting the generated script into an audio file using a voice synthesis engine" is a function for converting the generated voice script into actual voice data using voice synthesis technology.
[1496] The "means for transmitting the audio file to the user's terminal" is a function for transmitting the generated audio file to the terminal used by the user via the Internet or other communication means.
[1497] The "means for playing transmitted audio files" is a function for playing audio files stored in the user's terminal, allowing the user to listen to audio information.
[1498] "Means for analyzing data obtained from a camera or microphone and recognizing the user's emotions" refers to a function that recognizes the user's emotional state by capturing and analyzing the user's facial expressions and voice using a camera or microphone installed on the device.
[1499] The "means for adjusting information or voice based on recognized emotional information" is a function for presenting information in an optimal form or adjusting the tone or content of voice based on the acquired emotional information of the user.
[1500] A system for implementing this invention allows a user to select the information categories they wish to receive, collects and filters relevant information based on those categories, generates audio scripts and plays audio files, and recognizes the user's emotions and adjusts the information and audio based on those emotions.
[1501] 1. User Settings
[1502] When the user starts the system for the first time, they set the category of information they want to receive (news, weather, schedule, etc.), their name, and a friendly greeting. These settings allow the system to retrieve and provide information tailored to the user.
[1503] 2. Emotion recognition
[1504] The device uses a camera and microphone to recognize emotions in real time from the user's facial expressions and voice. Specifically, it uses an emotion engine to analyze facial expressions using images captured by the camera and emotions from the voice. A library called DeepFace is used for emotion analysis, and a voice recognition API is used for voice analysis.
[1505] 3. Information gathering
[1506] The server collects the latest data from various information sources, such as a news API, weather information service, and calendar service, based on the user's settings and emotion information. For example, it obtains the latest articles from the news API, the local weather forecast from the weather information service, and schedule information from the calendar service.
[1507] 4. Information analysis and filtering
[1508] The server analyzes the collected information using natural language processing technology to extract only the information relevant to the user. At the same time, based on the emotional information recognized by the emotion engine, it prioritizes and selects the information that best suits the user's current emotional state.
[1509] 5. Voice script generation
[1510] The server generates an original voice script based on the filtered information, which includes the user's name, a friendly greeting, the selected information category, and phrases with adjusted tone and speed depending on the emotional information.
[1511] 6. Speech synthesis
[1512] The generated script is sent to a speech synthesis engine (e.g., a text-to-speech engine) and converted into an audio file, which generates an audio file in a format that is easy for users to listen to.
[1513] 7. Sending and playing audio files
[1514] The generated audio file is sent to the user's device, which receives it and saves it in the appropriate folder. The user can immediately listen to the audio information by pressing the play button.
[1515] Specific examples
[1516] For example, if a user named "Sato" specifies that they would like to receive news, weather, and today's schedule information, and the emotion engine recognizes that the current emotion is "happy," the server retrieves the latest news articles from the news API and uses natural language processing technology to extract only those articles that are highly relevant to Mr. Sato. It also retrieves weather information related to Mr. Sato's location from the weather information service and lists today's schedule from the calendar service. The emotion engine synthesizes voice in a positive tone that matches the user's "happy" state.
[1517] Example prompts for generative AI models
[1518] "The user's name is Sato. I'm happy to hear from Sato. Please let me know about new products that might interest Sato."
[1519] This makes it possible to efficiently provide appropriate information in a manner that is in line with the customer's feelings, thereby improving customer satisfaction.
[1520] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1521] Step 1: User Setup
[1522] Users launch the application installed on their device, select the category of information they want to receive (news, weather, schedule, etc.), and enter their name or a friendly greeting.
[1523] Input (user): Information category, name, call
[1524] Output (terminal): Save as configuration information
[1525] Step 2: Emotion Recognition
[1526] The device uses a camera and microphone to capture the user's facial expressions and voice in real time, and then analyzes the captured data using an emotion engine (e.g., the DeepFace library or a voice recognition API) to recognize the user's emotions.
[1527] Input (terminal): Camera video, audio data
[1528] Output (terminal): Emotion recognition results (e.g., joy, sadness, anger)
[1529] Step 3: Sending preferences and emotions
[1530] The device sends the setting information entered by the user and the recognized emotion information to the server, where they are stored in a database as a user profile.
[1531] Input (device): setting information, emotion recognition results
[1532] Output (Server): Save as user profile
[1533] Step 4: Gather information
[1534] Based on the user profile, the server collects the latest data from news APIs, weather information services, calendar services, etc. For example, it obtains the latest articles from the news API, the local weather forecast from the weather information service, and today's schedule from the calendar service.
[1535] Input (server): User profile
[1536] Output (server): Collected information (news articles, weather forecasts, schedules)
[1537] Step 5: Information analysis and filtering
[1538] The server analyzes the collected information using natural language processing technology to extract only the information relevant to the user. At the same time, it prioritizes and selects the information that best suits the user's current emotional state based on the emotional information provided by the emotion engine.
[1539] Input (server): Collected information, emotional information
[1540] Output (server): Filtered information
[1541] Step 6: Generate voice script
[1542] The server generates an original voice script based on the filtered information, including a user-friendly call and selected information categories, and adjusts the tone and speed of the voice based on the emotional information.
[1543] Input (server): filtered information, emotional information
[1544] Output (server): Voice script
[1545] Step 7: Text-to-Speech
[1546] The server sends the generated script to a speech synthesis engine and converts it into an audio file (e.g., MP3 format), which generates an audio file in a format that is easy for users to listen to.
[1547] Input (server): Voice script
[1548] Output (server): Audio file
[1549] Step 8: Send and save the audio file
[1550] The server sends the generated audio file to the user's terminal, which receives it and saves it in an appropriate folder.
[1551] Input (server): Audio file
[1552] Output (Device): Saved audio file
[1553] Step 9: Play the audio file
[1554] The user plays the audio file stored on the device and can listen to the information by pressing the play button.
[1555] Input (user): Playback instructions
[1556] Output (terminal): Playback of audio information
[1557] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1558] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1559] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1560] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1561] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1562] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1563] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1564] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1565] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1566] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1567] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1568] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1569] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1570] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1571] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1572] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1573] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1574] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1575] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1576] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1577] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1578] The following is further disclosed regarding the above embodiment.
[1579] (Claim 1)
[1580] a means for the user to select the categories of information they wish to receive;
[1581] means for collecting relevant information from the Internet based on the selected information categories;
[1582] a means for filtering the collected information to extract only information relevant to the user;
[1583] A means for generating an original voice script based on the extracted information;
[1584] a means for converting the generated script into an audio file using a speech synthesis engine;
[1585] means for transmitting the audio file to a user's terminal;
[1586] means for playing the transmitted audio file;
[1587] A system including:
[1588] (Claim 2)
[1589] 10. The system of claim 1, further comprising means for including a user name and a friendly greeting in the voice script.
[1590] (Claim 3)
[1591] 2. The system according to claim 1, wherein the information gathering includes means for acquiring news, weather information, and schedule information.
[1592] "Example 1"
[1593] (Claim 1)
[1594] a means for the user to select the categories of information they wish to receive;
[1595] means for collecting relevant information from the network based on the selected information category;
[1596] A means for filtering the collected information using natural language processing technology and extracting only information relevant to the user;
[1597] a means for generating a friendly voice script using a generative AI model based on the extracted information; and
[1598] a means for converting the generated script into an audio file using a speech synthesis engine;
[1599] means for transmitting the audio file to the user's device;
[1600] means for playing the transmitted audio file;
[1601] A system including:
[1602] (Claim 2)
[1603] 10. The system of claim 1, wherein the voice script includes a user name and a friendly greeting.
[1604] (Claim 3)
[1605] 2. The system according to claim 1, wherein the information gathering includes means for obtaining news, weather forecasts, and schedule information.
[1606] "Application Example 1"
[1607] (Claim 1)
[1608] a means for the user to select the categories of information they wish to receive;
[1609] means for collecting relevant information from the Internet based on the selected information categories;
[1610] a means for filtering the collected information to extract only information relevant to the user;
[1611] A means for generating an original voice script based on the extracted information;
[1612] a means for converting the generated script into an audio file using a speech synthesis engine;
[1613] means for transmitting the audio file to a user's terminal;
[1614] means for playing the transmitted audio file;
[1615] A means for a user to select a desired product category and to be provided with information on products in the virtual store by voice;
[1616] A means for automatically playing detailed product information by voice when approaching a product in the virtual store;
[1617] A system including:
[1618] (Claim 2)
[1619] 10. The system of claim 1, further comprising means for including a user name and a friendly greeting in the voice script.
[1620] (Claim 3)
[1621] 2. The system according to claim 1, wherein the information gathering comprises means for acquiring news, weather information, schedule information, and product information.
[1622] "Example 2: Combining Emotion Engines"
[1623] (Claim 1)
[1624] a means for the user to select the categories of information they wish to receive;
[1625] means for collecting relevant information from a data network based on the selected information categories;
[1626] A means for analyzing the collected information using natural language processing technology and extracting only information relevant to the user;
[1627] means for recognizing an emotional state of a user and adjusting information based on the emotional state;
[1628] a means for generating an original voice script based on the extracted and adjusted information;
[1629] a means for converting the generated script into an audio file using a speech synthesis engine;
[1630] means for transmitting the audio file to a user's terminal;
[1631] means for playing the transmitted audio file;
[1632] A system including:
[1633] (Claim 2)
[1634] 10. The system of claim 1, further comprising means for including a user name and a friendly greeting in the voice script.
[1635] (Claim 3)
[1636] 2. The system according to claim 1, wherein the information gathering includes means for acquiring news information, weather information, and schedule information.
[1637] "Application example 2 when combining emotion engines"
[1638] (Claim 1)
[1639] a means for the user to select the categories of information they wish to receive;
[1640] means for collecting relevant information from the Internet based on the selected information categories;
[1641] a means for filtering the collected information to extract only information relevant to the user;
[1642] A means for generating an original voice script based on the extracted information;
[1643] a means for converting the generated script into an audio file using a speech synthesis engine;
[1644] means for transmitting the audio file to a user's terminal;
[1645] means for playing the transmitted audio file;
[1646] A means of analyzing data obtained from a camera or microphone and recognizing the user's emotions;
[1647] means for adjusting information or audio based on the recognized emotion information;
[1648] A system including:
[1649] (Claim 2)
[1650] 10. The system of claim 1, further comprising means for including a user name and a friendly greeting in the voice script.
[1651] (Claim 3)
[1652] 2. The system according to claim 1, wherein the information gathering includes means for acquiring news, weather information, and schedule information. [Explanation of symbols]
[1653] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for the user to select the categories of information they wish to receive; means for collecting relevant information from the Internet based on the selected information categories; a means for filtering the collected information to extract only information relevant to the user; A means for generating an original voice script based on the extracted information; a means for converting the generated script into an audio file using a speech synthesis engine; means for transmitting the audio file to a user's terminal; means for playing the transmitted audio file; A system including:
2. 10. The system of claim 1, further comprising means for including a user name and a friendly greeting in the voice script.
3. 2. The system according to claim 1, wherein the information gathering includes means for acquiring news, weather information, and schedule information.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A