System

The system addresses the limitations of existing search systems by enabling efficient information retrieval through voice and text interaction, summarizing search history, and delivering personalized results, enhancing user experience.

JP2026019740APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121488
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing search systems lack the ability to efficiently provide users with interactive information retrieval, manage question history, and deliver personalized search results, making it difficult for users to obtain information quickly and effectively.

Method used

A system incorporating voice and text input means, recognition and response means, a summary generation means, and a notification means, along with a personalized database, to facilitate efficient information search and retrieval, including voice and text question history summarization.

Benefits of technology

Enables users to efficiently obtain information and easily access summarized past search history, improving the user experience by providing personalized and interactive search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019740000001_ABST
    Figure 2026019740000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a voice input unit for inputting a question by voice, a voice recognition unit for converting voice data into text data, a search unit for performing an Internet search based on the text data, a voice response unit for converting a search result into voice data, a voice output unit for outputting the voice data to a user, a text input unit for inputting a text question, a search unit for performing an Internet search based on the text question, a text response unit for providing a method for returning a search result of the text question to the user by text, a summary generation unit for generating a summary of a voice and text question history, and a notification unit for notifying the user of a summary generation result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Information search has become a part of everyday activities in modern life. However, existing search systems do not adequately provide users with the means to quickly and efficiently obtain information. Furthermore, many voice assistants and chatbots only provide one-way information and lack interaction. Furthermore, it is difficult to manage question history and quickly retrieve information from that history, limiting the user's search experience. This makes information retrieval cumbersome, making it difficult for users to obtain the information they need in a timely manner. [Means for solving the problem]

[0005] The present invention provides a system including a voice input means, a voice recognition means, a search means, a voice response means, a voice output means, a character input means, a character response means, a summary generation means, and a notification means. A user can use the voice input means to ask a question by voice. The voice recognition means converts the voice question into text data, and the search means performs an Internet search based on the text data. The voice response means converts search results obtained by the search means into voice data, and the voice output means provides the user with the voice results. The user can also use the character input means to ask a text question, which then performs an Internet search, and the text response means provides the search results in text. Furthermore, the system summarizes the voice and text question history using the summary generation means and notifies the user via the notification means. This allows the user to efficiently obtain information and easily access a summarized version of their past search history, improving their information search experience.

[0006] "Voice input means" refers to a device or software that allows a user to input a question by voice.

[0007] "Speech recognition means" refers to a device or software for converting voice data into text data.

[0008] "Search means" refers to a device or software for conducting an Internet search based on text data.

[0009] "Voice response means" refers to a device or software for converting search results into voice data.

[0010] "Audio output means" refers to a device or software for outputting audio data to a user.

[0011] "Text input means" refers to a device or software that allows a user to input a question in text.

[0012] "Text response means" refers to a device or software for returning search results based on character input to the user in text.

[0013] "Summary generator" refers to a device or software for summarizing audio and text question histories.

[0014] "Notification means" refers to a device or software for notifying the user of the summary generated results.

[0015] A "personalized database" refers to a database that stores search history and preference information for each user and provides personalized search results. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The voice response and chat-linked search system of the present invention includes a voice input means installed in a smart device, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the converted text data, and a voice response means for converting the search results into voice data and providing it to the user. The user can also use the input means of the smartphone to submit text questions, and an Internet search can be performed based on the text input, with the text response means providing the results to the user in text. The voice and text question history is summarized by a summary generation means and notified to the user by a notification means.

[0038] Natural language description of the program

[0039] 1. Initial Setup:

[0040] The server stores the user's account information and search history in a personalized database, and also configures the search API and message API integration.

[0041] 2. Audio question reception:

[0042] The device is in standby mode and starts voice input when the user speaks a trigger phrase. The user then speaks a question.

[0043] 3. Speech Recognition:

[0044] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine in the cloud to convert the voice data into text data.

[0045] 4. Information Search:

[0046] The server analyzes the converted text data, generates an appropriate search query, and sends the generated search query to the search API to retrieve search results.

[0047] 5. Voice response generation:

[0048] The server analyzes the search results, selects the most relevant information, and generates natural language text based on the search results.The natural language text is converted into voice data using a voice response means and sent to the terminal.

[0049] 6. Audio Answer:

[0050] The terminal plays back the received voice data and provides the user with a voice response.

[0051] 7. Chat questions:

[0052] The user types in a text question on their smartphone and sends it.

[0053] 8. Handling chat questions:

[0054] The server analyzes the received text question, generates an appropriate search query using a search API, retrieves search results, analyzes the retrieved search results, organizes them as natural language text, and sends a text reply to the user.

[0055] 9. Question History Summary:

[0056] The server periodically transmits the question history of voice and text to the summary generating means, which automatically generates a summary, and notifies the user of the summary via the notification means.

[0057] Specific examples

[0058] Examples of voice questions:

[0059] User: "OK, Assistant, what's the weather like right now?"

[0060] 1. The device captures the audio data and sends it to the server.

[0061] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[0062] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[0063] 4. The server generates the answer text "Currently, the weather in Tokyo is sunny."

[0064] 5. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[0065] 6. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[0066] Examples of chat questions:

[0067] User (asking in chat app): "What's the weather like tomorrow?"

[0068] 1. The server receives this request through the chat API.

[0069] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[0070] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[0071] Example of a question history summary:

[0072] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[0073] 1. The server sends these search histories to the summary generator.

[0074] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0075] 3. The server notifies the user of the generated summary via a notification means.

[0076] The above is an embodiment of the invention and its specific examples. This system allows users to efficiently obtain information and easily refer to their past search history.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] The server sets up a personalized database containing the user's account information and search history, and also sets up integration between the search API and the messaging API.

[0080] Step 2:

[0081] The device goes into standby mode, waiting for the user to say the trigger phrase.

[0082] Step 3:

[0083] When the user says a trigger phrase (e.g., "OK, Assistant"), the device switches to voice input mode and records the user's question.

[0084] Step 4:

[0085] The device sends the recorded voice data to a cloud service, where it is processed by a voice recognition engine.

[0086] Step 5:

[0087] The server uses a voice recognition engine in the cloud to convert the voice data into text data.

[0088] Step 6:

[0089] The server analyzes the converted text data and generates an appropriate search query.

[0090] Step 7:

[0091] The server sends the generated search query to the search API and retrieves the search results.

[0092] Step 8:

[0093] The server analyzes the search results and selects the most relevant information.

[0094] Step 9:

[0095] The server generates a voice response text in a natural language based on the selected information.

[0096] Step 10:

[0097] The server uses a text-to-speech engine to convert the natural language text into audio data.

[0098] Step 11:

[0099] The server transmits the generated voice data to the terminal.

[0100] Step 12:

[0101] The terminal plays back the received voice data and provides the user with a voice response.

[0102] Step 13:

[0103] A user sends a text question via a chat app on their smartphone.

[0104] Step 14:

[0105] The server receives text questions through a message API.

[0106] Step 15:

[0107] The server analyzes the received text query and generates an appropriate search query.

[0108] Step 16:

[0109] The server sends the generated search query to the search API and retrieves the search results.

[0110] Step 17:

[0111] The server analyzes the search results and selects the most relevant information.

[0112] Step 18:

[0113] The server generates the selected information as natural language text and returns the text to the user.

[0114] Step 19:

[0115] The server sends the audio and text question history to the summary generator.

[0116] Step 20:

[0117] The server analyzes the summary received from the generation AI and notifies the user of the summary via a notification means.

[0118] This is the specific processing flow of the entire system of the present invention. This system allows users to efficiently obtain information while also checking past question history all at once.

[0119] Example 1

[0120] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0121] In today's information search systems using smart devices and cloud services, users need a way to efficiently ask questions by voice or text and quickly obtain results. However, existing systems often lack the accuracy of voice recognition and personalized search results, hindering the user experience. Furthermore, there are no established methods for effectively utilizing search history or properly notifying search results. Furthermore, relying on cloud services to process voice and text data can make integration complex.

[0122] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0123] In this invention, the server includes a voice input means for a user to input a question by voice, a voice recognition means for converting the voice data into text data, a search means for performing an information search based on the text data, a voice response means for converting the search results into voice data, a voice output means for outputting the voice data to the user, a character input means for inputting a text question, a search means for performing an information search based on the text question, a text response means for returning the search results of the text question to the user in text, a summary generation means for generating a summary of the voice and text question history, a notification means for notifying the user of the generated summary, and a means for transmitting the voice and text data to a cloud service and performing voice recognition and search using the cloud service. This enables a user to efficiently ask questions by voice and text, quickly obtain personalized search results, effectively utilize their search history, and receive appropriate notification of the results.

[0124] A "user" is an entity that uses the system to enter questions and obtain information.

[0125] "Voice input means" refers to a device or function that allows a user to input a question by voice.

[0126] "Speech recognition means" refers to the technology or engine that converts input voice data into text data.

[0127] "Search means" refers to a device or function for searching for information based on text data.

[0128] "Voice response means" refers to a device or technology that converts text data obtained as a search result into voice data.

[0129] "Audio output means" refers to a device or function for reproducing audio data to the user.

[0130] "Character input means" refers to a device or function that allows a user to input a question in text.

[0131] A "text response means" is a device or technology that responds to a user in text with the search results of a text question.

[0132] A "summary generator" is a device or technique that summarizes the audio and text question history and generates summarized information.

[0133] The "notification means" refers to a device or function for notifying the user of the results of the generated summary.

[0134] A "personalized database" is a database that stores each user's search history and account information and responds to them individually.

[0135] "Cloud services" are remote computing resources used to perform voice recognition and search processing.

[0136] The voice response and chat-linked search system according to the present invention is mainly composed of a server, a terminal, and a user. The specific processing of these components and the hardware and software used will be described below.

[0137] Server Roles and Configuration

[0138] The server stores the user's account information and search history in a personalization database. The server also configures search and messaging APIs, such as Google's search and messaging APIs (e.g., Twilio). The server then sends the received voice data to a cloud service for speech recognition. For this purpose, the Google Cloud Speech-to-Text API is used.

[0139] Terminal roles and configuration

[0140] The device accepts voice and text input from the user. When voice input is performed, the device recognizes trigger phrases and captures voice data. The recorded voice data is sent to a server for speech recognition. The device has the ability to generate voice data using the Google Cloud Text-to-Speech API and play it back to the user.

[0141] User Actions

[0142] The user inputs a question by voice or text. For example, if the user asks by voice, "OK, Assistant, what's the weather like today?", the device captures this and sends it to the server. The server performs speech recognition processing and converts it into text data. The server then calls the search API based on the query "What's the weather like today?" and retrieves search results. The server generates a text answer, "Currently, the weather in Tokyo is sunny," converts this into audio data, and sends it to the device. The device then plays back the audio, "Currently, the weather in Tokyo is sunny."

[0143] Use of cloud services

[0144] The voice and text data are sent to a cloud service, where voice recognition and information search are performed within the cloud. This enables highly accurate voice recognition and rapid search results. The server also periodically sends the voice and text question history to a summary generator, which automatically generates a summary. This summary is then notified to the user via a notification device.

[0145] Specific examples

[0146] Examples of voice questions:

[0147] User: "OK, Assistant, what's the weather like right now?"

[0148] 1. The device captures the audio data and sends it to the server.

[0149] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[0150] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[0151] 4. The server generates the answer text "Currently, the weather in Tokyo is sunny."

[0152] 5. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[0153] 6. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[0154] Examples of chat questions:

[0155] User (asking in chat app): "What's the weather like tomorrow?"

[0156] 1. The server receives this request through the chat API.

[0157] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[0158] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[0159] Example of a question history summary:

[0160] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[0161] 1. The server sends these search histories to the summary generator.

[0162] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0163] 3. The server notifies the user of the generated summary via a notification means.

[0164] This system configuration allows users to efficiently retrieve information and easily refer to past search history.The system provides an advanced user experience by combining a voice recognition engine, search API, and text processing engine.

[0165] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0166] Step 1:

[0167] The server saves the user's account information and search history in a personalized database. This is done when a new account is registered. The input is account information and search history, and the output is saved in the personalized database. Specifically, the user's name, email address, past search history, etc. are stored in the database.

[0168] Step 2:

[0169] The server configures the search API and message API integration. The input is authentication information, and the output is completion of API integration. Specifically, it obtains authentication tokens for Google's search API and general message APIs and includes them in the server configuration.

[0170] Step 3:

[0171] The device is in standby mode, and when the user utters a trigger phrase, voice input begins. The input is a trigger phrase such as "OK, Assistant," and the output is a transition to voice input mode. Specifically, the device switches from standby mode to voice input mode and becomes ready to record the user's question.

[0172] Step 4:

[0173] The user inputs a question by voice. The input is a voice question such as "What's the weather like now?", and the output is recorded voice data. Specifically, the user says "What's the weather like now?", and the voice is recorded by the device.

[0174] Step 5:

[0175] The device sends the recorded voice data to a cloud service. The input is the recorded voice data, and the output is the transmission of the voice data to the cloud service. Specifically, the recorded data is sent to a server over the Internet (e.g., Google Cloud Speech-to-Text).

[0176] Step 6:

[0177] The server uses a speech recognition engine in the cloud to convert the voice data into text data. The input is the voice data received from the cloud service, and the output is the converted text data. Specifically, the Google Cloud Speech-to-Text API is used to convert the voice data into text data such as "What's the weather like today?"

[0178] Step 7:

[0179] The server analyzes the converted text data and generates an appropriate search query. The input is text data and the output is a search query. Specifically, it analyzes the text "What's the weather like now?" to generate a search query such as "Current weather."

[0180] Step 8:

[0181] The server sends the generated search query to the search API and retrieves the search results. The input is the search query and the output is the search results. Specifically, the server sends the query "current weather" to Google's search API and receives the search results.

[0182] Step 9:

[0183] The server analyzes the search results, selects the most relevant information, and generates natural language text. The input is the search results, and the output is the natural language text. Specifically, the generated text is, "Currently, the weather in Tokyo is sunny."

[0184] Step 10:

[0185] The server uses a voice response means to convert the generated natural language text into voice data and send it to the terminal. The input is natural language text and the output is voice data. Specifically, the Google Cloud Text-to-Speech API is used to generate the voice "Currently, the weather in Tokyo is sunny," and this is sent to the terminal.

[0186] Step 11:

[0187] The device plays the received voice data and provides the user with a voice response. The input is the voice data, and the output is the voice that is played back. Specifically, the device plays back the voice, "Currently, the weather in Tokyo is sunny."

[0188] Step 12:

[0189] The user inputs a text question into their smartphone and sends it. The input is a text question such as "What's the weather like tomorrow?", and the output is the text data sent.

[0190] Step 13:

[0191] The server analyzes the received text question, generates an appropriate search query using the search API, and retrieves search results. The input is text data, and the output is search results. Specifically, the server analyzes the text "What's the weather like tomorrow?", generates the search query "Tomorrow's weather," and sends it to Google's search API to retrieve search results.

[0192] Step 14:

[0193] The server analyzes the search results, organizes them as natural language text, and sends a text reply to the user. The input is the search results, and the output is natural language text. Specifically, the server generates the text "The weather in Tokyo will be cloudy tomorrow," and sends it back to the user via chat.

[0194] Step 15:

[0195] The server periodically sends the audio and text question history to the summary generation means, which then automatically generates a summary. The input is the question history, and the output is the summary. Specifically, the server sends the search history from the past 24 hours to the summary generation means, and generates a summary such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0196] Step 16:

[0197] The server notifies the user of the generated summary via a notification means. The input is the summary, and the output is a notification to the user. Specifically, the server notifies the user of the generated summary via email or the notification function of their smartphone.

[0198] (Application example 1)

[0199] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0200] In order to enable users to easily order and search for menus using voice or text, food delivery services must simplify conventional operations and provide a more efficient and intuitive interface. They also need a system that can personalize order and search histories to make optimal suggestions to users.

[0201] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0202] In this invention, the server includes voice input means for a user to input orders or menu searches by voice, voice recognition means for converting voice data into text data, search means for performing an internet search and menu information search based on the text data, voice response means for converting search results into voice data, voice output means for outputting the voice data to the user, character input means for inputting orders or menu searches using character input means, search means for performing an internet search and menu information search based on the character input, character response means for providing a method for returning search results of the character input to the user in text, summary generation means for generating a summary of the voice and character input history, and notification means for notifying the user of the generated summary results. This allows users to efficiently and intuitively order or search by voice or character input and receive personalized suggestions based on their past history.

[0203] "Voice input means" refers to a means by which a user can input information using voice.

[0204] The "voice recognition means" is a means for converting voice data into text data.

[0205] The "search means" is a means for searching the Internet or menu information based on text data.

[0206] The "voice response means" is a means for converting search results into voice data and providing it to the user.

[0207] The "audio output means" is a means for outputting audio data to the user.

[0208] "Character input means" refers to a means by which a user inputs information using characters.

[0209] The "text response means" is a means for returning search results based on a text question to a user in text form.

[0210] The "summary generation means" is a means for generating summaries of the audio and text question history.

[0211] The "notification means" is a means for notifying the user of the results of the generated summary.

[0212] A "personalized database" is a database that stores a user's input history and provides personalized information.

[0213] The present invention provides a system that allows efficient and intuitive operation of a food delivery service. Specific embodiments for realizing this system are described below.

[0214] System configuration

[0215] The system consists of a user, a device, and a server. Users use devices such as smartphones and smart speakers to place orders and search for menus by voice or text. The device accepts both voice and text input.

[0216] Hardware and software used

[0217] Hardware: smartphones, smart speakers, microphones

[0218] Software: Speech recognition engine (e.g., Google Cloud Speech-to-Text API), speech synthesis engine (e.g., Google Text-to-Speech API)

[0219] Data processing and calculation

[0220] Voice input means

[0221] When a user orders or searches the menu by voice, the device captures the voice data, which is then sent to a cloud service where it is converted into text using a voice recognition engine.

[0222] Voice recognition means

[0223] The voice recognition engine in the cloud service converts voice data into text data. For example, a voice saying "I want to order a pizza" is converted into text "I want to order a pizza."

[0224] Search methods

[0225] The server analyzes the converted text data and generates appropriate queries to search the Internet and retrieve menu information, and the search results are retrieved on the server side.

[0226] Voice response means

[0227] The server analyzes the search results and selects the most relevant information, which is then generated as natural language text and converted into voice data using a speech synthesis engine.

[0228] Audio output means

[0229] The device plays audio data to the user and provides information such as search results and order confirmations.

[0230] Character input means and character response means

[0231] Users can also enter text on their smartphone screens. This text data is sent to the server, which uses search tools to generate appropriate search queries. The search results are organized as natural language text and provided to users in text format.

[0232] Summary generation means and notification means

[0233] The server transmits the user's voice and text input history to a summary generation means, which generates a summary such as "Today's order: pizza, salad" and notifies the user via a notification means.

[0234] Specific examples

[0235] For example, if a user says, "I'd like to order sushi," the device captures this speech and sends it to the cloud service. The speech recognition engine converts it into text data, and the server analyzes the text data to generate an appropriate search query. The search result, "The recommended sushi is tuna nigiri," is converted into audio data by the speech synthesis engine and played on the device.

[0236] Examples of prompts used by generative AI models include:

[0237] The user says, "I want to order sushi." Generate an appropriate response.

[0238] In this way, the present invention provides a system that allows users to efficiently and intuitively order and search using voice or text input.

[0239] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0240] Step 1:

[0241] Users can input orders and menu searches by voice. Users can input "I want to order pizza" into a device such as a smartphone or smart speaker.

[0242] Step 2:

[0243] The device captures the user's voice and generates audio data, which is then sent to a cloud service.

[0244] Step 3:

[0245] The server uses the cloud service's voice recognition engine to convert the voice data into text data. For example, the voice saying "I'd like to order a pizza" is converted into text data saying "I'd like to order a pizza."

[0246] Step 4:

[0247] The server analyzes the converted text data and generates an appropriate search query. Specifically, based on the input text data "I want to order pizza," it generates the query "pizza menu" to search for menu information.

[0248] Step 5:

[0249] The server performs an internet search and menu information search based on the generated search query, and obtains the relevant menu information using the search means.

[0250] Step 6:

[0251] The server analyzes the search results and selects the most relevant information, for example, "The recommended pizza is Margherita."

[0252] Step 7:

[0253] The server generates the selected information as natural language text and converts it into voice data using a speech synthesis engine. This is the process of converting the text "The recommended pizza is Margherita" into voice data.

[0254] Step 8:

[0255] The terminal plays back the voice data and provides the search results to the user. By playing back the voice data "The recommended pizza is Margherita" to the user, the terminal confirms and suggests the order.

[0256] Step 9:

[0257] When a user inputs a text question on the screen of a smartphone, the user uses the text input means to place an order or search for a menu item. The input text data is sent to the server.

[0258] Step 10:

[0259] The server performs an internet search and menu information search based on the received text data. For example, if the text data "I want to order sushi" is received, the server generates the search query "sushi menu."

[0260] Step 11:

[0261] The server retrieves the search results, organizes them as natural language text, and returns the text to the user, providing the user with the text, "The recommended sushi is tuna nigiri."

[0262] Step 12:

[0263] The server transmits the input history of voice and text to the summary generation means, which analyzes the input history and automatically generates a summary.

[0264] Step 13:

[0265] The server notifies the user of the generated summary via the notification means. For example, the server notifies the user of the summary "Today's order: pizza, sushi."

[0266] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0267] The voice response and chat-linked search system of the present invention includes a voice input means installed in a smart device, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the converted text data, and a voice response means for converting the search results into voice data and providing it to the user. The user can also use the smartphone's input means to submit text questions, and an Internet search can be performed based on the text input, with the text response means providing the results to the user in text. The voice and text question history is summarized by a summary generation means and notified to the user by a notification means. Furthermore, the system is equipped with an emotion engine that recognizes emotions from the user's voice data, and can adjust the response content based on the emotional information.

[0268] Natural language description of the program

[0269] 1. Initial Setup:

[0270] The server sets up a personalized database containing the user's account information and search history, as well as the search API, message API, and emotion engine.

[0271] 2. Audio question reception:

[0272] The device goes into standby mode, waiting for the user to say the trigger phrase.

[0273] 3. Voice input:

[0274] When the user says a trigger phrase (e.g., "OK, Assistant"), the device switches to voice input mode and records the user's question.

[0275] 4. Speech Recognition:

[0276] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine in the cloud to convert the voice data into text data.

[0277] 5. Information Search:

[0278] The server analyzes the converted text data, generates an appropriate search query, and sends the generated search query to the search API to retrieve search results.

[0279] 6. Emotion recognition:

[0280] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[0281] 7. Tailor your response:

[0282] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[0283] 8. Voice response generation:

[0284] The server analyzes the retrieved search results and emotion information, selects the most relevant information, and generates a natural language voice response text. The natural language text is converted into voice data using the voice response means and sent to the terminal.

[0285] 9. Audio Answer:

[0286] The terminal plays back the received voice data and provides the user with a voice response.

[0287] 10. Chat questions:

[0288] The user types in a text question on their smartphone and sends it.

[0289] 11. Handling chat questions:

[0290] The server receives text questions through the message API, analyzes the received text questions, generates appropriate search queries using the search API, retrieves search results, analyzes the retrieved search results, organizes them as natural language text, and sends a text reply to the user.

[0291] 12. Question History Summary:

[0292] The server sends the audio and text question history to the summary generation means, which automatically generates a summary. The summary is then notified to the user via the notification means. The server also adjusts the summary content and notification format based on the emotion information.

[0293] Specific examples

[0294] Examples of voice questions:

[0295] User: "OK, Assistant, what's the weather like right now?"

[0296] 1. The device captures the audio data and sends it to the server.

[0297] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[0298] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[0299] 4. The server simultaneously sends the voice data to the emotion engine to recognize the user's emotion.

[0300] 5. The server generates a response text based on the emotion information: "Currently, the weather in Tokyo is sunny."

[0301] 6. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[0302] 7. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[0303] Examples of chat questions:

[0304] User (asking in chat app): "What's the weather like tomorrow?"

[0305] 1. The server receives a text question through the message API.

[0306] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[0307] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[0308] Example of a question history summary:

[0309] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[0310] 1. The server sends these search histories to the summary generator.

[0311] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0312] 3. The server notifies the user of the generated summary via a notification method, and adjusts the summary content and notification format if the user's emotions are recognized.

[0313] The above is a specific embodiment of the system of the present invention that combines an emotion engine. This system allows users to obtain information efficiently and in a way that takes emotion into consideration, and also allows users to check their past question history all at once.

[0314] The processing flow will be explained below.

[0315] Step 1:

[0316] The server sets up a personalized database containing the user's account information and search history, as well as the search API, message API, and emotion engine integration.

[0317] Step 2:

[0318] The device enters standby mode and prepares to switch to voice input mode when the user speaks the trigger phrase.

[0319] Step 3:

[0320] The user says the trigger phrase (e.g., "OK, Assistant").

[0321] Step 4:

[0322] The device detects the user's trigger phrase, switches to voice input mode, and records the user's question.

[0323] Step 5:

[0324] The device sends the recorded audio data to a cloud service.

[0325] Step 6:

[0326] The server uses a voice recognition engine in the cloud to convert the voice data into text data.

[0327] Step 7:

[0328] The server analyzes the converted text data and generates an appropriate search query.

[0329] Step 8:

[0330] The server sends the generated search query to the search API and retrieves the search results.

[0331] Step 9:

[0332] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[0333] Step 10:

[0334] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[0335] Step 11:

[0336] The server generates a voice response text in a natural language based on the adjusted response content.

[0337] Step 12:

[0338] The server uses a text-to-speech engine to convert the natural language text into audio data.

[0339] Step 13:

[0340] The server transmits the generated voice data to the terminal.

[0341] Step 14:

[0342] The terminal plays back the received voice data and provides the user with a voice response.

[0343] Step 15:

[0344] A user sends a text question via a chat app on their smartphone.

[0345] Step 16:

[0346] The server receives text questions through a message API.

[0347] Step 17:

[0348] The server analyzes the received text query and generates an appropriate search query.

[0349] Step 18:

[0350] The server sends the generated search query to the search API and retrieves the search results.

[0351] Step 19:

[0352] The server analyzes the search results and selects the most relevant information.

[0353] Step 20:

[0354] The server generates the selected information as natural language text and replies to the user via chat.

[0355] Step 21:

[0356] The server sends the audio and text question history to the summary generator.

[0357] Step 22:

[0358] The summary generator summarizes the audio and text question history and automatically generates a summary.

[0359] Step 23:

[0360] The summary generator sends the summary to the notification unit. If emotion information is present, the summary content and notification format are adjusted based on that information.

[0361] Step 24:

[0362] The server notifies the user of the generated summary via the notification means.

[0363] In this way, the system of the present invention can efficiently process voice or text questions from users and provide answers that take the user's feelings into consideration. In addition, by summarizing and notifying the user of the question history, the system can further facilitate the user's information acquisition.

[0364] Example 2

[0365] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0366] Conventional search systems that use voice and text input are unable to consider user sentiment when providing search results, limiting their ability to provide optimal search results. Furthermore, they do not effectively utilize voice and text question history, and lack a means to easily review past questions. This makes it difficult for users to quickly and accurately obtain the information they need.

[0367] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0368] In this invention, the server includes a summary generation unit that generates summaries of the voice and text question history, a response content adjustment unit that adjusts the response content of search results based on emotion information, and an emotion engine that performs emotion recognition. This makes it possible to provide appropriate and personalized search results while taking into consideration the user's emotions. In addition, by automatically summarizing the past question history and notifying the user, efficient information management and confirmation is possible.

[0369] "Voice input means" refers to a device or interface that allows a user to input a question by voice.

[0370] "Speech recognition means" refers to a technology or module for converting input voice data into text data.

[0371] The "search means" is a technology or module for conducting an internet search based on the converted text data and obtaining appropriate search results.

[0372] "Voice response means" refers to a technology or module for converting search results into voice data and providing it to the user.

[0373] "Audio output means" refers to a device or interface for outputting audio data to a user and conveying information.

[0374] "Character input means" refers to a device or interface that allows a user to input characters.

[0375] A "text response means" is a technology or module for conducting a search based on character input and returning search results to the user in text.

[0376] The "summary generation means" is a technology or module for summarizing the audio and text question history and automatically generating a summary.

[0377] "Notification means" refers to a technology or module for notifying the user of the summary generated results.

[0378] An "emotion engine" is a technology or module for recognizing emotions from a user's voice data.

[0379] The "response content adjustment means" is a technology or module for adjusting the response content of search results based on the emotion information obtained from the emotion engine.

[0380] A "personalized database" is a database that stores each user's account information and search history and provides personalized search results.

[0381] A "cloud service" is a service that uses remote computing resources and storage provided via the Internet.

[0382] The voice response and chat-linked search system of the present invention is composed of multiple means installed in a smart device, which processes users' voice and text questions, provides appropriate responses, and provides more personalized services through emotion recognition.

[0383] This system includes a voice input means for the user to input a question by voice, a voice recognition means for converting voice data into text data, a search means for performing an internet search based on the text data, a voice response means for converting search results into voice data, a voice output means for outputting the voice data to the user, a text input means for inputting a text question, a text response means for returning search results to the user in text, a summary generation means for generating a summary of the voice and text question history, and a notification means for notifying the user of the generated summary. Furthermore, it includes an emotion engine for recognizing emotions from the user's voice data, and a response content adjustment means for adjusting the response content of the search results based on the emotion information. Each part of this system can be implemented using an ordinary smartphone or a cloud server.

[0384] The smartphone's microphone is used as the voice input means. When the user's voice input is recognized as a trigger phrase (e.g., "OK, Assistant"), the voice recognition means sends the voice to a cloud service, where it is converted into text data by a voice recognition engine (e.g., Google Cloud Speech-to-Text API). The search means then searches the converted text data using an internet search API (e.g., Google Custom Search API).

[0385] The search results are converted by the voice response means and notified to the user by voice through the voice output means. In this series of steps, an emotion recognition engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotional state, and the response content adjustment means determines an appropriate response based on this.

[0386] The smartphone keyboard is used as the text input means. The user types and sends a question in text. For example, if the user types "What's the weather going to be like tomorrow?", the question is sent to the server via a messaging API (e.g., Twilio API), which then performs an internet search. The retrieved search results are returned to the user in text via the text response means.

[0387] The question history is stored in a question history database. The summary generation means summarizes this history and notifies the user as necessary. For example, if a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?", the summary generation means generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy" and notifies the user via the notification means.

[0388] Specific examples

[0389] Examples of voice questions:

[0390] User: "OK, Assistant, what's the weather like right now?"

[0391] 1. The device captures the audio data and sends it to the server.

[0392] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[0393] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[0394] 4. The server simultaneously sends the voice data to the emotion engine to recognize the user's emotion.

[0395] 5. The server generates a response text based on the emotion information: "Currently, the weather in Tokyo is sunny."

[0396] 6. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[0397] 7. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[0398] Examples of chat questions:

[0399] User (asking in chat app): "What's the weather like tomorrow?"

[0400] 1. The server receives a text question through the message API.

[0401] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[0402] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[0403] Example of a question history summary:

[0404] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[0405] 1. The server sends these search histories to the summary generator.

[0406] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0407] 3. The server notifies the user of the generated summary via a notification means.

[0408] The system of the present invention allows users to obtain information efficiently and sensitively, and also allows them to review their past question history in one place, significantly improving user convenience and enabling more personalized services.

[0409] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0410] Step 1: Initial Setup

[0411] The server sets up a personalized database that stores user account information and past search history, and also establishes integration with the search API, message API, and emotion engine.

[0412] Input: User information, past search history

[0413] Output: Personalized database, API and engine integration settings

[0414] Step 2: Voice inquiry reception

[0415] The device will enter standby mode and wait for the user to say a trigger phrase, such as "OK, Assistant."

[0416] Input: Trigger phrase

[0417] Output: Activate voice input mode

[0418] Step 3: Voice Input

[0419] When the user utters the trigger phrase, the device switches to voice input mode and records the user's question, for example, "What's the weather like today?"

[0420] Input: User's voice question

[0421] Output: Recorded audio data

[0422] Step 4: Voice Recognition

[0423] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine to convert the voice data into text data.

[0424] Input: Recorded audio data

[0425] Output: Converted text data

[0426] Step 5: Information Search

[0427] The server analyzes the converted text data and generates an appropriate search query, which it then sends to an Internet search API to retrieve search results.

[0428] Input: Text data

[0429] Output: Search result data

[0430] Step 6: Emotion Recognition

[0431] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[0432] Input: Audio data

[0433] Output: Emotional information

[0434] Step 7: Tailor your response

[0435] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[0436] Input: Emotion information, search result data

[0437] Output: Adjusted response content

[0438] Step 8: Generate voice response

[0439] The server integrates the search results and emotional information to generate a natural language voice response text, converts the text into voice data, and sends it to the device.

[0440] Input: Adjusted response content

[0441] Output: Audio data

[0442] Step 9: Answer by voice

[0443] The terminal plays back the received voice data and provides the user with a voice response.

[0444] Input: Audio data

[0445] Output: A spoken response to the user

[0446] Step 10: Chat with your questions

[0447] A user types a text question into a smartphone and sends it, for example, "What's the weather like tomorrow?" in a chat app.

[0448] Input: User's character question

[0449] output: Character data to be sent

[0450] Step 11: Handling chat questions

[0451] The server receives text questions through the message API, generates search queries based on the questions using the search API, retrieves search results, organizes the results into natural language text, and sends a text reply to the user.

[0452] Input: User's character question

[0453] Output: Search result text

[0454] Step 12: Question History Summary

[0455] The server sends the audio and text question history to the summary generation means, which automatically generates a summary. The summary is then notified to the user via the notification means. The server also adjusts the summary content and notification format based on the emotion information.

[0456] Input: Voice question history, text question history

[0457] Output: Summary, Notification

[0458] (Application example 2)

[0459] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0460] In today's world, it's important to provide users with a way to quickly and efficiently obtain the information they need. However, traditional voice response and chat-based search systems are unable to adjust responses based on the user's emotional state, which can result in a poor user experience. Furthermore, they are unable to summarize multiple query histories and notify users, making it difficult to easily review past search results. Another problem is the inability to provide personalized results based on voice or text queries.

[0461] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice input means for a user to input a question by voice, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the text data, a voice response means for converting the search results into voice data, a voice output means for outputting the voice data to the user, a character input means for inputting a text question, a search means for performing an Internet search based on the text question, a text response means for providing the user with a method for replying to the search results of the text question in text form, a summary generation means for generating a summary of the voice and text question history, a notification means for notifying the user of the generated summary results, an emotion recognition means for acquiring emotion information, and a response adjustment means for adapting the search results based on the emotion information. This makes it possible to provide personalized search results according to the user's emotions, improve the quality of the user experience, and make it easy to check past question history.

[0462] The "voice input means" is a means for the user to input a question by voice.

[0463] The "voice recognition means" is a means for converting voice data into text data.

[0464] "Search means" refers to a means for conducting an Internet search based on text data.

[0465] The "voice response means" is a means for converting search results into voice data.

[0466] The "audio output means" is a means for outputting audio data to the user.

[0467] The "text input means" is a means for the user to input a text question.

[0468] The "text response means" is a means for returning search results for a text question to the user in text form.

[0469] The "summary generation means" is a means for generating summaries of the audio and text question history.

[0470] The "notification means" is a means for notifying the user of the results of the generated summary.

[0471] "Emotion recognition means" is a means for acquiring emotion information.

[0472] A "response adjustment means" is a means for adapting search results based on emotional information.

[0473] A "personalized database" is a database that stores a user's question history and customizes search results individually.

[0474] A "cloud service" is a service that processes voice data and text data on an online server and performs voice recognition and search.

[0475] To implement this invention, the following hardware and software are primarily used. A user uses a smart device (e.g., a smartphone) to ask a question by voice or text and obtain search results corresponding to the question. The following describes the functions and processing flow of the system that realizes this application example.

[0476] Required Hardware and Software

[0477] 1. Smart devices (e.g. smartphones, tablets): Provide an interface with users. Devices that allow voice input, voice output, and text input.

[0478] 2. Cloud Services: Online platforms for processing voice and text data. They perform speech recognition, search, and emotion recognition. They use services such as Google Speech-to-Text API, IBM Watson Tone Analyzer, and Amazon Polly.

[0479] 3. Server: Manages the user's personalized database and generates search results. Provides tailored responses based on the user's question history and sentiment information.

[0480] System Operation Overview

[0481] 1. Audio question reception:

[0482] The user speaks a voice trigger phrase (e.g., "OK, Assistant") into the smart device. When the device switches to voice input mode, the user asks, "What movies are recommended today?"

[0483] 2. Speech and Emotion Recognition:

[0484] The smart device sends the voice data to a cloud service, which uses the Google Speech-to-Text API to convert the voice data into text, and simultaneously sends the voice data to the IBM Watson Tone Analyzer to recognize the user's emotions.

[0485] 3. Information Search:

[0486] The server generates a search query based on the converted text data and sends it to the appropriate search API (e.g., a movie database). Based on the user's emotional information, the server can recommend movies.

[0487] 4. Tailor your response:

[0488] The server selects the most appropriate information from the search results based on emotional information and generates a response in natural language. For example, if the user feels like relaxing, it generates a response such as, "The recommended relaxing movie is 'Inception.'"

[0489] 5. Voice response generation:

[0490] The generated response is converted into voice data using Amazon Polly and sent to the smart device, which then outputs the converted voice data to the user and provides a response.

[0491] 6. Question History Summary:

[0492] The server summarizes the user's past question history and periodically notifies the user of the summarized information. The summary information is presented to the user in the form of, for example, "Past 7 days' viewing history: 'Inception', 'The Matrix'."

[0493] Examples and prompts

[0494] Specific examples

[0495] When a user says "What are the best movies tonight?", the following happens:

[0496] 1. The voice data is sent to a cloud service and converted into text data by a speech recognition engine. A query such as "Today's recommended movies" is generated.

[0497] 2. The emotion recognition engine detects the user's desire to relax from their voice.

[0498] 3. The server searches a movie database and applies a filter that recommends "relaxing movies."

[0499] 4. Generate a response sentence: "A recommended relaxing movie is 'Inception'." and convert it into speech.

[0500] 5. The smart device plays this audio to the user.

[0501] Example prompts

[0502] Here are some example prompts for using generative AI models:

[0503] User asks: "What movie do you recommend today?"

[0504] Hardware: Smartphone

[0505] Software Settings:

[0506] Speech Recognition: Google Speech-to-Text API

[0507] Emotion recognition: IBM Watson Tone Analyzer

[0508] Speech synthesis: Amazon Polly

[0509] Specific processing:

[0510] 1. The voice data is sent to the speech recognition engine and converted into text.

[0511] 2. Send the text data to the movie database API and get the search results.

[0512] 3. Send the voice data to the emotion engine and obtain the emotion information.

[0513] 4. Based on the emotional information, the system selects an appropriate movie from the search results and generates a response in natural language.

[0514] 5. The response is converted into audio data and played on the smartphone.

[0515] Expected response: "A movie recommendation to fit your current mood is 'Inception'."

[0516] The system provided by this invention allows users to obtain personalized search results based on their emotions and easily check their past question history. This system improves the quality of the user experience and provides a means to efficiently obtain information.

[0517] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0518] Step 1:

[0519] Voice inquiry reception

[0520] The device enters standby mode and waits for the user to speak a trigger phrase. The input is a voice trigger phrase spoken by the user (e.g., "OK, Assistant"), and the device captures the voice and switches to voice input mode to accept the user's question.

[0521] Step 2:

[0522] Voice input

[0523] The user asks the device, "What movies do you recommend today?" The device records the user's voice and sends the voice data to the cloud service. The input is the recorded voice data, and the output is the voice data sent to the cloud service.

[0524] Step 3:

[0525] Voice Recognition

[0526] The cloud service converts the transmitted voice data into text data using the Google Speech-to-Text API. The input is the voice data, and the output is the text data "What movies do you recommend today?" The cloud service analyzes the voice data and generates the corresponding text.

[0527] Step 4:

[0528] emotion recognition

[0529] The cloud service sends the voice data to the IBM Watson Tone Analyzer along with the text data to recognize the user's emotions. The input is the voice data, and the output is the user's emotional information. The cloud service analyzes the voice data and recognizes the user's emotional state.

[0530] Step 5:

[0531] Information Search

[0532] The server analyzes the text data obtained from speech recognition and generates an appropriate search query. For example, it sends a query such as "Today's recommended movies" to a movie database API and retrieves relevant search results. The input is the converted text data, and the output is a list of search results.

[0533] Step 6:

[0534] Tailoring response content

[0535] The server determines the optimal response based on the search results and emotional information, and generates a response in natural language. For example, if the user is in the mood to relax, it generates the text "A recommended relaxing movie is 'Inception.'" The input is the search results and emotional information, and the output is a response in natural language.

[0536] Step 7:

[0537] Voice Response Generation

[0538] The server converts the generated response text into voice data using Amazon Polly and sends it to the device. The input is text data and the output is voice data. The server analyzes the response text and generates voice data.

[0539] Step 8:

[0540] Answer by voice

[0541] The device plays the received voice data and provides a voice response to the user. The input is the voice data sent from the server, and the output is the voice played from the device. The device analyzes the voice data and conveys information to the user by voice.

[0542] Step 9:

[0543] Question History Summary

[0544] The server sends the user's past question history to the summary generation means, which then summarizes the question history. For example, it generates a summary result such as "Past 7 days' viewing history: 'Inception', 'The Matrix'." The input is the past question history data, and the output is summarized text data. The server analyzes the question history and generates summary information.

[0545] Step 10:

[0546] Summary Notification

[0547] The server notifies the user of the generated summarized question history via a notification means. The input is the summarized text data, and the output is notification information for the user. The server analyzes the summarized data and notifies the user in an appropriate format.

[0548] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0549] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0550] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0551] [Second embodiment]

[0552] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0553] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0554] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0555] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0556] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0557] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0558] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0559] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0560] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0561] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0562] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0563] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0564] The voice response and chat-linked search system of the present invention includes a voice input means installed in a smart device, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the converted text data, and a voice response means for converting the search results into voice data and providing it to the user. The user can also use the input means of the smartphone to submit text questions, and an Internet search can be performed based on the text input, with the text response means providing the results to the user in text. The voice and text question history is summarized by a summary generation means and notified to the user by a notification means.

[0565] Natural language description of the program

[0566] 1. Initial Setup:

[0567] The server stores the user's account information and search history in a personalized database, and also configures the search API and message API integration.

[0568] 2. Audio question reception:

[0569] The device is in standby mode and starts voice input when the user speaks a trigger phrase. The user then speaks a question.

[0570] 3. Speech Recognition:

[0571] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine in the cloud to convert the voice data into text data.

[0572] 4. Information Search:

[0573] The server analyzes the converted text data, generates an appropriate search query, and sends the generated search query to the search API to retrieve search results.

[0574] 5. Voice response generation:

[0575] The server analyzes the search results, selects the most relevant information, and generates natural language text based on the search results.The natural language text is converted into voice data using a voice response means and sent to the terminal.

[0576] 6. Audio Answer:

[0577] The terminal plays back the received voice data and provides the user with a voice response.

[0578] 7. Chat questions:

[0579] The user types in a text question on their smartphone and sends it.

[0580] 8. Handling chat questions:

[0581] The server analyzes the received text question, generates an appropriate search query using a search API, retrieves search results, analyzes the retrieved search results, organizes them as natural language text, and sends a text reply to the user.

[0582] 9. Question History Summary:

[0583] The server periodically transmits the question history of voice and text to the summary generating means, which automatically generates a summary, and notifies the user of the summary via the notification means.

[0584] Specific examples

[0585] Examples of voice questions:

[0586] User: "OK, Assistant, what's the weather like right now?"

[0587] 1. The device captures the audio data and sends it to the server.

[0588] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[0589] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[0590] 4. The server generates the answer text "Currently, the weather in Tokyo is sunny."

[0591] 5. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[0592] 6. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[0593] Examples of chat questions:

[0594] User (asking in chat app): "What's the weather like tomorrow?"

[0595] 1. The server receives this request through the chat API.

[0596] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[0597] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[0598] Example of a question history summary:

[0599] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[0600] 1. The server sends these search histories to the summary generator.

[0601] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0602] 3. The server notifies the user of the generated summary via a notification means.

[0603] The above is an embodiment of the invention and its specific examples. This system allows users to efficiently obtain information and easily refer to their past search history.

[0604] The processing flow will be explained below.

[0605] Step 1:

[0606] The server sets up a personalized database containing the user's account information and search history, and also sets up integration between the search API and the messaging API.

[0607] Step 2:

[0608] The device goes into standby mode, waiting for the user to say the trigger phrase.

[0609] Step 3:

[0610] When the user says a trigger phrase (e.g., "OK, Assistant"), the device switches to voice input mode and records the user's question.

[0611] Step 4:

[0612] The device sends the recorded voice data to a cloud service, where it is processed by a voice recognition engine.

[0613] Step 5:

[0614] The server uses a voice recognition engine in the cloud to convert the voice data into text data.

[0615] Step 6:

[0616] The server analyzes the converted text data and generates an appropriate search query.

[0617] Step 7:

[0618] The server sends the generated search query to the search API and retrieves the search results.

[0619] Step 8:

[0620] The server analyzes the search results and selects the most relevant information.

[0621] Step 9:

[0622] The server generates a voice response text in a natural language based on the selected information.

[0623] Step 10:

[0624] The server uses a text-to-speech engine to convert the natural language text into audio data.

[0625] Step 11:

[0626] The server transmits the generated voice data to the terminal.

[0627] Step 12:

[0628] The terminal plays back the received voice data and provides the user with a voice response.

[0629] Step 13:

[0630] A user sends a text question via a chat app on their smartphone.

[0631] Step 14:

[0632] The server receives text questions through a message API.

[0633] Step 15:

[0634] The server analyzes the received text query and generates an appropriate search query.

[0635] Step 16:

[0636] The server sends the generated search query to the search API and retrieves the search results.

[0637] Step 17:

[0638] The server analyzes the search results and selects the most relevant information.

[0639] Step 18:

[0640] The server generates the selected information as natural language text and returns the text to the user.

[0641] Step 19:

[0642] The server sends the audio and text question history to the summary generator.

[0643] Step 20:

[0644] The server analyzes the summary received from the generation AI and notifies the user of the summary via a notification means.

[0645] This is the specific processing flow of the entire system of the present invention. This system allows users to efficiently obtain information while also checking past question history all at once.

[0646] Example 1

[0647] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0648] In today's information search systems using smart devices and cloud services, users need a way to efficiently ask questions by voice or text and quickly obtain results. However, existing systems often lack the accuracy of voice recognition and personalized search results, hindering the user experience. Furthermore, there are no established methods for effectively utilizing search history or properly notifying search results. Furthermore, relying on cloud services to process voice and text data can make integration complex.

[0649] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0650] In this invention, the server includes a voice input means for a user to input a question by voice, a voice recognition means for converting the voice data into text data, a search means for performing an information search based on the text data, a voice response means for converting the search results into voice data, a voice output means for outputting the voice data to the user, a character input means for inputting a text question, a search means for performing an information search based on the text question, a text response means for returning the search results of the text question to the user in text, a summary generation means for generating a summary of the voice and text question history, a notification means for notifying the user of the generated summary, and a means for transmitting the voice and text data to a cloud service and performing voice recognition and search using the cloud service. This enables a user to efficiently ask questions by voice and text, quickly obtain personalized search results, effectively utilize their search history, and receive appropriate notification of the results.

[0651] A "user" is an entity that uses the system to enter questions and obtain information.

[0652] "Voice input means" refers to a device or function that allows a user to input a question by voice.

[0653] "Speech recognition means" refers to the technology or engine that converts input voice data into text data.

[0654] "Search means" refers to a device or function for searching for information based on text data.

[0655] "Voice response means" refers to a device or technology that converts text data obtained as a search result into voice data.

[0656] "Audio output means" refers to a device or function for reproducing audio data to the user.

[0657] "Character input means" refers to a device or function that allows a user to input a question in text.

[0658] A "text response means" is a device or technology that responds to a user in text with the search results of a text question.

[0659] A "summary generator" is a device or technique that summarizes the audio and text question history and generates summarized information.

[0660] The "notification means" refers to a device or function for notifying the user of the results of the generated summary.

[0661] A "personalized database" is a database that stores each user's search history and account information and responds to them individually.

[0662] "Cloud services" are remote computing resources used to perform voice recognition and search processing.

[0663] The voice response and chat-linked search system according to the present invention is mainly composed of a server, a terminal, and a user. The specific processing of these components and the hardware and software used will be described below.

[0664] Server Roles and Configuration

[0665] The server stores the user's account information and search history in a personalization database. The server also configures search and messaging APIs, such as Google's search and messaging APIs (e.g., Twilio). The server then sends the received voice data to a cloud service for speech recognition. For this purpose, the Google Cloud Speech-to-Text API is used.

[0666] Terminal roles and configuration

[0667] The device accepts voice and text input from the user. When voice input is performed, the device recognizes trigger phrases and captures voice data. The recorded voice data is sent to a server for speech recognition. The device has the ability to generate voice data using the Google Cloud Text-to-Speech API and play it back to the user.

[0668] User Actions

[0669] The user inputs a question by voice or text. For example, if the user asks by voice, "OK, Assistant, what's the weather like today?", the device captures this and sends it to the server. The server performs speech recognition processing and converts it into text data. The server then calls the search API based on the query "What's the weather like today?" and retrieves search results. The server generates a text answer, "Currently, the weather in Tokyo is sunny," converts this into audio data, and sends it to the device. The device then plays back the audio, "Currently, the weather in Tokyo is sunny."

[0670] Use of cloud services

[0671] The voice and text data are sent to a cloud service, where voice recognition and information search are performed within the cloud. This enables highly accurate voice recognition and rapid search results. The server also periodically sends the voice and text question history to a summary generator, which automatically generates a summary. This summary is then notified to the user via a notification device.

[0672] Specific examples

[0673] Examples of voice questions:

[0674] User: "OK, Assistant, what's the weather like right now?"

[0675] 1. The device captures the audio data and sends it to the server.

[0676] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[0677] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[0678] 4. The server generates the answer text "Currently, the weather in Tokyo is sunny."

[0679] 5. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[0680] 6. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[0681] Examples of chat questions:

[0682] User (asking in chat app): "What's the weather like tomorrow?"

[0683] 1. The server receives this request through the chat API.

[0684] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[0685] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[0686] Example of a question history summary:

[0687] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[0688] 1. The server sends these search histories to the summary generator.

[0689] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0690] 3. The server notifies the user of the generated summary via a notification means.

[0691] This system configuration allows users to efficiently retrieve information and easily refer to past search history.The system provides an advanced user experience by combining a voice recognition engine, search API, and text processing engine.

[0692] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0693] Step 1:

[0694] The server saves the user's account information and search history in a personalized database. This is done when a new account is registered. The input is account information and search history, and the output is saved in the personalized database. Specifically, the user's name, email address, past search history, etc. are stored in the database.

[0695] Step 2:

[0696] The server configures the search API and message API integration. The input is authentication information, and the output is completion of API integration. Specifically, it obtains authentication tokens for Google's search API and general message APIs and includes them in the server configuration.

[0697] Step 3:

[0698] The device is in standby mode, and when the user utters a trigger phrase, voice input begins. The input is a trigger phrase such as "OK, Assistant," and the output is a transition to voice input mode. Specifically, the device switches from standby mode to voice input mode and becomes ready to record the user's question.

[0699] Step 4:

[0700] The user inputs a question by voice. The input is a voice question such as "What's the weather like now?", and the output is recorded voice data. Specifically, the user says "What's the weather like now?", and the voice is recorded by the device.

[0701] Step 5:

[0702] The device sends the recorded voice data to a cloud service. The input is the recorded voice data, and the output is the transmission of the voice data to the cloud service. Specifically, the recorded data is sent to a server over the Internet (e.g., Google Cloud Speech-to-Text).

[0703] Step 6:

[0704] The server uses a speech recognition engine in the cloud to convert the voice data into text data. The input is the voice data received from the cloud service, and the output is the converted text data. Specifically, the Google Cloud Speech-to-Text API is used to convert the voice data into text data such as "What's the weather like today?"

[0705] Step 7:

[0706] The server analyzes the converted text data and generates an appropriate search query. The input is text data and the output is a search query. Specifically, it analyzes the text "What's the weather like now?" to generate a search query such as "Current weather."

[0707] Step 8:

[0708] The server sends the generated search query to the search API and retrieves the search results. The input is the search query and the output is the search results. Specifically, the server sends the query "current weather" to Google's search API and receives the search results.

[0709] Step 9:

[0710] The server analyzes the search results, selects the most relevant information, and generates natural language text. The input is the search results, and the output is the natural language text. Specifically, the generated text is, "Currently, the weather in Tokyo is sunny."

[0711] Step 10:

[0712] The server uses a voice response means to convert the generated natural language text into voice data and send it to the terminal. The input is natural language text and the output is voice data. Specifically, the Google Cloud Text-to-Speech API is used to generate the voice "Currently, the weather in Tokyo is sunny," and this is sent to the terminal.

[0713] Step 11:

[0714] The device plays the received voice data and provides the user with a voice response. The input is the voice data, and the output is the voice that is played back. Specifically, the device plays back the voice, "Currently, the weather in Tokyo is sunny."

[0715] Step 12:

[0716] The user inputs a text question into their smartphone and sends it. The input is a text question such as "What's the weather like tomorrow?", and the output is the text data sent.

[0717] Step 13:

[0718] The server analyzes the received text question, generates an appropriate search query using the search API, and retrieves search results. The input is text data, and the output is search results. Specifically, the server analyzes the text "What's the weather like tomorrow?", generates the search query "Tomorrow's weather," and sends it to Google's search API to retrieve search results.

[0719] Step 14:

[0720] The server analyzes the search results, organizes them as natural language text, and sends a text reply to the user. The input is the search results, and the output is natural language text. Specifically, the server generates the text "The weather in Tokyo will be cloudy tomorrow," and sends it back to the user via chat.

[0721] Step 15:

[0722] The server periodically sends the audio and text question history to the summary generation means, which then automatically generates a summary. The input is the question history, and the output is the summary. Specifically, the server sends the search history from the past 24 hours to the summary generation means, and generates a summary such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0723] Step 16:

[0724] The server notifies the user of the generated summary via a notification means. The input is the summary, and the output is a notification to the user. Specifically, the server notifies the user of the generated summary via email or the notification function of their smartphone.

[0725] (Application example 1)

[0726] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0727] In order to enable users to easily order and search for menus using voice or text, food delivery services must simplify conventional operations and provide a more efficient and intuitive interface. They also need a system that can personalize order and search histories to make optimal suggestions to users.

[0728] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0729] In this invention, the server includes voice input means for a user to input orders or menu searches by voice, voice recognition means for converting voice data into text data, search means for performing an internet search and menu information search based on the text data, voice response means for converting search results into voice data, voice output means for outputting the voice data to the user, character input means for inputting orders or menu searches using character input means, search means for performing an internet search and menu information search based on the character input, character response means for providing a method for returning search results of the character input to the user in text, summary generation means for generating a summary of the voice and character input history, and notification means for notifying the user of the generated summary results. This allows users to efficiently and intuitively order or search by voice or character input and receive personalized suggestions based on their past history.

[0730] "Voice input means" refers to a means by which a user can input information using voice.

[0731] The "voice recognition means" is a means for converting voice data into text data.

[0732] The "search means" is a means for searching the Internet or menu information based on text data.

[0733] The "voice response means" is a means for converting search results into voice data and providing it to the user.

[0734] The "audio output means" is a means for outputting audio data to the user.

[0735] "Character input means" refers to a means by which a user inputs information using characters.

[0736] The "text response means" is a means for returning search results based on a text question to a user in text form.

[0737] The "summary generation means" is a means for generating summaries of the audio and text question history.

[0738] The "notification means" is a means for notifying the user of the results of the generated summary.

[0739] A "personalized database" is a database that stores a user's input history and provides personalized information.

[0740] The present invention provides a system that allows efficient and intuitive operation of a food delivery service. Specific embodiments for realizing this system are described below.

[0741] System configuration

[0742] The system consists of a user, a device, and a server. Users use devices such as smartphones and smart speakers to place orders and search for menus by voice or text. The device accepts both voice and text input.

[0743] Hardware and software used

[0744] Hardware: smartphones, smart speakers, microphones

[0745] Software: Speech recognition engine (e.g., Google Cloud Speech-to-Text API), speech synthesis engine (e.g., Google Text-to-Speech API)

[0746] Data processing and calculation

[0747] Voice input means

[0748] When a user orders or searches the menu by voice, the device captures the voice data, which is then sent to a cloud service where it is converted into text using a voice recognition engine.

[0749] Voice recognition means

[0750] The voice recognition engine in the cloud service converts voice data into text data. For example, a voice saying "I want to order a pizza" is converted into text "I want to order a pizza."

[0751] Search methods

[0752] The server analyzes the converted text data and generates appropriate queries to search the Internet and retrieve menu information, and the search results are retrieved on the server side.

[0753] Voice response means

[0754] The server analyzes the search results and selects the most relevant information, which is then generated as natural language text and converted into voice data using a speech synthesis engine.

[0755] Audio output means

[0756] The device plays audio data to the user and provides information such as search results and order confirmations.

[0757] Character input means and character response means

[0758] Users can also enter text on their smartphone screens. This text data is sent to the server, which uses search tools to generate appropriate search queries. The search results are organized as natural language text and provided to users in text format.

[0759] Summary generation means and notification means

[0760] The server transmits the user's voice and text input history to a summary generation means, which generates a summary such as "Today's order: pizza, salad" and notifies the user via a notification means.

[0761] Specific examples

[0762] For example, if a user says, "I'd like to order sushi," the device captures this speech and sends it to the cloud service. The speech recognition engine converts it into text data, and the server analyzes the text data to generate an appropriate search query. The search result, "The recommended sushi is tuna nigiri," is converted into audio data by the speech synthesis engine and played on the device.

[0763] Examples of prompts used by generative AI models include:

[0764] The user says, "I want to order sushi." Generate an appropriate response.

[0765] In this way, the present invention provides a system that allows users to efficiently and intuitively order and search using voice or text input.

[0766] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0767] Step 1:

[0768] Users can input orders and menu searches by voice. Users can input "I want to order pizza" into a device such as a smartphone or smart speaker.

[0769] Step 2:

[0770] The device captures the user's voice and generates audio data, which is then sent to a cloud service.

[0771] Step 3:

[0772] The server uses the cloud service's voice recognition engine to convert the voice data into text data. For example, the voice saying "I'd like to order a pizza" is converted into text data saying "I'd like to order a pizza."

[0773] Step 4:

[0774] The server analyzes the converted text data and generates an appropriate search query. Specifically, based on the input text data "I want to order pizza," it generates the query "pizza menu" to search for menu information.

[0775] Step 5:

[0776] The server performs an internet search and menu information search based on the generated search query, and obtains the relevant menu information using the search means.

[0777] Step 6:

[0778] The server analyzes the search results and selects the most relevant information, for example, "The recommended pizza is Margherita."

[0779] Step 7:

[0780] The server generates the selected information as natural language text and converts it into voice data using a speech synthesis engine. This is the process of converting the text "The recommended pizza is Margherita" into voice data.

[0781] Step 8:

[0782] The terminal plays back the voice data and provides the search results to the user. By playing back the voice data "The recommended pizza is Margherita" to the user, the terminal confirms and suggests the order.

[0783] Step 9:

[0784] When a user inputs a text question on the screen of a smartphone, the user uses the text input means to place an order or search for a menu item. The input text data is sent to the server.

[0785] Step 10:

[0786] The server performs an internet search and menu information search based on the received text data. For example, if the text data "I want to order sushi" is received, the server generates the search query "sushi menu."

[0787] Step 11:

[0788] The server retrieves the search results, organizes them as natural language text, and returns the text to the user, providing the user with the text, "The recommended sushi is tuna nigiri."

[0789] Step 12:

[0790] The server transmits the input history of voice and text to the summary generation means, which analyzes the input history and automatically generates a summary.

[0791] Step 13:

[0792] The server notifies the user of the generated summary via the notification means. For example, the server notifies the user of the summary "Today's order: pizza, sushi."

[0793] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0794] The voice response and chat-linked search system of the present invention includes a voice input means installed in a smart device, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the converted text data, and a voice response means for converting the search results into voice data and providing it to the user. The user can also use the smartphone's input means to submit text questions, and an Internet search can be performed based on the text input, with the text response means providing the results to the user in text. The voice and text question history is summarized by a summary generation means and notified to the user by a notification means. Furthermore, the system is equipped with an emotion engine that recognizes emotions from the user's voice data, and can adjust the response content based on the emotional information.

[0795] Natural language description of the program

[0796] 1. Initial Setup:

[0797] The server sets up a personalized database containing the user's account information and search history, as well as the search API, message API, and emotion engine.

[0798] 2. Audio question reception:

[0799] The device goes into standby mode, waiting for the user to say the trigger phrase.

[0800] 3. Voice input:

[0801] When the user says a trigger phrase (e.g., "OK, Assistant"), the device switches to voice input mode and records the user's question.

[0802] 4. Speech Recognition:

[0803] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine in the cloud to convert the voice data into text data.

[0804] 5. Information Search:

[0805] The server analyzes the converted text data, generates an appropriate search query, and sends the generated search query to the search API to retrieve search results.

[0806] 6. Emotion recognition:

[0807] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[0808] 7. Tailor your response:

[0809] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[0810] 8. Voice response generation:

[0811] The server analyzes the retrieved search results and emotion information, selects the most relevant information, and generates a natural language voice response text. The natural language text is converted into voice data using the voice response means and sent to the terminal.

[0812] 9. Audio Answer:

[0813] The terminal plays back the received voice data and provides the user with a voice response.

[0814] 10. Chat questions:

[0815] The user types in a text question on their smartphone and sends it.

[0816] 11. Handling chat questions:

[0817] The server receives text questions through the message API, analyzes the received text questions, generates appropriate search queries using the search API, retrieves search results, analyzes the retrieved search results, organizes them as natural language text, and sends a text reply to the user.

[0818] 12. Question History Summary:

[0819] The server sends the audio and text question history to the summary generation means, which automatically generates a summary. The summary is then notified to the user via the notification means. The server also adjusts the summary content and notification format based on the emotion information.

[0820] Specific examples

[0821] Examples of voice questions:

[0822] User: "OK, Assistant, what's the weather like right now?"

[0823] 1. The device captures the audio data and sends it to the server.

[0824] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[0825] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[0826] 4. The server simultaneously sends the voice data to the emotion engine to recognize the user's emotion.

[0827] 5. The server generates a response text based on the emotion information: "Currently, the weather in Tokyo is sunny."

[0828] 6. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[0829] 7. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[0830] Examples of chat questions:

[0831] User (asking in chat app): "What's the weather like tomorrow?"

[0832] 1. The server receives a text question through the message API.

[0833] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[0834] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[0835] Example of a question history summary:

[0836] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[0837] 1. The server sends these search histories to the summary generator.

[0838] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0839] 3. The server notifies the user of the generated summary via a notification method, and adjusts the summary content and notification format if the user's emotions are recognized.

[0840] The above is a specific embodiment of the system of the present invention that combines an emotion engine. This system allows users to obtain information efficiently and in a way that takes emotion into consideration, and also allows users to check their past question history all at once.

[0841] The processing flow will be explained below.

[0842] Step 1:

[0843] The server sets up a personalized database containing the user's account information and search history, as well as the search API, message API, and emotion engine integration.

[0844] Step 2:

[0845] The device enters standby mode and prepares to switch to voice input mode when the user speaks the trigger phrase.

[0846] Step 3:

[0847] The user says the trigger phrase (e.g., "OK, Assistant").

[0848] Step 4:

[0849] The device detects the user's trigger phrase, switches to voice input mode, and records the user's question.

[0850] Step 5:

[0851] The device sends the recorded audio data to a cloud service.

[0852] Step 6:

[0853] The server uses a voice recognition engine in the cloud to convert the voice data into text data.

[0854] Step 7:

[0855] The server analyzes the converted text data and generates an appropriate search query.

[0856] Step 8:

[0857] The server sends the generated search query to the search API and retrieves the search results.

[0858] Step 9:

[0859] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[0860] Step 10:

[0861] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[0862] Step 11:

[0863] The server generates a voice response text in a natural language based on the adjusted response content.

[0864] Step 12:

[0865] The server uses a text-to-speech engine to convert the natural language text into audio data.

[0866] Step 13:

[0867] The server transmits the generated voice data to the terminal.

[0868] Step 14:

[0869] The terminal plays back the received voice data and provides the user with a voice response.

[0870] Step 15:

[0871] A user sends a text question via a chat app on their smartphone.

[0872] Step 16:

[0873] The server receives text questions through a message API.

[0874] Step 17:

[0875] The server analyzes the received text query and generates an appropriate search query.

[0876] Step 18:

[0877] The server sends the generated search query to the search API and retrieves the search results.

[0878] Step 19:

[0879] The server analyzes the search results and selects the most relevant information.

[0880] Step 20:

[0881] The server generates the selected information as natural language text and replies to the user via chat.

[0882] Step 21:

[0883] The server sends the audio and text question history to the summary generator.

[0884] Step 22:

[0885] The summary generator summarizes the audio and text question history and automatically generates a summary.

[0886] Step 23:

[0887] The summary generator sends the summary to the notification unit. If emotion information is present, the summary content and notification format are adjusted based on that information.

[0888] Step 24:

[0889] The server notifies the user of the generated summary via the notification means.

[0890] In this way, the system of the present invention can efficiently process voice or text questions from users and provide answers that take the user's feelings into consideration. In addition, by summarizing and notifying the user of the question history, the system can further facilitate the user's information acquisition.

[0891] Example 2

[0892] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0893] Conventional search systems that use voice and text input are unable to consider user sentiment when providing search results, limiting their ability to provide optimal search results. Furthermore, they do not effectively utilize voice and text question history, and lack a means to easily review past questions. This makes it difficult for users to quickly and accurately obtain the information they need.

[0894] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0895] In this invention, the server includes a summary generation unit that generates summaries of the voice and text question history, a response content adjustment unit that adjusts the response content of search results based on emotion information, and an emotion engine that performs emotion recognition. This makes it possible to provide appropriate and personalized search results while taking into consideration the user's emotions. In addition, by automatically summarizing the past question history and notifying the user, efficient information management and confirmation is possible.

[0896] "Voice input means" refers to a device or interface that allows a user to input a question by voice.

[0897] "Speech recognition means" refers to a technology or module for converting input voice data into text data.

[0898] The "search means" is a technology or module for conducting an internet search based on the converted text data and obtaining appropriate search results.

[0899] "Voice response means" refers to a technology or module for converting search results into voice data and providing it to the user.

[0900] "Audio output means" refers to a device or interface for outputting audio data to a user and conveying information.

[0901] "Character input means" refers to a device or interface that allows a user to input characters.

[0902] A "text response means" is a technology or module for conducting a search based on character input and returning search results to the user in text.

[0903] The "summary generation means" is a technology or module for summarizing the audio and text question history and automatically generating a summary.

[0904] "Notification means" refers to a technology or module for notifying the user of the summary generated results.

[0905] An "emotion engine" is a technology or module for recognizing emotions from a user's voice data.

[0906] The "response content adjustment means" is a technology or module for adjusting the response content of search results based on the emotion information obtained from the emotion engine.

[0907] A "personalized database" is a database that stores each user's account information and search history and provides personalized search results.

[0908] A "cloud service" is a service that uses remote computing resources and storage provided via the Internet.

[0909] The voice response and chat-linked search system of the present invention is composed of multiple means installed in a smart device, which processes users' voice and text questions, provides appropriate responses, and provides more personalized services through emotion recognition.

[0910] This system includes a voice input means for the user to input a question by voice, a voice recognition means for converting voice data into text data, a search means for performing an internet search based on the text data, a voice response means for converting search results into voice data, a voice output means for outputting the voice data to the user, a text input means for inputting a text question, a text response means for returning search results to the user in text, a summary generation means for generating a summary of the voice and text question history, and a notification means for notifying the user of the generated summary. Furthermore, it includes an emotion engine for recognizing emotions from the user's voice data, and a response content adjustment means for adjusting the response content of the search results based on the emotion information. Each part of this system can be implemented using an ordinary smartphone or a cloud server.

[0911] The smartphone's microphone is used as the voice input means. When the user's voice input is recognized as a trigger phrase (e.g., "OK, Assistant"), the voice recognition means sends the voice to a cloud service, where it is converted into text data by a voice recognition engine (e.g., Google Cloud Speech-to-Text API). The search means then searches the converted text data using an internet search API (e.g., Google Custom Search API).

[0912] The search results are converted by the voice response means and notified to the user by voice through the voice output means. In this series of steps, an emotion recognition engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotional state, and the response content adjustment means determines an appropriate response based on this.

[0913] The smartphone keyboard is used as the text input means. The user types and sends a question in text. For example, if the user types "What's the weather going to be like tomorrow?", the question is sent to the server via a messaging API (e.g., Twilio API), which then performs an internet search. The retrieved search results are returned to the user in text via the text response means.

[0914] The question history is stored in a question history database. The summary generation means summarizes this history and notifies the user as necessary. For example, if a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?", the summary generation means generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy" and notifies the user via the notification means.

[0915] Specific examples

[0916] Examples of voice questions:

[0917] User: "OK, Assistant, what's the weather like right now?"

[0918] 1. The device captures the audio data and sends it to the server.

[0919] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[0920] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[0921] 4. The server simultaneously sends the voice data to the emotion engine to recognize the user's emotion.

[0922] 5. The server generates a response text based on the emotion information: "Currently, the weather in Tokyo is sunny."

[0923] 6. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[0924] 7. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[0925] Examples of chat questions:

[0926] User (asking in chat app): "What's the weather like tomorrow?"

[0927] 1. The server receives a text question through the message API.

[0928] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[0929] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[0930] Example of a question history summary:

[0931] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[0932] 1. The server sends these search histories to the summary generator.

[0933] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[0934] 3. The server notifies the user of the generated summary via a notification means.

[0935] The system of the present invention allows users to obtain information efficiently and sensitively, and also allows them to review their past question history in one place, significantly improving user convenience and enabling more personalized services.

[0936] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0937] Step 1: Initial Setup

[0938] The server sets up a personalized database that stores user account information and past search history, and also establishes integration with the search API, message API, and emotion engine.

[0939] Input: User information, past search history

[0940] Output: Personalized database, API and engine integration settings

[0941] Step 2: Voice inquiry reception

[0942] The device will enter standby mode and wait for the user to say a trigger phrase, such as "OK, Assistant."

[0943] Input: Trigger phrase

[0944] Output: Activate voice input mode

[0945] Step 3: Voice Input

[0946] When the user utters the trigger phrase, the device switches to voice input mode and records the user's question, for example, "What's the weather like today?"

[0947] Input: User's voice question

[0948] Output: Recorded audio data

[0949] Step 4: Voice Recognition

[0950] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine to convert the voice data into text data.

[0951] Input: Recorded audio data

[0952] Output: Converted text data

[0953] Step 5: Information Search

[0954] The server analyzes the converted text data and generates an appropriate search query, which it then sends to an Internet search API to retrieve search results.

[0955] Input: Text data

[0956] Output: Search result data

[0957] Step 6: Emotion Recognition

[0958] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[0959] Input: Audio data

[0960] Output: Emotional information

[0961] Step 7: Tailor your response

[0962] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[0963] Input: Emotion information, search result data

[0964] Output: Adjusted response content

[0965] Step 8: Generate voice response

[0966] The server integrates the search results and emotional information to generate a natural language voice response text, converts the text into voice data, and sends it to the device.

[0967] Input: Adjusted response content

[0968] Output: Audio data

[0969] Step 9: Answer by voice

[0970] The terminal plays back the received voice data and provides the user with a voice response.

[0971] Input: Audio data

[0972] Output: A spoken response to the user

[0973] Step 10: Chat with your questions

[0974] A user types a text question into a smartphone and sends it, for example, "What's the weather like tomorrow?" in a chat app.

[0975] Input: User's character question

[0976] output: Character data to be sent

[0977] Step 11: Handling chat questions

[0978] The server receives text questions through the message API, generates search queries based on the questions using the search API, retrieves search results, organizes the results into natural language text, and sends a text reply to the user.

[0979] Input: User's character question

[0980] Output: Search result text

[0981] Step 12: Question History Summary

[0982] The server sends the audio and text question history to the summary generation means, which automatically generates a summary. The summary is then notified to the user via the notification means. The server also adjusts the summary content and notification format based on the emotion information.

[0983] Input: Voice question history, text question history

[0984] Output: Summary, Notification

[0985] (Application example 2)

[0986] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0987] In today's world, it's important to provide users with a way to quickly and efficiently obtain the information they need. However, traditional voice response and chat-based search systems are unable to adjust responses based on the user's emotional state, which can result in a poor user experience. Furthermore, they are unable to summarize multiple query histories and notify users, making it difficult to easily review past search results. Another problem is the inability to provide personalized results based on voice or text queries.

[0988] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice input means for a user to input a question by voice, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the text data, a voice response means for converting the search results into voice data, a voice output means for outputting the voice data to the user, a character input means for inputting a text question, a search means for performing an Internet search based on the text question, a text response means for providing the user with a method for replying to the search results of the text question in text form, a summary generation means for generating a summary of the voice and text question history, a notification means for notifying the user of the generated summary results, an emotion recognition means for acquiring emotion information, and a response adjustment means for adapting the search results based on the emotion information. This makes it possible to provide personalized search results according to the user's emotions, improve the quality of the user experience, and make it easy to check past question history.

[0989] The "voice input means" is a means for the user to input a question by voice.

[0990] The "voice recognition means" is a means for converting voice data into text data.

[0991] "Search means" refers to a means for conducting an Internet search based on text data.

[0992] The "voice response means" is a means for converting search results into voice data.

[0993] The "audio output means" is a means for outputting audio data to the user.

[0994] The "text input means" is a means for the user to input a text question.

[0995] The "text response means" is a means for returning search results for a text question to the user in text form.

[0996] The "summary generation means" is a means for generating summaries of the audio and text question history.

[0997] The "notification means" is a means for notifying the user of the results of the generated summary.

[0998] "Emotion recognition means" is a means for acquiring emotion information.

[0999] A "response adjustment means" is a means for adapting search results based on emotional information.

[1000] A "personalized database" is a database that stores a user's question history and customizes search results individually.

[1001] A "cloud service" is a service that processes voice data and text data on an online server and performs voice recognition and search.

[1002] To implement this invention, the following hardware and software are primarily used. A user uses a smart device (e.g., a smartphone) to ask a question by voice or text and obtain search results corresponding to the question. The following describes the functions and processing flow of the system that realizes this application example.

[1003] Required Hardware and Software

[1004] 1. Smart devices (e.g. smartphones, tablets): Provide an interface with users. Devices that allow voice input, voice output, and text input.

[1005] 2. Cloud Services: Online platforms for processing voice and text data. They perform speech recognition, search, and emotion recognition. They use services such as Google Speech-to-Text API, IBM Watson Tone Analyzer, and Amazon Polly.

[1006] 3. Server: Manages the user's personalized database and generates search results. Provides tailored responses based on the user's question history and sentiment information.

[1007] System Operation Overview

[1008] 1. Audio question reception:

[1009] The user speaks a voice trigger phrase (e.g., "OK, Assistant") into the smart device. When the device switches to voice input mode, the user asks, "What movies are recommended today?"

[1010] 2. Speech and Emotion Recognition:

[1011] The smart device sends the voice data to a cloud service, which uses the Google Speech-to-Text API to convert the voice data into text, and simultaneously sends the voice data to the IBM Watson Tone Analyzer to recognize the user's emotions.

[1012] 3. Information Search:

[1013] The server generates a search query based on the converted text data and sends it to the appropriate search API (e.g., a movie database). Based on the user's emotional information, the server can recommend movies.

[1014] 4. Tailor your response:

[1015] The server selects the most appropriate information from the search results based on emotional information and generates a response in natural language. For example, if the user feels like relaxing, it generates a response such as, "The recommended relaxing movie is 'Inception.'"

[1016] 5. Voice response generation:

[1017] The generated response is converted into voice data using Amazon Polly and sent to the smart device, which then outputs the converted voice data to the user and provides a response.

[1018] 6. Question History Summary:

[1019] The server summarizes the user's past question history and periodically notifies the user of the summarized information. The summary information is presented to the user in the form of, for example, "Past 7 days' viewing history: 'Inception', 'The Matrix'."

[1020] Examples and prompts

[1021] Specific examples

[1022] When a user says "What are the best movies tonight?", the following happens:

[1023] 1. The voice data is sent to a cloud service and converted into text data by a speech recognition engine. A query such as "Today's recommended movies" is generated.

[1024] 2. The emotion recognition engine detects the user's desire to relax from their voice.

[1025] 3. The server searches a movie database and applies a filter that recommends "relaxing movies."

[1026] 4. Generate a response sentence: "A recommended relaxing movie is 'Inception'." and convert it into speech.

[1027] 5. The smart device plays this audio to the user.

[1028] Example prompts

[1029] Here are some example prompts for using generative AI models:

[1030] User asks: "What movie do you recommend today?"

[1031] Hardware: Smartphone

[1032] Software Settings:

[1033] Speech Recognition: Google Speech-to-Text API

[1034] Emotion recognition: IBM Watson Tone Analyzer

[1035] Speech synthesis: Amazon Polly

[1036] Specific processing:

[1037] 1. The voice data is sent to the speech recognition engine and converted into text.

[1038] 2. Send the text data to the movie database API and get the search results.

[1039] 3. Send the voice data to the emotion engine and obtain the emotion information.

[1040] 4. Based on the emotional information, the system selects an appropriate movie from the search results and generates a response in natural language.

[1041] 5. The response is converted into audio data and played on the smartphone.

[1042] Expected response: "A movie recommendation to fit your current mood is 'Inception'."

[1043] The system provided by this invention allows users to obtain personalized search results based on their emotions and easily check their past question history. This system improves the quality of the user experience and provides a means to efficiently obtain information.

[1044] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1045] Step 1:

[1046] Voice inquiry reception

[1047] The device enters standby mode and waits for the user to speak a trigger phrase. The input is a voice trigger phrase spoken by the user (e.g., "OK, Assistant"), and the device captures the voice and switches to voice input mode to accept the user's question.

[1048] Step 2:

[1049] Voice input

[1050] The user asks the device, "What movies do you recommend today?" The device records the user's voice and sends the voice data to the cloud service. The input is the recorded voice data, and the output is the voice data sent to the cloud service.

[1051] Step 3:

[1052] Voice Recognition

[1053] The cloud service converts the transmitted voice data into text data using the Google Speech-to-Text API. The input is the voice data, and the output is the text data "What movies do you recommend today?" The cloud service analyzes the voice data and generates the corresponding text.

[1054] Step 4:

[1055] emotion recognition

[1056] The cloud service sends the voice data to the IBM Watson Tone Analyzer along with the text data to recognize the user's emotions. The input is the voice data, and the output is the user's emotional information. The cloud service analyzes the voice data and recognizes the user's emotional state.

[1057] Step 5:

[1058] Information Search

[1059] The server analyzes the text data obtained from speech recognition and generates an appropriate search query. For example, it sends a query such as "Today's recommended movies" to a movie database API and retrieves relevant search results. The input is the converted text data, and the output is a list of search results.

[1060] Step 6:

[1061] Tailoring response content

[1062] The server determines the optimal response based on the search results and emotional information, and generates a response in natural language. For example, if the user is in the mood to relax, it generates the text "A recommended relaxing movie is 'Inception.'" The input is the search results and emotional information, and the output is a response in natural language.

[1063] Step 7:

[1064] Voice Response Generation

[1065] The server converts the generated response text into voice data using Amazon Polly and sends it to the device. The input is text data and the output is voice data. The server analyzes the response text and generates voice data.

[1066] Step 8:

[1067] Answer by voice

[1068] The device plays the received voice data and provides a voice response to the user. The input is the voice data sent from the server, and the output is the voice played from the device. The device analyzes the voice data and conveys information to the user by voice.

[1069] Step 9:

[1070] Question History Summary

[1071] The server sends the user's past question history to the summary generation means, which then summarizes the question history. For example, it generates a summary result such as "Past 7 days' viewing history: 'Inception', 'The Matrix'." The input is the past question history data, and the output is summarized text data. The server analyzes the question history and generates summary information.

[1072] Step 10:

[1073] Summary Notification

[1074] The server notifies the user of the generated summarized question history via a notification means. The input is the summarized text data, and the output is notification information for the user. The server analyzes the summarized data and notifies the user in an appropriate format.

[1075] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1076] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1077] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1078] [Third embodiment]

[1079] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1080] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1081] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1082] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1083] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1084] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1085] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1086] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1087] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1088] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1089] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1090] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1091] The voice response and chat-linked search system of the present invention includes a voice input means installed in a smart device, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the converted text data, and a voice response means for converting the search results into voice data and providing it to the user. The user can also use the input means of the smartphone to submit text questions, and an Internet search can be performed based on the text input, with the text response means providing the results to the user in text. The voice and text question history is summarized by a summary generation means and notified to the user by a notification means.

[1092] Natural language description of the program

[1093] 1. Initial Setup:

[1094] The server stores the user's account information and search history in a personalized database, and also configures the search API and message API integration.

[1095] 2. Audio question reception:

[1096] The device is in standby mode and starts voice input when the user speaks a trigger phrase. The user then speaks a question.

[1097] 3. Speech Recognition:

[1098] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine in the cloud to convert the voice data into text data.

[1099] 4. Information Search:

[1100] The server analyzes the converted text data, generates an appropriate search query, and sends the generated search query to the search API to retrieve search results.

[1101] 5. Voice response generation:

[1102] The server analyzes the search results, selects the most relevant information, and generates natural language text based on the search results.The natural language text is converted into voice data using a voice response means and sent to the terminal.

[1103] 6. Audio Answer:

[1104] The terminal plays back the received voice data and provides the user with a voice response.

[1105] 7. Chat questions:

[1106] The user types in a text question on their smartphone and sends it.

[1107] 8. Handling chat questions:

[1108] The server analyzes the received text question, generates an appropriate search query using a search API, retrieves search results, analyzes the retrieved search results, organizes them as natural language text, and sends a text reply to the user.

[1109] 9. Question History Summary:

[1110] The server periodically transmits the question history of voice and text to the summary generating means, which automatically generates a summary, and notifies the user of the summary via the notification means.

[1111] Specific examples

[1112] Examples of voice questions:

[1113] User: "OK, Assistant, what's the weather like right now?"

[1114] 1. The device captures the audio data and sends it to the server.

[1115] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[1116] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[1117] 4. The server generates the answer text "Currently, the weather in Tokyo is sunny."

[1118] 5. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[1119] 6. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[1120] Examples of chat questions:

[1121] User (asking in chat app): "What's the weather like tomorrow?"

[1122] 1. The server receives this request through the chat API.

[1123] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[1124] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[1125] Example of a question history summary:

[1126] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[1127] 1. The server sends these search histories to the summary generator.

[1128] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1129] 3. The server notifies the user of the generated summary via a notification means.

[1130] The above is an embodiment of the invention and its specific examples. This system allows users to efficiently obtain information and easily refer to their past search history.

[1131] The processing flow will be explained below.

[1132] Step 1:

[1133] The server sets up a personalized database containing the user's account information and search history, and also sets up integration between the search API and the messaging API.

[1134] Step 2:

[1135] The device goes into standby mode, waiting for the user to say the trigger phrase.

[1136] Step 3:

[1137] When the user says a trigger phrase (e.g., "OK, Assistant"), the device switches to voice input mode and records the user's question.

[1138] Step 4:

[1139] The device sends the recorded voice data to a cloud service, where it is processed by a voice recognition engine.

[1140] Step 5:

[1141] The server uses a voice recognition engine in the cloud to convert the voice data into text data.

[1142] Step 6:

[1143] The server analyzes the converted text data and generates an appropriate search query.

[1144] Step 7:

[1145] The server sends the generated search query to the search API and retrieves the search results.

[1146] Step 8:

[1147] The server analyzes the search results and selects the most relevant information.

[1148] Step 9:

[1149] The server generates a voice response text in a natural language based on the selected information.

[1150] Step 10:

[1151] The server uses a text-to-speech engine to convert the natural language text into audio data.

[1152] Step 11:

[1153] The server transmits the generated voice data to the terminal.

[1154] Step 12:

[1155] The terminal plays back the received voice data and provides the user with a voice response.

[1156] Step 13:

[1157] A user sends a text question via a chat app on their smartphone.

[1158] Step 14:

[1159] The server receives text questions through a message API.

[1160] Step 15:

[1161] The server analyzes the received text query and generates an appropriate search query.

[1162] Step 16:

[1163] The server sends the generated search query to the search API and retrieves the search results.

[1164] Step 17:

[1165] The server analyzes the search results and selects the most relevant information.

[1166] Step 18:

[1167] The server generates the selected information as natural language text and returns the text to the user.

[1168] Step 19:

[1169] The server sends the audio and text question history to the summary generator.

[1170] Step 20:

[1171] The server analyzes the summary received from the generation AI and notifies the user of the summary via a notification means.

[1172] This is the specific processing flow of the entire system of the present invention. This system allows users to efficiently obtain information while also checking past question history all at once.

[1173] Example 1

[1174] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1175] In today's information search systems using smart devices and cloud services, users need a way to efficiently ask questions by voice or text and quickly obtain results. However, existing systems often lack the accuracy of voice recognition and personalized search results, hindering the user experience. Furthermore, there are no established methods for effectively utilizing search history or properly notifying search results. Furthermore, relying on cloud services to process voice and text data can make integration complex.

[1176] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1177] In this invention, the server includes a voice input means for a user to input a question by voice, a voice recognition means for converting the voice data into text data, a search means for performing an information search based on the text data, a voice response means for converting the search results into voice data, a voice output means for outputting the voice data to the user, a character input means for inputting a text question, a search means for performing an information search based on the text question, a text response means for returning the search results of the text question to the user in text, a summary generation means for generating a summary of the voice and text question history, a notification means for notifying the user of the generated summary, and a means for transmitting the voice and text data to a cloud service and performing voice recognition and search using the cloud service. This enables a user to efficiently ask questions by voice and text, quickly obtain personalized search results, effectively utilize their search history, and receive appropriate notification of the results.

[1178] A "user" is an entity that uses the system to enter questions and obtain information.

[1179] "Voice input means" refers to a device or function that allows a user to input a question by voice.

[1180] "Speech recognition means" refers to the technology or engine that converts input voice data into text data.

[1181] "Search means" refers to a device or function for searching for information based on text data.

[1182] "Voice response means" refers to a device or technology that converts text data obtained as a search result into voice data.

[1183] "Audio output means" refers to a device or function for reproducing audio data to the user.

[1184] "Character input means" refers to a device or function that allows a user to input a question in text.

[1185] A "text response means" is a device or technology that responds to a user in text with the search results of a text question.

[1186] A "summary generator" is a device or technique that summarizes the audio and text question history and generates summarized information.

[1187] The "notification means" refers to a device or function for notifying the user of the results of the generated summary.

[1188] A "personalized database" is a database that stores each user's search history and account information and responds to them individually.

[1189] "Cloud services" are remote computing resources used to perform voice recognition and search processing.

[1190] The voice response and chat-linked search system according to the present invention is mainly composed of a server, a terminal, and a user. The specific processing of these components and the hardware and software used will be described below.

[1191] Server Roles and Configuration

[1192] The server stores the user's account information and search history in a personalization database. The server also configures search and messaging APIs, such as Google's search and messaging APIs (e.g., Twilio). The server then sends the received voice data to a cloud service for speech recognition. For this purpose, the Google Cloud Speech-to-Text API is used.

[1193] Terminal roles and configuration

[1194] The device accepts voice and text input from the user. When voice input is performed, the device recognizes trigger phrases and captures voice data. The recorded voice data is sent to a server for speech recognition. The device has the ability to generate voice data using the Google Cloud Text-to-Speech API and play it back to the user.

[1195] User Actions

[1196] The user inputs a question by voice or text. For example, if the user asks by voice, "OK, Assistant, what's the weather like today?", the device captures this and sends it to the server. The server performs speech recognition processing and converts it into text data. The server then calls the search API based on the query "What's the weather like today?" and retrieves search results. The server generates a text answer, "Currently, the weather in Tokyo is sunny," converts this into audio data, and sends it to the device. The device then plays back the audio, "Currently, the weather in Tokyo is sunny."

[1197] Use of cloud services

[1198] The voice and text data are sent to a cloud service, where voice recognition and information search are performed within the cloud. This enables highly accurate voice recognition and rapid search results. The server also periodically sends the voice and text question history to a summary generator, which automatically generates a summary. This summary is then notified to the user via a notification device.

[1199] Specific examples

[1200] Examples of voice questions:

[1201] User: "OK, Assistant, what's the weather like right now?"

[1202] 1. The device captures the audio data and sends it to the server.

[1203] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[1204] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[1205] 4. The server generates the answer text "Currently, the weather in Tokyo is sunny."

[1206] 5. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[1207] 6. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[1208] Examples of chat questions:

[1209] User (asking in chat app): "What's the weather like tomorrow?"

[1210] 1. The server receives this request through the chat API.

[1211] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[1212] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[1213] Example of a question history summary:

[1214] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[1215] 1. The server sends these search histories to the summary generator.

[1216] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1217] 3. The server notifies the user of the generated summary via a notification means.

[1218] This system configuration allows users to efficiently retrieve information and easily refer to past search history.The system provides an advanced user experience by combining a voice recognition engine, search API, and text processing engine.

[1219] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1220] Step 1:

[1221] The server saves the user's account information and search history in a personalized database. This is done when a new account is registered. The input is account information and search history, and the output is saved in the personalized database. Specifically, the user's name, email address, past search history, etc. are stored in the database.

[1222] Step 2:

[1223] The server configures the search API and message API integration. The input is authentication information, and the output is completion of API integration. Specifically, it obtains authentication tokens for Google's search API and general message APIs and includes them in the server configuration.

[1224] Step 3:

[1225] The device is in standby mode, and when the user utters a trigger phrase, voice input begins. The input is a trigger phrase such as "OK, Assistant," and the output is a transition to voice input mode. Specifically, the device switches from standby mode to voice input mode and becomes ready to record the user's question.

[1226] Step 4:

[1227] The user inputs a question by voice. The input is a voice question such as "What's the weather like now?", and the output is recorded voice data. Specifically, the user says "What's the weather like now?", and the voice is recorded by the device.

[1228] Step 5:

[1229] The device sends the recorded voice data to a cloud service. The input is the recorded voice data, and the output is the transmission of the voice data to the cloud service. Specifically, the recorded data is sent to a server over the Internet (e.g., Google Cloud Speech-to-Text).

[1230] Step 6:

[1231] The server uses a speech recognition engine in the cloud to convert the voice data into text data. The input is the voice data received from the cloud service, and the output is the converted text data. Specifically, the Google Cloud Speech-to-Text API is used to convert the voice data into text data such as "What's the weather like today?"

[1232] Step 7:

[1233] The server analyzes the converted text data and generates an appropriate search query. The input is text data and the output is a search query. Specifically, it analyzes the text "What's the weather like now?" to generate a search query such as "Current weather."

[1234] Step 8:

[1235] The server sends the generated search query to the search API and retrieves the search results. The input is the search query and the output is the search results. Specifically, the server sends the query "current weather" to Google's search API and receives the search results.

[1236] Step 9:

[1237] The server analyzes the search results, selects the most relevant information, and generates natural language text. The input is the search results, and the output is the natural language text. Specifically, the generated text is, "Currently, the weather in Tokyo is sunny."

[1238] Step 10:

[1239] The server uses a voice response means to convert the generated natural language text into voice data and send it to the terminal. The input is natural language text and the output is voice data. Specifically, the Google Cloud Text-to-Speech API is used to generate the voice "Currently, the weather in Tokyo is sunny," and this is sent to the terminal.

[1240] Step 11:

[1241] The device plays the received voice data and provides the user with a voice response. The input is the voice data, and the output is the voice that is played back. Specifically, the device plays back the voice, "Currently, the weather in Tokyo is sunny."

[1242] Step 12:

[1243] The user inputs a text question into their smartphone and sends it. The input is a text question such as "What's the weather like tomorrow?", and the output is the text data sent.

[1244] Step 13:

[1245] The server analyzes the received text question, generates an appropriate search query using the search API, and retrieves search results. The input is text data, and the output is search results. Specifically, the server analyzes the text "What's the weather like tomorrow?", generates the search query "Tomorrow's weather," and sends it to Google's search API to retrieve search results.

[1246] Step 14:

[1247] The server analyzes the search results, organizes them as natural language text, and sends a text reply to the user. The input is the search results, and the output is natural language text. Specifically, the server generates the text "The weather in Tokyo will be cloudy tomorrow," and sends it back to the user via chat.

[1248] Step 15:

[1249] The server periodically sends the audio and text question history to the summary generation means, which then automatically generates a summary. The input is the question history, and the output is the summary. Specifically, the server sends the search history from the past 24 hours to the summary generation means, and generates a summary such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1250] Step 16:

[1251] The server notifies the user of the generated summary via a notification means. The input is the summary, and the output is a notification to the user. Specifically, the server notifies the user of the generated summary via email or the notification function of their smartphone.

[1252] (Application example 1)

[1253] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1254] In order to enable users to easily order and search for menus using voice or text, food delivery services must simplify conventional operations and provide a more efficient and intuitive interface. They also need a system that can personalize order and search histories to make optimal suggestions to users.

[1255] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1256] In this invention, the server includes voice input means for a user to input orders or menu searches by voice, voice recognition means for converting voice data into text data, search means for performing an internet search and menu information search based on the text data, voice response means for converting search results into voice data, voice output means for outputting the voice data to the user, character input means for inputting orders or menu searches using character input means, search means for performing an internet search and menu information search based on the character input, character response means for providing a method for returning search results of the character input to the user in text, summary generation means for generating a summary of the voice and character input history, and notification means for notifying the user of the generated summary results. This allows users to efficiently and intuitively order or search by voice or character input and receive personalized suggestions based on their past history.

[1257] "Voice input means" refers to a means by which a user can input information using voice.

[1258] The "voice recognition means" is a means for converting voice data into text data.

[1259] The "search means" is a means for searching the Internet or menu information based on text data.

[1260] The "voice response means" is a means for converting search results into voice data and providing it to the user.

[1261] The "audio output means" is a means for outputting audio data to the user.

[1262] "Character input means" refers to a means by which a user inputs information using characters.

[1263] The "text response means" is a means for returning search results based on a text question to a user in text form.

[1264] The "summary generation means" is a means for generating summaries of the audio and text question history.

[1265] The "notification means" is a means for notifying the user of the results of the generated summary.

[1266] A "personalized database" is a database that stores a user's input history and provides personalized information.

[1267] The present invention provides a system that allows efficient and intuitive operation of a food delivery service. Specific embodiments for realizing this system are described below.

[1268] System configuration

[1269] The system consists of a user, a device, and a server. Users use devices such as smartphones and smart speakers to place orders and search for menus by voice or text. The device accepts both voice and text input.

[1270] Hardware and software used

[1271] Hardware: smartphones, smart speakers, microphones

[1272] Software: Speech recognition engine (e.g., Google Cloud Speech-to-Text API), speech synthesis engine (e.g., Google Text-to-Speech API)

[1273] Data processing and calculation

[1274] Voice input means

[1275] When a user orders or searches the menu by voice, the device captures the voice data, which is then sent to a cloud service where it is converted into text using a voice recognition engine.

[1276] Voice recognition means

[1277] The voice recognition engine in the cloud service converts voice data into text data. For example, a voice saying "I want to order a pizza" is converted into text "I want to order a pizza."

[1278] Search methods

[1279] The server analyzes the converted text data and generates appropriate queries to search the Internet and retrieve menu information, and the search results are retrieved on the server side.

[1280] Voice response means

[1281] The server analyzes the search results and selects the most relevant information, which is then generated as natural language text and converted into voice data using a speech synthesis engine.

[1282] Audio output means

[1283] The device plays audio data to the user and provides information such as search results and order confirmations.

[1284] Character input means and character response means

[1285] Users can also enter text on their smartphone screens. This text data is sent to the server, which uses search tools to generate appropriate search queries. The search results are organized as natural language text and provided to users in text format.

[1286] Summary generation means and notification means

[1287] The server transmits the user's voice and text input history to a summary generation means, which generates a summary such as "Today's order: pizza, salad" and notifies the user via a notification means.

[1288] Specific examples

[1289] For example, if a user says, "I'd like to order sushi," the device captures this speech and sends it to the cloud service. The speech recognition engine converts it into text data, and the server analyzes the text data to generate an appropriate search query. The search result, "The recommended sushi is tuna nigiri," is converted into audio data by the speech synthesis engine and played on the device.

[1290] Examples of prompts used by generative AI models include:

[1291] The user says, "I want to order sushi." Generate an appropriate response.

[1292] In this way, the present invention provides a system that allows users to efficiently and intuitively order and search using voice or text input.

[1293] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1294] Step 1:

[1295] Users can input orders and menu searches by voice. Users can input "I want to order pizza" into a device such as a smartphone or smart speaker.

[1296] Step 2:

[1297] The device captures the user's voice and generates audio data, which is then sent to a cloud service.

[1298] Step 3:

[1299] The server uses the cloud service's voice recognition engine to convert the voice data into text data. For example, the voice saying "I'd like to order a pizza" is converted into text data saying "I'd like to order a pizza."

[1300] Step 4:

[1301] The server analyzes the converted text data and generates an appropriate search query. Specifically, based on the input text data "I want to order pizza," it generates the query "pizza menu" to search for menu information.

[1302] Step 5:

[1303] The server performs an internet search and menu information search based on the generated search query, and obtains the relevant menu information using the search means.

[1304] Step 6:

[1305] The server analyzes the search results and selects the most relevant information, for example, "The recommended pizza is Margherita."

[1306] Step 7:

[1307] The server generates the selected information as natural language text and converts it into voice data using a speech synthesis engine. This is the process of converting the text "The recommended pizza is Margherita" into voice data.

[1308] Step 8:

[1309] The terminal plays back the voice data and provides the search results to the user. By playing back the voice data "The recommended pizza is Margherita" to the user, the terminal confirms and suggests the order.

[1310] Step 9:

[1311] When a user inputs a text question on the screen of a smartphone, the user uses the text input means to place an order or search for a menu item. The input text data is sent to the server.

[1312] Step 10:

[1313] The server performs an internet search and menu information search based on the received text data. For example, if the text data "I want to order sushi" is received, the server generates the search query "sushi menu."

[1314] Step 11:

[1315] The server retrieves the search results, organizes them as natural language text, and returns the text to the user, providing the user with the text, "The recommended sushi is tuna nigiri."

[1316] Step 12:

[1317] The server transmits the input history of voice and text to the summary generation means, which analyzes the input history and automatically generates a summary.

[1318] Step 13:

[1319] The server notifies the user of the generated summary via the notification means. For example, the server notifies the user of the summary "Today's order: pizza, sushi."

[1320] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1321] The voice response and chat-linked search system of the present invention includes a voice input means installed in a smart device, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the converted text data, and a voice response means for converting the search results into voice data and providing it to the user. The user can also use the smartphone's input means to submit text questions, and an Internet search can be performed based on the text input, with the text response means providing the results to the user in text. The voice and text question history is summarized by a summary generation means and notified to the user by a notification means. Furthermore, the system is equipped with an emotion engine that recognizes emotions from the user's voice data, and can adjust the response content based on the emotional information.

[1322] Natural language description of the program

[1323] 1. Initial Setup:

[1324] The server sets up a personalized database containing the user's account information and search history, as well as the search API, message API, and emotion engine.

[1325] 2. Audio question reception:

[1326] The device goes into standby mode, waiting for the user to say the trigger phrase.

[1327] 3. Voice input:

[1328] When the user says a trigger phrase (e.g., "OK, Assistant"), the device switches to voice input mode and records the user's question.

[1329] 4. Speech Recognition:

[1330] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine in the cloud to convert the voice data into text data.

[1331] 5. Information Search:

[1332] The server analyzes the converted text data, generates an appropriate search query, and sends the generated search query to the search API to retrieve search results.

[1333] 6. Emotion recognition:

[1334] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[1335] 7. Tailor your response:

[1336] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[1337] 8. Voice response generation:

[1338] The server analyzes the retrieved search results and emotion information, selects the most relevant information, and generates a natural language voice response text. The natural language text is converted into voice data using the voice response means and sent to the terminal.

[1339] 9. Audio Answer:

[1340] The terminal plays back the received voice data and provides the user with a voice response.

[1341] 10. Chat questions:

[1342] The user types in a text question on their smartphone and sends it.

[1343] 11. Handling chat questions:

[1344] The server receives text questions through the message API, analyzes the received text questions, generates appropriate search queries using the search API, retrieves search results, analyzes the retrieved search results, organizes them as natural language text, and sends a text reply to the user.

[1345] 12. Question History Summary:

[1346] The server sends the audio and text question history to the summary generation means, which automatically generates a summary. The summary is then notified to the user via the notification means. The server also adjusts the summary content and notification format based on the emotion information.

[1347] Specific examples

[1348] Examples of voice questions:

[1349] User: "OK, Assistant, what's the weather like right now?"

[1350] 1. The device captures the audio data and sends it to the server.

[1351] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[1352] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[1353] 4. The server simultaneously sends the voice data to the emotion engine to recognize the user's emotion.

[1354] 5. The server generates a response text based on the emotion information: "Currently, the weather in Tokyo is sunny."

[1355] 6. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[1356] 7. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[1357] Examples of chat questions:

[1358] User (asking in chat app): "What's the weather like tomorrow?"

[1359] 1. The server receives a text question through the message API.

[1360] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[1361] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[1362] Example of a question history summary:

[1363] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[1364] 1. The server sends these search histories to the summary generator.

[1365] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1366] 3. The server notifies the user of the generated summary via a notification method, and adjusts the summary content and notification format if the user's emotions are recognized.

[1367] The above is a specific embodiment of the system of the present invention that combines an emotion engine. This system allows users to obtain information efficiently and in a way that takes emotion into consideration, and also allows users to check their past question history all at once.

[1368] The processing flow will be explained below.

[1369] Step 1:

[1370] The server sets up a personalized database containing the user's account information and search history, as well as the search API, message API, and emotion engine integration.

[1371] Step 2:

[1372] The device enters standby mode and prepares to switch to voice input mode when the user speaks the trigger phrase.

[1373] Step 3:

[1374] The user says the trigger phrase (e.g., "OK, Assistant").

[1375] Step 4:

[1376] The device detects the user's trigger phrase, switches to voice input mode, and records the user's question.

[1377] Step 5:

[1378] The device sends the recorded audio data to a cloud service.

[1379] Step 6:

[1380] The server uses a voice recognition engine in the cloud to convert the voice data into text data.

[1381] Step 7:

[1382] The server analyzes the converted text data and generates an appropriate search query.

[1383] Step 8:

[1384] The server sends the generated search query to the search API and retrieves the search results.

[1385] Step 9:

[1386] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[1387] Step 10:

[1388] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[1389] Step 11:

[1390] The server generates a voice response text in a natural language based on the adjusted response content.

[1391] Step 12:

[1392] The server uses a text-to-speech engine to convert the natural language text into audio data.

[1393] Step 13:

[1394] The server transmits the generated voice data to the terminal.

[1395] Step 14:

[1396] The terminal plays back the received voice data and provides the user with a voice response.

[1397] Step 15:

[1398] A user sends a text question via a chat app on their smartphone.

[1399] Step 16:

[1400] The server receives text questions through a message API.

[1401] Step 17:

[1402] The server analyzes the received text query and generates an appropriate search query.

[1403] Step 18:

[1404] The server sends the generated search query to the search API and retrieves the search results.

[1405] Step 19:

[1406] The server analyzes the search results and selects the most relevant information.

[1407] Step 20:

[1408] The server generates the selected information as natural language text and replies to the user via chat.

[1409] Step 21:

[1410] The server sends the audio and text question history to the summary generator.

[1411] Step 22:

[1412] The summary generator summarizes the audio and text question history and automatically generates a summary.

[1413] Step 23:

[1414] The summary generator sends the summary to the notification unit. If emotion information is present, the summary content and notification format are adjusted based on that information.

[1415] Step 24:

[1416] The server notifies the user of the generated summary via the notification means.

[1417] In this way, the system of the present invention can efficiently process voice or text questions from users and provide answers that take the user's feelings into consideration. In addition, by summarizing and notifying the user of the question history, the system can further facilitate the user's information acquisition.

[1418] Example 2

[1419] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1420] Conventional search systems that use voice and text input are unable to consider user sentiment when providing search results, limiting their ability to provide optimal search results. Furthermore, they do not effectively utilize voice and text question history, and lack a means to easily review past questions. This makes it difficult for users to quickly and accurately obtain the information they need.

[1421] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1422] In this invention, the server includes a summary generation unit that generates summaries of the voice and text question history, a response content adjustment unit that adjusts the response content of search results based on emotion information, and an emotion engine that performs emotion recognition. This makes it possible to provide appropriate and personalized search results while taking into consideration the user's emotions. In addition, by automatically summarizing the past question history and notifying the user, efficient information management and confirmation is possible.

[1423] "Voice input means" refers to a device or interface that allows a user to input a question by voice.

[1424] "Speech recognition means" refers to a technology or module for converting input voice data into text data.

[1425] The "search means" is a technology or module for conducting an internet search based on the converted text data and obtaining appropriate search results.

[1426] "Voice response means" refers to a technology or module for converting search results into voice data and providing it to the user.

[1427] "Audio output means" refers to a device or interface for outputting audio data to a user and conveying information.

[1428] "Character input means" refers to a device or interface that allows a user to input characters.

[1429] A "text response means" is a technology or module for conducting a search based on character input and returning search results to the user in text.

[1430] The "summary generation means" is a technology or module for summarizing the audio and text question history and automatically generating a summary.

[1431] "Notification means" refers to a technology or module for notifying the user of the summary generated results.

[1432] An "emotion engine" is a technology or module for recognizing emotions from a user's voice data.

[1433] The "response content adjustment means" is a technology or module for adjusting the response content of search results based on the emotion information obtained from the emotion engine.

[1434] A "personalized database" is a database that stores each user's account information and search history and provides personalized search results.

[1435] A "cloud service" is a service that uses remote computing resources and storage provided via the Internet.

[1436] The voice response and chat-linked search system of the present invention is composed of multiple means installed in a smart device, which processes users' voice and text questions, provides appropriate responses, and provides more personalized services through emotion recognition.

[1437] This system includes a voice input means for the user to input a question by voice, a voice recognition means for converting voice data into text data, a search means for performing an internet search based on the text data, a voice response means for converting search results into voice data, a voice output means for outputting the voice data to the user, a text input means for inputting a text question, a text response means for returning search results to the user in text, a summary generation means for generating a summary of the voice and text question history, and a notification means for notifying the user of the generated summary. Furthermore, it includes an emotion engine for recognizing emotions from the user's voice data, and a response content adjustment means for adjusting the response content of the search results based on the emotion information. Each part of this system can be implemented using an ordinary smartphone or a cloud server.

[1438] The smartphone's microphone is used as the voice input means. When the user's voice input is recognized as a trigger phrase (e.g., "OK, Assistant"), the voice recognition means sends the voice to a cloud service, where it is converted into text data by a voice recognition engine (e.g., Google Cloud Speech-to-Text API). The search means then searches the converted text data using an internet search API (e.g., Google Custom Search API).

[1439] The search results are converted by the voice response means and notified to the user by voice through the voice output means. In this series of steps, an emotion recognition engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotional state, and the response content adjustment means determines an appropriate response based on this.

[1440] The smartphone keyboard is used as the text input means. The user types and sends a question in text. For example, if the user types "What's the weather going to be like tomorrow?", the question is sent to the server via a messaging API (e.g., Twilio API), which then performs an internet search. The retrieved search results are returned to the user in text via the text response means.

[1441] The question history is stored in a question history database. The summary generation means summarizes this history and notifies the user as necessary. For example, if a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?", the summary generation means generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy" and notifies the user via the notification means.

[1442] Specific examples

[1443] Examples of voice questions:

[1444] User: "OK, Assistant, what's the weather like right now?"

[1445] 1. The device captures the audio data and sends it to the server.

[1446] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[1447] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[1448] 4. The server simultaneously sends the voice data to the emotion engine to recognize the user's emotion.

[1449] 5. The server generates a response text based on the emotion information: "Currently, the weather in Tokyo is sunny."

[1450] 6. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[1451] 7. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[1452] Examples of chat questions:

[1453] User (asking in chat app): "What's the weather like tomorrow?"

[1454] 1. The server receives a text question through the message API.

[1455] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[1456] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[1457] Example of a question history summary:

[1458] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[1459] 1. The server sends these search histories to the summary generator.

[1460] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1461] 3. The server notifies the user of the generated summary via a notification means.

[1462] The system of the present invention allows users to obtain information efficiently and sensitively, and also allows them to review their past question history in one place, significantly improving user convenience and enabling more personalized services.

[1463] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1464] Step 1: Initial Setup

[1465] The server sets up a personalized database that stores user account information and past search history, and also establishes integration with the search API, message API, and emotion engine.

[1466] Input: User information, past search history

[1467] Output: Personalized database, API and engine integration settings

[1468] Step 2: Voice inquiry reception

[1469] The device will enter standby mode and wait for the user to say a trigger phrase, such as "OK, Assistant."

[1470] Input: Trigger phrase

[1471] Output: Activate voice input mode

[1472] Step 3: Voice Input

[1473] When the user utters the trigger phrase, the device switches to voice input mode and records the user's question, for example, "What's the weather like today?"

[1474] Input: User's voice question

[1475] Output: Recorded audio data

[1476] Step 4: Voice Recognition

[1477] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine to convert the voice data into text data.

[1478] Input: Recorded audio data

[1479] Output: Converted text data

[1480] Step 5: Information Search

[1481] The server analyzes the converted text data and generates an appropriate search query, which it then sends to an Internet search API to retrieve search results.

[1482] Input: Text data

[1483] Output: Search result data

[1484] Step 6: Emotion Recognition

[1485] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[1486] Input: Audio data

[1487] Output: Emotional information

[1488] Step 7: Tailor your response

[1489] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[1490] Input: Emotion information, search result data

[1491] Output: Adjusted response content

[1492] Step 8: Generate voice response

[1493] The server integrates the search results and emotional information to generate a natural language voice response text, converts the text into voice data, and sends it to the device.

[1494] Input: Adjusted response content

[1495] Output: Audio data

[1496] Step 9: Answer by voice

[1497] The terminal plays back the received voice data and provides the user with a voice response.

[1498] Input: Audio data

[1499] Output: A spoken response to the user

[1500] Step 10: Chat with your questions

[1501] A user types a text question into a smartphone and sends it, for example, "What's the weather like tomorrow?" in a chat app.

[1502] Input: User's character question

[1503] output: Character data to be sent

[1504] Step 11: Handling chat questions

[1505] The server receives text questions through the message API, generates search queries based on the questions using the search API, retrieves search results, organizes the results into natural language text, and sends a text reply to the user.

[1506] Input: User's character question

[1507] Output: Search result text

[1508] Step 12: Question History Summary

[1509] The server sends the audio and text question history to the summary generation means, which automatically generates a summary. The summary is then notified to the user via the notification means. The server also adjusts the summary content and notification format based on the emotion information.

[1510] Input: Voice question history, text question history

[1511] Output: Summary, Notification

[1512] (Application example 2)

[1513] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1514] In today's world, it's important to provide users with a way to quickly and efficiently obtain the information they need. However, traditional voice response and chat-based search systems are unable to adjust responses based on the user's emotional state, which can result in a poor user experience. Furthermore, they are unable to summarize multiple query histories and notify users, making it difficult to easily review past search results. Another problem is the inability to provide personalized results based on voice or text queries.

[1515] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice input means for a user to input a question by voice, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the text data, a voice response means for converting the search results into voice data, a voice output means for outputting the voice data to the user, a character input means for inputting a text question, a search means for performing an Internet search based on the text question, a text response means for providing the user with a method for replying to the search results of the text question in text form, a summary generation means for generating a summary of the voice and text question history, a notification means for notifying the user of the generated summary results, an emotion recognition means for acquiring emotion information, and a response adjustment means for adapting the search results based on the emotion information. This makes it possible to provide personalized search results according to the user's emotions, improve the quality of the user experience, and make it easy to check past question history.

[1516] The "voice input means" is a means for the user to input a question by voice.

[1517] The "voice recognition means" is a means for converting voice data into text data.

[1518] "Search means" refers to a means for conducting an Internet search based on text data.

[1519] The "voice response means" is a means for converting search results into voice data.

[1520] The "audio output means" is a means for outputting audio data to the user.

[1521] The "text input means" is a means for the user to input a text question.

[1522] The "text response means" is a means for returning search results for a text question to the user in text form.

[1523] The "summary generation means" is a means for generating summaries of the audio and text question history.

[1524] The "notification means" is a means for notifying the user of the results of the generated summary.

[1525] "Emotion recognition means" is a means for acquiring emotion information.

[1526] A "response adjustment means" is a means for adapting search results based on emotional information.

[1527] A "personalized database" is a database that stores a user's question history and customizes search results individually.

[1528] A "cloud service" is a service that processes voice data and text data on an online server and performs voice recognition and search.

[1529] To implement this invention, the following hardware and software are primarily used. A user uses a smart device (e.g., a smartphone) to ask a question by voice or text and obtain search results corresponding to the question. The following describes the functions and processing flow of the system that realizes this application example.

[1530] Required Hardware and Software

[1531] 1. Smart devices (e.g. smartphones, tablets): Provide an interface with users. Devices that allow voice input, voice output, and text input.

[1532] 2. Cloud Services: Online platforms for processing voice and text data. They perform speech recognition, search, and emotion recognition. They use services such as Google Speech-to-Text API, IBM Watson Tone Analyzer, and Amazon Polly.

[1533] 3. Server: Manages the user's personalized database and generates search results. Provides tailored responses based on the user's question history and sentiment information.

[1534] System Operation Overview

[1535] 1. Audio question reception:

[1536] The user speaks a voice trigger phrase (e.g., "OK, Assistant") into the smart device. When the device switches to voice input mode, the user asks, "What movies are recommended today?"

[1537] 2. Speech and Emotion Recognition:

[1538] The smart device sends the voice data to a cloud service, which uses the Google Speech-to-Text API to convert the voice data into text, and simultaneously sends the voice data to the IBM Watson Tone Analyzer to recognize the user's emotions.

[1539] 3. Information Search:

[1540] The server generates a search query based on the converted text data and sends it to the appropriate search API (e.g., a movie database). Based on the user's emotional information, the server can recommend movies.

[1541] 4. Tailor your response:

[1542] The server selects the most appropriate information from the search results based on emotional information and generates a response in natural language. For example, if the user feels like relaxing, it generates a response such as, "The recommended relaxing movie is 'Inception.'"

[1543] 5. Voice response generation:

[1544] The generated response is converted into voice data using Amazon Polly and sent to the smart device, which then outputs the converted voice data to the user and provides a response.

[1545] 6. Question History Summary:

[1546] The server summarizes the user's past question history and periodically notifies the user of the summarized information. The summary information is presented to the user in the form of, for example, "Past 7 days' viewing history: 'Inception', 'The Matrix'."

[1547] Examples and prompts

[1548] Specific examples

[1549] When a user says "What are the best movies tonight?", the following happens:

[1550] 1. The voice data is sent to a cloud service and converted into text data by a speech recognition engine. A query such as "Today's recommended movies" is generated.

[1551] 2. The emotion recognition engine detects the user's desire to relax from their voice.

[1552] 3. The server searches a movie database and applies a filter that recommends "relaxing movies."

[1553] 4. Generate a response sentence: "A recommended relaxing movie is 'Inception'." and convert it into speech.

[1554] 5. The smart device plays this audio to the user.

[1555] Example prompts

[1556] Here are some example prompts for using generative AI models:

[1557] User asks: "What movie do you recommend today?"

[1558] Hardware: Smartphone

[1559] Software Settings:

[1560] Speech Recognition: Google Speech-to-Text API

[1561] Emotion recognition: IBM Watson Tone Analyzer

[1562] Speech synthesis: Amazon Polly

[1563] Specific processing:

[1564] 1. The voice data is sent to the speech recognition engine and converted into text.

[1565] 2. Send the text data to the movie database API and get the search results.

[1566] 3. Send the voice data to the emotion engine and obtain the emotion information.

[1567] 4. Based on the emotional information, the system selects an appropriate movie from the search results and generates a response in natural language.

[1568] 5. The response is converted into audio data and played on the smartphone.

[1569] Expected response: "A movie recommendation to fit your current mood is 'Inception'."

[1570] The system provided by this invention allows users to obtain personalized search results based on their emotions and easily check their past question history. This system improves the quality of the user experience and provides a means to efficiently obtain information.

[1571] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1572] Step 1:

[1573] Voice inquiry reception

[1574] The device enters standby mode and waits for the user to speak a trigger phrase. The input is a voice trigger phrase spoken by the user (e.g., "OK, Assistant"), and the device captures the voice and switches to voice input mode to accept the user's question.

[1575] Step 2:

[1576] Voice input

[1577] The user asks the device, "What movies do you recommend today?" The device records the user's voice and sends the voice data to the cloud service. The input is the recorded voice data, and the output is the voice data sent to the cloud service.

[1578] Step 3:

[1579] Voice Recognition

[1580] The cloud service converts the transmitted voice data into text data using the Google Speech-to-Text API. The input is the voice data, and the output is the text data "What movies do you recommend today?" The cloud service analyzes the voice data and generates the corresponding text.

[1581] Step 4:

[1582] emotion recognition

[1583] The cloud service sends the voice data to the IBM Watson Tone Analyzer along with the text data to recognize the user's emotions. The input is the voice data, and the output is the user's emotional information. The cloud service analyzes the voice data and recognizes the user's emotional state.

[1584] Step 5:

[1585] Information Search

[1586] The server analyzes the text data obtained from speech recognition and generates an appropriate search query. For example, it sends a query such as "Today's recommended movies" to a movie database API and retrieves relevant search results. The input is the converted text data, and the output is a list of search results.

[1587] Step 6:

[1588] Tailoring response content

[1589] The server determines the optimal response based on the search results and emotional information, and generates a response in natural language. For example, if the user is in the mood to relax, it generates the text "A recommended relaxing movie is 'Inception.'" The input is the search results and emotional information, and the output is a response in natural language.

[1590] Step 7:

[1591] Voice Response Generation

[1592] The server converts the generated response text into voice data using Amazon Polly and sends it to the device. The input is text data and the output is voice data. The server analyzes the response text and generates voice data.

[1593] Step 8:

[1594] Answer by voice

[1595] The device plays the received voice data and provides a voice response to the user. The input is the voice data sent from the server, and the output is the voice played from the device. The device analyzes the voice data and conveys information to the user by voice.

[1596] Step 9:

[1597] Question History Summary

[1598] The server sends the user's past question history to the summary generation means, which then summarizes the question history. For example, it generates a summary result such as "Past 7 days' viewing history: 'Inception', 'The Matrix'." The input is the past question history data, and the output is summarized text data. The server analyzes the question history and generates summary information.

[1599] Step 10:

[1600] Summary Notification

[1601] The server notifies the user of the generated summarized question history via a notification means. The input is the summarized text data, and the output is notification information for the user. The server analyzes the summarized data and notifies the user in an appropriate format.

[1602] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1603] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1604] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1605] [Fourth embodiment]

[1606] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1607] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1608] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1609] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1610] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1611] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1612] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1613] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1614] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1615] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1616] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1617] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1618] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1619] The voice response and chat-linked search system of the present invention includes a voice input means installed in a smart device, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the converted text data, and a voice response means for converting the search results into voice data and providing it to the user. The user can also use the input means of the smartphone to submit text questions, and an Internet search can be performed based on the text input, with the text response means providing the results to the user in text. The voice and text question history is summarized by a summary generation means and notified to the user by a notification means.

[1620] Natural language description of the program

[1621] 1. Initial Setup:

[1622] The server stores the user's account information and search history in a personalized database, and also configures the search API and message API integration.

[1623] 2. Audio question reception:

[1624] The device is in standby mode and starts voice input when the user speaks a trigger phrase. The user then speaks a question.

[1625] 3. Speech Recognition:

[1626] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine in the cloud to convert the voice data into text data.

[1627] 4. Information Search:

[1628] The server analyzes the converted text data, generates an appropriate search query, and sends the generated search query to the search API to retrieve search results.

[1629] 5. Voice response generation:

[1630] The server analyzes the search results, selects the most relevant information, and generates natural language text based on the search results.The natural language text is converted into voice data using a voice response means and sent to the terminal.

[1631] 6. Audio Answer:

[1632] The terminal plays back the received voice data and provides the user with a voice response.

[1633] 7. Chat questions:

[1634] The user types in a text question on their smartphone and sends it.

[1635] 8. Handling chat questions:

[1636] The server analyzes the received text question, generates an appropriate search query using a search API, retrieves search results, analyzes the retrieved search results, organizes them as natural language text, and sends a text reply to the user.

[1637] 9. Question History Summary:

[1638] The server periodically transmits the question history of voice and text to the summary generating means, which automatically generates a summary, and notifies the user of the summary via the notification means.

[1639] Specific examples

[1640] Examples of voice questions:

[1641] User: "OK, Assistant, what's the weather like right now?"

[1642] 1. The device captures the audio data and sends it to the server.

[1643] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[1644] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[1645] 4. The server generates the answer text "Currently, the weather in Tokyo is sunny."

[1646] 5. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[1647] 6. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[1648] Examples of chat questions:

[1649] User (asking in chat app): "What's the weather like tomorrow?"

[1650] 1. The server receives this request through the chat API.

[1651] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[1652] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[1653] Example of a question history summary:

[1654] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[1655] 1. The server sends these search histories to the summary generator.

[1656] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1657] 3. The server notifies the user of the generated summary via a notification means.

[1658] The above is an embodiment of the invention and its specific examples. This system allows users to efficiently obtain information and easily refer to their past search history.

[1659] The processing flow will be explained below.

[1660] Step 1:

[1661] The server sets up a personalized database containing the user's account information and search history, and also sets up integration between the search API and the messaging API.

[1662] Step 2:

[1663] The device goes into standby mode, waiting for the user to say the trigger phrase.

[1664] Step 3:

[1665] When the user says a trigger phrase (e.g., "OK, Assistant"), the device switches to voice input mode and records the user's question.

[1666] Step 4:

[1667] The device sends the recorded voice data to a cloud service, where it is processed by a voice recognition engine.

[1668] Step 5:

[1669] The server uses a voice recognition engine in the cloud to convert the voice data into text data.

[1670] Step 6:

[1671] The server analyzes the converted text data and generates an appropriate search query.

[1672] Step 7:

[1673] The server sends the generated search query to the search API and retrieves the search results.

[1674] Step 8:

[1675] The server analyzes the search results and selects the most relevant information.

[1676] Step 9:

[1677] The server generates a voice response text in a natural language based on the selected information.

[1678] Step 10:

[1679] The server uses a text-to-speech engine to convert the natural language text into audio data.

[1680] Step 11:

[1681] The server transmits the generated voice data to the terminal.

[1682] Step 12:

[1683] The terminal plays back the received voice data and provides the user with a voice response.

[1684] Step 13:

[1685] A user sends a text question via a chat app on their smartphone.

[1686] Step 14:

[1687] The server receives text questions through a message API.

[1688] Step 15:

[1689] The server analyzes the received text query and generates an appropriate search query.

[1690] Step 16:

[1691] The server sends the generated search query to the search API and retrieves the search results.

[1692] Step 17:

[1693] The server analyzes the search results and selects the most relevant information.

[1694] Step 18:

[1695] The server generates the selected information as natural language text and returns the text to the user.

[1696] Step 19:

[1697] The server sends the audio and text question history to the summary generator.

[1698] Step 20:

[1699] The server analyzes the summary received from the generation AI and notifies the user of the summary via a notification means.

[1700] This is the specific processing flow of the entire system of the present invention. This system allows users to efficiently obtain information while also checking past question history all at once.

[1701] Example 1

[1702] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1703] In today's information search systems using smart devices and cloud services, users need a way to efficiently ask questions by voice or text and quickly obtain results. However, existing systems often lack the accuracy of voice recognition and personalized search results, hindering the user experience. Furthermore, there are no established methods for effectively utilizing search history or properly notifying search results. Furthermore, relying on cloud services to process voice and text data can make integration complex.

[1704] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1705] In this invention, the server includes a voice input means for a user to input a question by voice, a voice recognition means for converting the voice data into text data, a search means for performing an information search based on the text data, a voice response means for converting the search results into voice data, a voice output means for outputting the voice data to the user, a character input means for inputting a text question, a search means for performing an information search based on the text question, a text response means for returning the search results of the text question to the user in text, a summary generation means for generating a summary of the voice and text question history, a notification means for notifying the user of the generated summary, and a means for transmitting the voice and text data to a cloud service and performing voice recognition and search using the cloud service. This enables a user to efficiently ask questions by voice and text, quickly obtain personalized search results, effectively utilize their search history, and receive appropriate notification of the results.

[1706] A "user" is an entity that uses the system to enter questions and obtain information.

[1707] "Voice input means" refers to a device or function that allows a user to input a question by voice.

[1708] "Speech recognition means" refers to the technology or engine that converts input voice data into text data.

[1709] "Search means" refers to a device or function for searching for information based on text data.

[1710] "Voice response means" refers to a device or technology that converts text data obtained as a search result into voice data.

[1711] "Audio output means" refers to a device or function for reproducing audio data to the user.

[1712] "Character input means" refers to a device or function that allows a user to input a question in text.

[1713] A "text response means" is a device or technology that responds to a user in text with the search results of a text question.

[1714] A "summary generator" is a device or technique that summarizes the audio and text question history and generates summarized information.

[1715] The "notification means" refers to a device or function for notifying the user of the results of the generated summary.

[1716] A "personalized database" is a database that stores each user's search history and account information and responds to them individually.

[1717] "Cloud services" are remote computing resources used to perform voice recognition and search processing.

[1718] The voice response and chat-linked search system according to the present invention is mainly composed of a server, a terminal, and a user. The specific processing of these components and the hardware and software used will be described below.

[1719] Server Roles and Configuration

[1720] The server stores the user's account information and search history in a personalization database. The server also configures search and messaging APIs, such as Google's search and messaging APIs (e.g., Twilio). The server then sends the received voice data to a cloud service for speech recognition. For this purpose, the Google Cloud Speech-to-Text API is used.

[1721] Terminal roles and configuration

[1722] The device accepts voice and text input from the user. When voice input is performed, the device recognizes trigger phrases and captures voice data. The recorded voice data is sent to a server for speech recognition. The device has the ability to generate voice data using the Google Cloud Text-to-Speech API and play it back to the user.

[1723] User Actions

[1724] The user inputs a question by voice or text. For example, if the user asks by voice, "OK, Assistant, what's the weather like today?", the device captures this and sends it to the server. The server performs speech recognition processing and converts it into text data. The server then calls the search API based on the query "What's the weather like today?" and retrieves search results. The server generates a text answer, "Currently, the weather in Tokyo is sunny," converts this into audio data, and sends it to the device. The device then plays back the audio, "Currently, the weather in Tokyo is sunny."

[1725] Use of cloud services

[1726] The voice and text data are sent to a cloud service, where voice recognition and information search are performed within the cloud. This enables highly accurate voice recognition and rapid search results. The server also periodically sends the voice and text question history to a summary generator, which automatically generates a summary. This summary is then notified to the user via a notification device.

[1727] Specific examples

[1728] Examples of voice questions:

[1729] User: "OK, Assistant, what's the weather like right now?"

[1730] 1. The device captures the audio data and sends it to the server.

[1731] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[1732] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[1733] 4. The server generates the answer text "Currently, the weather in Tokyo is sunny."

[1734] 5. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[1735] 6. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[1736] Examples of chat questions:

[1737] User (asking in chat app): "What's the weather like tomorrow?"

[1738] 1. The server receives this request through the chat API.

[1739] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[1740] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[1741] Example of a question history summary:

[1742] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[1743] 1. The server sends these search histories to the summary generator.

[1744] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1745] 3. The server notifies the user of the generated summary via a notification means.

[1746] This system configuration allows users to efficiently retrieve information and easily refer to past search history.The system provides an advanced user experience by combining a voice recognition engine, search API, and text processing engine.

[1747] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1748] Step 1:

[1749] The server saves the user's account information and search history in a personalized database. This is done when a new account is registered. The input is account information and search history, and the output is saved in the personalized database. Specifically, the user's name, email address, past search history, etc. are stored in the database.

[1750] Step 2:

[1751] The server configures the search API and message API integration. The input is authentication information, and the output is completion of API integration. Specifically, it obtains authentication tokens for Google's search API and general message APIs and includes them in the server configuration.

[1752] Step 3:

[1753] The device is in standby mode, and when the user utters a trigger phrase, voice input begins. The input is a trigger phrase such as "OK, Assistant," and the output is a transition to voice input mode. Specifically, the device switches from standby mode to voice input mode and becomes ready to record the user's question.

[1754] Step 4:

[1755] The user inputs a question by voice. The input is a voice question such as "What's the weather like now?", and the output is recorded voice data. Specifically, the user says "What's the weather like now?", and the voice is recorded by the device.

[1756] Step 5:

[1757] The device sends the recorded voice data to a cloud service. The input is the recorded voice data, and the output is the transmission of the voice data to the cloud service. Specifically, the recorded data is sent to a server over the Internet (e.g., Google Cloud Speech-to-Text).

[1758] Step 6:

[1759] The server uses a speech recognition engine in the cloud to convert the voice data into text data. The input is the voice data received from the cloud service, and the output is the converted text data. Specifically, the Google Cloud Speech-to-Text API is used to convert the voice data into text data such as "What's the weather like today?"

[1760] Step 7:

[1761] The server analyzes the converted text data and generates an appropriate search query. The input is text data and the output is a search query. Specifically, it analyzes the text "What's the weather like now?" to generate a search query such as "Current weather."

[1762] Step 8:

[1763] The server sends the generated search query to the search API and retrieves the search results. The input is the search query and the output is the search results. Specifically, the server sends the query "current weather" to Google's search API and receives the search results.

[1764] Step 9:

[1765] The server analyzes the search results, selects the most relevant information, and generates natural language text. The input is the search results, and the output is the natural language text. Specifically, the generated text is, "Currently, the weather in Tokyo is sunny."

[1766] Step 10:

[1767] The server uses a voice response means to convert the generated natural language text into voice data and send it to the terminal. The input is natural language text and the output is voice data. Specifically, the Google Cloud Text-to-Speech API is used to generate the voice "Currently, the weather in Tokyo is sunny," and this is sent to the terminal.

[1768] Step 11:

[1769] The device plays the received voice data and provides the user with a voice response. The input is the voice data, and the output is the voice that is played back. Specifically, the device plays back the voice, "Currently, the weather in Tokyo is sunny."

[1770] Step 12:

[1771] The user inputs a text question into their smartphone and sends it. The input is a text question such as "What's the weather like tomorrow?", and the output is the text data sent.

[1772] Step 13:

[1773] The server analyzes the received text question, generates an appropriate search query using the search API, and retrieves search results. The input is text data, and the output is search results. Specifically, the server analyzes the text "What's the weather like tomorrow?", generates the search query "Tomorrow's weather," and sends it to Google's search API to retrieve search results.

[1774] Step 14:

[1775] The server analyzes the search results, organizes them as natural language text, and sends a text reply to the user. The input is the search results, and the output is natural language text. Specifically, the server generates the text "The weather in Tokyo will be cloudy tomorrow," and sends it back to the user via chat.

[1776] Step 15:

[1777] The server periodically sends the audio and text question history to the summary generation means, which then automatically generates a summary. The input is the question history, and the output is the summary. Specifically, the server sends the search history from the past 24 hours to the summary generation means, and generates a summary such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1778] Step 16:

[1779] The server notifies the user of the generated summary via a notification means. The input is the summary, and the output is a notification to the user. Specifically, the server notifies the user of the generated summary via email or the notification function of their smartphone.

[1780] (Application example 1)

[1781] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1782] In order to enable users to easily order and search for menus using voice or text, food delivery services must simplify conventional operations and provide a more efficient and intuitive interface. They also need a system that can personalize order and search histories to make optimal suggestions to users.

[1783] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1784] In this invention, the server includes voice input means for a user to input orders or menu searches by voice, voice recognition means for converting voice data into text data, search means for performing an internet search and menu information search based on the text data, voice response means for converting search results into voice data, voice output means for outputting the voice data to the user, character input means for inputting orders or menu searches using character input means, search means for performing an internet search and menu information search based on the character input, character response means for providing a method for returning search results of the character input to the user in text, summary generation means for generating a summary of the voice and character input history, and notification means for notifying the user of the generated summary results. This allows users to efficiently and intuitively order or search by voice or character input and receive personalized suggestions based on their past history.

[1785] "Voice input means" refers to a means by which a user can input information using voice.

[1786] The "voice recognition means" is a means for converting voice data into text data.

[1787] The "search means" is a means for searching the Internet or menu information based on text data.

[1788] The "voice response means" is a means for converting search results into voice data and providing it to the user.

[1789] The "audio output means" is a means for outputting audio data to the user.

[1790] "Character input means" refers to a means by which a user inputs information using characters.

[1791] The "text response means" is a means for returning search results based on a text question to a user in text form.

[1792] The "summary generation means" is a means for generating summaries of the audio and text question history.

[1793] The "notification means" is a means for notifying the user of the results of the generated summary.

[1794] A "personalized database" is a database that stores a user's input history and provides personalized information.

[1795] The present invention provides a system that allows efficient and intuitive operation of a food delivery service. Specific embodiments for realizing this system are described below.

[1796] System configuration

[1797] The system consists of a user, a device, and a server. Users use devices such as smartphones and smart speakers to place orders and search for menus by voice or text. The device accepts both voice and text input.

[1798] Hardware and software used

[1799] Hardware: smartphones, smart speakers, microphones

[1800] Software: Speech recognition engine (e.g., Google Cloud Speech-to-Text API), speech synthesis engine (e.g., Google Text-to-Speech API)

[1801] Data processing and calculation

[1802] Voice input means

[1803] When a user orders or searches the menu by voice, the device captures the voice data, which is then sent to a cloud service where it is converted into text using a voice recognition engine.

[1804] Voice recognition means

[1805] The voice recognition engine in the cloud service converts voice data into text data. For example, a voice saying "I want to order a pizza" is converted into text "I want to order a pizza."

[1806] Search methods

[1807] The server analyzes the converted text data and generates appropriate queries to search the Internet and retrieve menu information, and the search results are retrieved on the server side.

[1808] Voice response means

[1809] The server analyzes the search results and selects the most relevant information, which is then generated as natural language text and converted into voice data using a speech synthesis engine.

[1810] Audio output means

[1811] The device plays audio data to the user and provides information such as search results and order confirmations.

[1812] Character input means and character response means

[1813] Users can also enter text on their smartphone screens. This text data is sent to the server, which uses search tools to generate appropriate search queries. The search results are organized as natural language text and provided to users in text format.

[1814] Summary generation means and notification means

[1815] The server transmits the user's voice and text input history to a summary generation means, which generates a summary such as "Today's order: pizza, salad" and notifies the user via a notification means.

[1816] Specific examples

[1817] For example, if a user says, "I'd like to order sushi," the device captures this speech and sends it to the cloud service. The speech recognition engine converts it into text data, and the server analyzes the text data to generate an appropriate search query. The search result, "The recommended sushi is tuna nigiri," is converted into audio data by the speech synthesis engine and played on the device.

[1818] Examples of prompts used by generative AI models include:

[1819] The user says, "I want to order sushi." Generate an appropriate response.

[1820] In this way, the present invention provides a system that allows users to efficiently and intuitively order and search using voice or text input.

[1821] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1822] Step 1:

[1823] Users can input orders and menu searches by voice. Users can input "I want to order pizza" into a device such as a smartphone or smart speaker.

[1824] Step 2:

[1825] The device captures the user's voice and generates audio data, which is then sent to a cloud service.

[1826] Step 3:

[1827] The server uses the cloud service's voice recognition engine to convert the voice data into text data. For example, the voice saying "I'd like to order a pizza" is converted into text data saying "I'd like to order a pizza."

[1828] Step 4:

[1829] The server analyzes the converted text data and generates an appropriate search query. Specifically, based on the input text data "I want to order pizza," it generates the query "pizza menu" to search for menu information.

[1830] Step 5:

[1831] The server performs an internet search and menu information search based on the generated search query, and obtains the relevant menu information using the search means.

[1832] Step 6:

[1833] The server analyzes the search results and selects the most relevant information, for example, "The recommended pizza is Margherita."

[1834] Step 7:

[1835] The server generates the selected information as natural language text and converts it into voice data using a speech synthesis engine. This is the process of converting the text "The recommended pizza is Margherita" into voice data.

[1836] Step 8:

[1837] The terminal plays back the voice data and provides the search results to the user. By playing back the voice data "The recommended pizza is Margherita" to the user, the terminal confirms and suggests the order.

[1838] Step 9:

[1839] When a user inputs a text question on the screen of a smartphone, the user uses the text input means to place an order or search for a menu item. The input text data is sent to the server.

[1840] Step 10:

[1841] The server performs an internet search and menu information search based on the received text data. For example, if the text data "I want to order sushi" is received, the server generates the search query "sushi menu."

[1842] Step 11:

[1843] The server retrieves the search results, organizes them as natural language text, and returns the text to the user, providing the user with the text, "The recommended sushi is tuna nigiri."

[1844] Step 12:

[1845] The server transmits the input history of voice and text to the summary generation means, which analyzes the input history and automatically generates a summary.

[1846] Step 13:

[1847] The server notifies the user of the generated summary via the notification means. For example, the server notifies the user of the summary "Today's order: pizza, sushi."

[1848] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1849] The voice response and chat-linked search system of the present invention includes a voice input means installed in a smart device, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the converted text data, and a voice response means for converting the search results into voice data and providing it to the user. The user can also use the smartphone's input means to submit text questions, and an Internet search can be performed based on the text input, with the text response means providing the results to the user in text. The voice and text question history is summarized by a summary generation means and notified to the user by a notification means. Furthermore, the system is equipped with an emotion engine that recognizes emotions from the user's voice data, and can adjust the response content based on the emotional information.

[1850] Natural language description of the program

[1851] 1. Initial Setup:

[1852] The server sets up a personalized database containing the user's account information and search history, as well as the search API, message API, and emotion engine.

[1853] 2. Audio question reception:

[1854] The device goes into standby mode, waiting for the user to say the trigger phrase.

[1855] 3. Voice input:

[1856] When the user says a trigger phrase (e.g., "OK, Assistant"), the device switches to voice input mode and records the user's question.

[1857] 4. Speech Recognition:

[1858] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine in the cloud to convert the voice data into text data.

[1859] 5. Information Search:

[1860] The server analyzes the converted text data, generates an appropriate search query, and sends the generated search query to the search API to retrieve search results.

[1861] 6. Emotion recognition:

[1862] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[1863] 7. Tailor your response:

[1864] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[1865] 8. Voice response generation:

[1866] The server analyzes the retrieved search results and emotion information, selects the most relevant information, and generates a natural language voice response text. The natural language text is converted into voice data using the voice response means and sent to the terminal.

[1867] 9. Audio Answer:

[1868] The terminal plays back the received voice data and provides the user with a voice response.

[1869] 10. Chat questions:

[1870] The user types in a text question on their smartphone and sends it.

[1871] 11. Handling chat questions:

[1872] The server receives text questions through the message API, analyzes the received text questions, generates appropriate search queries using the search API, retrieves search results, analyzes the retrieved search results, organizes them as natural language text, and sends a text reply to the user.

[1873] 12. Question History Summary:

[1874] The server sends the audio and text question history to the summary generation means, which automatically generates a summary. The summary is then notified to the user via the notification means. The server also adjusts the summary content and notification format based on the emotion information.

[1875] Specific examples

[1876] Examples of voice questions:

[1877] User: "OK, Assistant, what's the weather like right now?"

[1878] 1. The device captures the audio data and sends it to the server.

[1879] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[1880] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[1881] 4. The server simultaneously sends the voice data to the emotion engine to recognize the user's emotion.

[1882] 5. The server generates a response text based on the emotion information: "Currently, the weather in Tokyo is sunny."

[1883] 6. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[1884] 7. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[1885] Examples of chat questions:

[1886] User (asking in chat app): "What's the weather like tomorrow?"

[1887] 1. The server receives a text question through the message API.

[1888] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[1889] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[1890] Example of a question history summary:

[1891] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[1892] 1. The server sends these search histories to the summary generator.

[1893] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1894] 3. The server notifies the user of the generated summary via a notification method, and adjusts the summary content and notification format if the user's emotions are recognized.

[1895] The above is a specific embodiment of the system of the present invention that combines an emotion engine. This system allows users to obtain information efficiently and in a way that takes emotion into consideration, and also allows users to check their past question history all at once.

[1896] The processing flow will be explained below.

[1897] Step 1:

[1898] The server sets up a personalized database containing the user's account information and search history, as well as the search API, message API, and emotion engine integration.

[1899] Step 2:

[1900] The device enters standby mode and prepares to switch to voice input mode when the user speaks the trigger phrase.

[1901] Step 3:

[1902] The user says the trigger phrase (e.g., "OK, Assistant").

[1903] Step 4:

[1904] The device detects the user's trigger phrase, switches to voice input mode, and records the user's question.

[1905] Step 5:

[1906] The device sends the recorded audio data to a cloud service.

[1907] Step 6:

[1908] The server uses a voice recognition engine in the cloud to convert the voice data into text data.

[1909] Step 7:

[1910] The server analyzes the converted text data and generates an appropriate search query.

[1911] Step 8:

[1912] The server sends the generated search query to the search API and retrieves the search results.

[1913] Step 9:

[1914] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[1915] Step 10:

[1916] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[1917] Step 11:

[1918] The server generates a voice response text in a natural language based on the adjusted response content.

[1919] Step 12:

[1920] The server uses a text-to-speech engine to convert the natural language text into audio data.

[1921] Step 13:

[1922] The server transmits the generated voice data to the terminal.

[1923] Step 14:

[1924] The terminal plays back the received voice data and provides the user with a voice response.

[1925] Step 15:

[1926] A user sends a text question via a chat app on their smartphone.

[1927] Step 16:

[1928] The server receives text questions through a message API.

[1929] Step 17:

[1930] The server analyzes the received text query and generates an appropriate search query.

[1931] Step 18:

[1932] The server sends the generated search query to the search API and retrieves the search results.

[1933] Step 19:

[1934] The server analyzes the search results and selects the most relevant information.

[1935] Step 20:

[1936] The server generates the selected information as natural language text and replies to the user via chat.

[1937] Step 21:

[1938] The server sends the audio and text question history to the summary generator.

[1939] Step 22:

[1940] The summary generator summarizes the audio and text question history and automatically generates a summary.

[1941] Step 23:

[1942] The summary generator sends the summary to the notification unit. If emotion information is present, the summary content and notification format are adjusted based on that information.

[1943] Step 24:

[1944] The server notifies the user of the generated summary via the notification means.

[1945] In this way, the system of the present invention can efficiently process voice or text questions from users and provide answers that take the user's feelings into consideration. In addition, by summarizing and notifying the user of the question history, the system can further facilitate the user's information acquisition.

[1946] Example 2

[1947] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1948] Conventional search systems that use voice and text input are unable to consider user sentiment when providing search results, limiting their ability to provide optimal search results. Furthermore, they do not effectively utilize voice and text question history, and lack a means to easily review past questions. This makes it difficult for users to quickly and accurately obtain the information they need.

[1949] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1950] In this invention, the server includes a summary generation unit that generates summaries of the voice and text question history, a response content adjustment unit that adjusts the response content of search results based on emotion information, and an emotion engine that performs emotion recognition. This makes it possible to provide appropriate and personalized search results while taking into consideration the user's emotions. In addition, by automatically summarizing the past question history and notifying the user, efficient information management and confirmation is possible.

[1951] "Voice input means" refers to a device or interface that allows a user to input a question by voice.

[1952] "Speech recognition means" refers to a technology or module for converting input voice data into text data.

[1953] The "search means" is a technology or module for conducting an internet search based on the converted text data and obtaining appropriate search results.

[1954] "Voice response means" refers to a technology or module for converting search results into voice data and providing it to the user.

[1955] "Audio output means" refers to a device or interface for outputting audio data to a user and conveying information.

[1956] "Character input means" refers to a device or interface that allows a user to input characters.

[1957] A "text response means" is a technology or module for conducting a search based on character input and returning search results to the user in text.

[1958] The "summary generation means" is a technology or module for summarizing the audio and text question history and automatically generating a summary.

[1959] "Notification means" refers to a technology or module for notifying the user of the summary generated results.

[1960] An "emotion engine" is a technology or module for recognizing emotions from a user's voice data.

[1961] The "response content adjustment means" is a technology or module for adjusting the response content of search results based on the emotion information obtained from the emotion engine.

[1962] A "personalized database" is a database that stores each user's account information and search history and provides personalized search results.

[1963] A "cloud service" is a service that uses remote computing resources and storage provided via the Internet.

[1964] The voice response and chat-linked search system of the present invention is composed of multiple means installed in a smart device, which processes users' voice and text questions, provides appropriate responses, and provides more personalized services through emotion recognition.

[1965] This system includes a voice input means for the user to input a question by voice, a voice recognition means for converting voice data into text data, a search means for performing an internet search based on the text data, a voice response means for converting search results into voice data, a voice output means for outputting the voice data to the user, a text input means for inputting a text question, a text response means for returning search results to the user in text, a summary generation means for generating a summary of the voice and text question history, and a notification means for notifying the user of the generated summary. Furthermore, it includes an emotion engine for recognizing emotions from the user's voice data, and a response content adjustment means for adjusting the response content of the search results based on the emotion information. Each part of this system can be implemented using an ordinary smartphone or a cloud server.

[1966] The smartphone's microphone is used as the voice input means. When the user's voice input is recognized as a trigger phrase (e.g., "OK, Assistant"), the voice recognition means sends the voice to a cloud service, where it is converted into text data by a voice recognition engine (e.g., Google Cloud Speech-to-Text API). The search means then searches the converted text data using an internet search API (e.g., Google Custom Search API).

[1967] The search results are converted by the voice response means and notified to the user by voice through the voice output means. In this series of steps, an emotion recognition engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotional state, and the response content adjustment means determines an appropriate response based on this.

[1968] The smartphone keyboard is used as the text input means. The user types and sends a question in text. For example, if the user types "What's the weather going to be like tomorrow?", the question is sent to the server via a messaging API (e.g., Twilio API), which then performs an internet search. The retrieved search results are returned to the user in text via the text response means.

[1969] The question history is stored in a question history database. The summary generation means summarizes this history and notifies the user as necessary. For example, if a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?", the summary generation means generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy" and notifies the user via the notification means.

[1970] Specific examples

[1971] Examples of voice questions:

[1972] User: "OK, Assistant, what's the weather like right now?"

[1973] 1. The device captures the audio data and sends it to the server.

[1974] 2. The server performs speech recognition processing and converts the speech into text data such as "What's the weather like now?"

[1975] 3. The server sends the query "What's the weather like today?" to the search API and retrieves the search results.

[1976] 4. The server simultaneously sends the voice data to the emotion engine to recognize the user's emotion.

[1977] 5. The server generates a response text based on the emotion information: "Currently, the weather in Tokyo is sunny."

[1978] 6. The server converts "Currently, the weather in Tokyo is sunny." into voice data.

[1979] 7. The device will play a voice message saying, "Currently, the weather in Tokyo is sunny."

[1980] Examples of chat questions:

[1981] User (asking in chat app): "What's the weather like tomorrow?"

[1982] 1. The server receives a text question through the message API.

[1983] 2. The server uses the search API based on this question to send a search query for "tomorrow's weather" and retrieves the search results.

[1984] 3. The server generates the text "The weather in Tokyo will be cloudy tomorrow." and replies in the chat.

[1985] Example of a question history summary:

[1986] If a user asks the questions "What's the weather like now?" and "What's the weather like tomorrow?"

[1987] 1. The server sends these search histories to the summary generator.

[1988] 2. The summary generation means summarizes this information and generates a summary sentence such as "Today's weather: sunny, tomorrow's weather: cloudy."

[1989] 3. The server notifies the user of the generated summary via a notification means.

[1990] The system of the present invention allows users to obtain information efficiently and sensitively, and also allows them to review their past question history in one place, significantly improving user convenience and enabling more personalized services.

[1991] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1992] Step 1: Initial Setup

[1993] The server sets up a personalized database that stores user account information and past search history, and also establishes integration with the search API, message API, and emotion engine.

[1994] Input: User information, past search history

[1995] Output: Personalized database, API and engine integration settings

[1996] Step 2: Voice inquiry reception

[1997] The device will enter standby mode and wait for the user to say a trigger phrase, such as "OK, Assistant."

[1998] Input: Trigger phrase

[1999] Output: Activate voice input mode

[2000] Step 3: Voice Input

[2001] When the user utters the trigger phrase, the device switches to voice input mode and records the user's question, for example, "What's the weather like today?"

[2002] Input: User's voice question

[2003] Output: Recorded audio data

[2004] Step 4: Voice Recognition

[2005] The device sends the recorded voice data to a cloud service, and the server uses a voice recognition engine to convert the voice data into text data.

[2006] Input: Recorded audio data

[2007] Output: Converted text data

[2008] Step 5: Information Search

[2009] The server analyzes the converted text data and generates an appropriate search query, which it then sends to an Internet search API to retrieve search results.

[2010] Input: Text data

[2011] Output: Search result data

[2012] Step 6: Emotion Recognition

[2013] The server sends the recorded voice data to the emotion engine to recognize the user's emotions.

[2014] Input: Audio data

[2015] Output: Emotional information

[2016] Step 7: Tailor your response

[2017] The server adjusts the response content of the search results based on the emotional information obtained from the emotion engine.

[2018] Input: Emotion information, search result data

[2019] Output: Adjusted response content

[2020] Step 8: Generate voice response

[2021] The server integrates the search results and emotional information to generate a natural language voice response text, converts the text into voice data, and sends it to the device.

[2022] Input: Adjusted response content

[2023] Output: Audio data

[2024] Step 9: Answer by voice

[2025] The terminal plays back the received voice data and provides the user with a voice response.

[2026] Input: Audio data

[2027] Output: A spoken response to the user

[2028] Step 10: Chat with your questions

[2029] A user types a text question into a smartphone and sends it, for example, "What's the weather like tomorrow?" in a chat app.

[2030] Input: User's character question

[2031] output: Character data to be sent

[2032] Step 11: Handling chat questions

[2033] The server receives text questions through the message API, generates search queries based on the questions using the search API, retrieves search results, organizes the results into natural language text, and sends a text reply to the user.

[2034] Input: User's character question

[2035] Output: Search result text

[2036] Step 12: Question History Summary

[2037] The server sends the audio and text question history to the summary generation means, which automatically generates a summary. The summary is then notified to the user via the notification means. The server also adjusts the summary content and notification format based on the emotion information.

[2038] Input: Voice question history, text question history

[2039] Output: Summary, Notification

[2040] (Application example 2)

[2041] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2042] In today's world, it's important to provide users with a way to quickly and efficiently obtain the information they need. However, traditional voice response and chat-based search systems are unable to adjust responses based on the user's emotional state, which can result in a poor user experience. Furthermore, they are unable to summarize multiple query histories and notify users, making it difficult to easily review past search results. Another problem is the inability to provide personalized results based on voice or text queries.

[2043] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice input means for a user to input a question by voice, a voice recognition means for converting the voice data into text data, a search means for performing an Internet search based on the text data, a voice response means for converting the search results into voice data, a voice output means for outputting the voice data to the user, a character input means for inputting a text question, a search means for performing an Internet search based on the text question, a text response means for providing the user with a method for replying to the search results of the text question in text form, a summary generation means for generating a summary of the voice and text question history, a notification means for notifying the user of the generated summary results, an emotion recognition means for acquiring emotion information, and a response adjustment means for adapting the search results based on the emotion information. This makes it possible to provide personalized search results according to the user's emotions, improve the quality of the user experience, and make it easy to check past question history.

[2044] The "voice input means" is a means for the user to input a question by voice.

[2045] The "voice recognition means" is a means for converting voice data into text data.

[2046] "Search means" refers to a means for conducting an Internet search based on text data.

[2047] The "voice response means" is a means for converting search results into voice data.

[2048] The "audio output means" is a means for outputting audio data to the user.

[2049] The "text input means" is a means for the user to input a text question.

[2050] The "text response means" is a means for returning search results for a text question to the user in text form.

[2051] The "summary generation means" is a means for generating summaries of the audio and text question history.

[2052] The "notification means" is a means for notifying the user of the results of the generated summary.

[2053] "Emotion recognition means" is a means for acquiring emotion information.

[2054] A "response adjustment means" is a means for adapting search results based on emotional information.

[2055] A "personalized database" is a database that stores a user's question history and customizes search results individually.

[2056] A "cloud service" is a service that processes voice data and text data on an online server and performs voice recognition and search.

[2057] To implement this invention, the following hardware and software are primarily used. A user uses a smart device (e.g., a smartphone) to ask a question by voice or text and obtain search results corresponding to the question. The following describes the functions and processing flow of the system that realizes this application example.

[2058] Required Hardware and Software

[2059] 1. Smart devices (e.g. smartphones, tablets): Provide an interface with users. Devices that allow voice input, voice output, and text input.

[2060] 2. Cloud Services: Online platforms for processing voice and text data. They perform speech recognition, search, and emotion recognition. They use services such as Google Speech-to-Text API, IBM Watson Tone Analyzer, and Amazon Polly.

[2061] 3. Server: Manages the user's personalized database and generates search results. Provides tailored responses based on the user's question history and sentiment information.

[2062] System Operation Overview

[2063] 1. Audio question reception:

[2064] The user speaks a voice trigger phrase (e.g., "OK, Assistant") into the smart device. When the device switches to voice input mode, the user asks, "What movies are recommended today?"

[2065] 2. Speech and Emotion Recognition:

[2066] The smart device sends the voice data to a cloud service, which uses the Google Speech-to-Text API to convert the voice data into text, and simultaneously sends the voice data to the IBM Watson Tone Analyzer to recognize the user's emotions.

[2067] 3. Information Search:

[2068] The server generates a search query based on the converted text data and sends it to the appropriate search API (e.g., a movie database). Based on the user's emotional information, the server can recommend movies.

[2069] 4. Tailor your response:

[2070] The server selects the most appropriate information from the search results based on emotional information and generates a response in natural language. For example, if the user feels like relaxing, it generates a response such as, "The recommended relaxing movie is 'Inception.'"

[2071] 5. Voice response generation:

[2072] The generated response is converted into voice data using Amazon Polly and sent to the smart device, which then outputs the converted voice data to the user and provides a response.

[2073] 6. Question History Summary:

[2074] The server summarizes the user's past question history and periodically notifies the user of the summarized information. The summary information is presented to the user in the form of, for example, "Past 7 days' viewing history: 'Inception', 'The Matrix'."

[2075] Examples and prompts

[2076] Specific examples

[2077] When a user says "What are the best movies tonight?", the following happens:

[2078] 1. The voice data is sent to a cloud service and converted into text data by a speech recognition engine. A query such as "Today's recommended movies" is generated.

[2079] 2. The emotion recognition engine detects the user's desire to relax from their voice.

[2080] 3. The server searches a movie database and applies a filter that recommends "relaxing movies."

[2081] 4. Generate a response sentence: "A recommended relaxing movie is 'Inception'." and convert it into speech.

[2082] 5. The smart device plays this audio to the user.

[2083] Example prompts

[2084] Here are some example prompts for using generative AI models:

[2085] User asks: "What movie do you recommend today?"

[2086] Hardware: Smartphone

[2087] Software Settings:

[2088] Speech Recognition: Google Speech-to-Text API

[2089] Emotion recognition: IBM Watson Tone Analyzer

[2090] Speech synthesis: Amazon Polly

[2091] Specific processing:

[2092] 1. The voice data is sent to the speech recognition engine and converted into text.

[2093] 2. Send the text data to the movie database API and get the search results.

[2094] 3. Send the voice data to the emotion engine and obtain the emotion information.

[2095] 4. Based on the emotional information, the system selects an appropriate movie from the search results and generates a response in natural language.

[2096] 5. The response is converted into audio data and played on the smartphone.

[2097] Expected response: "A movie recommendation to fit your current mood is 'Inception'."

[2098] The system provided by this invention allows users to obtain personalized search results based on their emotions and easily check their past question history. This system improves the quality of the user experience and provides a means to efficiently obtain information.

[2099] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2100] Step 1:

[2101] Voice inquiry reception

[2102] The device enters standby mode and waits for the user to speak a trigger phrase. The input is a voice trigger phrase spoken by the user (e.g., "OK, Assistant"), and the device captures the voice and switches to voice input mode to accept the user's question.

[2103] Step 2:

[2104] Voice input

[2105] The user asks the device, "What movies do you recommend today?" The device records the user's voice and sends the voice data to the cloud service. The input is the recorded voice data, and the output is the voice data sent to the cloud service.

[2106] Step 3:

[2107] Voice Recognition

[2108] The cloud service converts the transmitted voice data into text data using the Google Speech-to-Text API. The input is the voice data, and the output is the text data "What movies do you recommend today?" The cloud service analyzes the voice data and generates the corresponding text.

[2109] Step 4:

[2110] emotion recognition

[2111] The cloud service sends the voice data to the IBM Watson Tone Analyzer along with the text data to recognize the user's emotions. The input is the voice data, and the output is the user's emotional information. The cloud service analyzes the voice data and recognizes the user's emotional state.

[2112] Step 5:

[2113] Information Search

[2114] The server analyzes the text data obtained from speech recognition and generates an appropriate search query. For example, it sends a query such as "Today's recommended movies" to a movie database API and retrieves relevant search results. The input is the converted text data, and the output is a list of search results.

[2115] Step 6:

[2116] Tailoring response content

[2117] The server determines the optimal response based on the search results and emotional information, and generates a response in natural language. For example, if the user is in the mood to relax, it generates the text "A recommended relaxing movie is 'Inception.'" The input is the search results and emotional information, and the output is a response in natural language.

[2118] Step 7:

[2119] Voice Response Generation

[2120] The server converts the generated response text into voice data using Amazon Polly and sends it to the device. The input is text data and the output is voice data. The server analyzes the response text and generates voice data.

[2121] Step 8:

[2122] Answer by voice

[2123] The device plays the received voice data and provides a voice response to the user. The input is the voice data sent from the server, and the output is the voice played from the device. The device analyzes the voice data and conveys information to the user by voice.

[2124] Step 9:

[2125] Question History Summary

[2126] The server sends the user's past question history to the summary generation means, which then summarizes the question history. For example, it generates a summary result such as "Past 7 days' viewing history: 'Inception', 'The Matrix'." The input is the past question history data, and the output is summarized text data. The server analyzes the question history and generates summary information.

[2127] Step 10:

[2128] Summary Notification

[2129] The server notifies the user of the generated summarized question history via a notification means. The input is the summarized text data, and the output is notification information for the user. The server analyzes the summarized data and notifies the user in an appropriate format.

[2130] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2131] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2132] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2133] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2134] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2135] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2136] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2137] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2138] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2139] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2140] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2141] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2142] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2143] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2144] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2145] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2146] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2147] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2148] Furthermore, the hardware structure of these various processo...

Claims

1. A voice input means for allowing a user to input a question by voice; a speech recognition means for converting speech data into text data; A search means for performing an internet search based on text data; a voice response means for converting the search results into voice data; audio output means for outputting audio data to a user; a character input means for inputting a text question; a search means for performing an internet search based on a text query; a text response means for providing a method for replying to the user in text form with search results for the text question; a summary generating means for generating a summary of the audio and text question history; a notification means for notifying a user of the summary generated result; A system including:

2. 2. The system of claim 1, further comprising means for storing the question history in a personalized database, thereby personalizing search results for each user.

3. The system of claim 1 , wherein the voice data and text data are transmitted to a cloud service and the cloud service is used for voice recognition and search.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A