System

A system using voice input, text conversion, and intent analysis allows visually impaired individuals to interact with the Internet effectively, addressing the limitations of existing technologies and enhancing their access to information and navigation.

JP2026017904APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118965
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

People with visual impairments face difficulties in fully utilizing the Internet due to insufficient web accessibility guidelines and conventional technologies, which fail to provide satisfactory support.

Method used

A system that includes voice input, conversion to text data, intent analysis, and voice feedback to enable visually impaired individuals to interact with the Internet using voice commands, facilitating actions like obtaining information and recognizing objects.

Benefits of technology

Enables visually impaired individuals to conveniently access and interact with the Internet through voice, improving their ability to obtain necessary information and navigate their surroundings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017904000001_ABST
    Figure 2026017904000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for inputting speech and generating speech data; means for transmitting the speech data to a server; means for converting the speech data to text data; means for parsing the text data and detecting user intent; means for performing an appropriate action based on the user intent; means for converting a result of the action to speech data; and means for transmitting the speech data to a terminal to provide speech feedback.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Currently, people who are blind or visually impaired have difficulty using the Internet to its full potential. To address this issue, various web accessibility guidelines have been established, but they are not sufficient on their own. Furthermore, with the aging of society, the number of people with visual impairments such as cataracts and glaucoma is also increasing. Conventional technologies have difficulty providing satisfactory support to these people. Therefore, the present invention aims to provide a new system that enables people with visual impairments to more easily use the Internet. [Means for solving the problem]

[0005] The present invention provides a system including: a means for inputting voice and generating voice data; a means for transmitting the voice data to a server; a means for converting the voice data into text data; a means for analyzing the text data and detecting a user's intent; a means for performing an appropriate action based on the user's intent; a means for converting the results of the action into voice data; and a means for transmitting the voice data to a terminal and providing voice feedback. This configuration enables visually impaired people to use the Internet solely through voice, making it easier to obtain the information they need. Furthermore, this system can perform a variety of actions, such as obtaining weather information, reading websites aloud, and recognizing surrounding objects, thereby significantly improving the lives of visually impaired people.

[0006] "Voice data" refers to data that has been converted from user speech or other sounds into digital signals and is in a format that can be handled by devices and systems.

[0007] "Text data" refers to data that has been converted from voice data into text information using voice recognition technology.

[0008] "Server" refers to a central computer system that receives, analyzes, and processes data sent from terminals.

[0009] A "terminal" is a device that is directly operated by a user and has the function of inputting voice and transmitting it to a server.

[0010] A "voice recognition engine" refers to software or algorithms used to convert voice data into text data.

[0011] A "natural language processing engine" refers to software or algorithms that analyze text data and understand its intent and meaning.

[0012] An "action" refers to a specific operation or process that a system performs based on a user's intentions or requests.

[0013] A "speech synthesis engine" refers to software or algorithms used to convert text data into speech data.

[0014] "Feedback" refers to the response or result that a system provides to a user in response to a user request, and is usually communicated to the user in the form of audio data. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] This invention is a system that utilizes voice recognition and generative AI to make the Internet easier for people with visual impairments. The system allows users to obtain necessary information or perform specific actions by giving voice commands.

[0037] Server Processing

[0038] The server receives the voice data sent from the device and converts it into text data. The converted text data is analyzed by a natural language processing engine to detect the user's intention. An appropriate action is taken based on the user's intention, and the results are generated as text data. The generated text data is then converted into voice data by a speech synthesis engine and sent to the device.

[0039] Terminal handling

[0040] The terminal has the function of inputting the user's voice and sending the voice data to the server. When the voice data is sent from the server, it is received and played back for the user. With the voice input and playback functions, it provides an interface for the user to interact with the system.

[0041] User operations

[0042] Users operate the system by issuing voice commands to the device. For example, when a user speaks to the device, "Please tell me the current weather," the voice data is sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing engine to detect the user's intent (a query about weather information). The server calls the weather information API to obtain current weather information and converts it into text data. The voice synthesis engine then converts the text data into voice data and sends it to the device. The device plays back the received voice data and tells the user, "The current weather is sunny."

[0043] Specific examples

[0044] For example, if a visually impaired person wants to know the weather information before going out, they can easily obtain the information by using this system. When the user speaks into the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data and obtains the current weather from the weather information API. If the weather information is "sunny," the voice synthesis engine generates voice data saying "The current weather is sunny" and sends it to the device. The device receives this voice data and plays it back to the user. This allows even visually impaired people to obtain the necessary weather information using only their voice.

[0045] As described above, the system of the present invention enables visually impaired people to conveniently use the Internet through a series of processes including voice data recognition, analysis, information acquisition, and voice feedback.

[0046] The processing flow will be explained below.

[0047] Step 1:

[0048] The user speaks into the terminal, "Please tell me the current weather." This voice is recorded by the microphone and saved as voice data.

[0049] Step 2:

[0050] The device sends the recorded voice data to the server. Specifically, the voice data is transferred to the server as packets via the network.

[0051] Step 3:

[0052] The server receives the voice data sent from the device, sends it to a voice analysis engine, and converts it into text data.

[0053] Step 4:

[0054] The server passes the text data converted by the speech analysis engine to a natural language processing engine to detect the user's intent. Specifically, it analyzes the intent, such as "I want weather information."

[0055] Step 5:

[0056] The server calls the appropriate API based on the user's intent. In this example, it calls the weather information API to obtain current weather information.

[0057] Step 6:

[0058] The server stores the acquired weather information in text format and then passes it to a speech synthesis engine, which converts it into voice data.

[0059] Step 7:

[0060] The server transmits the voice data generated by the voice synthesis engine to the terminal. Specifically, the server transfers the voice data to the terminal as packets via a network.

[0061] Step 8:

[0062] The terminal receives the voice data sent from the server and plays the voice data through the speaker, and the user receives voice feedback such as "The current weather is sunny."

[0063] The above is the specific flow of processing when a user inquires about weather information. This series of steps allows even visually impaired users to easily obtain the necessary information using only voice.

[0064] Example 1

[0065] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0066] There is a challenge to reduce the barriers to easy use of the Internet for many people, including the visually impaired, and enable smooth information acquisition using voice. With conventional technologies, visually impaired people often require special devices or skills to access information on the Internet, which poses a major barrier. To solve this challenge, a more intuitive and efficient method of acquiring information using voice input and generative AI models is needed.

[0067] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0068] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and detecting the user's intention, and means for executing an action, including appropriate data acquisition, based on the user's intention, thereby enabling users, including visually impaired people, to intuitively access information on the Internet through voice and efficiently acquire the information they need.

[0069] "Voice data" refers to data obtained by converting the contents of speech uttered by a user into a voice input device into digital format.

[0070] A "server device" is a device whose role is to receive voice data, convert it into text data, analyze the user's intentions, execute appropriate actions, convert the results into voice data, and send it to a terminal device.

[0071] "Text data" refers to data obtained by converting voice data into text format, and is analyzed by the server device.

[0072] "Analysis means" refers to a means for analyzing text data and detecting the user's intent, and can utilize a generative AI model.

[0073] An "action involving data acquisition" is a series of operations that involves acquiring data from an appropriate external source based on the user's intention.

[0074] The "voice input device" is a device that allows a user to input voice, and includes a microphone and the like.

[0075] A "generative AI model" is an artificial intelligence model that uses natural language processing to analyze text data.

[0076] "Voice feedback" refers to feedback in the form of voice that is provided as a response to the user by returning voice data transmitted from the server device to the terminal device.

[0077] A "terminal device" is a device that transmits voice data from a user to a server device, and receives and plays back voice data from the server device.

[0078] The present invention provides a system that enables many people, including the visually impaired, to efficiently obtain information on the Internet using voice. Specific embodiments for implementing this system are described below.

[0079] Server Processing

[0080] The server processes audio data using the following hardware and software:

[0081] Hardware: Regular server equipment (CPU, memory, storage)

[0082] Software: Google Cloud Speech-to-Text API, IBM Watson Speech to Text API, OpenAI GPT-4, Amazon Polly

[0083] When the server receives the voice data sent from the device, it converts the voice data into text data using the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API. The converted text data is analyzed using OpenAI GPT-4 to detect the user's intent. The appropriate data acquisition action is then executed based on the user's intent. For example, to obtain weather information, the weather information is obtained from the OpenWeatherMap API. The acquired data is again converted into voice data using Amazon Polly and sent to the device.

[0084] Terminal handling

[0085] The terminal provides an interface with the user using the following hardware and software.

[0086] Hardware: A terminal device with a microphone, speaker, and internet connection (e.g., a smartphone, Raspberry Pi, etc.)

[0087] Software: Voice input control, voice playback control, server communication interface

[0088] The device captures the user's voice with a microphone and sends the voice data to the server. When the voice data is sent from the server, it receives it and plays it on the speaker. The device plays the role of providing an interface for dialogue with the user.

[0089] User operations

[0090] The user operates the system by issuing voice commands to the device. For example, when the user says, "What is the current weather?", the voice data is sent to the server via the device. The server converts the voice data into text data and analyzes it using a generative AI model. Based on the analysis results, the server obtains weather information, converts it back into voice data, and sends it to the device. The device then plays the received voice data over its speaker, telling the user, "The current weather is sunny."

[0091] Specific examples

[0092] A specific example of the use of this system is when a visually impaired person wants to know the weather information before commuting to work in the morning. When the user speaks into the device, "What is the weather like now?", the voice data is sent to the server. The server converts the voice data into text data using the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API, and analyzes the "weather information" request using OpenAI GPT-4. Using the analysis results, it calls the weather information API to obtain current weather information. For example, if the weather information is "sunny," it uses Amazon Polly to generate voice data saying "The current weather is sunny" and sends it to the device. The device receives this voice data and plays it on its speaker to provide the user with weather information.

[0093] Prompt Sentence Examples

[0094] "Tell me the current weather."

[0095] "Where's the nearest restaurant?"

[0096] "Please tell me your plans for today."

[0097] As described above, this system is designed to enable users to intuitively obtain information through voice, through a series of processes including voice input, voice recognition, analysis of the generative AI model, information acquisition using an API, voice synthesis, and voice feedback.

[0098] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0099] Explanation of the program's processing steps

[0100] Terminal processing steps

[0101] Step 1:

[0102] The terminal uses an audio input device to capture the user's voice as digital audio data.

[0103] Input: User's voice

[0104] Output: Digital audio data

[0105] Specific operation: The device's microphone picks up the user's speech, "What's the weather like today?" and converts it into digital voice data.

[0106] Step 2:

[0107] The terminal transmits the captured audio data to a server via the Internet.

[0108] Input: Digital audio data

[0109] Output: HTTP request to the server

[0110] Specific operation: The device sends the audio data to the server as an HTTP POST request.

[0111] Step 3:

[0112] When the processing result of the audio data is returned from the server, it is received and the audio data is played back.

[0113] Input: Audio data from the server

[0114] Output: Audio playback from speakers

[0115] Specific operation: The device plays the audio data received from the server, "The current weather is sunny," on the speaker.

[0116] Server Processing Steps

[0117] Step 1:

[0118] The server receives the voice data transmitted from the terminal.

[0119] Input: Audio data from the device

[0120] Output: Received audio data

[0121] Specific behavior: The server receives the HTTP request and decodes the audio data.

[0122] Step 2:

[0123] The server calls a speech recognition API to convert the voice data into text data.

[0124] Input: Audio data

[0125] Output: Text data

[0126] Specific operation: The server uses the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API to convert the voice data into text data such as "What is the current weather?"

[0127] Step 3:

[0128] The server passes the converted text data to a generative AI model for analysis.

[0129] Input: Text data

[0130] Output: User intent analysis results

[0131] Specific operation: The server analyzes the text data using OpenAI GPT-4 and detects the user's intent, which is "inquire about weather information."

[0132] Step 4:

[0133] The server performs appropriate data retrieval actions based on the user's intent.

[0134] Input: User intent analysis results

[0135] Output: Data acquisition results

[0136] Specific operation: The server calls the weather information API and obtains the current weather information.

[0137] Step 5:

[0138] The server formats the acquired data as text data.

[0139] Input: Data acquisition results

[0140] Output: Formatted text data

[0141] Specific operation: The server converts the acquired weather information into text data such as "The current weather is sunny."

[0142] Step 6:

[0143] The server calls a speech synthesis API to convert the formatted text data into speech data.

[0144] Input: Formatted text data

[0145] Output: Audio data

[0146] Specific operation: The server uses a speech synthesis engine such as Amazon Polly to convert text data into speech data.

[0147] Step 7:

[0148] The server transmits the audio data to the terminal.

[0149] Input: Audio data

[0150] Output: HTTP response to the device

[0151] Specific operation: The server sends the generated audio data to the terminal as an HTTP response.

[0152] User operation steps

[0153] Step 1:

[0154] The user issues instructions by voice to the terminal.

[0155] Input: The user's intended spoken command

[0156] Output: Audio prompts

[0157] Specific behavior: The user speaks into the microphone, "What's the weather like today?"

[0158] Step 2:

[0159] The user receives audio feedback from the terminal.

[0160] Input: Audio data from the device

[0161] Output: Feedback to the user

[0162] Specific operation: The user listens to the audio played from the device and receives the information, "The current weather is sunny."

[0163] As described above, each processing step works seamlessly together, allowing the user to intuitively obtain the information they need through voice.

[0164] (Application example 1)

[0165] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0166] This project aims to solve the problem of visually impaired people having difficulty obtaining location information for the products and services they need when shopping in physical stores. It also aims to support more efficient shopping and quicker decision-making by obtaining information.

[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0168] In this invention, the server includes means for receiving voice input and generating voice data, means for transmitting the voice data to a server on a network, means for converting the voice data into text data, means for analyzing the text data and detecting a user's intention, means for performing an appropriate action based on the user's intention, means for converting the result of the action into voice data, means for transmitting the voice data to a terminal and providing voice feedback, and means for providing voice guidance of product and service location information within a physical store, thereby enabling visually impaired users to effectively obtain product and service location information and receive voice feedback within a physical store.

[0169] The "means for receiving voice input and generating voice data" refers to a device or system that receives a user's voice as input and converts it into digital voice data.

[0170] The "means for transmitting the voice data to a server on a network" refers to a device or software for transmitting the generated voice data to a server via the Internet.

[0171] The "means for converting the voice data into text data" refers to voice recognition technology or software that converts voice data into corresponding text data.

[0172] The "means for analyzing the text data and detecting the user's intent" refers to a natural language processing engine or algorithm for analyzing the text data and understanding what the user is looking for.

[0173] The "means for executing appropriate actions based on the user's intentions" refers to a system or process that performs specific operations or obtains information based on the analysis results.

[0174] The "means for converting the result of the action into voice data" is a voice synthesis engine that converts text data into voice data in order to verbally notify the user of the result of the executed action.

[0175] The "means for transmitting the voice data to the terminal and providing voice feedback" refers to a device or process for transmitting the generated voice data to the user's terminal and providing voice feedback.

[0176] "Means for providing audio guidance on the location information of products and services within a physical store" is a system for obtaining location information of products and services within a store and providing that information to the user via audio.

[0177] This invention provides a system that enables visually impaired people to obtain location information of products and services in a physical store by voice. The system mainly consists of a voice input device, a server, a voice recognition API, a natural language processing engine, a voice synthesis engine, and a user's terminal.

[0178] Hardware and Software Configuration

[0179] Voice input device: Smart glasses (e.g., "smart glasses" as a general term)

[0180] Server: Cloud server (e.g., "cloud server" as a general term)

[0181] Speech recognition API: An API that converts speech to text (e.g., the generic name "speech recognition API")

[0182] Natural language processing engine: An AI model that analyzes user intent (e.g., the generic term "natural language processing engine")

[0183] Speech synthesis engine: An API that converts text to speech (e.g., a generic term "speech synthesis API")

[0184] Devices: smart glasses, smartphones, etc.

[0185] Data processing and calculation flow

[0186] 1. Voice Input

[0187] The user speaks into the smart glasses, for example, "Where is the sugar?"

[0188] 2. Send

[0189] The smart glasses record this audio data and send it to a cloud server via the network.

[0190] 3. Voice Recognition

[0191] The cloud server uses a speech recognition API to convert the voice data into text data, which generates the text "Where is the sugar?"

[0192] 4. Natural Language Processing

[0193] The generated text data is passed to a natural language processing engine, which analyzes the user's intent and identifies it as "I want to know where the sugar is."

[0194] 5. Obtaining the necessary information

[0195] The server references the store's database and related APIs to obtain the current location of the sugar. For example, the server obtains location information such as "Sugar is on the third shelf on the right side."

[0196] 6. Speech Synthesis

[0197] The acquired location text data is sent to a speech synthesis API and converted into voice data such as "The sugar is on the third shelf on the right."

[0198] 7. Audio Feedback

[0199] The generated audio data is transmitted to the smart glasses and played as audio feedback to the user.

[0200] Examples of specific examples and prompts

[0201] Specific examples

[0202] A user wears smart glasses in a supermarket and asks aloud, "Where is the sugar?" This question is converted into text data and analyzed on a cloud server. The location information for sugar is retrieved from a product database, and the user is told aloud, "Sugar is on the third shelf on the right."

[0203] Prompt Sentence Examples

[0204] User: "Where's the sugar?"

[0205] Server: "The sugar is on the third shelf on the right."

[0206] This allows visually impaired users to effectively obtain location information for products and services in physical stores and receive voice feedback.

[0207] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0208] Step 1:

[0209] The user speaks to the smart glasses, for example, "Where is the sugar?" This speech is captured by the smart glasses' microphone and processed as digital audio data.

[0210] input:

[0211] User says "Where is the sugar?"

[0212] output:

[0213] Digital audio data

[0214] Specific behavior:

[0215] The user speaks into the smart glasses, and what is said is recorded as audio data.

[0216] Step 2:

[0217] The device processes the recorded audio data and sends it to a cloud server via the network.

[0218] input:

[0219] Digital audio data

[0220] output:

[0221] Audio data sent over the network

[0222] Specific behavior:

[0223] The smart glasses transmit audio data to a cloud server using Wi-Fi or mobile data.

[0224] Step 3:

[0225] The server uses a speech recognition API to convert the received voice data into text data. The speech recognition API analyzes the voice data and generates the text data "Where is the sugar?"

[0226] input:

[0227] Audio data

[0228] output:

[0229] Text data: "Where is the sugar?"

[0230] Specific behavior:

[0231] A speech recognition API analyzes the audio signal and converts it into corresponding text.

[0232] Step 4:

[0233] The server passes the converted text data to a natural language processing engine to analyze the user's intent. The generative AI model analyzes the text data and identifies the user's intent as "I want to know where the sugar is."

[0234] input:

[0235] Text data: "Where is the sugar?"

[0236] output:

[0237] Intent: "I want to know where the sugar is."

[0238] Specific behavior:

[0239] A generative AI model analyzes text data and understands the intent of the user's question.

[0240] Step 5:

[0241] The server executes the appropriate action based on the user's intention. Specifically, it references the store database and related APIs to obtain the sugar's current location information.

[0242] input:

[0243] Intent: "I want to know where the sugar is."

[0244] output:

[0245] Obtained location information: "Sugar is on the third shelf on the right."

[0246] Specific behavior:

[0247] The server accesses the store's database and searches for and identifies the location of the sugar.

[0248] Step 6:

[0249] The server sends the acquired location information to the speech synthesis API, which converts it into voice data. The speech synthesis API then converts the text "Sugar is on the third shelf on the right" into voice data.

[0250] input:

[0251] Text data: "Sugar is on the third shelf on the right."

[0252] output:

[0253] Audio data

[0254] Specific behavior:

[0255] A text-to-speech API converts text to speech.

[0256] Step 7:

[0257] The server sends the generated voice data to the device, which plays it back to the user. The smart glasses receive the voice data and provide voice feedback to the user, such as "The sugar is on the third shelf on the right."

[0258] input:

[0259] Audio data

[0260] output:

[0261] Voice feedback: "Sugar is on the third shelf on the right."

[0262] Specific behavior:

[0263] The smart glasses receive the audio data and output the audio to the user through the speakers.

[0264] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0265] This invention is a system that integrates speech recognition, generative AI, and emotion recognition to enable people with visual impairments to use the Internet more easily. This system is able to understand the user's intentions and emotions and provide appropriate feedback.

[0266] Server Processing

[0267] The server receives the voice data sent from the terminal and converts the voice data into text data. The converted text data is passed to a natural language processing engine for analysis, and the user's intention is detected. An emotion recognition engine is then used to identify the user's emotional state from the analyzed voice data. An appropriate action is taken based on the user's intention and emotion, and the result is generated as text data. The generated text data is then converted into voice data by a speech synthesis engine and transmitted to the terminal.

[0268] Terminal handling

[0269] The terminal has the function of inputting the user's voice and sending the voice data to the server. When the voice data is sent from the server, it is received and played back for the user. With the voice input and playback functions, the terminal provides an interface for the user to interact with the system, and also appropriately adjusts feedback based on the user's emotional state.

[0270] User operations

[0271] Users operate the system by issuing voice commands to their device. For example, if a user says to their device, "What's the weather like now?", the voice data is sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing engine to detect the user's intention. At the same time, an emotion recognition engine identifies the user's emotional state. The server calls a weather information API to obtain current weather information and converts it into text data. A speech synthesis engine converts the text data into voice data and sends it to the device. The device plays back the received voice data and tells the user, "The weather is currently sunny." If the emotion recognition engine identifies the user's emotion as "anxiety," it will adjust the feedback, such as by softening the tone.

[0272] Specific examples

[0273] For example, if a visually impaired person wants to know the weather before going out, they can easily obtain the information by using this system. When the user speaks into the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data to detect the user's intention and uses an emotion recognition engine to identify that the user is feeling "anxious." Here, the server obtains the current weather information from the weather information API, and if the weather is "sunny," converts it into voice data saying, "The current weather is sunny." The voice data is generated in a gentle tone and sent to the device. The device plays this voice data, providing the user with reassuring feedback.

[0274] In this way, the system of the present invention enables visually impaired people to use the Internet conveniently through a series of processes including voice data recognition, analysis, information acquisition, emotion recognition, and voice feedback. Furthermore, emotion recognition makes it possible to provide optimal feedback to users.

[0275] The processing flow will be explained below.

[0276] Step 1:

[0277] The user speaks to the device, saying, "What is the weather like today?" The device records this speech through the microphone and saves it as audio data.

[0278] Step 2:

[0279] The device sends the recorded audio data to the server, which transfers the audio data to the server in the form of network packets.

[0280] Step 3:

[0281] The server receives the voice data sent from the terminal, passes it to a voice recognition engine, and converts it into text data.

[0282] Step 4:

[0283] The server passes the text data to a natural language processing engine to analyze the user's intent, for example, to detect that the user is looking for weather information.

[0284] Step 5:

[0285] The server passes the voice data to an emotion recognition engine to analyze the user's emotional state, where the engine identifies emotions such as "anxiety" or "joy."

[0286] Step 6:

[0287] The server determines the appropriate action based on the user's intention and emotional state, for example, by calling a weather information API to get the current weather information.

[0288] Step 7:

[0289] The acquired weather information is stored in text format on the server, and then the text data is passed to a speech synthesis engine and converted into voice data.

[0290] Step 8:

[0291] The emotion recognition engine will adjust the tone and accent of the voice based on the user's emotional state. For example, if the user is feeling anxious, the voice will be produced in a gentler tone.

[0292] Step 9:

[0293] The server sends the generated audio data to the device, which is transferred in network packets.

[0294] Step 10:

[0295] The device receives the voice data sent from the server and plays it through the speaker. The user receives voice feedback such as "The current weather is sunny." This feedback is provided in an emotionally sensitive tone.

[0296] The above is the specific flow of processing in response to a user's inquiry about weather information in a system that combines an emotion engine. This system accurately recognizes the user's intentions and emotions and provides optimal feedback based on them, greatly improving convenience for people with visual impairments.

[0297] Example 2

[0298] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0299] Voice input and output is essential to enable people with visual impairments to easily use the Internet. However, conventional systems have difficulty accurately understanding the user's intent and grasping the user's emotional state to provide appropriate feedback. In particular, there is a need for improved accuracy in voice recognition and natural language processing, as well as the integration of user emotion recognition. A new system is needed to solve these problems.

[0300] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0301] In this invention, the server includes means for receiving voice input and generating voice data, means for transmitting the voice data to the server, means for converting the voice data into text data, means for analyzing the text data and detecting a user's intention, means for analyzing the voice data based on an emotional state, means for performing an appropriate action based on the user's intention and emotional state, means for converting the results of the action into voice data, and means for transmitting the voice data to a terminal and providing voice feedback, thereby enabling visually impaired users to easily obtain Internet information via voice and receive optimal feedback according to their emotional state.

[0302] "Means for inputting voice and generating voice data" refers to a device or software that takes in a user's speech and converts it into digital voice data.

[0303] "Means for transmitting the voice data to the server" refers to the function of a device or software that transmits collected voice data to a server via a network.

[0304] "Means for converting the voice data into text data" refers to software or algorithms that use voice recognition technology to convert voice data into text information.

[0305] "Means for analyzing the text data and detecting the user's intent" refers to algorithms or software that use natural language processing technology to understand the user's requests and instructions from the text data.

[0306] "Means for analyzing the voice data based on emotional state" refers to algorithms or software that analyzes the characteristics of the voice data to identify the user's emotions.

[0307] "Means for taking appropriate action based on the user's intentions and emotional state" refers to a control device or software that acquires appropriate information and generates responses based on the analysis results.

[0308] The "means for converting the results of the action into voice data" refers to software or a device that converts the acquired information or generated response into voice data using a voice synthesis algorithm.

[0309] The "means for transmitting the voice data to the terminal and providing voice feedback" refers to a function for transmitting the generated voice data to the terminal and providing information to the user by voice.

[0310] This invention is a system that enables people with visual impairments to use the Internet more easily. This system integrates speech recognition, generative AI, and emotion recognition to understand the user's intentions and emotions and provide appropriate feedback. The hardware and software used include a smartphone or voice assistant device, Google API, speech recognition software, a natural language processing engine (e.g., OpenAI's GPT-3), an emotion recognition engine (e.g., IBM Watson), and a speech synthesis engine (e.g., Amazon Polly).

[0311] Server Processing

[0312] The server receives the voice data sent from the device and converts it into text using voice recognition software such as the Google API. The converted text data is then passed to a natural language processing engine to analyze the user's intent. At the same time, the voice data is passed through an emotion recognition engine to identify the user's emotional state. For example, if the voice recognition software recognizes the voice data as "What is the current weather?" and the emotion recognition engine identifies it as "anxious," the server will take appropriate action based on these analysis results. Specifically, it calls a weather information API to obtain current weather information, converts it into text data, and then converts it into voice data using a voice synthesis engine. This voice data is generated in a gentle tone and sent to the device.

[0313] Terminal handling

[0314] The device has the function of inputting the user's voice and sending the voice data to the server. It also has the function of capturing the user's voice through a microphone and sending the voice data to the server. It also has the function of receiving the voice data sent from the server and playing it back to the user. For example, smartphones and voice-assisted devices provide these functions. This allows the user to interact with the system through a voice interface and receive feedback. The feedback is adjusted appropriately based on the user's emotional state.

[0315] User operations

[0316] The user operates the system by issuing voice commands to the device. For example, when the user speaks to the device, "Please tell me the current weather," the voice data is sent from the device to the server. The server analyzes the voice data and identifies the user's intent as "I want to know the weather" and their emotion as "anxiety." Based on this, the server obtains current weather information from a weather information API and generates text data such as "The current weather is sunny." This is converted into voice data using a voice synthesis engine and sent to the device. The device then plays the received voice data, telling the user, "The current weather is sunny." The tone of the feedback is adjusted gently according to the recognized emotional state.

[0317] Specific examples

[0318] For example, if a visually impaired person wants to know the weather before going out, they can easily obtain the information by using this system. When the user speaks to the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data to detect the user's intention, and an emotion recognition engine identifies that the user is feeling "anxious." In this case, the server obtains the current weather information from the weather information API, and if the weather is "sunny," converts it into voice data saying, "The current weather is sunny." The voice data is generated in a gentle tone and sent to the device. The device plays back this voice data, providing feedback that gives the user a sense of security.

[0319] Prompt Sentence Examples

[0320] Examples of prompts to be input to a generative AI model include:

[0321] Weather feedback in a gentle tone used when the user is feeling anxious:

[0322] "The user asks, 'What's the weather like today?' He's feeling anxious. Give him voice feedback in a gentle tone, saying, 'The weather is sunny right now. It's okay.'"

[0323] In this way, the system of the present invention enables visually impaired people to use the Internet conveniently through a series of processes including voice data recognition, analysis, information acquisition, emotion recognition, and voice feedback. Furthermore, emotion recognition makes it possible to provide optimal feedback to users.

[0324] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0325] Step 1:

[0326] The user gives a voice command

[0327] Specific operation: The user speaks to the terminal, "Please tell me the current weather."

[0328] Input: User's voice

[0329] Output: Audio data

[0330] Step 2:

[0331] The device collects voice data and sends it to the server.

[0332] Specific operation: The device collects the user's voice through the built-in microphone and transmits the voice data to the server via the network.

[0333] Input: Audio data

[0334] Output: Audio data sent to the server

[0335] Step 3:

[0336] The server converts the audio data into text data.

[0337] What happens: The server uses Google APIs or other speech recognition software to convert the voice data into text data, for example, "What is the weather like today?"

[0338] Input: Audio data

[0339] Output: Text data

[0340] Step 4:

[0341] Analysis using natural language processing engine and emotion recognition engine

[0342] Specific operation: The server passes the text data to a natural language processing engine (such as OpenAI's GPT-3) to analyze the user's intent. It also passes the voice data to an emotion recognition engine (such as IBM Watson) to identify the user's emotional state. For example, an intent such as "I want to know the current weather" can be distinguished from an emotion such as "anxiety."

[0343] Input: Text data, audio data

[0344] Output: User's intention and emotional state

[0345] Step 5:

[0346] The server gets the information it needs

[0347] Specific operation: The server retrieves the necessary information based on the user's intention. Specifically, it calls the weather information API and retrieves the current weather information, such as "sunny."

[0348] Input: User intent

[0349] Output: Information (e.g., current weather)

[0350] Step 6:

[0351] The server converts the text data into audio data.

[0352] Specific operation: The server passes the acquired information to a speech synthesis engine (such as Amazon Polly) and converts it into voice data. At this time, the tone of the feedback is adjusted according to the user's emotional state as identified by the emotion recognition engine. For example, it may generate a gentle tone saying, "The current weather is sunny."

[0353] Input: Text data, user's emotional state

[0354] Output: Audio data

[0355] Step 7:

[0356] The server sends the audio data to the device.

[0357] Specific operation: The server transmits the generated voice data to the terminal via the network.

[0358] Input: Audio data

[0359] Output: Audio data sent to the device

[0360] Step 8:

[0361] The device plays the audio data.

[0362] Specific operation: The device plays the received voice data through the speaker and provides feedback to the user, for example, "The current weather is sunny" in a gentle tone.

[0363] Input: Audio data

[0364] Output: The audio the user hears

[0365] (Application example 2)

[0366] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0367] Visually impaired users have difficulty obtaining information on the Internet, especially in virtual stores, where it is difficult to access product information. To solve this problem, a system that not only recognizes voice but also understands the user's emotions and provides appropriate feedback is needed.

[0368] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0369] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data to detect the user's intention, and means for identifying the user's emotion from the analyzed voice data, thereby appropriately adjusting voice feedback based on the user's emotional state and enabling visually impaired users to comfortably access product information in a virtual store.

[0370] "Voice data" refers to data that is a digital recording of human speech.

[0371] A "server" is a computer system that provides services to multiple terminals via a network.

[0372] "Text data" is digital data expressed as a string of characters.

[0373] "User intent" refers to the goals and requests analyzed based on the voice data entered by the user.

[0374] "Emotion identification" is the process of identifying a user's psychological state from their voice data.

[0375] An "action" is a specific operation or process that the system executes based on the user's intention.

[0376] "Voice feedback" is a function that provides responses to the user in the form of voice sent from the system.

[0377] A "terminal" is a device that is directly operated by a user, and is hardware that is capable of inputting and outputting audio.

[0378] A "natural language processing engine" is software that analyzes text data to understand its meaning and intent.

[0379] The present invention is a system that helps visually impaired people easily obtain information on the Internet. This system will be specifically described as a virtual store application that provides product information based on the user's voice input.

[0380] System Configuration Overview

[0381] The system mainly consists of the following hardware and software:

[0382] Audio input device (microphone)

[0383] server

[0384] Natural Language Processing Engine

[0385] Emotion Recognition Engine

[0386] Product information API

[0387] Text-to-Speech Engine (TTS)

[0388] Mobile devices (smartphones)

[0389] Server Processing

[0390] The server receives the voice data sent by the user and processes it as follows:

[0391] 1. Use a speech recognition engine to convert voice data into text data.

[0392] 2. Use a natural language processing engine to analyze user intent from text data.

[0393] 3. Identifying the user's emotional state from the analyzed voice data using an emotion identification engine.

[0394] 4. Call the product information API based on the user's intent and obtain the necessary product information.

[0395] 5. Generate the acquired information as text data.

[0396] 6. A speech synthesis engine is used to convert text data into speech data in order to generate speech feedback according to the emotional state.

[0397] 7. Send the audio data to the mobile device.

[0398] Mobile device processing

[0399] The mobile device works in conjunction with the voice input device to:

[0400] 1. The user's voice instructions are input through a microphone and voice data is generated.

[0401] 2. Send the audio data to the server.

[0402] 3. Receives the audio data sent from the server and plays it back to the user.

[0403] 4. Provide emotionally appropriate audio feedback.

[0404] User operations

[0405] Users operate the system by voice and obtain product information. For example, a user might say, "Please tell me about my new smartphone," and the voice data is sent to the server. The server converts the voice data into text data and uses a natural language processing engine to analyze the user's intention. An emotion recognition engine identifies the user's emotion and obtains smartphone information from the product information API. The obtained information is then converted into text data, and a voice synthesis engine is used to generate gentle-toned voice data, which is played back to the user.

[0406] Examples of specific examples and prompts

[0407] Examples:

[0408] When a user says, "Please tell me about the new smartphone," the system analyzes the voice data and understands that the user is looking for smartphone information. If the emotion recognition engine detects stress in the user's voice, the voice feedback is provided in a calmer tone.

[0409] Example prompt sentence:

[0410] "Please tell me the product information for the new smartphone."

[0411] "Tell me about your recent promotions."

[0412] In this way, the system of the present invention not only provides visually impaired users with voice-based information on the Internet, but also provides feedback that takes into account the user's emotional state.

[0413] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0414] Step 1:

[0415] The user issues a command by voice.

[0416] Specific operation: The user speaks into the microphone, "Please tell me about my new smartphone."

[0417] Input: Audio data

[0418] Output: Audio data

[0419] Step 2:

[0420] Acquire audio data on the device.

[0421] Specific operation: The microphone captures the user's speech as voice data and sends that data to the terminal.

[0422] Input: User's voice

[0423] Output: Audio data

[0424] Step 3:

[0425] The terminal transmits the voice data to the server.

[0426] Specific operation: The voice data acquired by the terminal is sent to the server via the network.

[0427] Input: Audio data

[0428] Output: Audio data sent to the server

[0429] Step 4:

[0430] The server converts the voice data into text data.

[0431] Specific operation: The server uses a speech recognition engine to convert the voice data into text data in a string format.

[0432] Input: Audio data

[0433] Output: Text data

[0434] Step 5:

[0435] Analyze text data and detect user intent.

[0436] Specific operation: The server uses a natural language processing engine to analyze the text data and detect the information the user is looking for (in this case, "smartphone information").

[0437] Input: Text data

[0438] Output: User intent

[0439] Step 6:

[0440] Identify user emotions from voice data.

[0441] Specific Operation: The server uses an emotion identification engine to identify the user's emotional state from the voice data.

[0442] Input: Audio data

[0443] Output: User's emotional state

[0444] Step 7:

[0445] Take appropriate action based on user intent.

[0446] Specific operation: Based on the user's intention (to obtain information from the smartphone), the server calls the product information API to obtain the necessary information.

[0447] Input: User intent, product information API

[0448] Output: Product information

[0449] Step 8:

[0450] Convert the result of the action into audio data.

[0451] Specific operation: The server converts the acquired product information into text data, and then converts the text data into voice data using a voice synthesis engine.

[0452] Input: Product information (text format)

[0453] Output: Audio data

[0454] Step 9:

[0455] Audio data is sent to the terminal to provide audio feedback.

[0456] Specific operation: The server transmits the generated voice data over the network to the terminal, which then plays the voice data. Feedback is provided in the form of tones based on the user's emotional state.

[0457] Input: Voice data, user's emotional state

[0458] Output: Audio data to be played

[0459] These steps create a system that allows users to obtain information solely through voice and receive feedback based on their emotional state at the time.

[0460] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0461] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0462] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0463] [Second embodiment]

[0464] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0465] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0466] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0467] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0468] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0469] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0470] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0471] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0472] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0473] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0474] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0475] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0476] This invention is a system that utilizes voice recognition and generative AI to make the Internet easier for people with visual impairments. The system allows users to obtain necessary information or perform specific actions by giving voice commands.

[0477] Server Processing

[0478] The server receives the voice data sent from the device and converts it into text data. The converted text data is analyzed by a natural language processing engine to detect the user's intention. An appropriate action is taken based on the user's intention, and the results are generated as text data. The generated text data is then converted into voice data by a speech synthesis engine and sent to the device.

[0479] Terminal handling

[0480] The terminal has the function of inputting the user's voice and sending the voice data to the server. When the voice data is sent from the server, it is received and played back for the user. With the voice input and playback functions, it provides an interface for the user to interact with the system.

[0481] User operations

[0482] Users operate the system by issuing voice commands to the device. For example, when a user speaks to the device, "Please tell me the current weather," the voice data is sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing engine to detect the user's intent (a query about weather information). The server calls the weather information API to obtain current weather information and converts it into text data. The voice synthesis engine then converts the text data into voice data and sends it to the device. The device plays back the received voice data and tells the user, "The current weather is sunny."

[0483] Specific examples

[0484] For example, if a visually impaired person wants to know the weather information before going out, they can easily obtain the information by using this system. When the user speaks into the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data and obtains the current weather from the weather information API. If the weather information is "sunny," the voice synthesis engine generates voice data saying "The current weather is sunny" and sends it to the device. The device receives this voice data and plays it back to the user. This allows even visually impaired people to obtain the necessary weather information using only their voice.

[0485] As described above, the system of the present invention enables visually impaired people to conveniently use the Internet through a series of processes including voice data recognition, analysis, information acquisition, and voice feedback.

[0486] The processing flow will be explained below.

[0487] Step 1:

[0488] The user speaks into the terminal, "Please tell me the current weather." This voice is recorded by the microphone and saved as voice data.

[0489] Step 2:

[0490] The device sends the recorded voice data to the server. Specifically, the voice data is transferred to the server as packets via the network.

[0491] Step 3:

[0492] The server receives the voice data sent from the device, sends it to a voice analysis engine, and converts it into text data.

[0493] Step 4:

[0494] The server passes the text data converted by the speech analysis engine to a natural language processing engine to detect the user's intent. Specifically, it analyzes the intent, such as "I want weather information."

[0495] Step 5:

[0496] The server calls the appropriate API based on the user's intent. In this example, it calls the weather information API to obtain current weather information.

[0497] Step 6:

[0498] The server stores the acquired weather information in text format and then passes it to a speech synthesis engine, which converts it into voice data.

[0499] Step 7:

[0500] The server transmits the voice data generated by the voice synthesis engine to the terminal. Specifically, the server transfers the voice data to the terminal as packets via a network.

[0501] Step 8:

[0502] The terminal receives the voice data sent from the server and plays the voice data through the speaker, and the user receives voice feedback such as "The current weather is sunny."

[0503] The above is the specific flow of processing when a user inquires about weather information. This series of steps allows even visually impaired users to easily obtain the necessary information using only voice.

[0504] Example 1

[0505] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0506] There is a challenge to reduce the barriers to easy use of the Internet for many people, including the visually impaired, and enable smooth information acquisition using voice. With conventional technologies, visually impaired people often require special devices or skills to access information on the Internet, which poses a major barrier. To solve this challenge, a more intuitive and efficient method of acquiring information using voice input and generative AI models is needed.

[0507] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0508] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and detecting the user's intention, and means for executing an action, including appropriate data acquisition, based on the user's intention, thereby enabling users, including visually impaired people, to intuitively access information on the Internet through voice and efficiently acquire the information they need.

[0509] "Voice data" refers to data obtained by converting the contents of speech uttered by a user into a voice input device into digital format.

[0510] A "server device" is a device whose role is to receive voice data, convert it into text data, analyze the user's intentions, execute appropriate actions, convert the results into voice data, and send it to a terminal device.

[0511] "Text data" refers to data obtained by converting voice data into text format, and is analyzed by the server device.

[0512] "Analysis means" refers to a means for analyzing text data and detecting the user's intent, and can utilize a generative AI model.

[0513] An "action involving data acquisition" is a series of operations that involves acquiring data from an appropriate external source based on the user's intention.

[0514] The "voice input device" is a device that allows a user to input voice, and includes a microphone and the like.

[0515] A "generative AI model" is an artificial intelligence model that uses natural language processing to analyze text data.

[0516] "Voice feedback" refers to feedback in the form of voice that is provided as a response to the user by returning voice data transmitted from the server device to the terminal device.

[0517] A "terminal device" is a device that transmits voice data from a user to a server device, and receives and plays back voice data from the server device.

[0518] The present invention provides a system that enables many people, including the visually impaired, to efficiently obtain information on the Internet using voice. Specific embodiments for implementing this system are described below.

[0519] Server Processing

[0520] The server processes audio data using the following hardware and software:

[0521] Hardware: Regular server equipment (CPU, memory, storage)

[0522] Software: Google Cloud Speech-to-Text API, IBM Watson Speech to Text API, OpenAI GPT-4, Amazon Polly

[0523] When the server receives the voice data sent from the device, it converts the voice data into text data using the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API. The converted text data is analyzed using OpenAI GPT-4 to detect the user's intent. The appropriate data acquisition action is then executed based on the user's intent. For example, to obtain weather information, the weather information is obtained from the OpenWeatherMap API. The acquired data is again converted into voice data using Amazon Polly and sent to the device.

[0524] Terminal handling

[0525] The terminal provides an interface with the user using the following hardware and software.

[0526] Hardware: A terminal device with a microphone, speaker, and internet connection (e.g., a smartphone, Raspberry Pi, etc.)

[0527] Software: Voice input control, voice playback control, server communication interface

[0528] The device captures the user's voice with a microphone and sends the voice data to the server. When the voice data is sent from the server, it receives it and plays it on the speaker. The device plays the role of providing an interface for dialogue with the user.

[0529] User operations

[0530] The user operates the system by issuing voice commands to the device. For example, when the user says, "What is the current weather?", the voice data is sent to the server via the device. The server converts the voice data into text data and analyzes it using a generative AI model. Based on the analysis results, the server obtains weather information, converts it back into voice data, and sends it to the device. The device then plays the received voice data over its speaker, telling the user, "The current weather is sunny."

[0531] Specific examples

[0532] A specific example of the use of this system is when a visually impaired person wants to know the weather information before commuting to work in the morning. When the user speaks into the device, "What is the weather like now?", the voice data is sent to the server. The server converts the voice data into text data using the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API, and analyzes the "weather information" request using OpenAI GPT-4. Using the analysis results, it calls the weather information API to obtain current weather information. For example, if the weather information is "sunny," it uses Amazon Polly to generate voice data saying "The current weather is sunny" and sends it to the device. The device receives this voice data and plays it on its speaker to provide the user with weather information.

[0533] Prompt Sentence Examples

[0534] "Tell me the current weather."

[0535] "Where's the nearest restaurant?"

[0536] "Please tell me your plans for today."

[0537] As described above, this system is designed to enable users to intuitively obtain information through voice, through a series of processes including voice input, voice recognition, analysis of the generative AI model, information acquisition using an API, voice synthesis, and voice feedback.

[0538] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0539] Explanation of the program's processing steps

[0540] Terminal processing steps

[0541] Step 1:

[0542] The terminal uses an audio input device to capture the user's voice as digital audio data.

[0543] Input: User's voice

[0544] Output: Digital audio data

[0545] Specific operation: The device's microphone picks up the user's speech, "What's the weather like today?" and converts it into digital voice data.

[0546] Step 2:

[0547] The terminal transmits the captured audio data to a server via the Internet.

[0548] Input: Digital audio data

[0549] Output: HTTP request to the server

[0550] Specific operation: The device sends the audio data to the server as an HTTP POST request.

[0551] Step 3:

[0552] When the processing result of the audio data is returned from the server, it is received and the audio data is played back.

[0553] Input: Audio data from the server

[0554] Output: Audio playback from speakers

[0555] Specific operation: The device plays the audio data received from the server, "The current weather is sunny," on the speaker.

[0556] Server Processing Steps

[0557] Step 1:

[0558] The server receives the voice data transmitted from the terminal.

[0559] Input: Audio data from the device

[0560] Output: Received audio data

[0561] Specific behavior: The server receives the HTTP request and decodes the audio data.

[0562] Step 2:

[0563] The server calls a speech recognition API to convert the voice data into text data.

[0564] Input: Audio data

[0565] Output: Text data

[0566] Specific operation: The server uses the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API to convert the voice data into text data such as "What is the current weather?"

[0567] Step 3:

[0568] The server passes the converted text data to a generative AI model for analysis.

[0569] Input: Text data

[0570] Output: User intent analysis results

[0571] Specific operation: The server analyzes the text data using OpenAI GPT-4 and detects the user's intent, which is "inquire about weather information."

[0572] Step 4:

[0573] The server performs appropriate data retrieval actions based on the user's intent.

[0574] Input: User intent analysis results

[0575] Output: Data acquisition results

[0576] Specific operation: The server calls the weather information API and obtains the current weather information.

[0577] Step 5:

[0578] The server formats the acquired data as text data.

[0579] Input: Data acquisition results

[0580] Output: Formatted text data

[0581] Specific operation: The server converts the acquired weather information into text data such as "The current weather is sunny."

[0582] Step 6:

[0583] The server calls a speech synthesis API to convert the formatted text data into speech data.

[0584] Input: Formatted text data

[0585] Output: Audio data

[0586] Specific operation: The server uses a speech synthesis engine such as Amazon Polly to convert text data into speech data.

[0587] Step 7:

[0588] The server transmits the audio data to the terminal.

[0589] Input: Audio data

[0590] Output: HTTP response to the device

[0591] Specific operation: The server sends the generated audio data to the terminal as an HTTP response.

[0592] User operation steps

[0593] Step 1:

[0594] The user issues instructions by voice to the terminal.

[0595] Input: The user's intended spoken command

[0596] Output: Audio prompts

[0597] Specific behavior: The user speaks into the microphone, "What's the weather like today?"

[0598] Step 2:

[0599] The user receives audio feedback from the terminal.

[0600] Input: Audio data from the device

[0601] Output: Feedback to the user

[0602] Specific operation: The user listens to the audio played from the device and receives the information, "The current weather is sunny."

[0603] As described above, each processing step works seamlessly together, allowing the user to intuitively obtain the information they need through voice.

[0604] (Application example 1)

[0605] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0606] This project aims to solve the problem of visually impaired people having difficulty obtaining location information for the products and services they need when shopping in physical stores. It also aims to support more efficient shopping and quicker decision-making by obtaining information.

[0607] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0608] In this invention, the server includes means for receiving voice input and generating voice data, means for transmitting the voice data to a server on a network, means for converting the voice data into text data, means for analyzing the text data and detecting a user's intention, means for performing an appropriate action based on the user's intention, means for converting the result of the action into voice data, means for transmitting the voice data to a terminal and providing voice feedback, and means for providing voice guidance of product and service location information within a physical store, thereby enabling visually impaired users to effectively obtain product and service location information and receive voice feedback within a physical store.

[0609] The "means for receiving voice input and generating voice data" refers to a device or system that receives a user's voice as input and converts it into digital voice data.

[0610] The "means for transmitting the voice data to a server on a network" refers to a device or software for transmitting the generated voice data to a server via the Internet.

[0611] The "means for converting the voice data into text data" refers to voice recognition technology or software that converts voice data into corresponding text data.

[0612] The "means for analyzing the text data and detecting the user's intent" refers to a natural language processing engine or algorithm for analyzing the text data and understanding what the user is looking for.

[0613] The "means for executing appropriate actions based on the user's intentions" refers to a system or process that performs specific operations or obtains information based on the analysis results.

[0614] The "means for converting the result of the action into voice data" is a voice synthesis engine that converts text data into voice data in order to verbally notify the user of the result of the executed action.

[0615] The "means for transmitting the voice data to the terminal and providing voice feedback" refers to a device or process for transmitting the generated voice data to the user's terminal and providing voice feedback.

[0616] "Means for providing audio guidance on the location information of products and services within a physical store" is a system for obtaining location information of products and services within a store and providing that information to the user via audio.

[0617] This invention provides a system that enables visually impaired people to obtain location information of products and services in a physical store by voice. The system mainly consists of a voice input device, a server, a voice recognition API, a natural language processing engine, a voice synthesis engine, and a user's terminal.

[0618] Hardware and Software Configuration

[0619] Voice input device: Smart glasses (e.g., "smart glasses" as a general term)

[0620] Server: Cloud server (e.g., "cloud server" as a general term)

[0621] Speech recognition API: An API that converts speech to text (e.g., the generic name "speech recognition API")

[0622] Natural language processing engine: An AI model that analyzes user intent (e.g., the generic term "natural language processing engine")

[0623] Speech synthesis engine: An API that converts text to speech (e.g., a generic term "speech synthesis API")

[0624] Devices: smart glasses, smartphones, etc.

[0625] Data processing and calculation flow

[0626] 1. Voice Input

[0627] The user speaks into the smart glasses, for example, "Where is the sugar?"

[0628] 2. Send

[0629] The smart glasses record this audio data and send it to a cloud server via the network.

[0630] 3. Voice Recognition

[0631] The cloud server uses a speech recognition API to convert the voice data into text data, which generates the text "Where is the sugar?"

[0632] 4. Natural Language Processing

[0633] The generated text data is passed to a natural language processing engine, which analyzes the user's intent and identifies it as "I want to know where the sugar is."

[0634] 5. Obtaining the necessary information

[0635] The server references the store's database and related APIs to obtain the current location of the sugar. For example, the server obtains location information such as "Sugar is on the third shelf on the right side."

[0636] 6. Speech Synthesis

[0637] The acquired location text data is sent to a speech synthesis API and converted into voice data such as "The sugar is on the third shelf on the right."

[0638] 7. Audio Feedback

[0639] The generated audio data is transmitted to the smart glasses and played as audio feedback to the user.

[0640] Examples of specific examples and prompts

[0641] Specific examples

[0642] A user wears smart glasses in a supermarket and asks aloud, "Where is the sugar?" This question is converted into text data and analyzed on a cloud server. The location information for sugar is retrieved from a product database, and the user is told aloud, "Sugar is on the third shelf on the right."

[0643] Prompt Sentence Examples

[0644] User: "Where's the sugar?"

[0645] Server: "The sugar is on the third shelf on the right."

[0646] This allows visually impaired users to effectively obtain location information for products and services in physical stores and receive voice feedback.

[0647] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0648] Step 1:

[0649] The user speaks to the smart glasses, for example, "Where is the sugar?" This speech is captured by the smart glasses' microphone and processed as digital audio data.

[0650] input:

[0651] User says "Where is the sugar?"

[0652] output:

[0653] Digital audio data

[0654] Specific behavior:

[0655] The user speaks into the smart glasses, and what is said is recorded as audio data.

[0656] Step 2:

[0657] The device processes the recorded audio data and sends it to a cloud server via the network.

[0658] input:

[0659] Digital audio data

[0660] output:

[0661] Audio data sent over the network

[0662] Specific behavior:

[0663] The smart glasses transmit audio data to a cloud server using Wi-Fi or mobile data.

[0664] Step 3:

[0665] The server uses a speech recognition API to convert the received voice data into text data. The speech recognition API analyzes the voice data and generates the text data "Where is the sugar?"

[0666] input:

[0667] Audio data

[0668] output:

[0669] Text data: "Where is the sugar?"

[0670] Specific behavior:

[0671] A speech recognition API analyzes the audio signal and converts it into corresponding text.

[0672] Step 4:

[0673] The server passes the converted text data to a natural language processing engine to analyze the user's intent. The generative AI model analyzes the text data and identifies the user's intent as "I want to know where the sugar is."

[0674] input:

[0675] Text data: "Where is the sugar?"

[0676] output:

[0677] Intent: "I want to know where the sugar is."

[0678] Specific behavior:

[0679] A generative AI model analyzes text data and understands the intent of the user's question.

[0680] Step 5:

[0681] The server executes the appropriate action based on the user's intention. Specifically, it references the store database and related APIs to obtain the sugar's current location information.

[0682] input:

[0683] Intent: "I want to know where the sugar is."

[0684] output:

[0685] Obtained location information: "Sugar is on the third shelf on the right."

[0686] Specific behavior:

[0687] The server accesses the store's database and searches for and identifies the location of the sugar.

[0688] Step 6:

[0689] The server sends the acquired location information to the speech synthesis API, which converts it into voice data. The speech synthesis API then converts the text "Sugar is on the third shelf on the right" into voice data.

[0690] input:

[0691] Text data: "Sugar is on the third shelf on the right."

[0692] output:

[0693] Audio data

[0694] Specific behavior:

[0695] A text-to-speech API converts text to speech.

[0696] Step 7:

[0697] The server sends the generated voice data to the device, which plays it back to the user. The smart glasses receive the voice data and provide voice feedback to the user, such as "The sugar is on the third shelf on the right."

[0698] input:

[0699] Audio data

[0700] output:

[0701] Voice feedback: "Sugar is on the third shelf on the right."

[0702] Specific behavior:

[0703] The smart glasses receive the audio data and output the audio to the user through the speakers.

[0704] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0705] This invention is a system that integrates speech recognition, generative AI, and emotion recognition to enable people with visual impairments to use the Internet more easily. This system is able to understand the user's intentions and emotions and provide appropriate feedback.

[0706] Server Processing

[0707] The server receives the voice data sent from the terminal and converts the voice data into text data. The converted text data is passed to a natural language processing engine for analysis, and the user's intention is detected. An emotion recognition engine is then used to identify the user's emotional state from the analyzed voice data. An appropriate action is taken based on the user's intention and emotion, and the result is generated as text data. The generated text data is then converted into voice data by a speech synthesis engine and transmitted to the terminal.

[0708] Terminal handling

[0709] The terminal has the function of inputting the user's voice and sending the voice data to the server. When the voice data is sent from the server, it is received and played back for the user. With the voice input and playback functions, the terminal provides an interface for the user to interact with the system, and also appropriately adjusts feedback based on the user's emotional state.

[0710] User operations

[0711] Users operate the system by issuing voice commands to their device. For example, if a user says to their device, "What's the weather like now?", the voice data is sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing engine to detect the user's intention. At the same time, an emotion recognition engine identifies the user's emotional state. The server calls a weather information API to obtain current weather information and converts it into text data. A speech synthesis engine converts the text data into voice data and sends it to the device. The device plays back the received voice data and tells the user, "The weather is currently sunny." If the emotion recognition engine identifies the user's emotion as "anxiety," it will adjust the feedback, such as by softening the tone.

[0712] Specific examples

[0713] For example, if a visually impaired person wants to know the weather before going out, they can easily obtain the information by using this system. When the user speaks into the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data to detect the user's intention and uses an emotion recognition engine to identify that the user is feeling "anxious." Here, the server obtains the current weather information from the weather information API, and if the weather is "sunny," converts it into voice data saying, "The current weather is sunny." The voice data is generated in a gentle tone and sent to the device. The device plays this voice data, providing the user with reassuring feedback.

[0714] In this way, the system of the present invention enables visually impaired people to use the Internet conveniently through a series of processes including voice data recognition, analysis, information acquisition, emotion recognition, and voice feedback. Furthermore, emotion recognition makes it possible to provide optimal feedback to users.

[0715] The processing flow will be explained below.

[0716] Step 1:

[0717] The user speaks to the device, saying, "What is the weather like today?" The device records this speech through the microphone and saves it as audio data.

[0718] Step 2:

[0719] The device sends the recorded audio data to the server, which transfers the audio data to the server in the form of network packets.

[0720] Step 3:

[0721] The server receives the voice data sent from the terminal, passes it to a voice recognition engine, and converts it into text data.

[0722] Step 4:

[0723] The server passes the text data to a natural language processing engine to analyze the user's intent, for example, to detect that the user is looking for weather information.

[0724] Step 5:

[0725] The server passes the voice data to an emotion recognition engine to analyze the user's emotional state, where the engine identifies emotions such as "anxiety" or "joy."

[0726] Step 6:

[0727] The server determines the appropriate action based on the user's intention and emotional state, for example, by calling a weather information API to get the current weather information.

[0728] Step 7:

[0729] The acquired weather information is stored in text format on the server, and then the text data is passed to a speech synthesis engine and converted into voice data.

[0730] Step 8:

[0731] The emotion recognition engine will adjust the tone and accent of the voice based on the user's emotional state. For example, if the user is feeling anxious, the voice will be produced in a gentler tone.

[0732] Step 9:

[0733] The server sends the generated audio data to the device, which is transferred in network packets.

[0734] Step 10:

[0735] The device receives the voice data sent from the server and plays it through the speaker. The user receives voice feedback such as "The current weather is sunny." This feedback is provided in an emotionally sensitive tone.

[0736] The above is the specific flow of processing in response to a user's inquiry about weather information in a system that combines an emotion engine. This system accurately recognizes the user's intentions and emotions and provides optimal feedback based on them, greatly improving convenience for people with visual impairments.

[0737] Example 2

[0738] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0739] Voice input and output is essential to enable people with visual impairments to easily use the Internet. However, conventional systems have difficulty accurately understanding the user's intent and grasping the user's emotional state to provide appropriate feedback. In particular, there is a need for improved accuracy in voice recognition and natural language processing, as well as the integration of user emotion recognition. A new system is needed to solve these problems.

[0740] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0741] In this invention, the server includes means for receiving voice input and generating voice data, means for transmitting the voice data to the server, means for converting the voice data into text data, means for analyzing the text data and detecting a user's intention, means for analyzing the voice data based on an emotional state, means for performing an appropriate action based on the user's intention and emotional state, means for converting the results of the action into voice data, and means for transmitting the voice data to a terminal and providing voice feedback, thereby enabling visually impaired users to easily obtain Internet information via voice and receive optimal feedback according to their emotional state.

[0742] "Means for inputting voice and generating voice data" refers to a device or software that takes in a user's speech and converts it into digital voice data.

[0743] "Means for transmitting the voice data to the server" refers to the function of a device or software that transmits collected voice data to a server via a network.

[0744] "Means for converting the voice data into text data" refers to software or algorithms that use voice recognition technology to convert voice data into text information.

[0745] "Means for analyzing the text data and detecting the user's intent" refers to algorithms or software that use natural language processing technology to understand the user's requests and instructions from the text data.

[0746] "Means for analyzing the voice data based on emotional state" refers to algorithms or software that analyzes the characteristics of the voice data to identify the user's emotions.

[0747] "Means for taking appropriate action based on the user's intentions and emotional state" refers to a control device or software that acquires appropriate information and generates responses based on the analysis results.

[0748] The "means for converting the results of the action into voice data" refers to software or a device that converts the acquired information or generated response into voice data using a voice synthesis algorithm.

[0749] The "means for transmitting the voice data to the terminal and providing voice feedback" refers to a function for transmitting the generated voice data to the terminal and providing information to the user by voice.

[0750] This invention is a system that enables people with visual impairments to use the Internet more easily. This system integrates speech recognition, generative AI, and emotion recognition to understand the user's intentions and emotions and provide appropriate feedback. The hardware and software used include a smartphone or voice assistant device, Google API, speech recognition software, a natural language processing engine (e.g., OpenAI's GPT-3), an emotion recognition engine (e.g., IBM Watson), and a speech synthesis engine (e.g., Amazon Polly).

[0751] Server Processing

[0752] The server receives the voice data sent from the device and converts it into text using voice recognition software such as the Google API. The converted text data is then passed to a natural language processing engine to analyze the user's intent. At the same time, the voice data is passed through an emotion recognition engine to identify the user's emotional state. For example, if the voice recognition software recognizes the voice data as "What is the current weather?" and the emotion recognition engine identifies it as "anxious," the server will take appropriate action based on these analysis results. Specifically, it calls a weather information API to obtain current weather information, converts it into text data, and then converts it into voice data using a voice synthesis engine. This voice data is generated in a gentle tone and sent to the device.

[0753] Terminal handling

[0754] The device has the function of inputting the user's voice and sending the voice data to the server. It also has the function of capturing the user's voice through a microphone and sending the voice data to the server. It also has the function of receiving the voice data sent from the server and playing it back to the user. For example, smartphones and voice-assisted devices provide these functions. This allows the user to interact with the system through a voice interface and receive feedback. The feedback is adjusted appropriately based on the user's emotional state.

[0755] User operations

[0756] The user operates the system by issuing voice commands to the device. For example, when the user speaks to the device, "Please tell me the current weather," the voice data is sent from the device to the server. The server analyzes the voice data and identifies the user's intent as "I want to know the weather" and their emotion as "anxiety." Based on this, the server obtains current weather information from a weather information API and generates text data such as "The current weather is sunny." This is converted into voice data using a voice synthesis engine and sent to the device. The device then plays the received voice data, telling the user, "The current weather is sunny." The tone of the feedback is adjusted gently according to the recognized emotional state.

[0757] Specific examples

[0758] For example, if a visually impaired person wants to know the weather before going out, they can easily obtain the information by using this system. When the user speaks to the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data to detect the user's intention, and an emotion recognition engine identifies that the user is feeling "anxious." In this case, the server obtains the current weather information from the weather information API, and if the weather is "sunny," converts it into voice data saying, "The current weather is sunny." The voice data is generated in a gentle tone and sent to the device. The device plays back this voice data, providing feedback that gives the user a sense of security.

[0759] Prompt Sentence Examples

[0760] Examples of prompts to be input to a generative AI model include:

[0761] Weather feedback in a gentle tone used when the user is feeling anxious:

[0762] "The user asks, 'What's the weather like today?' He's feeling anxious. Give him voice feedback in a gentle tone, saying, 'The weather is sunny right now. It's okay.'"

[0763] In this way, the system of the present invention enables visually impaired people to use the Internet conveniently through a series of processes including voice data recognition, analysis, information acquisition, emotion recognition, and voice feedback. Furthermore, emotion recognition makes it possible to provide optimal feedback to users.

[0764] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0765] Step 1:

[0766] The user gives a voice command

[0767] Specific operation: The user speaks to the terminal, "Please tell me the current weather."

[0768] Input: User's voice

[0769] Output: Audio data

[0770] Step 2:

[0771] The device collects voice data and sends it to the server.

[0772] Specific operation: The device collects the user's voice through the built-in microphone and transmits the voice data to the server via the network.

[0773] Input: Audio data

[0774] Output: Audio data sent to the server

[0775] Step 3:

[0776] The server converts the audio data into text data.

[0777] What happens: The server uses Google APIs or other speech recognition software to convert the voice data into text data, for example, "What is the weather like today?"

[0778] Input: Audio data

[0779] Output: Text data

[0780] Step 4:

[0781] Analysis using natural language processing engine and emotion recognition engine

[0782] Specific operation: The server passes the text data to a natural language processing engine (such as OpenAI's GPT-3) to analyze the user's intent. It also passes the voice data to an emotion recognition engine (such as IBM Watson) to identify the user's emotional state. For example, an intent such as "I want to know the current weather" can be distinguished from an emotion such as "anxiety."

[0783] Input: Text data, audio data

[0784] Output: User's intention and emotional state

[0785] Step 5:

[0786] The server gets the information it needs

[0787] Specific operation: The server retrieves the necessary information based on the user's intention. Specifically, it calls the weather information API and retrieves the current weather information, such as "sunny."

[0788] Input: User intent

[0789] Output: Information (e.g., current weather)

[0790] Step 6:

[0791] The server converts the text data into audio data.

[0792] Specific operation: The server passes the acquired information to a speech synthesis engine (such as Amazon Polly) and converts it into voice data. At this time, the tone of the feedback is adjusted according to the user's emotional state as identified by the emotion recognition engine. For example, it may generate a gentle tone saying, "The current weather is sunny."

[0793] Input: Text data, user's emotional state

[0794] Output: Audio data

[0795] Step 7:

[0796] The server sends the audio data to the device.

[0797] Specific operation: The server transmits the generated voice data to the terminal via the network.

[0798] Input: Audio data

[0799] Output: Audio data sent to the device

[0800] Step 8:

[0801] The device plays the audio data.

[0802] Specific operation: The device plays the received voice data through the speaker and provides feedback to the user, for example, "The current weather is sunny" in a gentle tone.

[0803] Input: Audio data

[0804] Output: The audio the user hears

[0805] (Application example 2)

[0806] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0807] Visually impaired users have difficulty obtaining information on the Internet, especially in virtual stores, where it is difficult to access product information. To solve this problem, a system that not only recognizes voice but also understands the user's emotions and provides appropriate feedback is needed.

[0808] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0809] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data to detect the user's intention, and means for identifying the user's emotion from the analyzed voice data, thereby appropriately adjusting voice feedback based on the user's emotional state and enabling visually impaired users to comfortably access product information in a virtual store.

[0810] "Voice data" refers to data that is a digital recording of human speech.

[0811] A "server" is a computer system that provides services to multiple terminals via a network.

[0812] "Text data" is digital data expressed as a string of characters.

[0813] "User intent" refers to the goals and requests analyzed based on the voice data entered by the user.

[0814] "Emotion identification" is the process of identifying a user's psychological state from their voice data.

[0815] An "action" is a specific operation or process that the system executes based on the user's intention.

[0816] "Voice feedback" is a function that provides responses to the user in the form of voice sent from the system.

[0817] A "terminal" is a device that is directly operated by a user, and is hardware that is capable of inputting and outputting audio.

[0818] A "natural language processing engine" is software that analyzes text data to understand its meaning and intent.

[0819] The present invention is a system that helps visually impaired people easily obtain information on the Internet. This system will be specifically described as a virtual store application that provides product information based on the user's voice input.

[0820] System Configuration Overview

[0821] The system mainly consists of the following hardware and software:

[0822] Audio input device (microphone)

[0823] server

[0824] Natural Language Processing Engine

[0825] Emotion Recognition Engine

[0826] Product information API

[0827] Text-to-Speech Engine (TTS)

[0828] Mobile devices (smartphones)

[0829] Server Processing

[0830] The server receives the voice data sent by the user and processes it as follows:

[0831] 1. Use a speech recognition engine to convert voice data into text data.

[0832] 2. Use a natural language processing engine to analyze user intent from text data.

[0833] 3. Identifying the user's emotional state from the analyzed voice data using an emotion identification engine.

[0834] 4. Call the product information API based on the user's intent and obtain the necessary product information.

[0835] 5. Generate the acquired information as text data.

[0836] 6. A speech synthesis engine is used to convert text data into speech data in order to generate speech feedback according to the emotional state.

[0837] 7. Send the audio data to the mobile device.

[0838] Mobile device processing

[0839] The mobile device works in conjunction with the voice input device to:

[0840] 1. The user's voice instructions are input through a microphone and voice data is generated.

[0841] 2. Send the audio data to the server.

[0842] 3. Receives the audio data sent from the server and plays it back to the user.

[0843] 4. Provide emotionally appropriate audio feedback.

[0844] User operations

[0845] Users operate the system by voice and obtain product information. For example, a user might say, "Please tell me about my new smartphone," and the voice data is sent to the server. The server converts the voice data into text data and uses a natural language processing engine to analyze the user's intention. An emotion recognition engine identifies the user's emotion and obtains smartphone information from the product information API. The obtained information is then converted into text data, and a voice synthesis engine is used to generate gentle-toned voice data, which is played back to the user.

[0846] Examples of specific examples and prompts

[0847] Examples:

[0848] When a user says, "Please tell me about the new smartphone," the system analyzes the voice data and understands that the user is looking for smartphone information. If the emotion recognition engine detects stress in the user's voice, the voice feedback is provided in a calmer tone.

[0849] Example prompt sentence:

[0850] "Please tell me the product information for the new smartphone."

[0851] "Tell me about your recent promotions."

[0852] In this way, the system of the present invention not only provides visually impaired users with voice-based information on the Internet, but also provides feedback that takes into account the user's emotional state.

[0853] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0854] Step 1:

[0855] The user issues a command by voice.

[0856] Specific operation: The user speaks into the microphone, "Please tell me about my new smartphone."

[0857] Input: Audio data

[0858] Output: Audio data

[0859] Step 2:

[0860] Acquire audio data on the device.

[0861] Specific operation: The microphone captures the user's speech as voice data and sends that data to the terminal.

[0862] Input: User's voice

[0863] Output: Audio data

[0864] Step 3:

[0865] The terminal transmits the voice data to the server.

[0866] Specific operation: The voice data acquired by the terminal is sent to the server via the network.

[0867] Input: Audio data

[0868] Output: Audio data sent to the server

[0869] Step 4:

[0870] The server converts the voice data into text data.

[0871] Specific operation: The server uses a speech recognition engine to convert the voice data into text data in a string format.

[0872] Input: Audio data

[0873] Output: Text data

[0874] Step 5:

[0875] Analyze text data and detect user intent.

[0876] Specific operation: The server uses a natural language processing engine to analyze the text data and detect the information the user is looking for (in this case, "smartphone information").

[0877] Input: Text data

[0878] Output: User intent

[0879] Step 6:

[0880] Identify user emotions from voice data.

[0881] Specific Operation: The server uses an emotion identification engine to identify the user's emotional state from the voice data.

[0882] Input: Audio data

[0883] Output: User's emotional state

[0884] Step 7:

[0885] Take appropriate action based on user intent.

[0886] Specific operation: Based on the user's intention (to obtain information from the smartphone), the server calls the product information API to obtain the necessary information.

[0887] Input: User intent, product information API

[0888] Output: Product information

[0889] Step 8:

[0890] Convert the result of the action into audio data.

[0891] Specific operation: The server converts the acquired product information into text data, and then converts the text data into voice data using a voice synthesis engine.

[0892] Input: Product information (text format)

[0893] Output: Audio data

[0894] Step 9:

[0895] Audio data is sent to the terminal to provide audio feedback.

[0896] Specific operation: The server transmits the generated voice data over the network to the terminal, which then plays the voice data. Feedback is provided in the form of tones based on the user's emotional state.

[0897] Input: Voice data, user's emotional state

[0898] Output: Audio data to be played

[0899] These steps create a system that allows users to obtain information solely through voice and receive feedback based on their emotional state at the time.

[0900] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0901] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0902] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0903] [Third embodiment]

[0904] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0905] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0906] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0907] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0908] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0909] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0910] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0911] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0912] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0913] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0914] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0915] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0916] This invention is a system that utilizes voice recognition and generative AI to make the Internet easier for people with visual impairments. The system allows users to obtain necessary information or perform specific actions by giving voice commands.

[0917] Server Processing

[0918] The server receives the voice data sent from the device and converts it into text data. The converted text data is analyzed by a natural language processing engine to detect the user's intention. An appropriate action is taken based on the user's intention, and the results are generated as text data. The generated text data is then converted into voice data by a speech synthesis engine and sent to the device.

[0919] Terminal handling

[0920] The terminal has the function of inputting the user's voice and sending the voice data to the server. When the voice data is sent from the server, it is received and played back for the user. With the voice input and playback functions, it provides an interface for the user to interact with the system.

[0921] User operations

[0922] Users operate the system by issuing voice commands to the device. For example, when a user speaks to the device, "Please tell me the current weather," the voice data is sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing engine to detect the user's intent (a query about weather information). The server calls the weather information API to obtain current weather information and converts it into text data. The voice synthesis engine then converts the text data into voice data and sends it to the device. The device plays back the received voice data and tells the user, "The current weather is sunny."

[0923] Specific examples

[0924] For example, if a visually impaired person wants to know the weather information before going out, they can easily obtain the information by using this system. When the user speaks into the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data and obtains the current weather from the weather information API. If the weather information is "sunny," the voice synthesis engine generates voice data saying "The current weather is sunny" and sends it to the device. The device receives this voice data and plays it back to the user. This allows even visually impaired people to obtain the necessary weather information using only their voice.

[0925] As described above, the system of the present invention enables visually impaired people to conveniently use the Internet through a series of processes including voice data recognition, analysis, information acquisition, and voice feedback.

[0926] The processing flow will be explained below.

[0927] Step 1:

[0928] The user speaks into the terminal, "Please tell me the current weather." This voice is recorded by the microphone and saved as voice data.

[0929] Step 2:

[0930] The device sends the recorded voice data to the server. Specifically, the voice data is transferred to the server as packets via the network.

[0931] Step 3:

[0932] The server receives the voice data sent from the device, sends it to a voice analysis engine, and converts it into text data.

[0933] Step 4:

[0934] The server passes the text data converted by the speech analysis engine to a natural language processing engine to detect the user's intent. Specifically, it analyzes the intent, such as "I want weather information."

[0935] Step 5:

[0936] The server calls the appropriate API based on the user's intent. In this example, it calls the weather information API to obtain current weather information.

[0937] Step 6:

[0938] The server stores the acquired weather information in text format and then passes it to a speech synthesis engine, which converts it into voice data.

[0939] Step 7:

[0940] The server transmits the voice data generated by the voice synthesis engine to the terminal. Specifically, the server transfers the voice data to the terminal as packets via a network.

[0941] Step 8:

[0942] The terminal receives the voice data sent from the server and plays the voice data through the speaker, and the user receives voice feedback such as "The current weather is sunny."

[0943] The above is the specific flow of processing when a user inquires about weather information. This series of steps allows even visually impaired users to easily obtain the necessary information using only voice.

[0944] Example 1

[0945] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0946] There is a challenge to reduce the barriers to easy use of the Internet for many people, including the visually impaired, and enable smooth information acquisition using voice. With conventional technologies, visually impaired people often require special devices or skills to access information on the Internet, which poses a major barrier. To solve this challenge, a more intuitive and efficient method of acquiring information using voice input and generative AI models is needed.

[0947] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0948] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and detecting the user's intention, and means for executing an action, including appropriate data acquisition, based on the user's intention, thereby enabling users, including visually impaired people, to intuitively access information on the Internet through voice and efficiently acquire the information they need.

[0949] "Voice data" refers to data obtained by converting the contents of speech uttered by a user into a voice input device into digital format.

[0950] A "server device" is a device whose role is to receive voice data, convert it into text data, analyze the user's intentions, execute appropriate actions, convert the results into voice data, and send it to a terminal device.

[0951] "Text data" refers to data obtained by converting voice data into text format, and is analyzed by the server device.

[0952] "Analysis means" refers to a means for analyzing text data and detecting the user's intent, and can utilize a generative AI model.

[0953] An "action involving data acquisition" is a series of operations that involves acquiring data from an appropriate external source based on the user's intention.

[0954] The "voice input device" is a device that allows a user to input voice, and includes a microphone and the like.

[0955] A "generative AI model" is an artificial intelligence model that uses natural language processing to analyze text data.

[0956] "Voice feedback" refers to feedback in the form of voice that is provided as a response to the user by returning voice data transmitted from the server device to the terminal device.

[0957] A "terminal device" is a device that transmits voice data from a user to a server device, and receives and plays back voice data from the server device.

[0958] The present invention provides a system that enables many people, including the visually impaired, to efficiently obtain information on the Internet using voice. Specific embodiments for implementing this system are described below.

[0959] Server Processing

[0960] The server processes audio data using the following hardware and software:

[0961] Hardware: Regular server equipment (CPU, memory, storage)

[0962] Software: Google Cloud Speech-to-Text API, IBM Watson Speech to Text API, OpenAI GPT-4, Amazon Polly

[0963] When the server receives the voice data sent from the device, it converts the voice data into text data using the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API. The converted text data is analyzed using OpenAI GPT-4 to detect the user's intent. The appropriate data acquisition action is then executed based on the user's intent. For example, to obtain weather information, the weather information is obtained from the OpenWeatherMap API. The acquired data is again converted into voice data using Amazon Polly and sent to the device.

[0964] Terminal handling

[0965] The terminal provides an interface with the user using the following hardware and software.

[0966] Hardware: A terminal device with a microphone, speaker, and internet connection (e.g., a smartphone, Raspberry Pi, etc.)

[0967] Software: Voice input control, voice playback control, server communication interface

[0968] The device captures the user's voice with a microphone and sends the voice data to the server. When the voice data is sent from the server, it receives it and plays it on the speaker. The device plays the role of providing an interface for dialogue with the user.

[0969] User operations

[0970] The user operates the system by issuing voice commands to the device. For example, when the user says, "What is the current weather?", the voice data is sent to the server via the device. The server converts the voice data into text data and analyzes it using a generative AI model. Based on the analysis results, the server obtains weather information, converts it back into voice data, and sends it to the device. The device then plays the received voice data over its speaker, telling the user, "The current weather is sunny."

[0971] Specific examples

[0972] A specific example of the use of this system is when a visually impaired person wants to know the weather information before commuting to work in the morning. When the user speaks into the device, "What is the weather like now?", the voice data is sent to the server. The server converts the voice data into text data using the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API, and analyzes the "weather information" request using OpenAI GPT-4. Using the analysis results, it calls the weather information API to obtain current weather information. For example, if the weather information is "sunny," it uses Amazon Polly to generate voice data saying "The current weather is sunny" and sends it to the device. The device receives this voice data and plays it on its speaker to provide the user with weather information.

[0973] Prompt Sentence Examples

[0974] "Tell me the current weather."

[0975] "Where's the nearest restaurant?"

[0976] "Please tell me your plans for today."

[0977] As described above, this system is designed to enable users to intuitively obtain information through voice, through a series of processes including voice input, voice recognition, analysis of the generative AI model, information acquisition using an API, voice synthesis, and voice feedback.

[0978] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0979] Explanation of the program's processing steps

[0980] Terminal processing steps

[0981] Step 1:

[0982] The terminal uses an audio input device to capture the user's voice as digital audio data.

[0983] Input: User's voice

[0984] Output: Digital audio data

[0985] Specific operation: The device's microphone picks up the user's speech, "What's the weather like today?" and converts it into digital voice data.

[0986] Step 2:

[0987] The terminal transmits the captured audio data to a server via the Internet.

[0988] Input: Digital audio data

[0989] Output: HTTP request to the server

[0990] Specific operation: The device sends the audio data to the server as an HTTP POST request.

[0991] Step 3:

[0992] When the processing result of the audio data is returned from the server, it is received and the audio data is played back.

[0993] Input: Audio data from the server

[0994] Output: Audio playback from speakers

[0995] Specific operation: The device plays the audio data received from the server, "The current weather is sunny," on the speaker.

[0996] Server Processing Steps

[0997] Step 1:

[0998] The server receives the voice data transmitted from the terminal.

[0999] Input: Audio data from the device

[1000] Output: Received audio data

[1001] Specific behavior: The server receives the HTTP request and decodes the audio data.

[1002] Step 2:

[1003] The server calls a speech recognition API to convert the voice data into text data.

[1004] Input: Audio data

[1005] Output: Text data

[1006] Specific operation: The server uses the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API to convert the voice data into text data such as "What is the current weather?"

[1007] Step 3:

[1008] The server passes the converted text data to a generative AI model for analysis.

[1009] Input: Text data

[1010] Output: User intent analysis results

[1011] Specific operation: The server analyzes the text data using OpenAI GPT-4 and detects the user's intent, which is "inquire about weather information."

[1012] Step 4:

[1013] The server performs appropriate data retrieval actions based on the user's intent.

[1014] Input: User intent analysis results

[1015] Output: Data acquisition results

[1016] Specific operation: The server calls the weather information API and obtains the current weather information.

[1017] Step 5:

[1018] The server formats the acquired data as text data.

[1019] Input: Data acquisition results

[1020] Output: Formatted text data

[1021] Specific operation: The server converts the acquired weather information into text data such as "The current weather is sunny."

[1022] Step 6:

[1023] The server calls a speech synthesis API to convert the formatted text data into speech data.

[1024] Input: Formatted text data

[1025] Output: Audio data

[1026] Specific operation: The server uses a speech synthesis engine such as Amazon Polly to convert text data into speech data.

[1027] Step 7:

[1028] The server transmits the audio data to the terminal.

[1029] Input: Audio data

[1030] Output: HTTP response to the device

[1031] Specific operation: The server sends the generated audio data to the terminal as an HTTP response.

[1032] User operation steps

[1033] Step 1:

[1034] The user issues instructions by voice to the terminal.

[1035] Input: The user's intended spoken command

[1036] Output: Audio prompts

[1037] Specific behavior: The user speaks into the microphone, "What's the weather like today?"

[1038] Step 2:

[1039] The user receives audio feedback from the terminal.

[1040] Input: Audio data from the device

[1041] Output: Feedback to the user

[1042] Specific operation: The user listens to the audio played from the device and receives the information, "The current weather is sunny."

[1043] As described above, each processing step works seamlessly together, allowing the user to intuitively obtain the information they need through voice.

[1044] (Application example 1)

[1045] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1046] This project aims to solve the problem of visually impaired people having difficulty obtaining location information for the products and services they need when shopping in physical stores. It also aims to support more efficient shopping and quicker decision-making by obtaining information.

[1047] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1048] In this invention, the server includes means for receiving voice input and generating voice data, means for transmitting the voice data to a server on a network, means for converting the voice data into text data, means for analyzing the text data and detecting a user's intention, means for performing an appropriate action based on the user's intention, means for converting the result of the action into voice data, means for transmitting the voice data to a terminal and providing voice feedback, and means for providing voice guidance of product and service location information within a physical store, thereby enabling visually impaired users to effectively obtain product and service location information and receive voice feedback within a physical store.

[1049] The "means for receiving voice input and generating voice data" refers to a device or system that receives a user's voice as input and converts it into digital voice data.

[1050] The "means for transmitting the voice data to a server on a network" refers to a device or software for transmitting the generated voice data to a server via the Internet.

[1051] The "means for converting the voice data into text data" refers to voice recognition technology or software that converts voice data into corresponding text data.

[1052] The "means for analyzing the text data and detecting the user's intent" refers to a natural language processing engine or algorithm for analyzing the text data and understanding what the user is looking for.

[1053] The "means for executing appropriate actions based on the user's intentions" refers to a system or process that performs specific operations or obtains information based on the analysis results.

[1054] The "means for converting the result of the action into voice data" is a voice synthesis engine that converts text data into voice data in order to verbally notify the user of the result of the executed action.

[1055] The "means for transmitting the voice data to the terminal and providing voice feedback" refers to a device or process for transmitting the generated voice data to the user's terminal and providing voice feedback.

[1056] "Means for providing audio guidance on the location information of products and services within a physical store" is a system for obtaining location information of products and services within a store and providing that information to the user via audio.

[1057] This invention provides a system that enables visually impaired people to obtain location information of products and services in a physical store by voice. The system mainly consists of a voice input device, a server, a voice recognition API, a natural language processing engine, a voice synthesis engine, and a user's terminal.

[1058] Hardware and Software Configuration

[1059] Voice input device: Smart glasses (e.g., "smart glasses" as a general term)

[1060] Server: Cloud server (e.g., "cloud server" as a general term)

[1061] Speech recognition API: An API that converts speech to text (e.g., the generic name "speech recognition API")

[1062] Natural language processing engine: An AI model that analyzes user intent (e.g., the generic term "natural language processing engine")

[1063] Speech synthesis engine: An API that converts text to speech (e.g., a generic term "speech synthesis API")

[1064] Devices: smart glasses, smartphones, etc.

[1065] Data processing and calculation flow

[1066] 1. Voice Input

[1067] The user speaks into the smart glasses, for example, "Where is the sugar?"

[1068] 2. Send

[1069] The smart glasses record this audio data and send it to a cloud server via the network.

[1070] 3. Voice Recognition

[1071] The cloud server uses a speech recognition API to convert the voice data into text data, which generates the text "Where is the sugar?"

[1072] 4. Natural Language Processing

[1073] The generated text data is passed to a natural language processing engine, which analyzes the user's intent and identifies it as "I want to know where the sugar is."

[1074] 5. Obtaining the necessary information

[1075] The server references the store's database and related APIs to obtain the current location of the sugar. For example, the server obtains location information such as "Sugar is on the third shelf on the right side."

[1076] 6. Speech Synthesis

[1077] The acquired location text data is sent to a speech synthesis API and converted into voice data such as "The sugar is on the third shelf on the right."

[1078] 7. Audio Feedback

[1079] The generated audio data is transmitted to the smart glasses and played as audio feedback to the user.

[1080] Examples of specific examples and prompts

[1081] Specific examples

[1082] A user wears smart glasses in a supermarket and asks aloud, "Where is the sugar?" This question is converted into text data and analyzed on a cloud server. The location information for sugar is retrieved from a product database, and the user is told aloud, "Sugar is on the third shelf on the right."

[1083] Prompt Sentence Examples

[1084] User: "Where's the sugar?"

[1085] Server: "The sugar is on the third shelf on the right."

[1086] This allows visually impaired users to effectively obtain location information for products and services in physical stores and receive voice feedback.

[1087] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1088] Step 1:

[1089] The user speaks to the smart glasses, for example, "Where is the sugar?" This speech is captured by the smart glasses' microphone and processed as digital audio data.

[1090] input:

[1091] User says "Where is the sugar?"

[1092] output:

[1093] Digital audio data

[1094] Specific behavior:

[1095] The user speaks into the smart glasses, and what is said is recorded as audio data.

[1096] Step 2:

[1097] The device processes the recorded audio data and sends it to a cloud server via the network.

[1098] input:

[1099] Digital audio data

[1100] output:

[1101] Audio data sent over the network

[1102] Specific behavior:

[1103] The smart glasses transmit audio data to a cloud server using Wi-Fi or mobile data.

[1104] Step 3:

[1105] The server uses a speech recognition API to convert the received voice data into text data. The speech recognition API analyzes the voice data and generates the text data "Where is the sugar?"

[1106] input:

[1107] Audio data

[1108] output:

[1109] Text data: "Where is the sugar?"

[1110] Specific behavior:

[1111] A speech recognition API analyzes the audio signal and converts it into corresponding text.

[1112] Step 4:

[1113] The server passes the converted text data to a natural language processing engine to analyze the user's intent. The generative AI model analyzes the text data and identifies the user's intent as "I want to know where the sugar is."

[1114] input:

[1115] Text data: "Where is the sugar?"

[1116] output:

[1117] Intent: "I want to know where the sugar is."

[1118] Specific behavior:

[1119] A generative AI model analyzes text data and understands the intent of the user's question.

[1120] Step 5:

[1121] The server executes the appropriate action based on the user's intention. Specifically, it references the store database and related APIs to obtain the sugar's current location information.

[1122] input:

[1123] Intent: "I want to know where the sugar is."

[1124] output:

[1125] Obtained location information: "Sugar is on the third shelf on the right."

[1126] Specific behavior:

[1127] The server accesses the store's database and searches for and identifies the location of the sugar.

[1128] Step 6:

[1129] The server sends the acquired location information to the speech synthesis API, which converts it into voice data. The speech synthesis API then converts the text "Sugar is on the third shelf on the right" into voice data.

[1130] input:

[1131] Text data: "Sugar is on the third shelf on the right."

[1132] output:

[1133] Audio data

[1134] Specific behavior:

[1135] A text-to-speech API converts text to speech.

[1136] Step 7:

[1137] The server sends the generated voice data to the device, which plays it back to the user. The smart glasses receive the voice data and provide voice feedback to the user, such as "The sugar is on the third shelf on the right."

[1138] input:

[1139] Audio data

[1140] output:

[1141] Voice feedback: "Sugar is on the third shelf on the right."

[1142] Specific behavior:

[1143] The smart glasses receive the audio data and output the audio to the user through the speakers.

[1144] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1145] This invention is a system that integrates speech recognition, generative AI, and emotion recognition to enable people with visual impairments to use the Internet more easily. This system is able to understand the user's intentions and emotions and provide appropriate feedback.

[1146] Server Processing

[1147] The server receives the voice data sent from the terminal and converts the voice data into text data. The converted text data is passed to a natural language processing engine for analysis, and the user's intention is detected. An emotion recognition engine is then used to identify the user's emotional state from the analyzed voice data. An appropriate action is taken based on the user's intention and emotion, and the result is generated as text data. The generated text data is then converted into voice data by a speech synthesis engine and transmitted to the terminal.

[1148] Terminal handling

[1149] The terminal has the function of inputting the user's voice and sending the voice data to the server. When the voice data is sent from the server, it is received and played back for the user. With the voice input and playback functions, the terminal provides an interface for the user to interact with the system, and also appropriately adjusts feedback based on the user's emotional state.

[1150] User operations

[1151] Users operate the system by issuing voice commands to their device. For example, if a user says to their device, "What's the weather like now?", the voice data is sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing engine to detect the user's intention. At the same time, an emotion recognition engine identifies the user's emotional state. The server calls a weather information API to obtain current weather information and converts it into text data. A speech synthesis engine converts the text data into voice data and sends it to the device. The device plays back the received voice data and tells the user, "The weather is currently sunny." If the emotion recognition engine identifies the user's emotion as "anxiety," it will adjust the feedback, such as by softening the tone.

[1152] Specific examples

[1153] For example, if a visually impaired person wants to know the weather before going out, they can easily obtain the information by using this system. When the user speaks into the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data to detect the user's intention and uses an emotion recognition engine to identify that the user is feeling "anxious." Here, the server obtains the current weather information from the weather information API, and if the weather is "sunny," converts it into voice data saying, "The current weather is sunny." The voice data is generated in a gentle tone and sent to the device. The device plays this voice data, providing the user with reassuring feedback.

[1154] In this way, the system of the present invention enables visually impaired people to use the Internet conveniently through a series of processes including voice data recognition, analysis, information acquisition, emotion recognition, and voice feedback. Furthermore, emotion recognition makes it possible to provide optimal feedback to users.

[1155] The processing flow will be explained below.

[1156] Step 1:

[1157] The user speaks to the device, saying, "What is the weather like today?" The device records this speech through the microphone and saves it as audio data.

[1158] Step 2:

[1159] The device sends the recorded audio data to the server, which transfers the audio data to the server in the form of network packets.

[1160] Step 3:

[1161] The server receives the voice data sent from the terminal, passes it to a voice recognition engine, and converts it into text data.

[1162] Step 4:

[1163] The server passes the text data to a natural language processing engine to analyze the user's intent, for example, to detect that the user is looking for weather information.

[1164] Step 5:

[1165] The server passes the voice data to an emotion recognition engine to analyze the user's emotional state, where the engine identifies emotions such as "anxiety" or "joy."

[1166] Step 6:

[1167] The server determines the appropriate action based on the user's intention and emotional state, for example, by calling a weather information API to get the current weather information.

[1168] Step 7:

[1169] The acquired weather information is stored in text format on the server, and then the text data is passed to a speech synthesis engine and converted into voice data.

[1170] Step 8:

[1171] The emotion recognition engine will adjust the tone and accent of the voice based on the user's emotional state. For example, if the user is feeling anxious, the voice will be produced in a gentler tone.

[1172] Step 9:

[1173] The server sends the generated audio data to the device, which is transferred in network packets.

[1174] Step 10:

[1175] The device receives the voice data sent from the server and plays it through the speaker. The user receives voice feedback such as "The current weather is sunny." This feedback is provided in an emotionally sensitive tone.

[1176] The above is the specific flow of processing in response to a user's inquiry about weather information in a system that combines an emotion engine. This system accurately recognizes the user's intentions and emotions and provides optimal feedback based on them, greatly improving convenience for people with visual impairments.

[1177] Example 2

[1178] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1179] Voice input and output is essential to enable people with visual impairments to easily use the Internet. However, conventional systems have difficulty accurately understanding the user's intent and grasping the user's emotional state to provide appropriate feedback. In particular, there is a need for improved accuracy in voice recognition and natural language processing, as well as the integration of user emotion recognition. A new system is needed to solve these problems.

[1180] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1181] In this invention, the server includes means for receiving voice input and generating voice data, means for transmitting the voice data to the server, means for converting the voice data into text data, means for analyzing the text data and detecting a user's intention, means for analyzing the voice data based on an emotional state, means for performing an appropriate action based on the user's intention and emotional state, means for converting the results of the action into voice data, and means for transmitting the voice data to a terminal and providing voice feedback, thereby enabling visually impaired users to easily obtain Internet information via voice and receive optimal feedback according to their emotional state.

[1182] "Means for inputting voice and generating voice data" refers to a device or software that takes in a user's speech and converts it into digital voice data.

[1183] "Means for transmitting the voice data to the server" refers to the function of a device or software that transmits collected voice data to a server via a network.

[1184] "Means for converting the voice data into text data" refers to software or algorithms that use voice recognition technology to convert voice data into text information.

[1185] "Means for analyzing the text data and detecting the user's intent" refers to algorithms or software that use natural language processing technology to understand the user's requests and instructions from the text data.

[1186] "Means for analyzing the voice data based on emotional state" refers to algorithms or software that analyzes the characteristics of the voice data to identify the user's emotions.

[1187] "Means for taking appropriate action based on the user's intentions and emotional state" refers to a control device or software that acquires appropriate information and generates responses based on the analysis results.

[1188] The "means for converting the results of the action into voice data" refers to software or a device that converts the acquired information or generated response into voice data using a voice synthesis algorithm.

[1189] The "means for transmitting the voice data to the terminal and providing voice feedback" refers to a function for transmitting the generated voice data to the terminal and providing information to the user by voice.

[1190] This invention is a system that enables people with visual impairments to use the Internet more easily. This system integrates speech recognition, generative AI, and emotion recognition to understand the user's intentions and emotions and provide appropriate feedback. The hardware and software used include a smartphone or voice assistant device, Google API, speech recognition software, a natural language processing engine (e.g., OpenAI's GPT-3), an emotion recognition engine (e.g., IBM Watson), and a speech synthesis engine (e.g., Amazon Polly).

[1191] Server Processing

[1192] The server receives the voice data sent from the device and converts it into text using voice recognition software such as the Google API. The converted text data is then passed to a natural language processing engine to analyze the user's intent. At the same time, the voice data is passed through an emotion recognition engine to identify the user's emotional state. For example, if the voice recognition software recognizes the voice data as "What is the current weather?" and the emotion recognition engine identifies it as "anxious," the server will take appropriate action based on these analysis results. Specifically, it calls a weather information API to obtain current weather information, converts it into text data, and then converts it into voice data using a voice synthesis engine. This voice data is generated in a gentle tone and sent to the device.

[1193] Terminal handling

[1194] The device has the function of inputting the user's voice and sending the voice data to the server. It also has the function of capturing the user's voice through a microphone and sending the voice data to the server. It also has the function of receiving the voice data sent from the server and playing it back to the user. For example, smartphones and voice-assisted devices provide these functions. This allows the user to interact with the system through a voice interface and receive feedback. The feedback is adjusted appropriately based on the user's emotional state.

[1195] User operations

[1196] The user operates the system by issuing voice commands to the device. For example, when the user speaks to the device, "Please tell me the current weather," the voice data is sent from the device to the server. The server analyzes the voice data and identifies the user's intent as "I want to know the weather" and their emotion as "anxiety." Based on this, the server obtains current weather information from a weather information API and generates text data such as "The current weather is sunny." This is converted into voice data using a voice synthesis engine and sent to the device. The device then plays the received voice data, telling the user, "The current weather is sunny." The tone of the feedback is adjusted gently according to the recognized emotional state.

[1197] Specific examples

[1198] For example, if a visually impaired person wants to know the weather before going out, they can easily obtain the information by using this system. When the user speaks to the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data to detect the user's intention, and an emotion recognition engine identifies that the user is feeling "anxious." In this case, the server obtains the current weather information from the weather information API, and if the weather is "sunny," converts it into voice data saying, "The current weather is sunny." The voice data is generated in a gentle tone and sent to the device. The device plays back this voice data, providing feedback that gives the user a sense of security.

[1199] Prompt Sentence Examples

[1200] Examples of prompts to be input to a generative AI model include:

[1201] Weather feedback in a gentle tone used when the user is feeling anxious:

[1202] "The user asks, 'What's the weather like today?' He's feeling anxious. Give him voice feedback in a gentle tone, saying, 'The weather is sunny right now. It's okay.'"

[1203] In this way, the system of the present invention enables visually impaired people to use the Internet conveniently through a series of processes including voice data recognition, analysis, information acquisition, emotion recognition, and voice feedback. Furthermore, emotion recognition makes it possible to provide optimal feedback to users.

[1204] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1205] Step 1:

[1206] The user gives a voice command

[1207] Specific operation: The user speaks to the terminal, "Please tell me the current weather."

[1208] Input: User's voice

[1209] Output: Audio data

[1210] Step 2:

[1211] The device collects voice data and sends it to the server.

[1212] Specific operation: The device collects the user's voice through the built-in microphone and transmits the voice data to the server via the network.

[1213] Input: Audio data

[1214] Output: Audio data sent to the server

[1215] Step 3:

[1216] The server converts the audio data into text data.

[1217] What happens: The server uses Google APIs or other speech recognition software to convert the voice data into text data, for example, "What is the weather like today?"

[1218] Input: Audio data

[1219] Output: Text data

[1220] Step 4:

[1221] Analysis using natural language processing engine and emotion recognition engine

[1222] Specific operation: The server passes the text data to a natural language processing engine (such as OpenAI's GPT-3) to analyze the user's intent. It also passes the voice data to an emotion recognition engine (such as IBM Watson) to identify the user's emotional state. For example, an intent such as "I want to know the current weather" can be distinguished from an emotion such as "anxiety."

[1223] Input: Text data, audio data

[1224] Output: User's intention and emotional state

[1225] Step 5:

[1226] The server gets the information it needs

[1227] Specific operation: The server retrieves the necessary information based on the user's intention. Specifically, it calls the weather information API and retrieves the current weather information, such as "sunny."

[1228] Input: User intent

[1229] Output: Information (e.g., current weather)

[1230] Step 6:

[1231] The server converts the text data into audio data.

[1232] Specific operation: The server passes the acquired information to a speech synthesis engine (such as Amazon Polly) and converts it into voice data. At this time, the tone of the feedback is adjusted according to the user's emotional state as identified by the emotion recognition engine. For example, it may generate a gentle tone saying, "The current weather is sunny."

[1233] Input: Text data, user's emotional state

[1234] Output: Audio data

[1235] Step 7:

[1236] The server sends the audio data to the device.

[1237] Specific operation: The server transmits the generated voice data to the terminal via the network.

[1238] Input: Audio data

[1239] Output: Audio data sent to the device

[1240] Step 8:

[1241] The device plays the audio data.

[1242] Specific operation: The device plays the received voice data through the speaker and provides feedback to the user, for example, "The current weather is sunny" in a gentle tone.

[1243] Input: Audio data

[1244] Output: The audio the user hears

[1245] (Application example 2)

[1246] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1247] Visually impaired users have difficulty obtaining information on the Internet, especially in virtual stores, where it is difficult to access product information. To solve this problem, a system that not only recognizes voice but also understands the user's emotions and provides appropriate feedback is needed.

[1248] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1249] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data to detect the user's intention, and means for identifying the user's emotion from the analyzed voice data, thereby appropriately adjusting voice feedback based on the user's emotional state and enabling visually impaired users to comfortably access product information in a virtual store.

[1250] "Voice data" refers to data that is a digital recording of human speech.

[1251] A "server" is a computer system that provides services to multiple terminals via a network.

[1252] "Text data" is digital data expressed as a string of characters.

[1253] "User intent" refers to the goals and requests analyzed based on the voice data entered by the user.

[1254] "Emotion identification" is the process of identifying a user's psychological state from their voice data.

[1255] An "action" is a specific operation or process that the system executes based on the user's intention.

[1256] "Voice feedback" is a function that provides responses to the user in the form of voice sent from the system.

[1257] A "terminal" is a device that is directly operated by a user, and is hardware that is capable of inputting and outputting audio.

[1258] A "natural language processing engine" is software that analyzes text data to understand its meaning and intent.

[1259] The present invention is a system that helps visually impaired people easily obtain information on the Internet. This system will be specifically described as a virtual store application that provides product information based on the user's voice input.

[1260] System Configuration Overview

[1261] The system mainly consists of the following hardware and software:

[1262] Audio input device (microphone)

[1263] server

[1264] Natural Language Processing Engine

[1265] Emotion Recognition Engine

[1266] Product information API

[1267] Text-to-Speech Engine (TTS)

[1268] Mobile devices (smartphones)

[1269] Server Processing

[1270] The server receives the voice data sent by the user and processes it as follows:

[1271] 1. Use a speech recognition engine to convert voice data into text data.

[1272] 2. Use a natural language processing engine to analyze user intent from text data.

[1273] 3. Identifying the user's emotional state from the analyzed voice data using an emotion identification engine.

[1274] 4. Call the product information API based on the user's intent and obtain the necessary product information.

[1275] 5. Generate the acquired information as text data.

[1276] 6. A speech synthesis engine is used to convert text data into speech data in order to generate speech feedback according to the emotional state.

[1277] 7. Send the audio data to the mobile device.

[1278] Mobile device processing

[1279] The mobile device works in conjunction with the voice input device to:

[1280] 1. The user's voice instructions are input through a microphone and voice data is generated.

[1281] 2. Send the audio data to the server.

[1282] 3. Receives the audio data sent from the server and plays it back to the user.

[1283] 4. Provide emotionally appropriate audio feedback.

[1284] User operations

[1285] Users operate the system by voice and obtain product information. For example, a user might say, "Please tell me about my new smartphone," and the voice data is sent to the server. The server converts the voice data into text data and uses a natural language processing engine to analyze the user's intention. An emotion recognition engine identifies the user's emotion and obtains smartphone information from the product information API. The obtained information is then converted into text data, and a voice synthesis engine is used to generate gentle-toned voice data, which is played back to the user.

[1286] Examples of specific examples and prompts

[1287] Examples:

[1288] When a user says, "Please tell me about the new smartphone," the system analyzes the voice data and understands that the user is looking for smartphone information. If the emotion recognition engine detects stress in the user's voice, the voice feedback is provided in a calmer tone.

[1289] Example prompt sentence:

[1290] "Please tell me the product information for the new smartphone."

[1291] "Tell me about your recent promotions."

[1292] In this way, the system of the present invention not only provides visually impaired users with voice-based information on the Internet, but also provides feedback that takes into account the user's emotional state.

[1293] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1294] Step 1:

[1295] The user issues a command by voice.

[1296] Specific operation: The user speaks into the microphone, "Please tell me about my new smartphone."

[1297] Input: Audio data

[1298] Output: Audio data

[1299] Step 2:

[1300] Acquire audio data on the device.

[1301] Specific operation: The microphone captures the user's speech as voice data and sends that data to the terminal.

[1302] Input: User's voice

[1303] Output: Audio data

[1304] Step 3:

[1305] The terminal transmits the voice data to the server.

[1306] Specific operation: The voice data acquired by the terminal is sent to the server via the network.

[1307] Input: Audio data

[1308] Output: Audio data sent to the server

[1309] Step 4:

[1310] The server converts the voice data into text data.

[1311] Specific operation: The server uses a speech recognition engine to convert the voice data into text data in a string format.

[1312] Input: Audio data

[1313] Output: Text data

[1314] Step 5:

[1315] Analyze text data and detect user intent.

[1316] Specific operation: The server uses a natural language processing engine to analyze the text data and detect the information the user is looking for (in this case, "smartphone information").

[1317] Input: Text data

[1318] Output: User intent

[1319] Step 6:

[1320] Identify user emotions from voice data.

[1321] Specific Operation: The server uses an emotion identification engine to identify the user's emotional state from the voice data.

[1322] Input: Audio data

[1323] Output: User's emotional state

[1324] Step 7:

[1325] Take appropriate action based on user intent.

[1326] Specific operation: Based on the user's intention (to obtain information from the smartphone), the server calls the product information API to obtain the necessary information.

[1327] Input: User intent, product information API

[1328] Output: Product information

[1329] Step 8:

[1330] Convert the result of the action into audio data.

[1331] Specific operation: The server converts the acquired product information into text data, and then converts the text data into voice data using a voice synthesis engine.

[1332] Input: Product information (text format)

[1333] Output: Audio data

[1334] Step 9:

[1335] Audio data is sent to the terminal to provide audio feedback.

[1336] Specific operation: The server transmits the generated voice data over the network to the terminal, which then plays the voice data. Feedback is provided in the form of tones based on the user's emotional state.

[1337] Input: Voice data, user's emotional state

[1338] Output: Audio data to be played

[1339] These steps create a system that allows users to obtain information solely through voice and receive feedback based on their emotional state at the time.

[1340] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1341] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1342] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1343] [Fourth embodiment]

[1344] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1345] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1346] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1347] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1348] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1349] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1350] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1351] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1352] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1353] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1354] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1355] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1356] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1357] This invention is a system that utilizes voice recognition and generative AI to make the Internet easier for people with visual impairments. The system allows users to obtain necessary information or perform specific actions by giving voice commands.

[1358] Server Processing

[1359] The server receives the voice data sent from the device and converts it into text data. The converted text data is analyzed by a natural language processing engine to detect the user's intention. An appropriate action is taken based on the user's intention, and the results are generated as text data. The generated text data is then converted into voice data by a speech synthesis engine and sent to the device.

[1360] Terminal handling

[1361] The terminal has the function of inputting the user's voice and sending the voice data to the server. When the voice data is sent from the server, it is received and played back for the user. With the voice input and playback functions, it provides an interface for the user to interact with the system.

[1362] User operations

[1363] Users operate the system by issuing voice commands to the device. For example, when a user speaks to the device, "Please tell me the current weather," the voice data is sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing engine to detect the user's intent (a query about weather information). The server calls the weather information API to obtain current weather information and converts it into text data. The voice synthesis engine then converts the text data into voice data and sends it to the device. The device plays back the received voice data and tells the user, "The current weather is sunny."

[1364] Specific examples

[1365] For example, if a visually impaired person wants to know the weather information before going out, they can easily obtain the information by using this system. When the user speaks into the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data and obtains the current weather from the weather information API. If the weather information is "sunny," the voice synthesis engine generates voice data saying "The current weather is sunny" and sends it to the device. The device receives this voice data and plays it back to the user. This allows even visually impaired people to obtain the necessary weather information using only their voice.

[1366] As described above, the system of the present invention enables visually impaired people to conveniently use the Internet through a series of processes including voice data recognition, analysis, information acquisition, and voice feedback.

[1367] The processing flow will be explained below.

[1368] Step 1:

[1369] The user speaks into the terminal, "Please tell me the current weather." This voice is recorded by the microphone and saved as voice data.

[1370] Step 2:

[1371] The device sends the recorded voice data to the server. Specifically, the voice data is transferred to the server as packets via the network.

[1372] Step 3:

[1373] The server receives the voice data sent from the device, sends it to a voice analysis engine, and converts it into text data.

[1374] Step 4:

[1375] The server passes the text data converted by the speech analysis engine to a natural language processing engine to detect the user's intent. Specifically, it analyzes the intent, such as "I want weather information."

[1376] Step 5:

[1377] The server calls the appropriate API based on the user's intent. In this example, it calls the weather information API to obtain current weather information.

[1378] Step 6:

[1379] The server stores the acquired weather information in text format and then passes it to a speech synthesis engine, which converts it into voice data.

[1380] Step 7:

[1381] The server transmits the voice data generated by the voice synthesis engine to the terminal. Specifically, the server transfers the voice data to the terminal as packets via a network.

[1382] Step 8:

[1383] The terminal receives the voice data sent from the server and plays the voice data through the speaker, and the user receives voice feedback such as "The current weather is sunny."

[1384] The above is the specific flow of processing when a user inquires about weather information. This series of steps allows even visually impaired users to easily obtain the necessary information using only voice.

[1385] Example 1

[1386] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1387] There is a challenge to reduce the barriers to easy use of the Internet for many people, including the visually impaired, and enable smooth information acquisition using voice. With conventional technologies, visually impaired people often require special devices or skills to access information on the Internet, which poses a major barrier. To solve this challenge, a more intuitive and efficient method of acquiring information using voice input and generative AI models is needed.

[1388] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1389] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and detecting the user's intention, and means for executing an action, including appropriate data acquisition, based on the user's intention, thereby enabling users, including visually impaired people, to intuitively access information on the Internet through voice and efficiently acquire the information they need.

[1390] "Voice data" refers to data obtained by converting the contents of speech uttered by a user into a voice input device into digital format.

[1391] A "server device" is a device whose role is to receive voice data, convert it into text data, analyze the user's intentions, execute appropriate actions, convert the results into voice data, and send it to a terminal device.

[1392] "Text data" refers to data obtained by converting voice data into text format, and is analyzed by the server device.

[1393] "Analysis means" refers to a means for analyzing text data and detecting the user's intent, and can utilize a generative AI model.

[1394] An "action involving data acquisition" is a series of operations that involves acquiring data from an appropriate external source based on the user's intention.

[1395] The "voice input device" is a device that allows a user to input voice, and includes a microphone and the like.

[1396] A "generative AI model" is an artificial intelligence model that uses natural language processing to analyze text data.

[1397] "Voice feedback" refers to feedback in the form of voice that is provided as a response to the user by returning voice data transmitted from the server device to the terminal device.

[1398] A "terminal device" is a device that transmits voice data from a user to a server device, and receives and plays back voice data from the server device.

[1399] The present invention provides a system that enables many people, including the visually impaired, to efficiently obtain information on the Internet using voice. Specific embodiments for implementing this system are described below.

[1400] Server Processing

[1401] The server processes audio data using the following hardware and software:

[1402] Hardware: Regular server equipment (CPU, memory, storage)

[1403] Software: Google Cloud Speech-to-Text API, IBM Watson Speech to Text API, OpenAI GPT-4, Amazon Polly

[1404] When the server receives the voice data sent from the device, it converts the voice data into text data using the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API. The converted text data is analyzed using OpenAI GPT-4 to detect the user's intent. The appropriate data acquisition action is then executed based on the user's intent. For example, to obtain weather information, the weather information is obtained from the OpenWeatherMap API. The acquired data is again converted into voice data using Amazon Polly and sent to the device.

[1405] Terminal handling

[1406] The terminal provides an interface with the user using the following hardware and software.

[1407] Hardware: A terminal device with a microphone, speaker, and internet connection (e.g., a smartphone, Raspberry Pi, etc.)

[1408] Software: Voice input control, voice playback control, server communication interface

[1409] The device captures the user's voice with a microphone and sends the voice data to the server. When the voice data is sent from the server, it receives it and plays it on the speaker. The device plays the role of providing an interface for dialogue with the user.

[1410] User operations

[1411] The user operates the system by issuing voice commands to the device. For example, when the user says, "What is the current weather?", the voice data is sent to the server via the device. The server converts the voice data into text data and analyzes it using a generative AI model. Based on the analysis results, the server obtains weather information, converts it back into voice data, and sends it to the device. The device then plays the received voice data over its speaker, telling the user, "The current weather is sunny."

[1412] Specific examples

[1413] A specific example of the use of this system is when a visually impaired person wants to know the weather information before commuting to work in the morning. When the user speaks into the device, "What is the weather like now?", the voice data is sent to the server. The server converts the voice data into text data using the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API, and analyzes the "weather information" request using OpenAI GPT-4. Using the analysis results, it calls the weather information API to obtain current weather information. For example, if the weather information is "sunny," it uses Amazon Polly to generate voice data saying "The current weather is sunny" and sends it to the device. The device receives this voice data and plays it on its speaker to provide the user with weather information.

[1414] Prompt Sentence Examples

[1415] "Tell me the current weather."

[1416] "Where's the nearest restaurant?"

[1417] "Please tell me your plans for today."

[1418] As described above, this system is designed to enable users to intuitively obtain information through voice, through a series of processes including voice input, voice recognition, analysis of the generative AI model, information acquisition using an API, voice synthesis, and voice feedback.

[1419] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1420] Explanation of the program's processing steps

[1421] Terminal processing steps

[1422] Step 1:

[1423] The terminal uses an audio input device to capture the user's voice as digital audio data.

[1424] Input: User's voice

[1425] Output: Digital audio data

[1426] Specific operation: The device's microphone picks up the user's speech, "What's the weather like today?" and converts it into digital voice data.

[1427] Step 2:

[1428] The terminal transmits the captured audio data to a server via the Internet.

[1429] Input: Digital audio data

[1430] Output: HTTP request to the server

[1431] Specific operation: The device sends the audio data to the server as an HTTP POST request.

[1432] Step 3:

[1433] When the processing result of the audio data is returned from the server, it is received and the audio data is played back.

[1434] Input: Audio data from the server

[1435] Output: Audio playback from speakers

[1436] Specific operation: The device plays the audio data received from the server, "The current weather is sunny," on the speaker.

[1437] Server Processing Steps

[1438] Step 1:

[1439] The server receives the voice data transmitted from the terminal.

[1440] Input: Audio data from the device

[1441] Output: Received audio data

[1442] Specific behavior: The server receives the HTTP request and decodes the audio data.

[1443] Step 2:

[1444] The server calls a speech recognition API to convert the voice data into text data.

[1445] Input: Audio data

[1446] Output: Text data

[1447] Specific operation: The server uses the Google Cloud Speech-to-Text API or IBM Watson Speech to Text API to convert the voice data into text data such as "What is the current weather?"

[1448] Step 3:

[1449] The server passes the converted text data to a generative AI model for analysis.

[1450] Input: Text data

[1451] Output: User intent analysis results

[1452] Specific operation: The server analyzes the text data using OpenAI GPT-4 and detects the user's intent, which is "inquire about weather information."

[1453] Step 4:

[1454] The server performs appropriate data retrieval actions based on the user's intent.

[1455] Input: User intent analysis results

[1456] Output: Data acquisition results

[1457] Specific operation: The server calls the weather information API and obtains the current weather information.

[1458] Step 5:

[1459] The server formats the acquired data as text data.

[1460] Input: Data acquisition results

[1461] Output: Formatted text data

[1462] Specific operation: The server converts the acquired weather information into text data such as "The current weather is sunny."

[1463] Step 6:

[1464] The server calls a speech synthesis API to convert the formatted text data into speech data.

[1465] Input: Formatted text data

[1466] Output: Audio data

[1467] Specific operation: The server uses a speech synthesis engine such as Amazon Polly to convert text data into speech data.

[1468] Step 7:

[1469] The server transmits the audio data to the terminal.

[1470] Input: Audio data

[1471] Output: HTTP response to the device

[1472] Specific operation: The server sends the generated audio data to the terminal as an HTTP response.

[1473] User operation steps

[1474] Step 1:

[1475] The user issues instructions by voice to the terminal.

[1476] Input: The user's intended spoken command

[1477] Output: Audio prompts

[1478] Specific behavior: The user speaks into the microphone, "What's the weather like today?"

[1479] Step 2:

[1480] The user receives audio feedback from the terminal.

[1481] Input: Audio data from the device

[1482] Output: Feedback to the user

[1483] Specific operation: The user listens to the audio played from the device and receives the information, "The current weather is sunny."

[1484] As described above, each processing step works seamlessly together, allowing the user to intuitively obtain the information they need through voice.

[1485] (Application example 1)

[1486] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1487] This project aims to solve the problem of visually impaired people having difficulty obtaining location information for the products and services they need when shopping in physical stores. It also aims to support more efficient shopping and quicker decision-making by obtaining information.

[1488] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1489] In this invention, the server includes means for receiving voice input and generating voice data, means for transmitting the voice data to a server on a network, means for converting the voice data into text data, means for analyzing the text data and detecting a user's intention, means for performing an appropriate action based on the user's intention, means for converting the result of the action into voice data, means for transmitting the voice data to a terminal and providing voice feedback, and means for providing voice guidance of product and service location information within a physical store, thereby enabling visually impaired users to effectively obtain product and service location information and receive voice feedback within a physical store.

[1490] The "means for receiving voice input and generating voice data" refers to a device or system that receives a user's voice as input and converts it into digital voice data.

[1491] The "means for transmitting the voice data to a server on a network" refers to a device or software for transmitting the generated voice data to a server via the Internet.

[1492] The "means for converting the voice data into text data" refers to voice recognition technology or software that converts voice data into corresponding text data.

[1493] The "means for analyzing the text data and detecting the user's intent" refers to a natural language processing engine or algorithm for analyzing the text data and understanding what the user is looking for.

[1494] The "means for executing appropriate actions based on the user's intentions" refers to a system or process that performs specific operations or obtains information based on the analysis results.

[1495] The "means for converting the result of the action into voice data" is a voice synthesis engine that converts text data into voice data in order to verbally notify the user of the result of the executed action.

[1496] The "means for transmitting the voice data to the terminal and providing voice feedback" refers to a device or process for transmitting the generated voice data to the user's terminal and providing voice feedback.

[1497] "Means for providing audio guidance on the location information of products and services within a physical store" is a system for obtaining location information of products and services within a store and providing that information to the user via audio.

[1498] This invention provides a system that enables visually impaired people to obtain location information of products and services in a physical store by voice. The system mainly consists of a voice input device, a server, a voice recognition API, a natural language processing engine, a voice synthesis engine, and a user's terminal.

[1499] Hardware and Software Configuration

[1500] Voice input device: Smart glasses (e.g., "smart glasses" as a general term)

[1501] Server: Cloud server (e.g., "cloud server" as a general term)

[1502] Speech recognition API: An API that converts speech to text (e.g., the generic name "speech recognition API")

[1503] Natural language processing engine: An AI model that analyzes user intent (e.g., the generic term "natural language processing engine")

[1504] Speech synthesis engine: An API that converts text to speech (e.g., a generic term "speech synthesis API")

[1505] Devices: smart glasses, smartphones, etc.

[1506] Data processing and calculation flow

[1507] 1. Voice Input

[1508] The user speaks into the smart glasses, for example, "Where is the sugar?"

[1509] 2. Send

[1510] The smart glasses record this audio data and send it to a cloud server via the network.

[1511] 3. Voice Recognition

[1512] The cloud server uses a speech recognition API to convert the voice data into text data, which generates the text "Where is the sugar?"

[1513] 4. Natural Language Processing

[1514] The generated text data is passed to a natural language processing engine, which analyzes the user's intent and identifies it as "I want to know where the sugar is."

[1515] 5. Obtaining the necessary information

[1516] The server references the store's database and related APIs to obtain the current location of the sugar. For example, the server obtains location information such as "Sugar is on the third shelf on the right side."

[1517] 6. Speech Synthesis

[1518] The acquired location text data is sent to a speech synthesis API and converted into voice data such as "The sugar is on the third shelf on the right."

[1519] 7. Audio Feedback

[1520] The generated audio data is transmitted to the smart glasses and played as audio feedback to the user.

[1521] Examples of specific examples and prompts

[1522] Specific examples

[1523] A user wears smart glasses in a supermarket and asks aloud, "Where is the sugar?" This question is converted into text data and analyzed on a cloud server. The location information for sugar is retrieved from a product database, and the user is told aloud, "Sugar is on the third shelf on the right."

[1524] Prompt Sentence Examples

[1525] User: "Where's the sugar?"

[1526] Server: "The sugar is on the third shelf on the right."

[1527] This allows visually impaired users to effectively obtain location information for products and services in physical stores and receive voice feedback.

[1528] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1529] Step 1:

[1530] The user speaks to the smart glasses, for example, "Where is the sugar?" This speech is captured by the smart glasses' microphone and processed as digital audio data.

[1531] input:

[1532] User says "Where is the sugar?"

[1533] output:

[1534] Digital audio data

[1535] Specific behavior:

[1536] The user speaks into the smart glasses, and what is said is recorded as audio data.

[1537] Step 2:

[1538] The device processes the recorded audio data and sends it to a cloud server via the network.

[1539] input:

[1540] Digital audio data

[1541] output:

[1542] Audio data sent over the network

[1543] Specific behavior:

[1544] The smart glasses transmit audio data to a cloud server using Wi-Fi or mobile data.

[1545] Step 3:

[1546] The server uses a speech recognition API to convert the received voice data into text data. The speech recognition API analyzes the voice data and generates the text data "Where is the sugar?"

[1547] input:

[1548] Audio data

[1549] output:

[1550] Text data: "Where is the sugar?"

[1551] Specific behavior:

[1552] A speech recognition API analyzes the audio signal and converts it into corresponding text.

[1553] Step 4:

[1554] The server passes the converted text data to a natural language processing engine to analyze the user's intent. The generative AI model analyzes the text data and identifies the user's intent as "I want to know where the sugar is."

[1555] input:

[1556] Text data: "Where is the sugar?"

[1557] output:

[1558] Intent: "I want to know where the sugar is."

[1559] Specific behavior:

[1560] A generative AI model analyzes text data and understands the intent of the user's question.

[1561] Step 5:

[1562] The server executes the appropriate action based on the user's intention. Specifically, it references the store database and related APIs to obtain the sugar's current location information.

[1563] input:

[1564] Intent: "I want to know where the sugar is."

[1565] output:

[1566] Obtained location information: "Sugar is on the third shelf on the right."

[1567] Specific behavior:

[1568] The server accesses the store's database and searches for and identifies the location of the sugar.

[1569] Step 6:

[1570] The server sends the acquired location information to the speech synthesis API, which converts it into voice data. The speech synthesis API then converts the text "Sugar is on the third shelf on the right" into voice data.

[1571] input:

[1572] Text data: "Sugar is on the third shelf on the right."

[1573] output:

[1574] Audio data

[1575] Specific behavior:

[1576] A text-to-speech API converts text to speech.

[1577] Step 7:

[1578] The server sends the generated voice data to the device, which plays it back to the user. The smart glasses receive the voice data and provide voice feedback to the user, such as "The sugar is on the third shelf on the right."

[1579] input:

[1580] Audio data

[1581] output:

[1582] Voice feedback: "Sugar is on the third shelf on the right."

[1583] Specific behavior:

[1584] The smart glasses receive the audio data and output the audio to the user through the speakers.

[1585] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1586] This invention is a system that integrates speech recognition, generative AI, and emotion recognition to enable people with visual impairments to use the Internet more easily. This system is able to understand the user's intentions and emotions and provide appropriate feedback.

[1587] Server Processing

[1588] The server receives the voice data sent from the terminal and converts the voice data into text data. The converted text data is passed to a natural language processing engine for analysis, and the user's intention is detected. An emotion recognition engine is then used to identify the user's emotional state from the analyzed voice data. An appropriate action is taken based on the user's intention and emotion, and the result is generated as text data. The generated text data is then converted into voice data by a speech synthesis engine and transmitted to the terminal.

[1589] Terminal handling

[1590] The terminal has the function of inputting the user's voice and sending the voice data to the server. When the voice data is sent from the server, it is received and played back for the user. With the voice input and playback functions, the terminal provides an interface for the user to interact with the system, and also appropriately adjusts feedback based on the user's emotional state.

[1591] User operations

[1592] Users operate the system by issuing voice commands to their device. For example, if a user says to their device, "What's the weather like now?", the voice data is sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing engine to detect the user's intention. At the same time, an emotion recognition engine identifies the user's emotional state. The server calls a weather information API to obtain current weather information and converts it into text data. A speech synthesis engine converts the text data into voice data and sends it to the device. The device plays back the received voice data and tells the user, "The weather is currently sunny." If the emotion recognition engine identifies the user's emotion as "anxiety," it will adjust the feedback, such as by softening the tone.

[1593] Specific examples

[1594] For example, if a visually impaired person wants to know the weather before going out, they can easily obtain the information by using this system. When the user speaks into the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data to detect the user's intention and uses an emotion recognition engine to identify that the user is feeling "anxious." Here, the server obtains the current weather information from the weather information API, and if the weather is "sunny," converts it into voice data saying, "The current weather is sunny." The voice data is generated in a gentle tone and sent to the device. The device plays this voice data, providing the user with reassuring feedback.

[1595] In this way, the system of the present invention enables visually impaired people to use the Internet conveniently through a series of processes including voice data recognition, analysis, information acquisition, emotion recognition, and voice feedback. Furthermore, emotion recognition makes it possible to provide optimal feedback to users.

[1596] The processing flow will be explained below.

[1597] Step 1:

[1598] The user speaks to the device, saying, "What is the weather like today?" The device records this speech through the microphone and saves it as audio data.

[1599] Step 2:

[1600] The device sends the recorded audio data to the server, which transfers the audio data to the server in the form of network packets.

[1601] Step 3:

[1602] The server receives the voice data sent from the terminal, passes it to a voice recognition engine, and converts it into text data.

[1603] Step 4:

[1604] The server passes the text data to a natural language processing engine to analyze the user's intent, for example, to detect that the user is looking for weather information.

[1605] Step 5:

[1606] The server passes the voice data to an emotion recognition engine to analyze the user's emotional state, where the engine identifies emotions such as "anxiety" or "joy."

[1607] Step 6:

[1608] The server determines the appropriate action based on the user's intention and emotional state, for example, by calling a weather information API to get the current weather information.

[1609] Step 7:

[1610] The acquired weather information is stored in text format on the server, and then the text data is passed to a speech synthesis engine and converted into voice data.

[1611] Step 8:

[1612] The emotion recognition engine will adjust the tone and accent of the voice based on the user's emotional state. For example, if the user is feeling anxious, the voice will be produced in a gentler tone.

[1613] Step 9:

[1614] The server sends the generated audio data to the device, which is transferred in network packets.

[1615] Step 10:

[1616] The device receives the voice data sent from the server and plays it through the speaker. The user receives voice feedback such as "The current weather is sunny." This feedback is provided in an emotionally sensitive tone.

[1617] The above is the specific flow of processing in response to a user's inquiry about weather information in a system that combines an emotion engine. This system accurately recognizes the user's intentions and emotions and provides optimal feedback based on them, greatly improving convenience for people with visual impairments.

[1618] Example 2

[1619] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1620] Voice input and output is essential to enable people with visual impairments to easily use the Internet. However, conventional systems have difficulty accurately understanding the user's intent and grasping the user's emotional state to provide appropriate feedback. In particular, there is a need for improved accuracy in voice recognition and natural language processing, as well as the integration of user emotion recognition. A new system is needed to solve these problems.

[1621] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1622] In this invention, the server includes means for receiving voice input and generating voice data, means for transmitting the voice data to the server, means for converting the voice data into text data, means for analyzing the text data and detecting a user's intention, means for analyzing the voice data based on an emotional state, means for performing an appropriate action based on the user's intention and emotional state, means for converting the results of the action into voice data, and means for transmitting the voice data to a terminal and providing voice feedback, thereby enabling visually impaired users to easily obtain Internet information via voice and receive optimal feedback according to their emotional state.

[1623] "Means for inputting voice and generating voice data" refers to a device or software that takes in a user's speech and converts it into digital voice data.

[1624] "Means for transmitting the voice data to the server" refers to the function of a device or software that transmits collected voice data to a server via a network.

[1625] "Means for converting the voice data into text data" refers to software or algorithms that use voice recognition technology to convert voice data into text information.

[1626] "Means for analyzing the text data and detecting the user's intent" refers to algorithms or software that use natural language processing technology to understand the user's requests and instructions from the text data.

[1627] "Means for analyzing the voice data based on emotional state" refers to algorithms or software that analyzes the characteristics of the voice data to identify the user's emotions.

[1628] "Means for taking appropriate action based on the user's intentions and emotional state" refers to a control device or software that acquires appropriate information and generates responses based on the analysis results.

[1629] The "means for converting the results of the action into voice data" refers to software or a device that converts the acquired information or generated response into voice data using a voice synthesis algorithm.

[1630] The "means for transmitting the voice data to the terminal and providing voice feedback" refers to a function for transmitting the generated voice data to the terminal and providing information to the user by voice.

[1631] This invention is a system that enables people with visual impairments to use the Internet more easily. This system integrates speech recognition, generative AI, and emotion recognition to understand the user's intentions and emotions and provide appropriate feedback. The hardware and software used include a smartphone or voice assistant device, Google API, speech recognition software, a natural language processing engine (e.g., OpenAI's GPT-3), an emotion recognition engine (e.g., IBM Watson), and a speech synthesis engine (e.g., Amazon Polly).

[1632] Server Processing

[1633] The server receives the voice data sent from the device and converts it into text using voice recognition software such as the Google API. The converted text data is then passed to a natural language processing engine to analyze the user's intent. At the same time, the voice data is passed through an emotion recognition engine to identify the user's emotional state. For example, if the voice recognition software recognizes the voice data as "What is the current weather?" and the emotion recognition engine identifies it as "anxious," the server will take appropriate action based on these analysis results. Specifically, it calls a weather information API to obtain current weather information, converts it into text data, and then converts it into voice data using a voice synthesis engine. This voice data is generated in a gentle tone and sent to the device.

[1634] Terminal handling

[1635] The device has the function of inputting the user's voice and sending the voice data to the server. It also has the function of capturing the user's voice through a microphone and sending the voice data to the server. It also has the function of receiving the voice data sent from the server and playing it back to the user. For example, smartphones and voice-assisted devices provide these functions. This allows the user to interact with the system through a voice interface and receive feedback. The feedback is adjusted appropriately based on the user's emotional state.

[1636] User operations

[1637] The user operates the system by issuing voice commands to the device. For example, when the user speaks to the device, "Please tell me the current weather," the voice data is sent from the device to the server. The server analyzes the voice data and identifies the user's intent as "I want to know the weather" and their emotion as "anxiety." Based on this, the server obtains current weather information from a weather information API and generates text data such as "The current weather is sunny." This is converted into voice data using a voice synthesis engine and sent to the device. The device then plays the received voice data, telling the user, "The current weather is sunny." The tone of the feedback is adjusted gently according to the recognized emotional state.

[1638] Specific examples

[1639] For example, if a visually impaired person wants to know the weather before going out, they can easily obtain the information by using this system. When the user speaks to the device, "Please tell me the current weather," the voice data is sent to the server. The server analyzes the voice data to detect the user's intention, and an emotion recognition engine identifies that the user is feeling "anxious." In this case, the server obtains the current weather information from the weather information API, and if the weather is "sunny," converts it into voice data saying, "The current weather is sunny." The voice data is generated in a gentle tone and sent to the device. The device plays back this voice data, providing feedback that gives the user a sense of security.

[1640] Prompt Sentence Examples

[1641] Examples of prompts to be input to a generative AI model include:

[1642] Weather feedback in a gentle tone used when the user is feeling anxious:

[1643] "The user asks, 'What's the weather like today?' He's feeling anxious. Give him voice feedback in a gentle tone, saying, 'The weather is sunny right now. It's okay.'"

[1644] In this way, the system of the present invention enables visually impaired people to use the Internet conveniently through a series of processes including voice data recognition, analysis, information acquisition, emotion recognition, and voice feedback. Furthermore, emotion recognition makes it possible to provide optimal feedback to users.

[1645] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1646] Step 1:

[1647] The user gives a voice command

[1648] Specific operation: The user speaks to the terminal, "Please tell me the current weather."

[1649] Input: User's voice

[1650] Output: Audio data

[1651] Step 2:

[1652] The device collects voice data and sends it to the server.

[1653] Specific operation: The device collects the user's voice through the built-in microphone and transmits the voice data to the server via the network.

[1654] Input: Audio data

[1655] Output: Audio data sent to the server

[1656] Step 3:

[1657] The server converts the audio data into text data.

[1658] What happens: The server uses Google APIs or other speech recognition software to convert the voice data into text data, for example, "What is the weather like today?"

[1659] Input: Audio data

[1660] Output: Text data

[1661] Step 4:

[1662] Analysis using natural language processing engine and emotion recognition engine

[1663] Specific operation: The server passes the text data to a natural language processing engine (such as OpenAI's GPT-3) to analyze the user's intent. It also passes the voice data to an emotion recognition engine (such as IBM Watson) to identify the user's emotional state. For example, an intent such as "I want to know the current weather" can be distinguished from an emotion such as "anxiety."

[1664] Input: Text data, audio data

[1665] Output: User's intention and emotional state

[1666] Step 5:

[1667] The server gets the information it needs

[1668] Specific operation: The server retrieves the necessary information based on the user's intention. Specifically, it calls the weather information API and retrieves the current weather information, such as "sunny."

[1669] Input: User intent

[1670] Output: Information (e.g., current weather)

[1671] Step 6:

[1672] The server converts the text data into audio data.

[1673] Specific operation: The server passes the acquired information to a speech synthesis engine (such as Amazon Polly) and converts it into voice data. At this time, the tone of the feedback is adjusted according to the user's emotional state as identified by the emotion recognition engine. For example, it may generate a gentle tone saying, "The current weather is sunny."

[1674] Input: Text data, user's emotional state

[1675] Output: Audio data

[1676] Step 7:

[1677] The server sends the audio data to the device.

[1678] Specific operation: The server transmits the generated voice data to the terminal via the network.

[1679] Input: Audio data

[1680] Output: Audio data sent to the device

[1681] Step 8:

[1682] The device plays the audio data.

[1683] Specific operation: The device plays the received voice data through the speaker and provides feedback to the user, for example, "The current weather is sunny" in a gentle tone.

[1684] Input: Audio data

[1685] Output: The audio the user hears

[1686] (Application example 2)

[1687] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1688] Visually impaired users have difficulty obtaining information on the Internet, especially in virtual stores, where it is difficult to access product information. To solve this problem, a system that not only recognizes voice but also understands the user's emotions and provides appropriate feedback is needed.

[1689] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1690] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data to detect the user's intention, and means for identifying the user's emotion from the analyzed voice data, thereby appropriately adjusting voice feedback based on the user's emotional state and enabling visually impaired users to comfortably access product information in a virtual store.

[1691] "Voice data" refers to data that is a digital recording of human speech.

[1692] A "server" is a computer system that provides services to multiple terminals via a network.

[1693] "Text data" is digital data expressed as a string of characters.

[1694] "User intent" refers to the goals and requests analyzed based on the voice data entered by the user.

[1695] "Emotion identification" is the process of identifying a user's psychological state from their voice data.

[1696] An "action" is a specific operation or process that the system executes based on the user's intention.

[1697] "Voice feedback" is a function that provides responses to the user in the form of voice sent from the system.

[1698] A "terminal" is a device that is directly operated by a user, and is hardware that is capable of inputting and outputting audio.

[1699] A "natural language processing engine" is software that analyzes text data to understand its meaning and intent.

[1700] The present invention is a system that helps visually impaired people easily obtain information on the Internet. This system will be specifically described as a virtual store application that provides product information based on the user's voice input.

[1701] System Configuration Overview

[1702] The system mainly consists of the following hardware and software:

[1703] Audio input device (microphone)

[1704] server

[1705] Natural Language Processing Engine

[1706] Emotion Recognition Engine

[1707] Product information API

[1708] Text-to-Speech Engine (TTS)

[1709] Mobile devices (smartphones)

[1710] Server Processing

[1711] The server receives the voice data sent by the user and processes it as follows:

[1712] 1. Use a speech recognition engine to convert voice data into text data.

[1713] 2. Use a natural language processing engine to analyze user intent from text data.

[1714] 3. Identifying the user's emotional state from the analyzed voice data using an emotion identification engine.

[1715] 4. Call the product information API based on the user's intent and obtain the necessary product information.

[1716] 5. Generate the acquired information as text data.

[1717] 6. A speech synthesis engine is used to convert text data into speech data in order to generate speech feedback according to the emotional state.

[1718] 7. Send the audio data to the mobile device.

[1719] Mobile device processing

[1720] The mobile device works in conjunction with the voice input device to:

[1721] 1. The user's voice instructions are input through a microphone and voice data is generated.

[1722] 2. Send the audio data to the server.

[1723] 3. Receives the audio data sent from the server and plays it back to the user.

[1724] 4. Provide emotionally appropriate audio feedback.

[1725] User operations

[1726] Users operate the system by voice and obtain product information. For example, a user might say, "Please tell me about my new smartphone," and the voice data is sent to the server. The server converts the voice data into text data and uses a natural language processing engine to analyze the user's intention. An emotion recognition engine identifies the user's emotion and obtains smartphone information from the product information API. The obtained information is then converted into text data, and a voice synthesis engine is used to generate gentle-toned voice data, which is played back to the user.

[1727] Examples of specific examples and prompts

[1728] Examples:

[1729] When a user says, "Please tell me about the new smartphone," the system analyzes the voice data and understands that the user is looking for smartphone information. If the emotion recognition engine detects stress in the user's voice, the voice feedback is provided in a calmer tone.

[1730] Example prompt sentence:

[1731] "Please tell me the product information for the new smartphone."

[1732] "Tell me about your recent promotions."

[1733] In this way, the system of the present invention not only provides visually impaired users with voice-based information on the Internet, but also provides feedback that takes into account the user's emotional state.

[1734] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1735] Step 1:

[1736] The user issues a command by voice.

[1737] Specific operation: The user speaks into the microphone, "Please tell me about my new smartphone."

[1738] Input: Audio data

[1739] Output: Audio data

[1740] Step 2:

[1741] Acquire audio data on the device.

[1742] Specific operation: The microphone captures the user's speech as voice data and sends that data to the terminal.

[1743] Input: User's voice

[1744] Output: Audio data

[1745] Step 3:

[1746] The terminal transmits the voice data to the server.

[1747] Specific operation: The voice data acquired by the terminal is sent to the server via the network.

[1748] Input: Audio data

[1749] Output: Audio data sent to the server

[1750] Step 4:

[1751] The server converts the voice data into text data.

[1752] Specific operation: The server uses a speech recognition engine to convert the voice data into text data in a string format.

[1753] Input: Audio data

[1754] Output: Text data

[1755] Step 5:

[1756] Analyze text data and detect user intent.

[1757] Specific operation: The server uses a natural language processing engine to analyze the text data and detect the information the user is looking for (in this case, "smartphone information").

[1758] Input: Text data

[1759] Output: User intent

[1760] Step 6:

[1761] Identify user emotions from voice data.

[1762] Specific Operation: The server uses an emotion identification engine to identify the user's emotional state from the voice data.

[1763] Input: Audio data

[1764] Output: User's emotional state

[1765] Step 7:

[1766] Take appropriate action based on user intent.

[1767] Specific operation: Based on the user's intention (to obtain information from the smartphone), the server calls the product information API to obtain the necessary information.

[1768] Input: User intent, product information API

[1769] Output: Product information

[1770] Step 8:

[1771] Convert the result of the action into audio data.

[1772] Specific operation: The server converts the acquired product information into text data, and then converts the text data into voice data using a voice synthesis engine.

[1773] Input: Product information (text format)

[1774] Output: Audio data

[1775] Step 9:

[1776] Audio data is sent to the terminal to provide audio feedback.

[1777] Specific operation: The server transmits the generated voice data over the network to the terminal, which then plays the voice data. Feedback is provided in the form of tones based on the user's emotional state.

[1778] Input: Voice data, user's emotional state

[1779] Output: Audio data to be played

[1780] These steps create a system that allows users to obtain information solely through voice and receive feedback based on their emotional state at the time.

[1781] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1782] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1783] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1784] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1785] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1786] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1787] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1788] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1789] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1790] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1791] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1792] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1793] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1794] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1795] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1796] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1797] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1798] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1799] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1800] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1801] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1802] The following is further disclosed regarding the above embodiment.

[1803] (Claim 1)

[1804] A means for receiving voice and generating voice data;

[1805] means for transmitting the voice data to a server;

[1806] means for converting the voice data into text data;

[1807] means for analyzing the text data and detecting a user's intention;

[1808] means for performing appropriate actions based on the user's intent;

[1809] means for converting the results of said actions into audio data;

[1810] means for transmitting the voice data to a terminal and providing voice feedback;

[1811] A system including:

[1812] (Claim 2)

[1813] 2. The system of claim 1, wherein said voice data input means uses a microphone.

[1814] (Claim 3)

[1815] 10. The system of claim 1, wherein the analyzing means uses a natural language processing engine.

[1816] (Claim 4)

[1817] 2. The system according to claim 1, wherein the action execution means calls an API for acquiring weather information.

[1818] (Claim 5)

[1819] 2. The system of claim 1, wherein the voice data conversion means uses a voice synthesis engine.

[1820] "Example 1"

[1821] (Claim 1)

[1822] A means for receiving voice and generating voice data;

[1823] means for transmitting the voice data to a server device;

[1824] means for converting the voice data into text data;

[1825] means for analyzing the text data and detecting a user's intention;

[1826] means for performing actions, including appropriate data acquisition, based on the user's intent;

[1827] means for converting the results of said data acquisition into audio data;

[1828] means for transmitting the voice data to a terminal device to provide voice feedback;

[1829] A system including:

[1830] (Claim 2)

[1831] 2. The system of claim 1, wherein said voice data input means uses a voice input device.

[1832] (Claim 3)

[1833] 10. The system of claim 1, wherein the analysis means uses a generative AI model.

[1834] "Application Example 1"

[1835] (Claim 1)

[1836] A means for receiving voice and generating voice data;

[1837] means for transmitting the voice data to a server on a network;

[1838] means for converting the voice data into text data;

[1839] means for analyzing the text data and detecting a user's intention;

[1840] means for performing appropriate actions based on the user's intent;

[1841] means for converting the results of said actions into audio data;

[1842] means for transmitting the voice data to a terminal and providing voice feedback;

[1843] A means of providing audio location information for products and services within a physical store;

[1844] A system including:

[1845] (Claim 2)

[1846] 2. The system of claim 1, wherein said voice data input means uses a voice input device.

[1847] (Claim 3)

[1848] 10. The system of claim 1, wherein the analysis means uses a generative AI model.

[1849] "Example 2: Combining Emotion Engines"

[1850] (Claim 1)

[1851] A means for receiving voice and generating voice data;

[1852] means for transmitting the voice data to a server;

[1853] means for converting the voice data into text data;

[1854] means for analyzing the text data and detecting a user's intention;

[1855] means for analyzing said voice data based on an emotional state;

[1856] means for performing appropriate actions based on the user's intent and emotional state;

[1857] means for converting the results of said actions into audio data;

[1858] means for transmitting the voice data to a terminal and providing voice feedback;

[1859] A system including:

[1860] (Claim 2)

[1861] 2. The system of claim 1, wherein said voice data input means uses a voice input / output device.

[1862] (Claim 3)

[1863] 10. The system of claim 1, wherein the analyzing means uses a natural language processing engine.

[1864] "Application example 2 when combining emotion engines"

[1865] (Claim 1)

[1866] A means for receiving voice and generating voice data;

[1867] means for transmitting the voice data to a server;

[1868] means for converting the voice data into text data;

[1869] means for analyzing the text data and detecting a user's intention;

[1870] means for identifying a user's emotion from the analyzed voice data;

[1871] means for performing appropriate actions based on the user's intentions and emotions;

[1872] means for converting the results of said actions into audio data;

[1873] means for transmitting the voice data to a terminal and providing voice feedback based on the user's emotional state;

[1874] A system including:

[1875] (Claim 2)

[1876] 2. The system of claim 1, wherein said voice data input means uses a microphone.

[1877] (Claim 3)

[1878] 10. The system of claim 1, wherein the analyzing means uses a natural language processing engine. [Explanation of symbols]

[1879] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving voice and generating voice data; means for transmitting the voice data to a server; means for converting the voice data into text data; means for analyzing the text data and detecting a user's intention; means for performing appropriate actions based on the user's intent; means for converting the results of said actions into audio data; means for transmitting the voice data to a terminal and providing voice feedback; A system including:

2. 2. The system of claim 1, wherein said voice data input means uses a microphone.

3. The system of claim 1 , wherein the analyzing means uses a natural language processing engine.

4. 2. The system according to claim 1, wherein the action execution means calls an API for acquiring weather information.

5. 2. The system of claim 1, wherein said voice data conversion means uses a voice synthesis engine.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A