System

The system allows for efficient and cost-effective creation and updating of voice guidance by analyzing text, generating voice data, and converting it into a usable format, addressing the limitations of conventional outsourcing methods.

JP2026024083APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126404
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Conventional methods for creating voice guidance require outsourcing to specialist companies, leading to high time, effort, and cost, and necessitate repeated efforts for content updates due to operational changes or revisions.

Method used

A system that includes means for analyzing generated text, generating voice data based on the analyzed text, converting the data into an outputtable format, and providing it to users, enabling quick and low-cost creation and updating of professional voice guidance.

Benefits of technology

Enables anyone to easily and quickly create and update high-quality voice guidance, reducing time, effort, and expense.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026024083000001_ABST
    Figure 2026024083000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for analyzing a generated sentence; means for generating predetermined voice data based on the analyzed sentence; means for converting the generated voice data into an outputtable format; and means for providing the generated voice data to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional methods for creating voice guidance require outsourcing to a specialist company, which requires a great deal of time, effort, and expense to complete. Furthermore, once the guidance is created, the same effort must be repeated whenever changes to the content occur due to operational changes or revisions. For this reason, there is a demand for a method that allows for the frequent creation and updating of high-quality voice guidance, quickly and at low cost. [Means for solving the problem]

[0005] The present invention provides a system including a means for analyzing generated text, a means for generating predetermined voice data based on the analyzed text, a means for converting the generated voice data into an outputtable format, and a means for providing the generated voice data to a user, thereby enabling anyone to easily and quickly create and update professional voice guidance at low cost, thereby saving a great deal of time, effort, and expense.

[0006] The "means for analyzing the generated text" is a method for analyzing the generated text based on the requirements provided by the user and understanding its contents.

[0007] The "means for generating predetermined voice data based on the analyzed sentence" is a method for generating voice data using a voice synthesis system based on the analyzed sentence.

[0008] The "means for converting the generated voice data into an outputtable format" is a method for converting the generated voice data into a format that can be easily used by the user.

[0009] The "means for providing the generated voice data to the user" is a method for providing the converted voice data to the user via the Internet.

[0010] "Voice setting information selected in advance by the user" is information in which the user specifies in advance voice characteristics such as voice type, intonation, and speed.

[0011] "User interface" refers to an interactive screen or input device that allows a user to input requirements for guidance and voice setting information and to issue a request to generate voice guidance. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0020] [First embodiment]

[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0033] Server-side processing

[0034] Request received

[0035] The server has the functionality to receive information such as the conditions for generating guidance messages and voice settings sent by the user. Specifically, it has an API that waits for HTTP requests and extracts the necessary information.

[0036] Sentence generation

[0037] Based on the received conditions, the server generates guidance text using a generative AI model (e.g., a general-purpose language model), sending requirements and keywords as prompts to the AI ​​model and receiving the generated text.

[0038] Speech synthesis

[0039] The generated text is passed to a speech synthesis system (e.g., a speech synthesis API) to generate speech data. The speech synthesis system operates based on the voice setting information (such as voice type, intonation, speed, etc.) specified by the user. The generated speech data is then returned to the server.

[0040] Providing audio data

[0041] The generated audio data is provided to the user, who then converts it to the appropriate format and generates a URL that the user can use to download or stream it.

[0042] Terminal side processing

[0043] Interface provided

[0044] The terminal provides the user with a flexible input form. This form provides an interface for entering conditions for generating guidance messages and voice setting information. For example, the user can select the content of the guidance message and the items for setting the voice through the screen of a web application.

[0045] Send request

[0046] Once the user has completed their input, the device sends this information to the server as an HTTP request, which is then formatted in JSON or other formats for efficient delivery to the server.

[0047] Receiving the results

[0048] The device receives the audio data returned from the server and displays it to the user. The user can then click a link to download the audio data or play it in their browser.

[0049] User-side processing

[0050] Setting input

[0051] The user uses the device interface to input the conditions for the announcement and voice settings. For example, when creating an announcement for a new exhibition, the user enters a sentence such as "Please create announcements for the start and end of the new exhibition" and voice settings such as "female voice, calm tone, normal speed."

[0052] Request confirmation

[0053] After checking the input, the user presses the send button to send a request to the server. The request is formatted by the terminal and passed to the server.

[0054] Check the results

[0055] Once the generated voice data is provided by the server, the user can review it through their device. If they are satisfied with the voice data, they can use it as is. If corrections or changes are required, they can re-enter the data and repeat the same process.

[0056] Specific examples

[0057] For example, consider the case of creating a guide for a new exhibition. The user enters the conditions "Please create a guide for the start and end of the new exhibition" and the voice settings "female voice, calm tone, normal speed" into a form on their device and submits it. The server receives this information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user can click the link to check the generated voice data for the guide. If necessary, they can submit a revision request again using the same procedure.

[0058] The processing flow will be explained below.

[0059] Server-side processing

[0060] Step 1: Receiving a request

[0061] The server waits for HTTP requests. When a request arrives, it analyzes the guidance generation conditions and voice setting information and decodes it into the appropriate format.

[0062] Step 2: Sentence generation

[0063] The server generates guidance text using the generative AI model, creates prompts based on the generation conditions, and passes them to the generative AI. The server receives the generated text and passes it on to the next process.

[0064] Step 3: Text-to-Speech

[0065] The server passes the generated text and voice setting information selected by the user in advance to the speech synthesis system, generates voice data, calls the speech synthesis API, and receives the generated results.

[0066] Step 4: Provide audio data

[0067] The server converts the generated audio data into an appropriate format, generates a download link or streaming URL for the user, and returns this link as an HTTP response.

[0068] Terminal side processing

[0069] Step 1: Provide an interface

[0070] The terminal displays a form for the user to input the conditions for generating guidance messages and voice settings. A user interface is provided using HTML, CSS, and JavaScript.

[0071] Step 2: Submitting a request

[0072] When a user enters information into the input form and presses the submit button, the terminal formats the information into JSON format and sends it to the server as an HTTP request.

[0073] Step 3: Receiving the results

[0074] The device receives the HTTP response from the server, obtains the download link for the audio data, and the URL for streaming, and displays this to the user.

[0075] User-side processing

[0076] Step 1: Enter settings

[0077] The user uses the terminal interface to input the conditions and voice setting information for the announcement text, for example, by writing "Announcement of the opening and closing of a new exhibition" in the input field.

[0078] Step 2: Request confirmation

[0079] The user checks the input and presses the send button, which causes the device to send the request to the server.

[0080] Step 3: Check the results

[0081] The user clicks on the download link or streaming URL for the audio data displayed on the device, plays back the generated audio guidance, and checks it. If necessary, the user can make corrections or make a new request.

[0082] Example 1

[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0084] Conventional voice guidance systems have limitations in the quality and flexibility of the text and voice data they generate, making it difficult to provide voice guidance customized to the user's needs. Furthermore, there is a need for a system that can quickly generate guidance text based on user-entered conditions and provide that content as voice data efficiently and accurately.

[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0086] In this invention, the server includes means for receiving a request including conditions from a user, means for generating sentences using a generative model based on the received conditions, means for synthesizing the generated sentences and converting them into voice data, means for converting the generated voice data into an outputtable format, and means for providing the generated voice data to the user, thereby enabling high-quality and customized voice guidance to be provided quickly.

[0087] A "request including conditions from the user" is a series of request contents including conditions for generating a guidance message and voice setting information specified by the user.

[0088] A "generative model" is an algorithm or system that uses natural language processing technology to generate sentences based on user requests.

[0089] The "means for generating a sentence" is a process for automatically generating a guidance sentence based on the received conditions using a generative model.

[0090] The "synthesis means for converting into voice data" is a process of converting the generated sentence into voice data using voice synthesis technology.

[0091] The "means for converting into an outputtable format" is a process for converting the generated audio data into a format that is easily accessible to the user and ready for provision.

[0092] The "means for providing to the user" refers to a method for making the generated audio data available to the user via download or streaming.

[0093] "Voice setting information" is parameter information such as the type of voice, intonation, and speed used during voice synthesis.

[0094] "Interface" is a general term for an input form or user interface that allows a user to input conditions for generating guidance messages and voice setting information.

[0095] The following describes an embodiment of the present invention. This system is composed of a server, a terminal, and a user. The specific operation of each element will be explained in detail.

[0096] Server-side processing

[0097] Request received

[0098] The server receives the guidance message generation conditions and voice setting information sent by the user. Specifically, an API that waits for requests is set at the endpoint using a web framework such as Flask or Django. When a request arrives, the server extracts the necessary information (for example, JSON data entered by the user) from the request object.

[0099] Sentence generation

[0100] The server generates guidance text using a generative AI model (for example, a model using natural language processing technology) based on the extracted conditions. The requirements and keywords are sent to the AI ​​model as prompts, and the generated text is obtained. A specific example of a generative AI model that can be used is a general-purpose natural language generation model.

[0101] Speech synthesis

[0102] The server passes the generated text to a synthesis system (e.g., a speech synthesis API) to convert it into speech data, and generates speech data. Voice setting information (voice type, intonation, speed, etc.) is sent along with the API request. The generated speech data is returned to the server.

[0103] Providing audio data

[0104] The server converts the generated audio data into a format that can be output, and generates a URL that the user can download or stream. The URL is sent back to the user as an HTTP response.

[0105] Terminal side processing

[0106] Interface provided

[0107] The terminal provides the user with an input form. Using HTML and JavaScript as a web front end, we create a form that allows users to enter conditions for generating guidance messages and voice setting information. For example, there is a text input field for the guidance message content and a drop-down menu for selecting the voice type.

[0108] Send request

[0109] Once the user has completed the input, the device formats the information into JSON format and sends it to the server as an HTTP request, using a JavaScript function such as fetch.

[0110] Receiving the results

[0111] The device receives the URL of the audio data returned by the server and displays it to the user, who can then click the provided link to download the audio data or play it in their browser.

[0112] User-side processing

[0113] Setting input

[0114] The user uses the terminal interface to input the conditions for the guide message and voice settings, such as "Please generate a guide message for a new exhibition" or "Female voice, calm tone, normal speed."

[0115] Request confirmation

[0116] After checking the input, the user presses the send button to send a request to the server. The device captures the click event of the send button, formats the information in JSON format, and passes it to the server.

[0117] Check the results

[0118] Once the generated voice data is provided by the server, the user can review it through their device. If they are satisfied with the voice data, they can use it as is. If corrections or changes are required, they can re-enter the data and repeat the same process.

[0119] Specific examples

[0120] For example, consider the case of creating a guide for a new exhibition. The user enters the conditions "Please create opening and closing guides for a new exhibition" and voice settings of "female voice, calm tone, normal speed" into a form on their device and submits it. The server receives the information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user clicks the link to check the generated voice data for the guide. If necessary, they can submit a revision request again using the same procedure.

[0121] Example prompt: "Generate a guide for a new exhibition in a female voice, with a calm tone and normal speed."

[0122] This enables the system to quickly provide high-quality, customized voice guidance.

[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0124] Step 1:

[0125] The user inputs the guidance conditions.

[0126] The user uses the device's interface to input the conditions for the announcement text and voice setting information. For example, they can input content such as "Please generate an announcement text for a new exhibition" and voice setting information such as "female voice, calm tone, normal speed." The input data is temporarily stored in the device's memory.

[0127] Input: User's guidance conditions and voice setting information

[0128] Output: Saved guidance conditions and voice setting information

[0129] Step 2:

[0130] The device sends a request to the server

[0131] The device formats the conditions and setting information entered by the user into JSON format and sends it to the server as an HTTP request using the JavaScript fetch function.

[0132] Input: Saved guidance conditions and voice setting information

[0133] Output: HTTP request sent to the server

[0134] Step 3:

[0135] The server receives the request

[0136] The server receives HTTP requests sent from devices. An API endpoint built using a web framework (e.g., Flask or Django) listens for the requests and extracts data from the request object when it is received.

[0137] Input: HTTP request sent from the terminal

[0138] Output: Extracted guidance conditions and voice setting information

[0139] Step 4:

[0140] The server generates the message

[0141] The server sends a prompt to the generative AI model based on the extracted conditions, generates a guidance message, and then makes a request to the generative model (e.g., a general-purpose language model) to obtain the generated guidance message from the API.

[0142] Input: Extracted guidance conditions

[0143] Output: Generated guidance text

[0144] Step 5:

[0145] The server synthesizes the voice data

[0146] The server sends the generated guidance message and voice setting information to a speech synthesis system to convert it into voice data. The speech synthesis system (e.g., a speech synthesis API) converts the text into voice data based on the specified voice settings and returns it to the server.

[0147] Input: Generated guidance text and voice setting information

[0148] Output: Generated audio data

[0149] Step 6:

[0150] The server provides the audio data

[0151] The server converts the generated audio data into a format that allows the user to download or stream it, provides the converted audio data as a URL that the user can access, and returns the URL to the user as an HTTP response.

[0152] Input: Generated audio data

[0153] Output: A URL of the audio data that can be accessed by the user

[0154] Step 7:

[0155] The terminal displays the results to the user

[0156] The device receives the URL of the audio data returned by the server and displays it to the user, who can then click the provided link to download the audio data or play it in their browser.

[0157] Input: URL of the audio data returned from the server

[0158] Output: Displayed as a link that the user can access

[0159] In this way, by detailing the specific operations performed at each step and the inputs and outputs, the processing flow of this system becomes clear.

[0160] (Application example 1)

[0161] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0162] Currently, many food delivery services rely on text-based notifications, which can create visual constraints and disrupt user flow. Therefore, a more intuitive and real-time means of providing information is needed. Additionally, systems that can provide voice guidance customized based on individual user preferences and settings are anticipated.

[0163] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0164] In this invention, the server includes a means for receiving a request from a user, a means for generating guidance sentences using a generative AI model based on the received request, and a means for analyzing the generated guidance sentences, thereby enabling the user to receive more personalized voice guidance in real time.

[0165] "User" refers to the person who uses the system to receive voice guidance.

[0166] A "means for receiving a request" is a component of the system that has the function of receiving information entered by a user.

[0167] "Generative AI models" refer to algorithms and techniques that use artificial intelligence to generate text.

[0168] The "means for generating guidance text" is the part of the system that automatically generates guidance text based on the received request.

[0169] The "means for analyzing the guidance text" is a system function that analyzes the generated text and obtains information for creating voice data based on the content of the text.

[0170] The "means for generating predetermined voice data" is the part of the system that generates voice data in a predetermined format based on the analyzed sentence.

[0171] A "means for converting to an outputtable format" is a system component that has the ability to convert the generated audio data into a format that can be used by the user.

[0172] The "means for providing voice data" is a function of the system for providing the generated voice data to the user.

[0173] "Voice setting information" refers to setting information related to the voice selected in advance by the user, such as the type of voice, intonation, and speed.

[0174] "Means operating through a user interface" means the part of the system that has an interface through which a user can directly interact with to input and send requests.

[0175] Server-side processing

[0176] The server receives a request from the user, generates a guide message based on the request, and creates and provides audio data. Specifically, the server follows the steps below.

[0177] First, the server receives a request from the user, including the conditions for generating the guidance text and voice setting information. At this time, the information is sent via an HTTP request, and the server uses an API to extract that information. Next, based on the received conditions, the generative AI model is used to generate the guidance text. At this time, the requirements and keywords are sent to the generative AI model as prompts, and the generated text is received. The generated text is passed to the speech synthesis system, which generates voice data. The speech synthesis system operates based on the voice setting information specified by the user (voice type, intonation, speed, etc.). The generated voice data is then returned to the server, converted into an appropriate format, and provided via a URL that the user can download or stream.

[0178] Terminal side processing

[0179] The device provides the user with an input form, formats the request, and sends it to the server. The user enters the conditions for generating the guidance text and audio setting information through the web application screen. Once the input is complete, the device formats this information in JSON format or similar and sends it to the server as an HTTP request. The device receives the audio data returned from the server and displays it to the user. The user can click the provided link to download the audio data or play the audio in their browser.

[0180] User-side processing

[0181] The user uses the device interface to input the conditions for the guidance text and voice settings. For example, in the case of a food delivery service, the user might input the condition "Please create a guidance message for the delivery status" and voice setting information such as "Male voice, calm tone, normal speed." After checking the input information, the user presses the send button to send a request to the server. The request is formatted by the device and passed to the server. When the generated voice data is provided by the server, the user checks it through the device. After checking the voice data, if the user is satisfied, they can use it as is. If corrections or changes are needed, the user can enter the information again and repeat the same process.

[0182] Specific examples

[0183] For example, consider a food delivery service where a user wants to provide delivery status information. The user enters a condition such as "Please create a delivery status information" and voice settings such as "male voice, calm tone, normal speed" into a form on their device and submits it. The server receives this information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user can click the link to check the generated voice data of the delivery status.

[0184] Prompt Sentence Examples

[0185] "The delivery person has left the restaurant. The estimated delivery time is 30 minutes."

[0186] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0187] Step 1:

[0188] The server receives guidance message generation conditions and voice setting information from the user.

[0189] Specifically, the information entered by the user using the device interface is sent to the server as an HTTP request, which is sent in JSON format and received by the server.

[0190] Input: Guidance message generation conditions, voice setting information (HTTP request sent in JSON format)

[0191] Output: Guidance message generation conditions, voice setting information (stored as variables within the server)

[0192] Step 2:

[0193] The server generates the guidance text using a generative AI model based on the received guidance text generation conditions.

[0194] For example, OpenAI's GPT-3 is used as the generative AI model. In this case, the conditions for generating the guidance sentence are sent to the AI ​​model as a prompt, and the generated sentence is received.

[0195] Input: Guidance sentence generation conditions (sent as prompt sentence to the AI ​​model)

[0196] Output: Generated guidance text

[0197] Step 3:

[0198] The server analyzes the generated guidance text and obtains information for generating predetermined voice data based on the text.

[0199] Specifically, the content of the guidance text is analyzed and the parameters necessary for the voice synthesis system are extracted.

[0200] Input: Generated guidance text

[0201] Output: Speech synthesis parameters (voice type, intonation, speed, etc.)

[0202] Step 4:

[0203] The server uses a voice synthesis system to generate predetermined voice data based on the voice synthesis parameters.

[0204] The Google Cloud Text-to-Speech API is used for voice synthesis. The generated text and voice setting information are sent to the voice synthesis system, which then generates the voice data.

[0205] Input: Generated guidance text, voice setting information

[0206] Output: Generated audio data

[0207] Step 5:

[0208] The server converts the generated audio data into an outputtable format.

[0209] This process involves encoding the audio data into MP3 format and converting it so that users can download or stream it.

[0210] Input: Generated audio data (internal format)

[0211] Output: Audio data in outputtable format (MP3 format)

[0212] Step 6:

[0213] The server provides the converted audio data to the user.

[0214] Specifically, it hosts the audio data and generates a URL that users can access, which is then provided to the user.

[0215] Input: Audio data in an outputtable format (MP3 format)

[0216] Output: Access URL for audio data

[0217] Step 7:

[0218] The terminal displays the access URL for the audio data received from the server to the user.

[0219] Users can click on the provided URL to view the generated audio data.

[0220] Input: Access URL for audio data

[0221] Output: Access URL for the displayed audio data (user interface)

[0222] The following are specific examples of prompt sentences:

[0223] Example prompt sentence:

[0224] "The delivery person has left the restaurant. The estimated delivery time is 30 minutes."

[0225] A guidance sentence is generated based on this prompt sentence, and the subsequent series of processes are executed.

[0226] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0227] Server-side processing

[0228] Request received

[0229] The server receives a request for generating voice guidance from the user. The request includes the conditions for generating guidance, voice setting information, and the user's emotional information. The server analyzes this information and decodes it into an appropriate format.

[0230] Emotion Recognition and Sentence Generation

[0231] The server uses an emotion engine to recognize the user's emotion based on the received user emotion information. The emotion engine analyzes emotions from the user's voice and text input and provides the recognition results to a generative AI model. The server then uses the generative AI model to generate guidance text based on the emotion information. In this text generation process, the tone and content of the text are adjusted according to the emotion information.

[0232] Speech synthesis

[0233] The server uses a speech synthesis system to generate voice data based on the analyzed text and voice setting information selected in advance by the user. The generated voice data is then returned to the server.

[0234] Providing audio data

[0235] The server converts the generated audio data into an appropriate format, generates a download link or streaming URL for the user, and returns this link as an HTTP response.

[0236] Terminal side processing

[0237] Interface provided

[0238] The terminal provides the user with a form for inputting the conditions for generating guidance messages, voice settings, and emotion information. The user interface uses HTML, CSS, and JavaScript, and allows voice and text input.

[0239] Send request

[0240] When a user enters information and presses the send button, the device formats the entered information into JSON format and sends it to the server as an HTTP request.

[0241] Receiving the results

[0242] The device receives the HTTP response from the server and obtains the download link and streaming URL for the audio data, which are then displayed to the user, allowing them to download or preview the data.

[0243] User-side processing

[0244] Setting input

[0245] The user uses the terminal interface to input the conditions for the guidance message, voice setting information, and emotional information, which includes information obtained from the user's facial expression and tone of voice, for example.

[0246] Request confirmation

[0247] The user checks the input and presses the send button to send a request to the server. The request is formatted by the terminal and passed to the server.

[0248] Check the results

[0249] Once the audio data is provided by the server, the user can check the download link or streaming URL of the audio data through their device, play the generated audio guidance to check the content, and resubmit the adjustment request if necessary.

[0250] Specific examples

[0251] For example, consider the case where a user wants to create a guide for a new exhibition. The user enters the condition "Please create opening and closing instructions for the new exhibition," along with voice setting information (female voice, calm tone, normal speed) and emotional information into a form on the device. The device sends this information to the server, which uses an emotion engine to analyze the user's emotions. Based on the analysis results, a generative AI model generates sentences, and a speech synthesis system generates voice data. The generated voice data is provided by the server, and the user can play it back on the device and request corrections as needed.

[0252] The processing flow will be explained below.

[0253] Server-side processing

[0254] Step 1:

[0255] The server waits for HTTP requests and receives requests from users, including conditions for generating announcements, voice settings, and emotion information. For example, the request may include content such as "Please create announcements for the start and end of a new exhibition" and voice settings such as "female voice, calm tone, normal speed."

[0256] Step 2:

[0257] The server extracts emotion information from the received request and passes it to the emotion engine. The emotion engine recognizes emotions from the user's input (text or voice) and provides the results to the generative AI model. For example, if the user is recognized as happy, the tone of the text will be adjusted to be more cheerful and positive.

[0258] Step 3:

[0259] The server uses a generative AI model to generate guidance text based on the emotional information returned by the emotion engine. For example, a more friendly tone of text is generated based on the emotional information. This text is then converted into audio data in the next step.

[0260] Step 4:

[0261] The server passes the generated text and user-specified voice settings to a speech synthesis system, which then generates voice data. The speech synthesis system then generates voice data with the specified voice type, intonation, and speed, for example. The generated voice data is then returned to the server.

[0262] Step 5:

[0263] The server converts the generated audio data into the appropriate format and generates a download link or streaming URL, which is sent back to the user as an HTTP response.

[0264] Terminal side processing

[0265] Step 1:

[0266] The terminal provides the user with a form for inputting the conditions for generating the announcement, voice setting information, and emotion information. For example, the user might write "Please create an announcement for the start and end of a new exhibition" in the input field, and select "female voice, calm tone, normal speed" as the voice setting information.

[0267] Step 2:

[0268] When a user inputs information and presses the send button, the device converts the information into JSON format and sends it to the server as an HTTP request. This request includes the text generation conditions, voice settings, and emotion information.

[0269] Step 3:

[0270] The device receives the HTTP response from the server and obtains the download link and streaming URL for the audio data, which are then displayed to the user, allowing them to download or preview the data.

[0271] User-side processing

[0272] Step 1:

[0273] The user inputs the conditions for the announcement, voice setting information, and emotional information using the terminal interface. For example, if the user is in a happy mood and wants to create an "announcement for the opening and closing of a new exhibition," the emotional information is also input at the same time.

[0274] Step 2:

[0275] The user checks the input and presses the send button, which causes the device to send the request to the server. The request is then formatted by the device and passed to the server.

[0276] Step 3:

[0277] Once the voice data is provided by the server, the user can check it through their device. The generated voice guidance is played back and the user can check the content. If they are satisfied, they can use it as is, and if necessary, they can send a request for adjustments again. For example, if they want the tone and intonation of the guidance to be a little brighter, they can set it again on their device and make another request.

[0278] Example 2

[0279] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0280] Conventional voice guidance generation systems generate voice data without considering the user's emotional information, which means they are unable to provide guidance that reflects the user's emotions or specific requirements. In addition, the tone and content of the generated voice data are fixed, making it difficult to flexibly respond to the diverse needs of users.

[0281] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0282] In this invention, the server includes means for analyzing emotional information received from a user, means for analyzing generated text, means for adjusting the tone and content of the generated text based on the analyzed emotional information, means for generating predetermined voice data based on the analyzed text, means for converting the generated voice data into an outputtable format, means for providing the generated voice data to the user, and user interface means for receiving and shaping input information from the user, thereby making it possible to provide voice guidance according to the user's emotions and specific requirements.

[0283] The "means for analyzing the generated text" is a means for analyzing the generated natural language text to understand the meaning and evaluate the content.

[0284] The "means for generating predetermined voice data" refers to a means for synthesizing voice data based on analyzed text. For example, this would be a voice synthesis system that generates voice from text.

[0285] The "means for converting the generated audio data into an outputtable format" refers to a means for converting the audio data into a format that can be used by the user (for example, a download link or a streaming URL).

[0286] The "means for providing the generated voice data to the user" refers to a means for converting the generated voice data into an appropriate format and transmitting it to the user. For example, this would be a process for sending a download link as an HTTP response.

[0287] The "means for analyzing emotional information received from a user" refers to a means for analyzing emotional information (such as voice or text) sent from a user and recognizing the emotional state of the user.

[0288] The "means for adjusting the tone and content of the generated sentence" refers to a means for adjusting the tone and content of the generated sentence based on the recognized emotion information, thereby generating an appropriate sentence according to the user's emotion and desires.

[0289] "User interface means for receiving and formatting user-input information" refers to means for providing an interface for collecting information entered by a user, formatting it appropriately, and sending it to a server. Examples of this include forms using HTML, CSS, and JavaScript.

[0290] The system of this invention generates and provides voice guidance to users according to their emotions. The system is mainly composed of three elements: a server, a terminal, and a user, each of which plays a specific role.

[0291] Server-side processing

[0292] The server receives a voice guidance generation request from the user. The request includes the guidance generation conditions, voice setting information, and user emotion information. Specifically, the server performs the following process.

[0293] Receiving and parsing the request

[0294] The server receives the HTTP request and parses the request body to extract the necessary information, using software such as Node.js.

[0295] emotion recognition

[0296] The server uses an emotion engine (e.g., EmotionRecognitionEngine) to analyze the emotion information provided by the user and recognize the user's emotional state based on the analysis results.

[0297] Sentence generation

[0298] The server provides prompts to a generative AI model (e.g., GPT-4) based on the recognized emotion information, and generates guidance text. The generative AI model adjusts the tone and content of the text according to the emotion.

[0299] Speech synthesis

[0300] The server generates voice data using a voice synthesis system (for example, Google Text-to-Speech) based on the generated text and voice setting information selected in advance by the user.

[0301] Converting and providing audio data

[0302] The generated audio data is converted into an appropriate format, a download link or streaming URL is generated to provide to the user, and the generated link is returned as an HTTP response.

[0303] Terminal side processing

[0304] The terminal provides a form for the user to input the conditions for generating the guidance message, voice setting information, and emotion information, using web technologies such as HTML, CSS, and JavaScript.

[0305] Sending user-entered information

[0306] When a user provides input information and presses the submit button, the device formats the information into JSON format and sends it to the server as an HTTP request.

[0307] Receiving and displaying results

[0308] The device receives the response from the server and displays a download link for the audio data and a streaming URL to the user.

[0309] User-side processing

[0310] Entering Settings

[0311] The user uses the terminal interface to input the conditions for the guidance message, voice setting information, and emotional information, which includes information obtained from the user's facial expression and tone of voice, for example.

[0312] Submitting and confirming your request

[0313] The user checks the input contents and presses the send button to send the request to the server.

[0314] Checking the results

[0315] The user can check the voice data provided by the server through the terminal, play the generated voice guidance, and resubmit the adjustment request if necessary.

[0316] Specific examples

[0317] For example, consider the case where a user wants to create a guide for a new exhibition. The user inputs conditions such as "Please create opening and closing guides for the new exhibition," and voice settings such as "female voice, calm tone, normal speed." In addition, the user's emotional information is also input. The device sends this information to the server, which uses an emotion engine to analyze the user's emotions. Based on the results of this analysis, a generative AI model generates sentences, and a speech synthesis system generates voice data. The generated voice data is provided by the server, and the user can play it back on their device.

[0318] Prompt Sentence Examples

[0319] "Please create opening and closing announcements for a new exhibition. Please use a female voice, calm tone, and normal speed."

[0320] This system makes it possible to easily generate and provide voice guidance that responds to the user's emotions and specific requirements.

[0321] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0322] Step 1:

[0323] Providing user input information

[0324] The user uses the terminal to input the information generation conditions, voice setting information, and emotion information. For example, the user inputs the information generation conditions such as "Please create an announcement for the start and end of a new exhibition," the voice setting information such as "female voice, calm tone, normal speed," and emotion information. The input information is collected by a form on the terminal.

[0325] Step 2:

[0326] Sending input information

[0327] The device formats the information received from the user into JSON format. The formatted data is sent to the server as an HTTP request. The input is the guidance generation conditions, voice setting information, and emotion information provided by the user, and the output is the request data sent to the server.

[0328] Step 3:

[0329] The server receives and analyzes the request

[0330] The server receives an HTTP request from the device. The server parses the received data and extracts the message generation conditions, voice setting information, and emotion information. For example, the request body is parsed using Node.js. The input is the request data sent from the device, and the output is the analyzed message generation conditions, voice setting information, and emotion information.

[0331] Step 4:

[0332] Emotional information analysis

[0333] The server analyzes the user's emotional information using an emotion engine. The emotion engine receives the user's voice or text as emotional information, analyzes it, and recognizes the emotional state. For example, an EmotionRecognitionEngine is used. The input is the analyzed emotional information, and the output is the recognized emotional state.

[0334] Step 5:

[0335] Sentence generation

[0336] The server provides a prompt to the generative AI model based on the recognized emotional information to generate a guidance sentence. For example, using a generative AI model (GPT-4), it generates a prompt such as, "User's emotion: [emotional state]. Please generate a guidance sentence based on the following conditions: [guidance sentence generation conditions]." The input is the recognized emotional state and the guidance sentence generation conditions, and the output is the generated guidance sentence.

[0337] Step 6:

[0338] Speech synthesis

[0339] The server generates voice data using a voice synthesis system based on the generated text and voice setting information selected in advance by the user. For example, Google Text-to-Speech is used. The input is the generated guidance text and voice setting information, and the output is the generated voice data.

[0340] Step 7:

[0341] Converting and providing audio data

[0342] The server converts the generated audio data into an appropriate format and generates a download link or streaming URL to provide it to the user. For example, the server stores the audio data on the server and generates a download link for it. The input is the generated audio data, and the output is a download link or streaming URL.

[0343] Step 8:

[0344] Receiving and displaying results

[0345] The terminal receives the HTTP response from the server and displays a download link or streaming URL for the audio data to the user. The user can then check the generated audio guidance and play it back as needed. The input is the download link or streaming URL provided by the server, and the output is the user's confirmation and playback of the audio guidance.

[0346] (Application example 2)

[0347] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0348] Conventional voice guidance systems are unable to take into account the emotional state of the user, making it difficult to provide flexible guidance tailored to the passenger's emotions and situation, and making it difficult to reduce passenger stress and anxiety.Autonomous vehicles are required to provide an environment in which passengers can travel in a relaxed and comfortable manner, and current systems are unable to meet this need.

[0349] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0350] In this invention, the server includes means for analyzing the generated sentence, means for generating predetermined voice data based on the analyzed sentence, means for converting the generated voice data into an outputtable format, means for providing the generated voice data to a user, and means for analyzing the user's emotional information and generating a sentence according to the emotional information using a generative AI model based on the analysis results, thereby making it possible to provide flexible and appropriate voice guidance according to the emotional state of passengers.

[0351] A generative AI model is a type of artificial intelligence that generates sentences in natural language based on input data. Specifically, it generates sentences according to prompts and outputs text that reflects the user's intentions and emotions.

[0352] A prompt is an input that instructs a generative AI model on how to generate a sentence. This prompt guides the theme and tone of the sentence that the model generates.

[0353] (Emotional information) is data obtained from the user's facial expressions, tone of voice, behavior, etc., and is an indicator of the user's current emotional state.

[0354] (Voice setting information) is information about voice characteristics that the user selects in advance. Specifically, it refers to settings such as the gender, tone, and speed of the voice.

[0355] An in-vehicle display is a display device installed in an autonomous vehicle, which is used by users to input information and visually confirm guidance.

[0356] A microphone is a device for collecting sound, and in particular, its role is to obtain the user's voice as input.

[0357] A (camera) is a device that captures video and is used to collect visual data such as a user's facial expressions.

[0358] A speech synthesis system is a system that converts text data into speech data and is responsible for generating natural-sounding speech.

[0359] A user interface is an interface through which a user inputs information into a system, and is designed to improve operability.

[0360] An HTTP request is a data request sent from a client to a server, and is a basic communication method in web communication.

[0361] An HTTP response is a response from a server to a client, containing the results of a request.

[0362] The present invention aims to provide a system that provides optimal voice guidance based on the user's emotional information. This system is particularly effective in autonomous vehicles, and aims to provide a comfortable travel experience by providing flexible guidance according to the passenger's emotional state.

[0363] System configuration

[0364] The system of the present invention comprises the following main components:

[0365] 1. Server

[0366] The server receives requests from users, analyzes them, generates sentences, synthesizes speech, and provides speech data. Specifically, it uses the following software components:

[0367] Emotion engine: An open-source library for analyzing user emotional information (e.g., Librosa).

[0368] Generative AI model: GPT-4 for generating natural language sentences.

[0369] Speech synthesis system: Google Text-to-Speech API that converts text into audio data.

[0370] 2. Terminal

[0371] The terminal is equipped with an in-vehicle display, microphone, and camera, and has the ability to collect user input information and communicate with a server.

[0372] In-vehicle display: A display device that allows passengers to input information via touch.

[0373] Microphone: An input device for capturing the user's voice.

[0374] Camera: A device for capturing the user's facial expressions and acquiring emotional information.

[0375] 3. Users

[0376] Users can operate the system to receive the necessary guidance. For example, passengers can request guidance by operating the touch panel or by voice while on board.

[0377] Processing flow

[0378] When the server receives a request from a user, it performs the following process.

[0379] 1. Request Analysis

[0380] The server analyzes the request and extracts the user's emotional information, voice setting information, and guidance conditions.

[0381] 2. Emotional Information Analysis

[0382] An emotion engine is used to analyze emotional information from the user's tone of voice and facial expressions.

[0383] 3. Sentence generation

[0384] Based on the analysis results, prompts are input to a generative AI model (GPT-4) to generate guidance sentences that match the user's emotions. For example, if the user seems anxious, the following prompts are used:

[0385] "The user seems anxious. Please generate gentle, relaxing instructions."

[0386] 4. Speech Synthesis

[0387] The generated text is input into a speech synthesis system (Google Text-to-Speech API) to generate voice data.

[0388] 5. Provision of audio data

[0389] The generated voice data is converted into an appropriate format and provided to the user. Specifically, a URL link is generated and sent to the user's device.

[0390] Specific examples

[0391] For example, if a passenger is stuck in traffic and requests, "Tell me how to have fun in traffic jams," the system will process the request as follows:

[0392] The server analyzes the request and detects anxiety from the user's tone of voice, then feeds the following prompt to the generative AI model:

[0393] "Passengers seem anxious in traffic. Please generate a guide to make them feel happy."

[0394] The text generated by the AI ​​model is input into the Google Text-to-Speech API to generate audio data, and a URL link to that audio data is provided to the device so that passengers can play it back.

[0395] The above is an embodiment of the present invention. The present invention makes it possible to provide flexible and appropriate guidance according to the emotional state of passengers, thereby realizing a comfortable travel experience.

[0396] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0397] Program processing steps

[0398] Step 1:

[0399] The device receives user input information (guidance conditions, voice settings, and emotion information). This input information is collected through touch operations, voice input, facial recognition, etc. After collecting the input information, the device formats it into JSON format and sends it to the server as an HTTP request.

[0400] Step 2:

[0401] The server receives a JSON-formatted HTTP request sent from the device. It analyzes the information contained in the request (guidance conditions, voice settings, and emotion information) and decodes it into an appropriate format. The analyzed data is used in the next emotion analysis step.

[0402] Step 3:

[0403] The server uses an emotion engine to recognize the user's emotion based on the received user emotion information. Specifically, it analyzes the voice and text data to identify the user's emotional state (e.g., anxiety, relief, enjoyment, etc.). The analysis results are used in the next step.

[0404] Step 4:

[0405] The server inputs the emotional information obtained from the emotion engine into a generative AI model (GPT-4) to create a prompt. This prompt specifically indicates the theme and tone of the text to be generated. For example, when providing guidance to an anxious passenger, the prompt might say, "The user seems anxious. Please generate a gentle, relaxing message." This prompt is used to have the generative AI model generate the text.

[0406] Step 5:

[0407] The server receives the text output from the generative AI model and then inputs it into a speech synthesis system. The speech synthesis system (e.g., Google Text-to-Speech API) converts the text data into audio data. This audio data is based on the text generated by the generative AI model and is adjusted according to the voice settings previously set by the user.

[0408] Step 6:

[0409] The server receives the audio data generated by the speech synthesis system and converts it into a format that can be played on the user's device (e.g., MP3 format). The converted audio data is then generated as a download link or streaming URL.

[0410] Step 7:

[0411] The server sends an HTTP response to the device, which includes a download link or streaming URL for the generated audio data. The device receives this response and provides the user with an option to play the audio data.

[0412] Step 8:

[0413] The terminal provides an interface for the user to play the audio data, allowing the user to listen to the guidance, and displays the interface again so that the user can make further requests as necessary.

[0414] The above steps make it possible to provide optimal voice guidance based on the user's emotional information. As a specific example of operation, if a user inputs "Please tell me how to have fun even in traffic jams," this information is sent to the server, and appropriate guidance is provided using the generative AI model and speech synthesis system.

[0415] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0416] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0417] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0418] [Second embodiment]

[0419] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0420] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0421] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0422] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0423] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0424] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0425] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0426] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0427] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0428] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0429] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0430] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0431] Server-side processing

[0432] Request received

[0433] The server has the functionality to receive information such as the conditions for generating guidance messages and voice settings sent by the user. Specifically, it has an API that waits for HTTP requests and extracts the necessary information.

[0434] Sentence generation

[0435] Based on the received conditions, the server generates guidance text using a generative AI model (e.g., a general-purpose language model), sending requirements and keywords as prompts to the AI ​​model and receiving the generated text.

[0436] Speech synthesis

[0437] The generated text is passed to a speech synthesis system (e.g., a speech synthesis API) to generate speech data. The speech synthesis system operates based on the voice setting information (such as voice type, intonation, speed, etc.) specified by the user. The generated speech data is then returned to the server.

[0438] Providing audio data

[0439] The generated audio data is provided to the user, who then converts it to the appropriate format and generates a URL that the user can use to download or stream it.

[0440] Terminal side processing

[0441] Interface provided

[0442] The terminal provides the user with a flexible input form. This form provides an interface for entering conditions for generating guidance messages and voice setting information. For example, the user can select the content of the guidance message and the items for setting the voice through the screen of a web application.

[0443] Send request

[0444] Once the user has completed their input, the device sends this information to the server as an HTTP request, which is then formatted in JSON or other formats for efficient delivery to the server.

[0445] Receiving the results

[0446] The device receives the audio data returned from the server and displays it to the user. The user can then click a link to download the audio data or play it in their browser.

[0447] User-side processing

[0448] Setting input

[0449] The user uses the device interface to input the conditions for the announcement and voice settings. For example, when creating an announcement for a new exhibition, the user enters a sentence such as "Please create announcements for the start and end of the new exhibition" and voice settings such as "female voice, calm tone, normal speed."

[0450] Request confirmation

[0451] After checking the input, the user presses the send button to send a request to the server. The request is formatted by the terminal and passed to the server.

[0452] Check the results

[0453] Once the generated voice data is provided by the server, the user can review it through their device. If they are satisfied with the voice data, they can use it as is. If corrections or changes are required, they can re-enter the data and repeat the same process.

[0454] Specific examples

[0455] For example, consider the case of creating a guide for a new exhibition. The user enters the conditions "Please create a guide for the start and end of the new exhibition" and the voice settings "female voice, calm tone, normal speed" into a form on their device and submits it. The server receives this information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user can click the link to check the generated voice data for the guide. If necessary, they can submit a revision request again using the same procedure.

[0456] The processing flow will be explained below.

[0457] Server-side processing

[0458] Step 1: Receiving a request

[0459] The server waits for HTTP requests. When a request arrives, it analyzes the guidance generation conditions and voice setting information and decodes it into the appropriate format.

[0460] Step 2: Sentence generation

[0461] The server generates guidance text using the generative AI model, creates prompts based on the generation conditions, and passes them to the generative AI. The server receives the generated text and passes it on to the next process.

[0462] Step 3: Text-to-Speech

[0463] The server passes the generated text and voice setting information selected by the user in advance to the speech synthesis system, generates voice data, calls the speech synthesis API, and receives the generated results.

[0464] Step 4: Provide audio data

[0465] The server converts the generated audio data into an appropriate format, generates a download link or streaming URL for the user, and returns this link as an HTTP response.

[0466] Terminal side processing

[0467] Step 1: Provide an interface

[0468] The terminal displays a form for the user to input the conditions for generating guidance messages and voice settings. A user interface is provided using HTML, CSS, and JavaScript.

[0469] Step 2: Submitting a request

[0470] When a user enters information into the input form and presses the submit button, the terminal formats the information into JSON format and sends it to the server as an HTTP request.

[0471] Step 3: Receiving the results

[0472] The device receives the HTTP response from the server, obtains the download link for the audio data, and the URL for streaming, and displays this to the user.

[0473] User-side processing

[0474] Step 1: Enter settings

[0475] The user uses the terminal interface to input the conditions and voice setting information for the announcement text, for example, by writing "Announcement of the opening and closing of a new exhibition" in the input field.

[0476] Step 2: Request confirmation

[0477] The user checks the input and presses the send button, which causes the device to send the request to the server.

[0478] Step 3: Check the results

[0479] The user clicks on the download link or streaming URL for the audio data displayed on the device, plays back the generated audio guidance, and checks it. If necessary, the user can make corrections or make a new request.

[0480] Example 1

[0481] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0482] Conventional voice guidance systems have limitations in the quality and flexibility of the text and voice data they generate, making it difficult to provide voice guidance customized to the user's needs. Furthermore, there is a need for a system that can quickly generate guidance text based on user-entered conditions and provide that content as voice data efficiently and accurately.

[0483] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0484] In this invention, the server includes means for receiving a request including conditions from a user, means for generating sentences using a generative model based on the received conditions, means for synthesizing the generated sentences and converting them into voice data, means for converting the generated voice data into an outputtable format, and means for providing the generated voice data to the user, thereby enabling high-quality and customized voice guidance to be provided quickly.

[0485] A "request including conditions from the user" is a series of request contents including conditions for generating a guidance message and voice setting information specified by the user.

[0486] A "generative model" is an algorithm or system that uses natural language processing technology to generate sentences based on user requests.

[0487] The "means for generating a sentence" is a process for automatically generating a guidance sentence based on the received conditions using a generative model.

[0488] The "synthesis means for converting into voice data" is a process of converting the generated sentence into voice data using voice synthesis technology.

[0489] The "means for converting into an outputtable format" is a process for converting the generated audio data into a format that is easily accessible to the user and ready for provision.

[0490] The "means for providing to the user" refers to a method for making the generated audio data available to the user via download or streaming.

[0491] "Voice setting information" is parameter information such as the type of voice, intonation, and speed used during voice synthesis.

[0492] "Interface" is a general term for an input form or user interface that allows a user to input conditions for generating guidance messages and voice setting information.

[0493] The following describes an embodiment of the present invention. This system is composed of a server, a terminal, and a user. The specific operation of each element will be explained in detail.

[0494] Server-side processing

[0495] Request received

[0496] The server receives the guidance message generation conditions and voice setting information sent by the user. Specifically, an API that waits for requests is set at the endpoint using a web framework such as Flask or Django. When a request arrives, the server extracts the necessary information (for example, JSON data entered by the user) from the request object.

[0497] Sentence generation

[0498] The server generates guidance text using a generative AI model (for example, a model using natural language processing technology) based on the extracted conditions. The requirements and keywords are sent to the AI ​​model as prompts, and the generated text is obtained. A specific example of a generative AI model that can be used is a general-purpose natural language generation model.

[0499] Speech synthesis

[0500] The server passes the generated text to a synthesis system (e.g., a speech synthesis API) to convert it into speech data, and generates speech data. Voice setting information (voice type, intonation, speed, etc.) is sent along with the API request. The generated speech data is returned to the server.

[0501] Providing audio data

[0502] The server converts the generated audio data into a format that can be output, and generates a URL that the user can download or stream. The URL is sent back to the user as an HTTP response.

[0503] Terminal side processing

[0504] Interface provided

[0505] The terminal provides the user with an input form. Using HTML and JavaScript as a web front end, we create a form that allows users to enter conditions for generating guidance messages and voice setting information. For example, there is a text input field for the guidance message content and a drop-down menu for selecting the voice type.

[0506] Send request

[0507] Once the user has completed the input, the device formats the information into JSON format and sends it to the server as an HTTP request, using a JavaScript function such as fetch.

[0508] Receiving the results

[0509] The device receives the URL of the audio data returned by the server and displays it to the user, who can then click the provided link to download the audio data or play it in their browser.

[0510] User-side processing

[0511] Setting input

[0512] The user uses the terminal interface to input the conditions for the guide message and voice settings, such as "Please generate a guide message for a new exhibition" or "Female voice, calm tone, normal speed."

[0513] Request confirmation

[0514] After checking the input, the user presses the send button to send a request to the server. The device captures the click event of the send button, formats the information in JSON format, and passes it to the server.

[0515] Check the results

[0516] Once the generated voice data is provided by the server, the user can review it through their device. If they are satisfied with the voice data, they can use it as is. If corrections or changes are required, they can re-enter the data and repeat the same process.

[0517] Specific examples

[0518] For example, consider the case of creating a guide for a new exhibition. The user enters the conditions "Please create opening and closing guides for a new exhibition" and voice settings of "female voice, calm tone, normal speed" into a form on their device and submits it. The server receives the information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user clicks the link to check the generated voice data for the guide. If necessary, they can submit a revision request again using the same procedure.

[0519] Example prompt: "Generate a guide for a new exhibition in a female voice, with a calm tone and normal speed."

[0520] This enables the system to quickly provide high-quality, customized voice guidance.

[0521] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0522] Step 1:

[0523] The user inputs the guidance conditions.

[0524] The user uses the device's interface to input the conditions for the announcement text and voice setting information. For example, they can input content such as "Please generate an announcement text for a new exhibition" and voice setting information such as "female voice, calm tone, normal speed." The input data is temporarily stored in the device's memory.

[0525] Input: User's guidance conditions and voice setting information

[0526] Output: Saved guidance conditions and voice setting information

[0527] Step 2:

[0528] The device sends a request to the server

[0529] The device formats the conditions and setting information entered by the user into JSON format and sends it to the server as an HTTP request using the JavaScript fetch function.

[0530] Input: Saved guidance conditions and voice setting information

[0531] Output: HTTP request sent to the server

[0532] Step 3:

[0533] The server receives the request

[0534] The server receives HTTP requests sent from devices. An API endpoint built using a web framework (e.g., Flask or Django) listens for the requests and extracts data from the request object when it is received.

[0535] Input: HTTP request sent from the terminal

[0536] Output: Extracted guidance conditions and voice setting information

[0537] Step 4:

[0538] The server generates the message

[0539] The server sends a prompt to the generative AI model based on the extracted conditions, generates a guidance message, and then makes a request to the generative model (e.g., a general-purpose language model) to obtain the generated guidance message from the API.

[0540] Input: Extracted guidance conditions

[0541] Output: Generated guidance text

[0542] Step 5:

[0543] The server synthesizes the voice data

[0544] The server sends the generated guidance message and voice setting information to a speech synthesis system to convert it into voice data. The speech synthesis system (e.g., a speech synthesis API) converts the text into voice data based on the specified voice settings and returns it to the server.

[0545] Input: Generated guidance text and voice setting information

[0546] Output: Generated audio data

[0547] Step 6:

[0548] The server provides the audio data

[0549] The server converts the generated audio data into a format that allows the user to download or stream it, provides the converted audio data as a URL that the user can access, and returns the URL to the user as an HTTP response.

[0550] Input: Generated audio data

[0551] Output: A URL of the audio data that can be accessed by the user

[0552] Step 7:

[0553] The terminal displays the results to the user

[0554] The device receives the URL of the audio data returned by the server and displays it to the user, who can then click the provided link to download the audio data or play it in their browser.

[0555] Input: URL of the audio data returned from the server

[0556] Output: Displayed as a link that the user can access

[0557] In this way, by detailing the specific operations performed at each step and the inputs and outputs, the processing flow of this system becomes clear.

[0558] (Application example 1)

[0559] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0560] Currently, many food delivery services rely on text-based notifications, which can create visual constraints and disrupt user flow. Therefore, a more intuitive and real-time means of providing information is needed. Additionally, systems that can provide voice guidance customized based on individual user preferences and settings are anticipated.

[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0562] In this invention, the server includes a means for receiving a request from a user, a means for generating guidance sentences using a generative AI model based on the received request, and a means for analyzing the generated guidance sentences, thereby enabling the user to receive more personalized voice guidance in real time.

[0563] "User" refers to the person who uses the system to receive voice guidance.

[0564] A "means for receiving a request" is a component of the system that has the function of receiving information entered by a user.

[0565] "Generative AI models" refer to algorithms and techniques that use artificial intelligence to generate text.

[0566] The "means for generating guidance text" is the part of the system that automatically generates guidance text based on the received request.

[0567] The "means for analyzing the guidance text" is a system function that analyzes the generated text and obtains information for creating voice data based on the content of the text.

[0568] The "means for generating predetermined voice data" is the part of the system that generates voice data in a predetermined format based on the analyzed sentence.

[0569] A "means for converting to an outputtable format" is a system component that has the ability to convert the generated audio data into a format that can be used by the user.

[0570] The "means for providing voice data" is a function of the system for providing the generated voice data to the user.

[0571] "Voice setting information" refers to setting information related to the voice selected in advance by the user, such as the type of voice, intonation, and speed.

[0572] "Means operating through a user interface" means the part of the system that has an interface through which a user can directly interact with to input and send requests.

[0573] Server-side processing

[0574] The server receives a request from the user, generates a guide message based on the request, and creates and provides audio data. Specifically, the server follows the steps below.

[0575] First, the server receives a request from the user, including the conditions for generating the guidance text and voice setting information. At this time, the information is sent via an HTTP request, and the server uses an API to extract that information. Next, based on the received conditions, the generative AI model is used to generate the guidance text. At this time, the requirements and keywords are sent to the generative AI model as prompts, and the generated text is received. The generated text is passed to the speech synthesis system, which generates voice data. The speech synthesis system operates based on the voice setting information specified by the user (voice type, intonation, speed, etc.). The generated voice data is then returned to the server, converted into an appropriate format, and provided via a URL that the user can download or stream.

[0576] Terminal side processing

[0577] The device provides the user with an input form, formats the request, and sends it to the server. The user enters the conditions for generating the guidance text and audio setting information through the web application screen. Once the input is complete, the device formats this information in JSON format or similar and sends it to the server as an HTTP request. The device receives the audio data returned from the server and displays it to the user. The user can click the provided link to download the audio data or play the audio in their browser.

[0578] User-side processing

[0579] The user uses the device interface to input the conditions for the guidance text and voice settings. For example, in the case of a food delivery service, the user might input the condition "Please create a guidance message for the delivery status" and voice setting information such as "Male voice, calm tone, normal speed." After checking the input information, the user presses the send button to send a request to the server. The request is formatted by the device and passed to the server. When the generated voice data is provided by the server, the user checks it through the device. After checking the voice data, if the user is satisfied, they can use it as is. If corrections or changes are needed, the user can enter the information again and repeat the same process.

[0580] Specific examples

[0581] For example, consider a food delivery service where a user wants to provide delivery status information. The user enters a condition such as "Please create a delivery status information" and voice settings such as "male voice, calm tone, normal speed" into a form on their device and submits it. The server receives this information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user can click the link to check the generated voice data of the delivery status.

[0582] Prompt Sentence Examples

[0583] "The delivery person has left the restaurant. The estimated delivery time is 30 minutes."

[0584] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0585] Step 1:

[0586] The server receives guidance message generation conditions and voice setting information from the user.

[0587] Specifically, the information entered by the user using the device interface is sent to the server as an HTTP request, which is sent in JSON format and received by the server.

[0588] Input: Guidance message generation conditions, voice setting information (HTTP request sent in JSON format)

[0589] Output: Guidance message generation conditions, voice setting information (stored as variables within the server)

[0590] Step 2:

[0591] The server generates the guidance text using a generative AI model based on the received guidance text generation conditions.

[0592] For example, OpenAI's GPT-3 is used as the generative AI model. In this case, the conditions for generating the guidance sentence are sent to the AI ​​model as a prompt, and the generated sentence is received.

[0593] Input: Guidance sentence generation conditions (sent as prompt sentence to the AI ​​model)

[0594] Output: Generated guidance text

[0595] Step 3:

[0596] The server analyzes the generated guidance text and obtains information for generating predetermined voice data based on the text.

[0597] Specifically, the content of the guidance text is analyzed and the parameters necessary for the voice synthesis system are extracted.

[0598] Input: Generated guidance text

[0599] Output: Speech synthesis parameters (voice type, intonation, speed, etc.)

[0600] Step 4:

[0601] The server uses a voice synthesis system to generate predetermined voice data based on the voice synthesis parameters.

[0602] The Google Cloud Text-to-Speech API is used for voice synthesis. The generated text and voice setting information are sent to the voice synthesis system, which then generates the voice data.

[0603] Input: Generated guidance text, voice setting information

[0604] Output: Generated audio data

[0605] Step 5:

[0606] The server converts the generated audio data into an outputtable format.

[0607] This process involves encoding the audio data into MP3 format and converting it so that users can download or stream it.

[0608] Input: Generated audio data (internal format)

[0609] Output: Audio data in outputtable format (MP3 format)

[0610] Step 6:

[0611] The server provides the converted audio data to the user.

[0612] Specifically, it hosts the audio data and generates a URL that users can access, which is then provided to the user.

[0613] Input: Audio data in an outputtable format (MP3 format)

[0614] Output: Access URL for audio data

[0615] Step 7:

[0616] The terminal displays the access URL for the audio data received from the server to the user.

[0617] Users can click on the provided URL to view the generated audio data.

[0618] Input: Access URL for audio data

[0619] Output: Access URL for the displayed audio data (user interface)

[0620] The following are specific examples of prompt sentences:

[0621] Example prompt sentence:

[0622] "The delivery person has left the restaurant. The estimated delivery time is 30 minutes."

[0623] A guidance sentence is generated based on this prompt sentence, and the subsequent series of processes are executed.

[0624] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0625] Server-side processing

[0626] Request received

[0627] The server receives a request for generating voice guidance from the user. The request includes the conditions for generating guidance, voice setting information, and the user's emotional information. The server analyzes this information and decodes it into an appropriate format.

[0628] Emotion Recognition and Sentence Generation

[0629] The server uses an emotion engine to recognize the user's emotion based on the received user emotion information. The emotion engine analyzes emotions from the user's voice and text input and provides the recognition results to a generative AI model. The server then uses the generative AI model to generate guidance text based on the emotion information. In this text generation process, the tone and content of the text are adjusted according to the emotion information.

[0630] Speech synthesis

[0631] The server uses a speech synthesis system to generate voice data based on the analyzed text and voice setting information selected in advance by the user. The generated voice data is then returned to the server.

[0632] Providing audio data

[0633] The server converts the generated audio data into an appropriate format, generates a download link or streaming URL for the user, and returns this link as an HTTP response.

[0634] Terminal side processing

[0635] Interface provided

[0636] The terminal provides the user with a form for inputting the conditions for generating guidance messages, voice settings, and emotion information. The user interface uses HTML, CSS, and JavaScript, and allows voice and text input.

[0637] Send request

[0638] When a user enters information and presses the send button, the device formats the entered information into JSON format and sends it to the server as an HTTP request.

[0639] Receiving the results

[0640] The device receives the HTTP response from the server and obtains the download link and streaming URL for the audio data, which are then displayed to the user, allowing them to download or preview the data.

[0641] User-side processing

[0642] Setting input

[0643] The user uses the terminal interface to input the conditions for the guidance message, voice setting information, and emotional information, which includes information obtained from the user's facial expression and tone of voice, for example.

[0644] Request confirmation

[0645] The user checks the input and presses the send button to send a request to the server. The request is formatted by the terminal and passed to the server.

[0646] Check the results

[0647] Once the audio data is provided by the server, the user can check the download link or streaming URL of the audio data through their device, play the generated audio guidance to check the content, and resubmit the adjustment request if necessary.

[0648] Specific examples

[0649] For example, consider the case where a user wants to create a guide for a new exhibition. The user enters the condition "Please create opening and closing instructions for the new exhibition," along with voice setting information (female voice, calm tone, normal speed) and emotional information into a form on the device. The device sends this information to the server, which uses an emotion engine to analyze the user's emotions. Based on the analysis results, a generative AI model generates sentences, and a speech synthesis system generates voice data. The generated voice data is provided by the server, and the user can play it back on the device and request corrections as needed.

[0650] The processing flow will be explained below.

[0651] Server-side processing

[0652] Step 1:

[0653] The server waits for HTTP requests and receives requests from users, including conditions for generating announcements, voice settings, and emotion information. For example, the request may include content such as "Please create announcements for the start and end of a new exhibition" and voice settings such as "female voice, calm tone, normal speed."

[0654] Step 2:

[0655] The server extracts emotion information from the received request and passes it to the emotion engine. The emotion engine recognizes emotions from the user's input (text or voice) and provides the results to the generative AI model. For example, if the user is recognized as happy, the tone of the text will be adjusted to be more cheerful and positive.

[0656] Step 3:

[0657] The server uses a generative AI model to generate guidance text based on the emotional information returned by the emotion engine. For example, a more friendly tone of text is generated based on the emotional information. This text is then converted into audio data in the next step.

[0658] Step 4:

[0659] The server passes the generated text and user-specified voice settings to a speech synthesis system, which then generates voice data. The speech synthesis system then generates voice data with the specified voice type, intonation, and speed, for example. The generated voice data is then returned to the server.

[0660] Step 5:

[0661] The server converts the generated audio data into the appropriate format and generates a download link or streaming URL, which is sent back to the user as an HTTP response.

[0662] Terminal side processing

[0663] Step 1:

[0664] The terminal provides the user with a form for inputting the conditions for generating the announcement, voice setting information, and emotion information. For example, the user might write "Please create an announcement for the start and end of a new exhibition" in the input field, and select "female voice, calm tone, normal speed" as the voice setting information.

[0665] Step 2:

[0666] When a user inputs information and presses the send button, the device converts the information into JSON format and sends it to the server as an HTTP request. This request includes the text generation conditions, voice settings, and emotion information.

[0667] Step 3:

[0668] The device receives the HTTP response from the server and obtains the download link and streaming URL for the audio data, which are then displayed to the user, allowing them to download or preview the data.

[0669] User-side processing

[0670] Step 1:

[0671] The user inputs the conditions for the announcement, voice setting information, and emotional information using the terminal interface. For example, if the user is in a happy mood and wants to create an "announcement for the opening and closing of a new exhibition," the emotional information is also input at the same time.

[0672] Step 2:

[0673] The user checks the input and presses the send button, which causes the device to send the request to the server. The request is then formatted by the device and passed to the server.

[0674] Step 3:

[0675] Once the voice data is provided by the server, the user can check it through their device. The generated voice guidance is played back and the user can check the content. If they are satisfied, they can use it as is, and if necessary, they can send a request for adjustments again. For example, if they want the tone and intonation of the guidance to be a little brighter, they can set it again on their device and make another request.

[0676] Example 2

[0677] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0678] Conventional voice guidance generation systems generate voice data without considering the user's emotional information, which means they are unable to provide guidance that reflects the user's emotions or specific requirements. In addition, the tone and content of the generated voice data are fixed, making it difficult to flexibly respond to the diverse needs of users.

[0679] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0680] In this invention, the server includes means for analyzing emotional information received from a user, means for analyzing generated text, means for adjusting the tone and content of the generated text based on the analyzed emotional information, means for generating predetermined voice data based on the analyzed text, means for converting the generated voice data into an outputtable format, means for providing the generated voice data to the user, and user interface means for receiving and shaping input information from the user, thereby making it possible to provide voice guidance according to the user's emotions and specific requirements.

[0681] The "means for analyzing the generated text" is a means for analyzing the generated natural language text to understand the meaning and evaluate the content.

[0682] The "means for generating predetermined voice data" refers to a means for synthesizing voice data based on analyzed text. For example, this would be a voice synthesis system that generates voice from text.

[0683] The "means for converting the generated audio data into an outputtable format" refers to a means for converting the audio data into a format that can be used by the user (for example, a download link or a streaming URL).

[0684] The "means for providing the generated voice data to the user" refers to a means for converting the generated voice data into an appropriate format and transmitting it to the user. For example, this would be a process for sending a download link as an HTTP response.

[0685] The "means for analyzing emotional information received from a user" refers to a means for analyzing emotional information (such as voice or text) sent from a user and recognizing the emotional state of the user.

[0686] The "means for adjusting the tone and content of the generated sentence" refers to a means for adjusting the tone and content of the generated sentence based on the recognized emotion information, thereby generating an appropriate sentence according to the user's emotion and desires.

[0687] "User interface means for receiving and formatting user-input information" refers to means for providing an interface for collecting information entered by a user, formatting it appropriately, and sending it to a server. Examples of this include forms using HTML, CSS, and JavaScript.

[0688] The system of this invention generates and provides voice guidance to users according to their emotions. The system is mainly composed of three elements: a server, a terminal, and a user, each of which plays a specific role.

[0689] Server-side processing

[0690] The server receives a voice guidance generation request from the user. The request includes the guidance generation conditions, voice setting information, and user emotion information. Specifically, the server performs the following process.

[0691] Receiving and parsing the request

[0692] The server receives the HTTP request and parses the request body to extract the necessary information, using software such as Node.js.

[0693] emotion recognition

[0694] The server uses an emotion engine (e.g., EmotionRecognitionEngine) to analyze the emotion information provided by the user and recognize the user's emotional state based on the analysis results.

[0695] Sentence generation

[0696] The server provides prompts to a generative AI model (e.g., GPT-4) based on the recognized emotion information, and generates guidance text. The generative AI model adjusts the tone and content of the text according to the emotion.

[0697] Speech synthesis

[0698] The server generates voice data using a voice synthesis system (for example, Google Text-to-Speech) based on the generated text and voice setting information selected in advance by the user.

[0699] Converting and providing audio data

[0700] The generated audio data is converted into an appropriate format, a download link or streaming URL is generated to provide to the user, and the generated link is returned as an HTTP response.

[0701] Terminal side processing

[0702] The terminal provides a form for the user to input the conditions for generating the guidance message, voice setting information, and emotion information, using web technologies such as HTML, CSS, and JavaScript.

[0703] Sending user-entered information

[0704] When a user provides input information and presses the submit button, the device formats the information into JSON format and sends it to the server as an HTTP request.

[0705] Receiving and displaying results

[0706] The device receives the response from the server and displays a download link for the audio data and a streaming URL to the user.

[0707] User-side processing

[0708] Entering Settings

[0709] The user uses the terminal interface to input the conditions for the guidance message, voice setting information, and emotional information, which includes information obtained from the user's facial expression and tone of voice, for example.

[0710] Submitting and confirming your request

[0711] The user checks the input contents and presses the send button to send the request to the server.

[0712] Checking the results

[0713] The user can check the voice data provided by the server through the terminal, play the generated voice guidance, and resubmit the adjustment request if necessary.

[0714] Specific examples

[0715] For example, consider the case where a user wants to create a guide for a new exhibition. The user inputs conditions such as "Please create opening and closing guides for the new exhibition," and voice settings such as "female voice, calm tone, normal speed." In addition, the user's emotional information is also input. The device sends this information to the server, which uses an emotion engine to analyze the user's emotions. Based on the results of this analysis, a generative AI model generates sentences, and a speech synthesis system generates voice data. The generated voice data is provided by the server, and the user can play it back on their device.

[0716] Prompt Sentence Examples

[0717] "Please create opening and closing announcements for a new exhibition. Please use a female voice, calm tone, and normal speed."

[0718] This system makes it possible to easily generate and provide voice guidance that responds to the user's emotions and specific requirements.

[0719] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0720] Step 1:

[0721] Providing user input information

[0722] The user uses the terminal to input the information generation conditions, voice setting information, and emotion information. For example, the user inputs the information generation conditions such as "Please create an announcement for the start and end of a new exhibition," the voice setting information such as "female voice, calm tone, normal speed," and emotion information. The input information is collected by a form on the terminal.

[0723] Step 2:

[0724] Sending input information

[0725] The device formats the information received from the user into JSON format. The formatted data is sent to the server as an HTTP request. The input is the guidance generation conditions, voice setting information, and emotion information provided by the user, and the output is the request data sent to the server.

[0726] Step 3:

[0727] The server receives and analyzes the request

[0728] The server receives an HTTP request from the device. The server parses the received data and extracts the message generation conditions, voice setting information, and emotion information. For example, the request body is parsed using Node.js. The input is the request data sent from the device, and the output is the analyzed message generation conditions, voice setting information, and emotion information.

[0729] Step 4:

[0730] Emotional information analysis

[0731] The server analyzes the user's emotional information using an emotion engine. The emotion engine receives the user's voice or text as emotional information, analyzes it, and recognizes the emotional state. For example, an EmotionRecognitionEngine is used. The input is the analyzed emotional information, and the output is the recognized emotional state.

[0732] Step 5:

[0733] Sentence generation

[0734] The server provides a prompt to the generative AI model based on the recognized emotional information to generate a guidance sentence. For example, using a generative AI model (GPT-4), it generates a prompt such as, "User's emotion: [emotional state]. Please generate a guidance sentence based on the following conditions: [guidance sentence generation conditions]." The input is the recognized emotional state and the guidance sentence generation conditions, and the output is the generated guidance sentence.

[0735] Step 6:

[0736] Speech synthesis

[0737] The server generates voice data using a voice synthesis system based on the generated text and voice setting information selected in advance by the user. For example, Google Text-to-Speech is used. The input is the generated guidance text and voice setting information, and the output is the generated voice data.

[0738] Step 7:

[0739] Converting and providing audio data

[0740] The server converts the generated audio data into an appropriate format and generates a download link or streaming URL to provide it to the user. For example, the server stores the audio data on the server and generates a download link for it. The input is the generated audio data, and the output is a download link or streaming URL.

[0741] Step 8:

[0742] Receiving and displaying results

[0743] The terminal receives the HTTP response from the server and displays a download link or streaming URL for the audio data to the user. The user can then check the generated audio guidance and play it back as needed. The input is the download link or streaming URL provided by the server, and the output is the user's confirmation and playback of the audio guidance.

[0744] (Application example 2)

[0745] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0746] Conventional voice guidance systems are unable to take into account the emotional state of the user, making it difficult to provide flexible guidance tailored to the passenger's emotions and situation, and making it difficult to reduce passenger stress and anxiety.Autonomous vehicles are required to provide an environment in which passengers can travel in a relaxed and comfortable manner, and current systems are unable to meet this need.

[0747] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0748] In this invention, the server includes means for analyzing the generated sentence, means for generating predetermined voice data based on the analyzed sentence, means for converting the generated voice data into an outputtable format, means for providing the generated voice data to a user, and means for analyzing the user's emotional information and generating a sentence according to the emotional information using a generative AI model based on the analysis results, thereby making it possible to provide flexible and appropriate voice guidance according to the emotional state of passengers.

[0749] A generative AI model is a type of artificial intelligence that generates sentences in natural language based on input data. Specifically, it generates sentences according to prompts and outputs text that reflects the user's intentions and emotions.

[0750] A prompt is an input that instructs a generative AI model on how to generate a sentence. This prompt guides the theme and tone of the sentence that the model generates.

[0751] (Emotional information) is data obtained from the user's facial expressions, tone of voice, behavior, etc., and is an indicator of the user's current emotional state.

[0752] (Voice setting information) is information about voice characteristics that the user selects in advance. Specifically, it refers to settings such as the gender, tone, and speed of the voice.

[0753] An in-vehicle display is a display device installed in an autonomous vehicle, which is used by users to input information and visually confirm guidance.

[0754] A microphone is a device for collecting sound, and in particular, its role is to obtain the user's voice as input.

[0755] A (camera) is a device that captures video and is used to collect visual data such as a user's facial expressions.

[0756] A speech synthesis system is a system that converts text data into speech data and is responsible for generating natural-sounding speech.

[0757] A user interface is an interface through which a user inputs information into a system, and is designed to improve operability.

[0758] An HTTP request is a data request sent from a client to a server, and is a basic communication method in web communication.

[0759] An HTTP response is a response from a server to a client, containing the results of a request.

[0760] The present invention aims to provide a system that provides optimal voice guidance based on the user's emotional information. This system is particularly effective in autonomous vehicles, and aims to provide a comfortable travel experience by providing flexible guidance according to the passenger's emotional state.

[0761] System configuration

[0762] The system of the present invention comprises the following main components:

[0763] 1. Server

[0764] The server receives requests from users, analyzes them, generates sentences, synthesizes speech, and provides speech data. Specifically, it uses the following software components:

[0765] Emotion engine: An open-source library for analyzing user emotional information (e.g., Librosa).

[0766] Generative AI model: GPT-4 for generating natural language sentences.

[0767] Speech synthesis system: Google Text-to-Speech API that converts text into audio data.

[0768] 2. Terminal

[0769] The terminal is equipped with an in-vehicle display, microphone, and camera, and has the ability to collect user input information and communicate with a server.

[0770] In-vehicle display: A display device that allows passengers to input information via touch.

[0771] Microphone: An input device for capturing the user's voice.

[0772] Camera: A device for capturing the user's facial expressions and acquiring emotional information.

[0773] 3. Users

[0774] Users can operate the system to receive the necessary guidance. For example, passengers can request guidance by operating the touch panel or by voice while on board.

[0775] Processing flow

[0776] When the server receives a request from a user, it performs the following process.

[0777] 1. Request Analysis

[0778] The server analyzes the request and extracts the user's emotional information, voice setting information, and guidance conditions.

[0779] 2. Emotional Information Analysis

[0780] An emotion engine is used to analyze emotional information from the user's tone of voice and facial expressions.

[0781] 3. Sentence generation

[0782] Based on the analysis results, prompts are input to a generative AI model (GPT-4) to generate guidance sentences that match the user's emotions. For example, if the user seems anxious, the following prompts are used:

[0783] "The user seems anxious. Please generate gentle, relaxing instructions."

[0784] 4. Speech Synthesis

[0785] The generated text is input into a speech synthesis system (Google Text-to-Speech API) to generate voice data.

[0786] 5. Provision of audio data

[0787] The generated voice data is converted into an appropriate format and provided to the user. Specifically, a URL link is generated and sent to the user's device.

[0788] Specific examples

[0789] For example, if a passenger is stuck in traffic and requests, "Tell me how to have fun in traffic jams," the system will process the request as follows:

[0790] The server analyzes the request and detects anxiety from the user's tone of voice, then feeds the following prompt to the generative AI model:

[0791] "Passengers seem anxious in traffic. Please generate a guide to make them feel happy."

[0792] The text generated by the AI ​​model is input into the Google Text-to-Speech API to generate audio data, and a URL link to that audio data is provided to the device so that passengers can play it back.

[0793] The above is an embodiment of the present invention. The present invention makes it possible to provide flexible and appropriate guidance according to the emotional state of passengers, thereby realizing a comfortable travel experience.

[0794] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0795] Program processing steps

[0796] Step 1:

[0797] The device receives user input information (guidance conditions, voice settings, and emotion information). This input information is collected through touch operations, voice input, facial recognition, etc. After collecting the input information, the device formats it into JSON format and sends it to the server as an HTTP request.

[0798] Step 2:

[0799] The server receives a JSON-formatted HTTP request sent from the device. It analyzes the information contained in the request (guidance conditions, voice settings, and emotion information) and decodes it into an appropriate format. The analyzed data is used in the next emotion analysis step.

[0800] Step 3:

[0801] The server uses an emotion engine to recognize the user's emotion based on the received user emotion information. Specifically, it analyzes the voice and text data to identify the user's emotional state (e.g., anxiety, relief, enjoyment, etc.). The analysis results are used in the next step.

[0802] Step 4:

[0803] The server inputs the emotional information obtained from the emotion engine into a generative AI model (GPT-4) to create a prompt. This prompt specifically indicates the theme and tone of the text to be generated. For example, when providing guidance to an anxious passenger, the prompt might say, "The user seems anxious. Please generate a gentle, relaxing message." This prompt is used to have the generative AI model generate the text.

[0804] Step 5:

[0805] The server receives the text output from the generative AI model and then inputs it into a speech synthesis system. The speech synthesis system (e.g., Google Text-to-Speech API) converts the text data into audio data. This audio data is based on the text generated by the generative AI model and is adjusted according to the voice settings previously set by the user.

[0806] Step 6:

[0807] The server receives the audio data generated by the speech synthesis system and converts it into a format that can be played on the user's device (e.g., MP3 format). The converted audio data is then generated as a download link or streaming URL.

[0808] Step 7:

[0809] The server sends an HTTP response to the device, which includes a download link or streaming URL for the generated audio data. The device receives this response and provides the user with an option to play the audio data.

[0810] Step 8:

[0811] The terminal provides an interface for the user to play the audio data, allowing the user to listen to the guidance, and displays the interface again so that the user can make further requests as necessary.

[0812] The above steps make it possible to provide optimal voice guidance based on the user's emotional information. As a specific example of operation, if a user inputs "Please tell me how to have fun even in traffic jams," this information is sent to the server, and appropriate guidance is provided using the generative AI model and speech synthesis system.

[0813] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0814] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0815] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0816] [Third embodiment]

[0817] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0818] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0819] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0820] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0821] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0822] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0823] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0824] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0825] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0826] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0827] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0828] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0829] Server-side processing

[0830] Request received

[0831] The server has the functionality to receive information such as the conditions for generating guidance messages and voice settings sent by the user. Specifically, it has an API that waits for HTTP requests and extracts the necessary information.

[0832] Sentence generation

[0833] Based on the received conditions, the server generates guidance text using a generative AI model (e.g., a general-purpose language model), sending requirements and keywords as prompts to the AI ​​model and receiving the generated text.

[0834] Speech synthesis

[0835] The generated text is passed to a speech synthesis system (e.g., a speech synthesis API) to generate speech data. The speech synthesis system operates based on the voice setting information (such as voice type, intonation, speed, etc.) specified by the user. The generated speech data is then returned to the server.

[0836] Providing audio data

[0837] The generated audio data is provided to the user, who then converts it to the appropriate format and generates a URL that the user can use to download or stream it.

[0838] Terminal side processing

[0839] Interface provided

[0840] The terminal provides the user with a flexible input form. This form provides an interface for entering conditions for generating guidance messages and voice setting information. For example, the user can select the content of the guidance message and the items for setting the voice through the screen of a web application.

[0841] Send request

[0842] Once the user has completed their input, the device sends this information to the server as an HTTP request, which is then formatted in JSON or other formats for efficient delivery to the server.

[0843] Receiving the results

[0844] The device receives the audio data returned from the server and displays it to the user. The user can then click a link to download the audio data or play it in their browser.

[0845] User-side processing

[0846] Setting input

[0847] The user uses the device interface to input the conditions for the announcement and voice settings. For example, when creating an announcement for a new exhibition, the user enters a sentence such as "Please create announcements for the start and end of the new exhibition" and voice settings such as "female voice, calm tone, normal speed."

[0848] Request confirmation

[0849] After checking the input, the user presses the send button to send a request to the server. The request is formatted by the terminal and passed to the server.

[0850] Check the results

[0851] Once the generated voice data is provided by the server, the user can review it through their device. If they are satisfied with the voice data, they can use it as is. If corrections or changes are required, they can re-enter the data and repeat the same process.

[0852] Specific examples

[0853] For example, consider the case of creating a guide for a new exhibition. The user enters the conditions "Please create a guide for the start and end of the new exhibition" and the voice settings "female voice, calm tone, normal speed" into a form on their device and submits it. The server receives this information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user can click the link to check the generated voice data for the guide. If necessary, they can submit a revision request again using the same procedure.

[0854] The processing flow will be explained below.

[0855] Server-side processing

[0856] Step 1: Receiving a request

[0857] The server waits for HTTP requests. When a request arrives, it analyzes the guidance generation conditions and voice setting information and decodes it into the appropriate format.

[0858] Step 2: Sentence generation

[0859] The server generates guidance text using the generative AI model, creates prompts based on the generation conditions, and passes them to the generative AI. The server receives the generated text and passes it on to the next process.

[0860] Step 3: Text-to-Speech

[0861] The server passes the generated text and voice setting information selected by the user in advance to the speech synthesis system, generates voice data, calls the speech synthesis API, and receives the generated results.

[0862] Step 4: Provide audio data

[0863] The server converts the generated audio data into an appropriate format, generates a download link or streaming URL for the user, and returns this link as an HTTP response.

[0864] Terminal side processing

[0865] Step 1: Provide an interface

[0866] The terminal displays a form for the user to input the conditions for generating guidance messages and voice settings. A user interface is provided using HTML, CSS, and JavaScript.

[0867] Step 2: Submitting a request

[0868] When a user enters information into the input form and presses the submit button, the terminal formats the information into JSON format and sends it to the server as an HTTP request.

[0869] Step 3: Receiving the results

[0870] The device receives the HTTP response from the server, obtains the download link for the audio data, and the URL for streaming, and displays this to the user.

[0871] User-side processing

[0872] Step 1: Enter settings

[0873] The user uses the terminal interface to input the conditions and voice setting information for the announcement text, for example, by writing "Announcement of the opening and closing of a new exhibition" in the input field.

[0874] Step 2: Request confirmation

[0875] The user checks the input and presses the send button, which causes the device to send the request to the server.

[0876] Step 3: Check the results

[0877] The user clicks on the download link or streaming URL for the audio data displayed on the device, plays back the generated audio guidance, and checks it. If necessary, the user can make corrections or make a new request.

[0878] Example 1

[0879] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0880] Conventional voice guidance systems have limitations in the quality and flexibility of the text and voice data they generate, making it difficult to provide voice guidance customized to the user's needs. Furthermore, there is a need for a system that can quickly generate guidance text based on user-entered conditions and provide that content as voice data efficiently and accurately.

[0881] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0882] In this invention, the server includes means for receiving a request including conditions from a user, means for generating sentences using a generative model based on the received conditions, means for synthesizing the generated sentences and converting them into voice data, means for converting the generated voice data into an outputtable format, and means for providing the generated voice data to the user, thereby enabling high-quality and customized voice guidance to be provided quickly.

[0883] A "request including conditions from the user" is a series of request contents including conditions for generating a guidance message and voice setting information specified by the user.

[0884] A "generative model" is an algorithm or system that uses natural language processing technology to generate sentences based on user requests.

[0885] The "means for generating a sentence" is a process for automatically generating a guidance sentence based on the received conditions using a generative model.

[0886] The "synthesis means for converting into voice data" is a process of converting the generated sentence into voice data using voice synthesis technology.

[0887] The "means for converting into an outputtable format" is a process for converting the generated audio data into a format that is easily accessible to the user and ready for provision.

[0888] The "means for providing to the user" refers to a method for making the generated audio data available to the user via download or streaming.

[0889] "Voice setting information" is parameter information such as the type of voice, intonation, and speed used during voice synthesis.

[0890] "Interface" is a general term for an input form or user interface that allows a user to input conditions for generating guidance messages and voice setting information.

[0891] The following describes an embodiment of the present invention. This system is composed of a server, a terminal, and a user. The specific operation of each element will be explained in detail.

[0892] Server-side processing

[0893] Request received

[0894] The server receives the guidance message generation conditions and voice setting information sent by the user. Specifically, an API that waits for requests is set at the endpoint using a web framework such as Flask or Django. When a request arrives, the server extracts the necessary information (for example, JSON data entered by the user) from the request object.

[0895] Sentence generation

[0896] The server generates guidance text using a generative AI model (for example, a model using natural language processing technology) based on the extracted conditions. The requirements and keywords are sent to the AI ​​model as prompts, and the generated text is obtained. A specific example of a generative AI model that can be used is a general-purpose natural language generation model.

[0897] Speech synthesis

[0898] The server passes the generated text to a synthesis system (e.g., a speech synthesis API) to convert it into speech data, and generates speech data. Voice setting information (voice type, intonation, speed, etc.) is sent along with the API request. The generated speech data is returned to the server.

[0899] Providing audio data

[0900] The server converts the generated audio data into a format that can be output, and generates a URL that the user can download or stream. The URL is sent back to the user as an HTTP response.

[0901] Terminal side processing

[0902] Interface provided

[0903] The terminal provides the user with an input form. Using HTML and JavaScript as a web front end, we create a form that allows users to enter conditions for generating guidance messages and voice setting information. For example, there is a text input field for the guidance message content and a drop-down menu for selecting the voice type.

[0904] Send request

[0905] Once the user has completed the input, the device formats the information into JSON format and sends it to the server as an HTTP request, using a JavaScript function such as fetch.

[0906] Receiving the results

[0907] The device receives the URL of the audio data returned by the server and displays it to the user, who can then click the provided link to download the audio data or play it in their browser.

[0908] User-side processing

[0909] Setting input

[0910] The user uses the terminal interface to input the conditions for the guide message and voice settings, such as "Please generate a guide message for a new exhibition" or "Female voice, calm tone, normal speed."

[0911] Request confirmation

[0912] After checking the input, the user presses the send button to send a request to the server. The device captures the click event of the send button, formats the information in JSON format, and passes it to the server.

[0913] Check the results

[0914] Once the generated voice data is provided by the server, the user can review it through their device. If they are satisfied with the voice data, they can use it as is. If corrections or changes are required, they can re-enter the data and repeat the same process.

[0915] Specific examples

[0916] For example, consider the case of creating a guide for a new exhibition. The user enters the conditions "Please create opening and closing guides for a new exhibition" and voice settings of "female voice, calm tone, normal speed" into a form on their device and submits it. The server receives the information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user clicks the link to check the generated voice data for the guide. If necessary, they can submit a revision request again using the same procedure.

[0917] Example prompt: "Generate a guide for a new exhibition in a female voice, with a calm tone and normal speed."

[0918] This enables the system to quickly provide high-quality, customized voice guidance.

[0919] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0920] Step 1:

[0921] The user inputs the guidance conditions.

[0922] The user uses the device's interface to input the conditions for the announcement text and voice setting information. For example, they can input content such as "Please generate an announcement text for a new exhibition" and voice setting information such as "female voice, calm tone, normal speed." The input data is temporarily stored in the device's memory.

[0923] Input: User's guidance conditions and voice setting information

[0924] Output: Saved guidance conditions and voice setting information

[0925] Step 2:

[0926] The device sends a request to the server

[0927] The device formats the conditions and setting information entered by the user into JSON format and sends it to the server as an HTTP request using the JavaScript fetch function.

[0928] Input: Saved guidance conditions and voice setting information

[0929] Output: HTTP request sent to the server

[0930] Step 3:

[0931] The server receives the request

[0932] The server receives HTTP requests sent from devices. An API endpoint built using a web framework (e.g., Flask or Django) listens for the requests and extracts data from the request object when it is received.

[0933] Input: HTTP request sent from the terminal

[0934] Output: Extracted guidance conditions and voice setting information

[0935] Step 4:

[0936] The server generates the message

[0937] The server sends a prompt to the generative AI model based on the extracted conditions, generates a guidance message, and then makes a request to the generative model (e.g., a general-purpose language model) to obtain the generated guidance message from the API.

[0938] Input: Extracted guidance conditions

[0939] Output: Generated guidance text

[0940] Step 5:

[0941] The server synthesizes the voice data

[0942] The server sends the generated guidance message and voice setting information to a speech synthesis system to convert it into voice data. The speech synthesis system (e.g., a speech synthesis API) converts the text into voice data based on the specified voice settings and returns it to the server.

[0943] Input: Generated guidance text and voice setting information

[0944] Output: Generated audio data

[0945] Step 6:

[0946] The server provides the audio data

[0947] The server converts the generated audio data into a format that allows the user to download or stream it, provides the converted audio data as a URL that the user can access, and returns the URL to the user as an HTTP response.

[0948] Input: Generated audio data

[0949] Output: A URL of the audio data that can be accessed by the user

[0950] Step 7:

[0951] The terminal displays the results to the user

[0952] The device receives the URL of the audio data returned by the server and displays it to the user, who can then click the provided link to download the audio data or play it in their browser.

[0953] Input: URL of the audio data returned from the server

[0954] Output: Displayed as a link that the user can access

[0955] In this way, by detailing the specific operations performed at each step and the inputs and outputs, the processing flow of this system becomes clear.

[0956] (Application example 1)

[0957] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0958] Currently, many food delivery services rely on text-based notifications, which can create visual constraints and disrupt user flow. Therefore, a more intuitive and real-time means of providing information is needed. Additionally, systems that can provide voice guidance customized based on individual user preferences and settings are anticipated.

[0959] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0960] In this invention, the server includes a means for receiving a request from a user, a means for generating guidance sentences using a generative AI model based on the received request, and a means for analyzing the generated guidance sentences, thereby enabling the user to receive more personalized voice guidance in real time.

[0961] "User" refers to the person who uses the system to receive voice guidance.

[0962] A "means for receiving a request" is a component of the system that has the function of receiving information entered by a user.

[0963] "Generative AI models" refer to algorithms and techniques that use artificial intelligence to generate text.

[0964] The "means for generating guidance text" is the part of the system that automatically generates guidance text based on the received request.

[0965] The "means for analyzing the guidance text" is a system function that analyzes the generated text and obtains information for creating voice data based on the content of the text.

[0966] The "means for generating predetermined voice data" is the part of the system that generates voice data in a predetermined format based on the analyzed sentence.

[0967] A "means for converting to an outputtable format" is a system component that has the ability to convert the generated audio data into a format that can be used by the user.

[0968] The "means for providing voice data" is a function of the system for providing the generated voice data to the user.

[0969] "Voice setting information" refers to setting information related to the voice selected in advance by the user, such as the type of voice, intonation, and speed.

[0970] "Means operating through a user interface" means the part of the system that has an interface through which a user can directly interact with to input and send requests.

[0971] Server-side processing

[0972] The server receives a request from the user, generates a guide message based on the request, and creates and provides audio data. Specifically, the server follows the steps below.

[0973] First, the server receives a request from the user, including the conditions for generating the guidance text and voice setting information. At this time, the information is sent via an HTTP request, and the server uses an API to extract that information. Next, based on the received conditions, the generative AI model is used to generate the guidance text. At this time, the requirements and keywords are sent to the generative AI model as prompts, and the generated text is received. The generated text is passed to the speech synthesis system, which generates voice data. The speech synthesis system operates based on the voice setting information specified by the user (voice type, intonation, speed, etc.). The generated voice data is then returned to the server, converted into an appropriate format, and provided via a URL that the user can download or stream.

[0974] Terminal side processing

[0975] The device provides the user with an input form, formats the request, and sends it to the server. The user enters the conditions for generating the guidance text and audio setting information through the web application screen. Once the input is complete, the device formats this information in JSON format or similar and sends it to the server as an HTTP request. The device receives the audio data returned from the server and displays it to the user. The user can click the provided link to download the audio data or play the audio in their browser.

[0976] User-side processing

[0977] The user uses the device interface to input the conditions for the guidance text and voice settings. For example, in the case of a food delivery service, the user might input the condition "Please create a guidance message for the delivery status" and voice setting information such as "Male voice, calm tone, normal speed." After checking the input information, the user presses the send button to send a request to the server. The request is formatted by the device and passed to the server. When the generated voice data is provided by the server, the user checks it through the device. After checking the voice data, if the user is satisfied, they can use it as is. If corrections or changes are needed, the user can enter the information again and repeat the same process.

[0978] Specific examples

[0979] For example, consider a food delivery service where a user wants to provide delivery status information. The user enters a condition such as "Please create a delivery status information" and voice settings such as "male voice, calm tone, normal speed" into a form on their device and submits it. The server receives this information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user can click the link to check the generated voice data of the delivery status.

[0980] Prompt Sentence Examples

[0981] "The delivery person has left the restaurant. The estimated delivery time is 30 minutes."

[0982] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0983] Step 1:

[0984] The server receives guidance message generation conditions and voice setting information from the user.

[0985] Specifically, the information entered by the user using the device interface is sent to the server as an HTTP request, which is sent in JSON format and received by the server.

[0986] Input: Guidance message generation conditions, voice setting information (HTTP request sent in JSON format)

[0987] Output: Guidance message generation conditions, voice setting information (stored as variables within the server)

[0988] Step 2:

[0989] The server generates the guidance text using a generative AI model based on the received guidance text generation conditions.

[0990] For example, OpenAI's GPT-3 is used as the generative AI model. In this case, the conditions for generating the guidance sentence are sent to the AI ​​model as a prompt, and the generated sentence is received.

[0991] Input: Guidance sentence generation conditions (sent as prompt sentence to the AI ​​model)

[0992] Output: Generated guidance text

[0993] Step 3:

[0994] The server analyzes the generated guidance text and obtains information for generating predetermined voice data based on the text.

[0995] Specifically, the content of the guidance text is analyzed and the parameters necessary for the voice synthesis system are extracted.

[0996] Input: Generated guidance text

[0997] Output: Speech synthesis parameters (voice type, intonation, speed, etc.)

[0998] Step 4:

[0999] The server uses a voice synthesis system to generate predetermined voice data based on the voice synthesis parameters.

[1000] The Google Cloud Text-to-Speech API is used for voice synthesis. The generated text and voice setting information are sent to the voice synthesis system, which then generates the voice data.

[1001] Input: Generated guidance text, voice setting information

[1002] Output: Generated audio data

[1003] Step 5:

[1004] The server converts the generated audio data into an outputtable format.

[1005] This process involves encoding the audio data into MP3 format and converting it so that users can download or stream it.

[1006] Input: Generated audio data (internal format)

[1007] Output: Audio data in outputtable format (MP3 format)

[1008] Step 6:

[1009] The server provides the converted audio data to the user.

[1010] Specifically, it hosts the audio data and generates a URL that users can access, which is then provided to the user.

[1011] Input: Audio data in an outputtable format (MP3 format)

[1012] Output: Access URL for audio data

[1013] Step 7:

[1014] The terminal displays the access URL for the audio data received from the server to the user.

[1015] Users can click on the provided URL to view the generated audio data.

[1016] Input: Access URL for audio data

[1017] Output: Access URL for the displayed audio data (user interface)

[1018] The following are specific examples of prompt sentences:

[1019] Example prompt sentence:

[1020] "The delivery person has left the restaurant. The estimated delivery time is 30 minutes."

[1021] A guidance sentence is generated based on this prompt sentence, and the subsequent series of processes are executed.

[1022] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1023] Server-side processing

[1024] Request received

[1025] The server receives a request for generating voice guidance from the user. The request includes the conditions for generating guidance, voice setting information, and the user's emotional information. The server analyzes this information and decodes it into an appropriate format.

[1026] Emotion Recognition and Sentence Generation

[1027] The server uses an emotion engine to recognize the user's emotion based on the received user emotion information. The emotion engine analyzes emotions from the user's voice and text input and provides the recognition results to a generative AI model. The server then uses the generative AI model to generate guidance text based on the emotion information. In this text generation process, the tone and content of the text are adjusted according to the emotion information.

[1028] Speech synthesis

[1029] The server uses a speech synthesis system to generate voice data based on the analyzed text and voice setting information selected in advance by the user. The generated voice data is then returned to the server.

[1030] Providing audio data

[1031] The server converts the generated audio data into an appropriate format, generates a download link or streaming URL for the user, and returns this link as an HTTP response.

[1032] Terminal side processing

[1033] Interface provided

[1034] The terminal provides the user with a form for inputting the conditions for generating guidance messages, voice settings, and emotion information. The user interface uses HTML, CSS, and JavaScript, and allows voice and text input.

[1035] Send request

[1036] When a user enters information and presses the send button, the device formats the entered information into JSON format and sends it to the server as an HTTP request.

[1037] Receiving the results

[1038] The device receives the HTTP response from the server and obtains the download link and streaming URL for the audio data, which are then displayed to the user, allowing them to download or preview the data.

[1039] User-side processing

[1040] Setting input

[1041] The user uses the terminal interface to input the conditions for the guidance message, voice setting information, and emotional information, which includes information obtained from the user's facial expression and tone of voice, for example.

[1042] Request confirmation

[1043] The user checks the input and presses the send button to send a request to the server. The request is formatted by the terminal and passed to the server.

[1044] Check the results

[1045] Once the audio data is provided by the server, the user can check the download link or streaming URL of the audio data through their device, play the generated audio guidance to check the content, and resubmit the adjustment request if necessary.

[1046] Specific examples

[1047] For example, consider the case where a user wants to create a guide for a new exhibition. The user enters the condition "Please create opening and closing instructions for the new exhibition," along with voice setting information (female voice, calm tone, normal speed) and emotional information into a form on the device. The device sends this information to the server, which uses an emotion engine to analyze the user's emotions. Based on the analysis results, a generative AI model generates sentences, and a speech synthesis system generates voice data. The generated voice data is provided by the server, and the user can play it back on the device and request corrections as needed.

[1048] The processing flow will be explained below.

[1049] Server-side processing

[1050] Step 1:

[1051] The server waits for HTTP requests and receives requests from users, including conditions for generating announcements, voice settings, and emotion information. For example, the request may include content such as "Please create announcements for the start and end of a new exhibition" and voice settings such as "female voice, calm tone, normal speed."

[1052] Step 2:

[1053] The server extracts emotion information from the received request and passes it to the emotion engine. The emotion engine recognizes emotions from the user's input (text or voice) and provides the results to the generative AI model. For example, if the user is recognized as happy, the tone of the text will be adjusted to be more cheerful and positive.

[1054] Step 3:

[1055] The server uses a generative AI model to generate guidance text based on the emotional information returned by the emotion engine. For example, a more friendly tone of text is generated based on the emotional information. This text is then converted into audio data in the next step.

[1056] Step 4:

[1057] The server passes the generated text and user-specified voice settings to a speech synthesis system, which then generates voice data. The speech synthesis system then generates voice data with the specified voice type, intonation, and speed, for example. The generated voice data is then returned to the server.

[1058] Step 5:

[1059] The server converts the generated audio data into the appropriate format and generates a download link or streaming URL, which is sent back to the user as an HTTP response.

[1060] Terminal side processing

[1061] Step 1:

[1062] The terminal provides the user with a form for inputting the conditions for generating the announcement, voice setting information, and emotion information. For example, the user might write "Please create an announcement for the start and end of a new exhibition" in the input field, and select "female voice, calm tone, normal speed" as the voice setting information.

[1063] Step 2:

[1064] When a user inputs information and presses the send button, the device converts the information into JSON format and sends it to the server as an HTTP request. This request includes the text generation conditions, voice settings, and emotion information.

[1065] Step 3:

[1066] The device receives the HTTP response from the server and obtains the download link and streaming URL for the audio data, which are then displayed to the user, allowing them to download or preview the data.

[1067] User-side processing

[1068] Step 1:

[1069] The user inputs the conditions for the announcement, voice setting information, and emotional information using the terminal interface. For example, if the user is in a happy mood and wants to create an "announcement for the opening and closing of a new exhibition," the emotional information is also input at the same time.

[1070] Step 2:

[1071] The user checks the input and presses the send button, which causes the device to send the request to the server. The request is then formatted by the device and passed to the server.

[1072] Step 3:

[1073] Once the voice data is provided by the server, the user can check it through their device. The generated voice guidance is played back and the user can check the content. If they are satisfied, they can use it as is, and if necessary, they can send a request for adjustments again. For example, if they want the tone and intonation of the guidance to be a little brighter, they can set it again on their device and make another request.

[1074] Example 2

[1075] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1076] Conventional voice guidance generation systems generate voice data without considering the user's emotional information, which means they are unable to provide guidance that reflects the user's emotions or specific requirements. In addition, the tone and content of the generated voice data are fixed, making it difficult to flexibly respond to the diverse needs of users.

[1077] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1078] In this invention, the server includes means for analyzing emotional information received from a user, means for analyzing generated text, means for adjusting the tone and content of the generated text based on the analyzed emotional information, means for generating predetermined voice data based on the analyzed text, means for converting the generated voice data into an outputtable format, means for providing the generated voice data to the user, and user interface means for receiving and shaping input information from the user, thereby making it possible to provide voice guidance according to the user's emotions and specific requirements.

[1079] The "means for analyzing the generated text" is a means for analyzing the generated natural language text to understand the meaning and evaluate the content.

[1080] The "means for generating predetermined voice data" refers to a means for synthesizing voice data based on analyzed text. For example, this would be a voice synthesis system that generates voice from text.

[1081] The "means for converting the generated audio data into an outputtable format" refers to a means for converting the audio data into a format that can be used by the user (for example, a download link or a streaming URL).

[1082] The "means for providing the generated voice data to the user" refers to a means for converting the generated voice data into an appropriate format and transmitting it to the user. For example, this would be a process for sending a download link as an HTTP response.

[1083] The "means for analyzing emotional information received from a user" refers to a means for analyzing emotional information (such as voice or text) sent from a user and recognizing the emotional state of the user.

[1084] The "means for adjusting the tone and content of the generated sentence" refers to a means for adjusting the tone and content of the generated sentence based on the recognized emotion information, thereby generating an appropriate sentence according to the user's emotion and desires.

[1085] "User interface means for receiving and formatting user-input information" refers to means for providing an interface for collecting information entered by a user, formatting it appropriately, and sending it to a server. Examples of this include forms using HTML, CSS, and JavaScript.

[1086] The system of this invention generates and provides voice guidance to users according to their emotions. The system is mainly composed of three elements: a server, a terminal, and a user, each of which plays a specific role.

[1087] Server-side processing

[1088] The server receives a voice guidance generation request from the user. The request includes the guidance generation conditions, voice setting information, and user emotion information. Specifically, the server performs the following process.

[1089] Receiving and parsing the request

[1090] The server receives the HTTP request and parses the request body to extract the necessary information, using software such as Node.js.

[1091] emotion recognition

[1092] The server uses an emotion engine (e.g., EmotionRecognitionEngine) to analyze the emotion information provided by the user and recognize the user's emotional state based on the analysis results.

[1093] Sentence generation

[1094] The server provides prompts to a generative AI model (e.g., GPT-4) based on the recognized emotion information, and generates guidance text. The generative AI model adjusts the tone and content of the text according to the emotion.

[1095] Speech synthesis

[1096] The server generates voice data using a voice synthesis system (for example, Google Text-to-Speech) based on the generated text and voice setting information selected in advance by the user.

[1097] Converting and providing audio data

[1098] The generated audio data is converted into an appropriate format, a download link or streaming URL is generated to provide to the user, and the generated link is returned as an HTTP response.

[1099] Terminal side processing

[1100] The terminal provides a form for the user to input the conditions for generating the guidance message, voice setting information, and emotion information, using web technologies such as HTML, CSS, and JavaScript.

[1101] Sending user-entered information

[1102] When a user provides input information and presses the submit button, the device formats the information into JSON format and sends it to the server as an HTTP request.

[1103] Receiving and displaying results

[1104] The device receives the response from the server and displays a download link for the audio data and a streaming URL to the user.

[1105] User-side processing

[1106] Entering Settings

[1107] The user uses the terminal interface to input the conditions for the guidance message, voice setting information, and emotional information, which includes information obtained from the user's facial expression and tone of voice, for example.

[1108] Submitting and confirming your request

[1109] The user checks the input contents and presses the send button to send the request to the server.

[1110] Checking the results

[1111] The user can check the voice data provided by the server through the terminal, play the generated voice guidance, and resubmit the adjustment request if necessary.

[1112] Specific examples

[1113] For example, consider the case where a user wants to create a guide for a new exhibition. The user inputs conditions such as "Please create opening and closing guides for the new exhibition," and voice settings such as "female voice, calm tone, normal speed." In addition, the user's emotional information is also input. The device sends this information to the server, which uses an emotion engine to analyze the user's emotions. Based on the results of this analysis, a generative AI model generates sentences, and a speech synthesis system generates voice data. The generated voice data is provided by the server, and the user can play it back on their device.

[1114] Prompt Sentence Examples

[1115] "Please create opening and closing announcements for a new exhibition. Please use a female voice, calm tone, and normal speed."

[1116] This system makes it possible to easily generate and provide voice guidance that responds to the user's emotions and specific requirements.

[1117] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1118] Step 1:

[1119] Providing user input information

[1120] The user uses the terminal to input the information generation conditions, voice setting information, and emotion information. For example, the user inputs the information generation conditions such as "Please create an announcement for the start and end of a new exhibition," the voice setting information such as "female voice, calm tone, normal speed," and emotion information. The input information is collected by a form on the terminal.

[1121] Step 2:

[1122] Sending input information

[1123] The device formats the information received from the user into JSON format. The formatted data is sent to the server as an HTTP request. The input is the guidance generation conditions, voice setting information, and emotion information provided by the user, and the output is the request data sent to the server.

[1124] Step 3:

[1125] The server receives and analyzes the request

[1126] The server receives an HTTP request from the device. The server parses the received data and extracts the message generation conditions, voice setting information, and emotion information. For example, the request body is parsed using Node.js. The input is the request data sent from the device, and the output is the analyzed message generation conditions, voice setting information, and emotion information.

[1127] Step 4:

[1128] Emotional information analysis

[1129] The server analyzes the user's emotional information using an emotion engine. The emotion engine receives the user's voice or text as emotional information, analyzes it, and recognizes the emotional state. For example, an EmotionRecognitionEngine is used. The input is the analyzed emotional information, and the output is the recognized emotional state.

[1130] Step 5:

[1131] Sentence generation

[1132] The server provides a prompt to the generative AI model based on the recognized emotional information to generate a guidance sentence. For example, using a generative AI model (GPT-4), it generates a prompt such as, "User's emotion: [emotional state]. Please generate a guidance sentence based on the following conditions: [guidance sentence generation conditions]." The input is the recognized emotional state and the guidance sentence generation conditions, and the output is the generated guidance sentence.

[1133] Step 6:

[1134] Speech synthesis

[1135] The server generates voice data using a voice synthesis system based on the generated text and voice setting information selected in advance by the user. For example, Google Text-to-Speech is used. The input is the generated guidance text and voice setting information, and the output is the generated voice data.

[1136] Step 7:

[1137] Converting and providing audio data

[1138] The server converts the generated audio data into an appropriate format and generates a download link or streaming URL to provide it to the user. For example, the server stores the audio data on the server and generates a download link for it. The input is the generated audio data, and the output is a download link or streaming URL.

[1139] Step 8:

[1140] Receiving and displaying results

[1141] The terminal receives the HTTP response from the server and displays a download link or streaming URL for the audio data to the user. The user can then check the generated audio guidance and play it back as needed. The input is the download link or streaming URL provided by the server, and the output is the user's confirmation and playback of the audio guidance.

[1142] (Application example 2)

[1143] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1144] Conventional voice guidance systems are unable to take into account the emotional state of the user, making it difficult to provide flexible guidance tailored to the passenger's emotions and situation, and making it difficult to reduce passenger stress and anxiety.Autonomous vehicles are required to provide an environment in which passengers can travel in a relaxed and comfortable manner, and current systems are unable to meet this need.

[1145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1146] In this invention, the server includes means for analyzing the generated sentence, means for generating predetermined voice data based on the analyzed sentence, means for converting the generated voice data into an outputtable format, means for providing the generated voice data to a user, and means for analyzing the user's emotional information and generating a sentence according to the emotional information using a generative AI model based on the analysis results, thereby making it possible to provide flexible and appropriate voice guidance according to the emotional state of passengers.

[1147] A generative AI model is a type of artificial intelligence that generates sentences in natural language based on input data. Specifically, it generates sentences according to prompts and outputs text that reflects the user's intentions and emotions.

[1148] A prompt is an input that instructs a generative AI model on how to generate a sentence. This prompt guides the theme and tone of the sentence that the model generates.

[1149] (Emotional information) is data obtained from the user's facial expressions, tone of voice, behavior, etc., and is an indicator of the user's current emotional state.

[1150] (Voice setting information) is information about voice characteristics that the user selects in advance. Specifically, it refers to settings such as the gender, tone, and speed of the voice.

[1151] An in-vehicle display is a display device installed in an autonomous vehicle, which is used by users to input information and visually confirm guidance.

[1152] A microphone is a device for collecting sound, and in particular, its role is to obtain the user's voice as input.

[1153] A (camera) is a device that captures video and is used to collect visual data such as a user's facial expressions.

[1154] A speech synthesis system is a system that converts text data into speech data and is responsible for generating natural-sounding speech.

[1155] A user interface is an interface through which a user inputs information into a system, and is designed to improve operability.

[1156] An HTTP request is a data request sent from a client to a server, and is a basic communication method in web communication.

[1157] An HTTP response is a response from a server to a client, containing the results of a request.

[1158] The present invention aims to provide a system that provides optimal voice guidance based on the user's emotional information. This system is particularly effective in autonomous vehicles, and aims to provide a comfortable travel experience by providing flexible guidance according to the passenger's emotional state.

[1159] System configuration

[1160] The system of the present invention comprises the following main components:

[1161] 1. Server

[1162] The server receives requests from users, analyzes them, generates sentences, synthesizes speech, and provides speech data. Specifically, it uses the following software components:

[1163] Emotion engine: An open-source library for analyzing user emotional information (e.g., Librosa).

[1164] Generative AI model: GPT-4 for generating natural language sentences.

[1165] Speech synthesis system: Google Text-to-Speech API that converts text into audio data.

[1166] 2. Terminal

[1167] The terminal is equipped with an in-vehicle display, microphone, and camera, and has the ability to collect user input information and communicate with a server.

[1168] In-vehicle display: A display device that allows passengers to input information via touch.

[1169] Microphone: An input device for capturing the user's voice.

[1170] Camera: A device for capturing the user's facial expressions and acquiring emotional information.

[1171] 3. Users

[1172] Users can operate the system to receive the necessary guidance. For example, passengers can request guidance by operating the touch panel or by voice while on board.

[1173] Processing flow

[1174] When the server receives a request from a user, it performs the following process.

[1175] 1. Request Analysis

[1176] The server analyzes the request and extracts the user's emotional information, voice setting information, and guidance conditions.

[1177] 2. Emotional Information Analysis

[1178] An emotion engine is used to analyze emotional information from the user's tone of voice and facial expressions.

[1179] 3. Sentence generation

[1180] Based on the analysis results, prompts are input to a generative AI model (GPT-4) to generate guidance sentences that match the user's emotions. For example, if the user seems anxious, the following prompts are used:

[1181] "The user seems anxious. Please generate gentle, relaxing instructions."

[1182] 4. Speech Synthesis

[1183] The generated text is input into a speech synthesis system (Google Text-to-Speech API) to generate voice data.

[1184] 5. Provision of audio data

[1185] The generated voice data is converted into an appropriate format and provided to the user. Specifically, a URL link is generated and sent to the user's device.

[1186] Specific examples

[1187] For example, if a passenger is stuck in traffic and requests, "Tell me how to have fun in traffic jams," the system will process the request as follows:

[1188] The server analyzes the request and detects anxiety from the user's tone of voice, then feeds the following prompt to the generative AI model:

[1189] "Passengers seem anxious in traffic. Please generate a guide to make them feel happy."

[1190] The text generated by the AI ​​model is input into the Google Text-to-Speech API to generate audio data, and a URL link to that audio data is provided to the device so that passengers can play it back.

[1191] The above is an embodiment of the present invention. The present invention makes it possible to provide flexible and appropriate guidance according to the emotional state of passengers, thereby realizing a comfortable travel experience.

[1192] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1193] Program processing steps

[1194] Step 1:

[1195] The device receives user input information (guidance conditions, voice settings, and emotion information). This input information is collected through touch operations, voice input, facial recognition, etc. After collecting the input information, the device formats it into JSON format and sends it to the server as an HTTP request.

[1196] Step 2:

[1197] The server receives a JSON-formatted HTTP request sent from the device. It analyzes the information contained in the request (guidance conditions, voice settings, and emotion information) and decodes it into an appropriate format. The analyzed data is used in the next emotion analysis step.

[1198] Step 3:

[1199] The server uses an emotion engine to recognize the user's emotion based on the received user emotion information. Specifically, it analyzes the voice and text data to identify the user's emotional state (e.g., anxiety, relief, enjoyment, etc.). The analysis results are used in the next step.

[1200] Step 4:

[1201] The server inputs the emotional information obtained from the emotion engine into a generative AI model (GPT-4) to create a prompt. This prompt specifically indicates the theme and tone of the text to be generated. For example, when providing guidance to an anxious passenger, the prompt might say, "The user seems anxious. Please generate a gentle, relaxing message." This prompt is used to have the generative AI model generate the text.

[1202] Step 5:

[1203] The server receives the text output from the generative AI model and then inputs it into a speech synthesis system. The speech synthesis system (e.g., Google Text-to-Speech API) converts the text data into audio data. This audio data is based on the text generated by the generative AI model and is adjusted according to the voice settings previously set by the user.

[1204] Step 6:

[1205] The server receives the audio data generated by the speech synthesis system and converts it into a format that can be played on the user's device (e.g., MP3 format). The converted audio data is then generated as a download link or streaming URL.

[1206] Step 7:

[1207] The server sends an HTTP response to the device, which includes a download link or streaming URL for the generated audio data. The device receives this response and provides the user with an option to play the audio data.

[1208] Step 8:

[1209] The terminal provides an interface for the user to play the audio data, allowing the user to listen to the guidance, and displays the interface again so that the user can make further requests as necessary.

[1210] The above steps make it possible to provide optimal voice guidance based on the user's emotional information. As a specific example of operation, if a user inputs "Please tell me how to have fun even in traffic jams," this information is sent to the server, and appropriate guidance is provided using the generative AI model and speech synthesis system.

[1211] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1212] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1213] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1214] [Fourth embodiment]

[1215] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1216] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1217] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1218] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1219] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1220] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1221] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1222] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1223] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1224] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1225] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1226] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1227] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1228] Server-side processing

[1229] Request received

[1230] The server has the functionality to receive information such as the conditions for generating guidance messages and voice settings sent by the user. Specifically, it has an API that waits for HTTP requests and extracts the necessary information.

[1231] Sentence generation

[1232] Based on the received conditions, the server generates guidance text using a generative AI model (e.g., a general-purpose language model), sending requirements and keywords as prompts to the AI ​​model and receiving the generated text.

[1233] Speech synthesis

[1234] The generated text is passed to a speech synthesis system (e.g., a speech synthesis API) to generate speech data. The speech synthesis system operates based on the voice setting information (such as voice type, intonation, speed, etc.) specified by the user. The generated speech data is then returned to the server.

[1235] Providing audio data

[1236] The generated audio data is provided to the user, who then converts it to the appropriate format and generates a URL that the user can use to download or stream it.

[1237] Terminal side processing

[1238] Interface provided

[1239] The terminal provides the user with a flexible input form. This form provides an interface for entering conditions for generating guidance messages and voice setting information. For example, the user can select the content of the guidance message and the items for setting the voice through the screen of a web application.

[1240] Send request

[1241] Once the user has completed their input, the device sends this information to the server as an HTTP request, which is then formatted in JSON or other formats for efficient delivery to the server.

[1242] Receiving the results

[1243] The device receives the audio data returned from the server and displays it to the user. The user can then click a link to download the audio data or play it in their browser.

[1244] User-side processing

[1245] Setting input

[1246] The user uses the device interface to input the conditions for the announcement and voice settings. For example, when creating an announcement for a new exhibition, the user enters a sentence such as "Please create announcements for the start and end of the new exhibition" and voice settings such as "female voice, calm tone, normal speed."

[1247] Request confirmation

[1248] After checking the input, the user presses the send button to send a request to the server. The request is formatted by the terminal and passed to the server.

[1249] Check the results

[1250] Once the generated voice data is provided by the server, the user can review it through their device. If they are satisfied with the voice data, they can use it as is. If corrections or changes are required, they can re-enter the data and repeat the same process.

[1251] Specific examples

[1252] For example, consider the case of creating a guide for a new exhibition. The user enters the conditions "Please create a guide for the start and end of the new exhibition" and the voice settings "female voice, calm tone, normal speed" into a form on their device and submits it. The server receives this information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user can click the link to check the generated voice data for the guide. If necessary, they can submit a revision request again using the same procedure.

[1253] The processing flow will be explained below.

[1254] Server-side processing

[1255] Step 1: Receiving a request

[1256] The server waits for HTTP requests. When a request arrives, it analyzes the guidance generation conditions and voice setting information and decodes it into the appropriate format.

[1257] Step 2: Sentence generation

[1258] The server generates guidance text using the generative AI model, creates prompts based on the generation conditions, and passes them to the generative AI. The server receives the generated text and passes it on to the next process.

[1259] Step 3: Text-to-Speech

[1260] The server passes the generated text and voice setting information selected by the user in advance to the speech synthesis system, generates voice data, calls the speech synthesis API, and receives the generated results.

[1261] Step 4: Provide audio data

[1262] The server converts the generated audio data into an appropriate format, generates a download link or streaming URL for the user, and returns this link as an HTTP response.

[1263] Terminal side processing

[1264] Step 1: Provide an interface

[1265] The terminal displays a form for the user to input the conditions for generating guidance messages and voice settings. A user interface is provided using HTML, CSS, and JavaScript.

[1266] Step 2: Submitting a request

[1267] When a user enters information into the input form and presses the submit button, the terminal formats the information into JSON format and sends it to the server as an HTTP request.

[1268] Step 3: Receiving the results

[1269] The device receives the HTTP response from the server, obtains the download link for the audio data, and the URL for streaming, and displays this to the user.

[1270] User-side processing

[1271] Step 1: Enter settings

[1272] The user uses the terminal interface to input the conditions and voice setting information for the announcement text, for example, by writing "Announcement of the opening and closing of a new exhibition" in the input field.

[1273] Step 2: Request confirmation

[1274] The user checks the input and presses the send button, which causes the device to send the request to the server.

[1275] Step 3: Check the results

[1276] The user clicks on the download link or streaming URL for the audio data displayed on the device, plays back the generated audio guidance, and checks it. If necessary, the user can make corrections or make a new request.

[1277] Example 1

[1278] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1279] Conventional voice guidance systems have limitations in the quality and flexibility of the text and voice data they generate, making it difficult to provide voice guidance customized to the user's needs. Furthermore, there is a need for a system that can quickly generate guidance text based on user-entered conditions and provide that content as voice data efficiently and accurately.

[1280] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1281] In this invention, the server includes means for receiving a request including conditions from a user, means for generating sentences using a generative model based on the received conditions, means for synthesizing the generated sentences and converting them into voice data, means for converting the generated voice data into an outputtable format, and means for providing the generated voice data to the user, thereby enabling high-quality and customized voice guidance to be provided quickly.

[1282] A "request including conditions from the user" is a series of request contents including conditions for generating a guidance message and voice setting information specified by the user.

[1283] A "generative model" is an algorithm or system that uses natural language processing technology to generate sentences based on user requests.

[1284] The "means for generating a sentence" is a process for automatically generating a guidance sentence based on the received conditions using a generative model.

[1285] The "synthesis means for converting into voice data" is a process of converting the generated sentence into voice data using voice synthesis technology.

[1286] The "means for converting into an outputtable format" is a process for converting the generated audio data into a format that is easily accessible to the user and ready for provision.

[1287] The "means for providing to the user" refers to a method for making the generated audio data available to the user via download or streaming.

[1288] "Voice setting information" is parameter information such as the type of voice, intonation, and speed used during voice synthesis.

[1289] "Interface" is a general term for an input form or user interface that allows a user to input conditions for generating guidance messages and voice setting information.

[1290] The following describes an embodiment of the present invention. This system is composed of a server, a terminal, and a user. The specific operation of each element will be explained in detail.

[1291] Server-side processing

[1292] Request received

[1293] The server receives the guidance message generation conditions and voice setting information sent by the user. Specifically, an API that waits for requests is set at the endpoint using a web framework such as Flask or Django. When a request arrives, the server extracts the necessary information (for example, JSON data entered by the user) from the request object.

[1294] Sentence generation

[1295] The server generates guidance text using a generative AI model (for example, a model using natural language processing technology) based on the extracted conditions. The requirements and keywords are sent to the AI ​​model as prompts, and the generated text is obtained. A specific example of a generative AI model that can be used is a general-purpose natural language generation model.

[1296] Speech synthesis

[1297] The server passes the generated text to a synthesis system (e.g., a speech synthesis API) to convert it into speech data, and generates speech data. Voice setting information (voice type, intonation, speed, etc.) is sent along with the API request. The generated speech data is returned to the server.

[1298] Providing audio data

[1299] The server converts the generated audio data into a format that can be output, and generates a URL that the user can download or stream. The URL is sent back to the user as an HTTP response.

[1300] Terminal side processing

[1301] Interface provided

[1302] The terminal provides the user with an input form. Using HTML and JavaScript as a web front end, we create a form that allows users to enter conditions for generating guidance messages and voice setting information. For example, there is a text input field for the guidance message content and a drop-down menu for selecting the voice type.

[1303] Send request

[1304] Once the user has completed the input, the device formats the information into JSON format and sends it to the server as an HTTP request, using a JavaScript function such as fetch.

[1305] Receiving the results

[1306] The device receives the URL of the audio data returned by the server and displays it to the user, who can then click the provided link to download the audio data or play it in their browser.

[1307] User-side processing

[1308] Setting input

[1309] The user uses the terminal interface to input the conditions for the guide message and voice settings, such as "Please generate a guide message for a new exhibition" or "Female voice, calm tone, normal speed."

[1310] Request confirmation

[1311] After checking the input, the user presses the send button to send a request to the server. The device captures the click event of the send button, formats the information in JSON format, and passes it to the server.

[1312] Check the results

[1313] Once the generated voice data is provided by the server, the user can review it through their device. If they are satisfied with the voice data, they can use it as is. If corrections or changes are required, they can re-enter the data and repeat the same process.

[1314] Specific examples

[1315] For example, consider the case of creating a guide for a new exhibition. The user enters the conditions "Please create opening and closing guides for a new exhibition" and voice settings of "female voice, calm tone, normal speed" into a form on their device and submits it. The server receives the information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user clicks the link to check the generated voice data for the guide. If necessary, they can submit a revision request again using the same procedure.

[1316] Example prompt: "Generate a guide for a new exhibition in a female voice, with a calm tone and normal speed."

[1317] This enables the system to quickly provide high-quality, customized voice guidance.

[1318] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1319] Step 1:

[1320] The user inputs the guidance conditions.

[1321] The user uses the device's interface to input the conditions for the announcement text and voice setting information. For example, they can input content such as "Please generate an announcement text for a new exhibition" and voice setting information such as "female voice, calm tone, normal speed." The input data is temporarily stored in the device's memory.

[1322] Input: User's guidance conditions and voice setting information

[1323] Output: Saved guidance conditions and voice setting information

[1324] Step 2:

[1325] The device sends a request to the server

[1326] The device formats the conditions and setting information entered by the user into JSON format and sends it to the server as an HTTP request using the JavaScript fetch function.

[1327] Input: Saved guidance conditions and voice setting information

[1328] Output: HTTP request sent to the server

[1329] Step 3:

[1330] The server receives the request

[1331] The server receives HTTP requests sent from devices. An API endpoint built using a web framework (e.g., Flask or Django) listens for the requests and extracts data from the request object when it is received.

[1332] Input: HTTP request sent from the terminal

[1333] Output: Extracted guidance conditions and voice setting information

[1334] Step 4:

[1335] The server generates the message

[1336] The server sends a prompt to the generative AI model based on the extracted conditions, generates a guidance message, and then makes a request to the generative model (e.g., a general-purpose language model) to obtain the generated guidance message from the API.

[1337] Input: Extracted guidance conditions

[1338] Output: Generated guidance text

[1339] Step 5:

[1340] The server synthesizes the voice data

[1341] The server sends the generated guidance message and voice setting information to a speech synthesis system to convert it into voice data. The speech synthesis system (e.g., a speech synthesis API) converts the text into voice data based on the specified voice settings and returns it to the server.

[1342] Input: Generated guidance text and voice setting information

[1343] Output: Generated audio data

[1344] Step 6:

[1345] The server provides the audio data

[1346] The server converts the generated audio data into a format that allows the user to download or stream it, provides the converted audio data as a URL that the user can access, and returns the URL to the user as an HTTP response.

[1347] Input: Generated audio data

[1348] Output: A URL of the audio data that can be accessed by the user

[1349] Step 7:

[1350] The terminal displays the results to the user

[1351] The device receives the URL of the audio data returned by the server and displays it to the user, who can then click the provided link to download the audio data or play it in their browser.

[1352] Input: URL of the audio data returned from the server

[1353] Output: Displayed as a link that the user can access

[1354] In this way, by detailing the specific operations performed at each step and the inputs and outputs, the processing flow of this system becomes clear.

[1355] (Application example 1)

[1356] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1357] Currently, many food delivery services rely on text-based notifications, which can create visual constraints and disrupt user flow. Therefore, a more intuitive and real-time means of providing information is needed. Additionally, systems that can provide voice guidance customized based on individual user preferences and settings are anticipated.

[1358] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1359] In this invention, the server includes a means for receiving a request from a user, a means for generating guidance sentences using a generative AI model based on the received request, and a means for analyzing the generated guidance sentences, thereby enabling the user to receive more personalized voice guidance in real time.

[1360] "User" refers to the person who uses the system to receive voice guidance.

[1361] A "means for receiving a request" is a component of the system that has the function of receiving information entered by a user.

[1362] "Generative AI models" refer to algorithms and techniques that use artificial intelligence to generate text.

[1363] The "means for generating guidance text" is the part of the system that automatically generates guidance text based on the received request.

[1364] The "means for analyzing the guidance text" is a system function that analyzes the generated text and obtains information for creating voice data based on the content of the text.

[1365] The "means for generating predetermined voice data" is the part of the system that generates voice data in a predetermined format based on the analyzed sentence.

[1366] A "means for converting to an outputtable format" is a system component that has the ability to convert the generated audio data into a format that can be used by the user.

[1367] The "means for providing voice data" is a function of the system for providing the generated voice data to the user.

[1368] "Voice setting information" refers to setting information related to the voice selected in advance by the user, such as the type of voice, intonation, and speed.

[1369] "Means operating through a user interface" means the part of the system that has an interface through which a user can directly interact with to input and send requests.

[1370] Server-side processing

[1371] The server receives a request from the user, generates a guide message based on the request, and creates and provides audio data. Specifically, the server follows the steps below.

[1372] First, the server receives a request from the user, including the conditions for generating the guidance text and voice setting information. At this time, the information is sent via an HTTP request, and the server uses an API to extract that information. Next, based on the received conditions, the generative AI model is used to generate the guidance text. At this time, the requirements and keywords are sent to the generative AI model as prompts, and the generated text is received. The generated text is passed to the speech synthesis system, which generates voice data. The speech synthesis system operates based on the voice setting information specified by the user (voice type, intonation, speed, etc.). The generated voice data is then returned to the server, converted into an appropriate format, and provided via a URL that the user can download or stream.

[1373] Terminal side processing

[1374] The device provides the user with an input form, formats the request, and sends it to the server. The user enters the conditions for generating the guidance text and audio setting information through the web application screen. Once the input is complete, the device formats this information in JSON format or similar and sends it to the server as an HTTP request. The device receives the audio data returned from the server and displays it to the user. The user can click the provided link to download the audio data or play the audio in their browser.

[1375] User-side processing

[1376] The user uses the device interface to input the conditions for the guidance text and voice settings. For example, in the case of a food delivery service, the user might input the condition "Please create a guidance message for the delivery status" and voice setting information such as "Male voice, calm tone, normal speed." After checking the input information, the user presses the send button to send a request to the server. The request is formatted by the device and passed to the server. When the generated voice data is provided by the server, the user checks it through the device. After checking the voice data, if the user is satisfied, they can use it as is. If corrections or changes are needed, the user can enter the information again and repeat the same process.

[1377] Specific examples

[1378] For example, consider a food delivery service where a user wants to provide delivery status information. The user enters a condition such as "Please create a delivery status information" and voice settings such as "male voice, calm tone, normal speed" into a form on their device and submits it. The server receives this information, generates text using a generative AI, and generates voice data using a voice synthesis system. The generated voice data is returned to the server, and a URL is provided that the user can easily access. The user can click the link to check the generated voice data of the delivery status.

[1379] Prompt Sentence Examples

[1380] "The delivery person has left the restaurant. The estimated delivery time is 30 minutes."

[1381] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1382] Step 1:

[1383] The server receives guidance message generation conditions and voice setting information from the user.

[1384] Specifically, the information entered by the user using the device interface is sent to the server as an HTTP request, which is sent in JSON format and received by the server.

[1385] Input: Guidance message generation conditions, voice setting information (HTTP request sent in JSON format)

[1386] Output: Guidance message generation conditions, voice setting information (stored as variables within the server)

[1387] Step 2:

[1388] The server generates the guidance text using a generative AI model based on the received guidance text generation conditions.

[1389] For example, OpenAI's GPT-3 is used as the generative AI model. In this case, the conditions for generating the guidance sentence are sent to the AI ​​model as a prompt, and the generated sentence is received.

[1390] Input: Guidance sentence generation conditions (sent as prompt sentence to the AI ​​model)

[1391] Output: Generated guidance text

[1392] Step 3:

[1393] The server analyzes the generated guidance text and obtains information for generating predetermined voice data based on the text.

[1394] Specifically, the content of the guidance text is analyzed and the parameters necessary for the voice synthesis system are extracted.

[1395] Input: Generated guidance text

[1396] Output: Speech synthesis parameters (voice type, intonation, speed, etc.)

[1397] Step 4:

[1398] The server uses a voice synthesis system to generate predetermined voice data based on the voice synthesis parameters.

[1399] The Google Cloud Text-to-Speech API is used for voice synthesis. The generated text and voice setting information are sent to the voice synthesis system, which then generates the voice data.

[1400] Input: Generated guidance text, voice setting information

[1401] Output: Generated audio data

[1402] Step 5:

[1403] The server converts the generated audio data into an outputtable format.

[1404] This process involves encoding the audio data into MP3 format and converting it so that users can download or stream it.

[1405] Input: Generated audio data (internal format)

[1406] Output: Audio data in outputtable format (MP3 format)

[1407] Step 6:

[1408] The server provides the converted audio data to the user.

[1409] Specifically, it hosts the audio data and generates a URL that users can access, which is then provided to the user.

[1410] Input: Audio data in an outputtable format (MP3 format)

[1411] Output: Access URL for audio data

[1412] Step 7:

[1413] The terminal displays the access URL for the audio data received from the server to the user.

[1414] Users can click on the provided URL to view the generated audio data.

[1415] Input: Access URL for audio data

[1416] Output: Access URL for the displayed audio data (user interface)

[1417] The following are specific examples of prompt sentences:

[1418] Example prompt sentence:

[1419] "The delivery person has left the restaurant. The estimated delivery time is 30 minutes."

[1420] A guidance sentence is generated based on this prompt sentence, and the subsequent series of processes are executed.

[1421] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1422] Server-side processing

[1423] Request received

[1424] The server receives a request for generating voice guidance from the user. The request includes the conditions for generating guidance, voice setting information, and the user's emotional information. The server analyzes this information and decodes it into an appropriate format.

[1425] Emotion Recognition and Sentence Generation

[1426] The server uses an emotion engine to recognize the user's emotion based on the received user emotion information. The emotion engine analyzes emotions from the user's voice and text input and provides the recognition results to a generative AI model. The server then uses the generative AI model to generate guidance text based on the emotion information. In this text generation process, the tone and content of the text are adjusted according to the emotion information.

[1427] Speech synthesis

[1428] The server uses a speech synthesis system to generate voice data based on the analyzed text and voice setting information selected in advance by the user. The generated voice data is then returned to the server.

[1429] Providing audio data

[1430] The server converts the generated audio data into an appropriate format, generates a download link or streaming URL for the user, and returns this link as an HTTP response.

[1431] Terminal side processing

[1432] Interface provided

[1433] The terminal provides the user with a form for inputting the conditions for generating guidance messages, voice settings, and emotion information. The user interface uses HTML, CSS, and JavaScript, and allows voice and text input.

[1434] Send request

[1435] When a user enters information and presses the send button, the device formats the entered information into JSON format and sends it to the server as an HTTP request.

[1436] Receiving the results

[1437] The device receives the HTTP response from the server and obtains the download link and streaming URL for the audio data, which are then displayed to the user, allowing them to download or preview the data.

[1438] User-side processing

[1439] Setting input

[1440] The user uses the terminal interface to input the conditions for the guidance message, voice setting information, and emotional information, which includes information obtained from the user's facial expression and tone of voice, for example.

[1441] Request confirmation

[1442] The user checks the input and presses the send button to send a request to the server. The request is formatted by the terminal and passed to the server.

[1443] Check the results

[1444] Once the audio data is provided by the server, the user can check the download link or streaming URL of the audio data through their device, play the generated audio guidance to check the content, and resubmit the adjustment request if necessary.

[1445] Specific examples

[1446] For example, consider the case where a user wants to create a guide for a new exhibition. The user enters the condition "Please create opening and closing instructions for the new exhibition," along with voice setting information (female voice, calm tone, normal speed) and emotional information into a form on the device. The device sends this information to the server, which uses an emotion engine to analyze the user's emotions. Based on the analysis results, a generative AI model generates sentences, and a speech synthesis system generates voice data. The generated voice data is provided by the server, and the user can play it back on the device and request corrections as needed.

[1447] The processing flow will be explained below.

[1448] Server-side processing

[1449] Step 1:

[1450] The server waits for HTTP requests and receives requests from users, including conditions for generating announcements, voice settings, and emotion information. For example, the request may include content such as "Please create announcements for the start and end of a new exhibition" and voice settings such as "female voice, calm tone, normal speed."

[1451] Step 2:

[1452] The server extracts emotion information from the received request and passes it to the emotion engine. The emotion engine recognizes emotions from the user's input (text or voice) and provides the results to the generative AI model. For example, if the user is recognized as happy, the tone of the text will be adjusted to be more cheerful and positive.

[1453] Step 3:

[1454] The server uses a generative AI model to generate guidance text based on the emotional information returned by the emotion engine. For example, a more friendly tone of text is generated based on the emotional information. This text is then converted into audio data in the next step.

[1455] Step 4:

[1456] The server passes the generated text and user-specified voice settings to a speech synthesis system, which then generates voice data. The speech synthesis system then generates voice data with the specified voice type, intonation, and speed, for example. The generated voice data is then returned to the server.

[1457] Step 5:

[1458] The server converts the generated audio data into the appropriate format and generates a download link or streaming URL, which is sent back to the user as an HTTP response.

[1459] Terminal side processing

[1460] Step 1:

[1461] The terminal provides the user with a form for inputting the conditions for generating the announcement, voice setting information, and emotion information. For example, the user might write "Please create an announcement for the start and end of a new exhibition" in the input field, and select "female voice, calm tone, normal speed" as the voice setting information.

[1462] Step 2:

[1463] When a user inputs information and presses the send button, the device converts the information into JSON format and sends it to the server as an HTTP request. This request includes the text generation conditions, voice settings, and emotion information.

[1464] Step 3:

[1465] The device receives the HTTP response from the server and obtains the download link and streaming URL for the audio data, which are then displayed to the user, allowing them to download or preview the data.

[1466] User-side processing

[1467] Step 1:

[1468] The user inputs the conditions for the announcement, voice setting information, and emotional information using the terminal interface. For example, if the user is in a happy mood and wants to create an "announcement for the opening and closing of a new exhibition," the emotional information is also input at the same time.

[1469] Step 2:

[1470] The user checks the input and presses the send button, which causes the device to send the request to the server. The request is then formatted by the device and passed to the server.

[1471] Step 3:

[1472] Once the voice data is provided by the server, the user can check it through their device. The generated voice guidance is played back and the user can check the content. If they are satisfied, they can use it as is, and if necessary, they can send a request for adjustments again. For example, if they want the tone and intonation of the guidance to be a little brighter, they can set it again on their device and make another request.

[1473] Example 2

[1474] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1475] Conventional voice guidance generation systems generate voice data without considering the user's emotional information, which means they are unable to provide guidance that reflects the user's emotions or specific requirements. In addition, the tone and content of the generated voice data are fixed, making it difficult to flexibly respond to the diverse needs of users.

[1476] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1477] In this invention, the server includes means for analyzing emotional information received from a user, means for analyzing generated text, means for adjusting the tone and content of the generated text based on the analyzed emotional information, means for generating predetermined voice data based on the analyzed text, means for converting the generated voice data into an outputtable format, means for providing the generated voice data to the user, and user interface means for receiving and shaping input information from the user, thereby making it possible to provide voice guidance according to the user's emotions and specific requirements.

[1478] The "means for analyzing the generated text" is a means for analyzing the generated natural language text to understand the meaning and evaluate the content.

[1479] The "means for generating predetermined voice data" refers to a means for synthesizing voice data based on analyzed text. For example, this would be a voice synthesis system that generates voice from text.

[1480] The "means for converting the generated audio data into an outputtable format" refers to a means for converting the audio data into a format that can be used by the user (for example, a download link or a streaming URL).

[1481] The "means for providing the generated voice data to the user" refers to a means for converting the generated voice data into an appropriate format and transmitting it to the user. For example, this would be a process for sending a download link as an HTTP response.

[1482] The "means for analyzing emotional information received from a user" refers to a means for analyzing emotional information (such as voice or text) sent from a user and recognizing the emotional state of the user.

[1483] The "means for adjusting the tone and content of the generated sentence" refers to a means for adjusting the tone and content of the generated sentence based on the recognized emotion information, thereby generating an appropriate sentence according to the user's emotion and desires.

[1484] "User interface means for receiving and formatting user-input information" refers to means for providing an interface for collecting information entered by a user, formatting it appropriately, and sending it to a server. Examples of this include forms using HTML, CSS, and JavaScript.

[1485] The system of this invention generates and provides voice guidance to users according to their emotions. The system is mainly composed of three elements: a server, a terminal, and a user, each of which plays a specific role.

[1486] Server-side processing

[1487] The server receives a voice guidance generation request from the user. The request includes the guidance generation conditions, voice setting information, and user emotion information. Specifically, the server performs the following process.

[1488] Receiving and parsing the request

[1489] The server receives the HTTP request and parses the request body to extract the necessary information, using software such as Node.js.

[1490] emotion recognition

[1491] The server uses an emotion engine (e.g., EmotionRecognitionEngine) to analyze the emotion information provided by the user and recognize the user's emotional state based on the analysis results.

[1492] Sentence generation

[1493] The server provides prompts to a generative AI model (e.g., GPT-4) based on the recognized emotion information, and generates guidance text. The generative AI model adjusts the tone and content of the text according to the emotion.

[1494] Speech synthesis

[1495] The server generates voice data using a voice synthesis system (for example, Google Text-to-Speech) based on the generated text and voice setting information selected in advance by the user.

[1496] Converting and providing audio data

[1497] The generated audio data is converted into an appropriate format, a download link or streaming URL is generated to provide to the user, and the generated link is returned as an HTTP response.

[1498] Terminal side processing

[1499] The terminal provides a form for the user to input the conditions for generating the guidance message, voice setting information, and emotion information, using web technologies such as HTML, CSS, and JavaScript.

[1500] Sending user-entered information

[1501] When a user provides input information and presses the submit button, the device formats the information into JSON format and sends it to the server as an HTTP request.

[1502] Receiving and displaying results

[1503] The device receives the response from the server and displays a download link for the audio data and a streaming URL to the user.

[1504] User-side processing

[1505] Entering Settings

[1506] The user uses the terminal interface to input the conditions for the guidance message, voice setting information, and emotional information, which includes information obtained from the user's facial expression and tone of voice, for example.

[1507] Submitting and confirming your request

[1508] The user checks the input contents and presses the send button to send the request to the server.

[1509] Checking the results

[1510] The user can check the voice data provided by the server through the terminal, play the generated voice guidance, and resubmit the adjustment request if necessary.

[1511] Specific examples

[1512] For example, consider the case where a user wants to create a guide for a new exhibition. The user inputs conditions such as "Please create opening and closing guides for the new exhibition," and voice settings such as "female voice, calm tone, normal speed." In addition, the user's emotional information is also input. The device sends this information to the server, which uses an emotion engine to analyze the user's emotions. Based on the results of this analysis, a generative AI model generates sentences, and a speech synthesis system generates voice data. The generated voice data is provided by the server, and the user can play it back on their device.

[1513] Prompt Sentence Examples

[1514] "Please create opening and closing announcements for a new exhibition. Please use a female voice, calm tone, and normal speed."

[1515] This system makes it possible to easily generate and provide voice guidance that responds to the user's emotions and specific requirements.

[1516] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1517] Step 1:

[1518] Providing user input information

[1519] The user uses the terminal to input the information generation conditions, voice setting information, and emotion information. For example, the user inputs the information generation conditions such as "Please create an announcement for the start and end of a new exhibition," the voice setting information such as "female voice, calm tone, normal speed," and emotion information. The input information is collected by a form on the terminal.

[1520] Step 2:

[1521] Sending input information

[1522] The device formats the information received from the user into JSON format. The formatted data is sent to the server as an HTTP request. The input is the guidance generation conditions, voice setting information, and emotion information provided by the user, and the output is the request data sent to the server.

[1523] Step 3:

[1524] The server receives and analyzes the request

[1525] The server receives an HTTP request from the device. The server parses the received data and extracts the message generation conditions, voice setting information, and emotion information. For example, the request body is parsed using Node.js. The input is the request data sent from the device, and the output is the analyzed message generation conditions, voice setting information, and emotion information.

[1526] Step 4:

[1527] Emotional information analysis

[1528] The server analyzes the user's emotional information using an emotion engine. The emotion engine receives the user's voice or text as emotional information, analyzes it, and recognizes the emotional state. For example, an EmotionRecognitionEngine is used. The input is the analyzed emotional information, and the output is the recognized emotional state.

[1529] Step 5:

[1530] Sentence generation

[1531] The server provides a prompt to the generative AI model based on the recognized emotional information to generate a guidance sentence. For example, using a generative AI model (GPT-4), it generates a prompt such as, "User's emotion: [emotional state]. Please generate a guidance sentence based on the following conditions: [guidance sentence generation conditions]." The input is the recognized emotional state and the guidance sentence generation conditions, and the output is the generated guidance sentence.

[1532] Step 6:

[1533] Speech synthesis

[1534] The server generates voice data using a voice synthesis system based on the generated text and voice setting information selected in advance by the user. For example, Google Text-to-Speech is used. The input is the generated guidance text and voice setting information, and the output is the generated voice data.

[1535] Step 7:

[1536] Converting and providing audio data

[1537] The server converts the generated audio data into an appropriate format and generates a download link or streaming URL to provide it to the user. For example, the server stores the audio data on the server and generates a download link for it. The input is the generated audio data, and the output is a download link or streaming URL.

[1538] Step 8:

[1539] Receiving and displaying results

[1540] The terminal receives the HTTP response from the server and displays a download link or streaming URL for the audio data to the user. The user can then check the generated audio guidance and play it back as needed. The input is the download link or streaming URL provided by the server, and the output is the user's confirmation and playback of the audio guidance.

[1541] (Application example 2)

[1542] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1543] Conventional voice guidance systems are unable to take into account the emotional state of the user, making it difficult to provide flexible guidance tailored to the passenger's emotions and situation, and making it difficult to reduce passenger stress and anxiety.Autonomous vehicles are required to provide an environment in which passengers can travel in a relaxed and comfortable manner, and current systems are unable to meet this need.

[1544] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1545] In this invention, the server includes means for analyzing the generated sentence, means for generating predetermined voice data based on the analyzed sentence, means for converting the generated voice data into an outputtable format, means for providing the generated voice data to a user, and means for analyzing the user's emotional information and generating a sentence according to the emotional information using a generative AI model based on the analysis results, thereby making it possible to provide flexible and appropriate voice guidance according to the emotional state of passengers.

[1546] A generative AI model is a type of artificial intelligence that generates sentences in natural language based on input data. Specifically, it generates sentences according to prompts and outputs text that reflects the user's intentions and emotions.

[1547] A prompt is an input that instructs a generative AI model on how to generate a sentence. This prompt guides the theme and tone of the sentence that the model generates.

[1548] (Emotional information) is data obtained from the user's facial expressions, tone of voice, behavior, etc., and is an indicator of the user's current emotional state.

[1549] (Voice setting information) is information about voice characteristics that the user selects in advance. Specifically, it refers to settings such as the gender, tone, and speed of the voice.

[1550] An in-vehicle display is a display device installed in an autonomous vehicle, which is used by users to input information and visually confirm guidance.

[1551] A microphone is a device for collecting sound, and in particular, its role is to obtain the user's voice as input.

[1552] A (camera) is a device that captures video and is used to collect visual data such as a user's facial expressions.

[1553] A speech synthesis system is a system that converts text data into speech data and is responsible for generating natural-sounding speech.

[1554] A user interface is an interface through which a user inputs information into a system, and is designed to improve operability.

[1555] An HTTP request is a data request sent from a client to a server, and is a basic communication method in web communication.

[1556] An HTTP response is a response from a server to a client, containing the results of a request.

[1557] The present invention aims to provide a system that provides optimal voice guidance based on the user's emotional information. This system is particularly effective in autonomous vehicles, and aims to provide a comfortable travel experience by providing flexible guidance according to the passenger's emotional state.

[1558] System configuration

[1559] The system of the present invention comprises the following main components:

[1560] 1. Server

[1561] The server receives requests from users, analyzes them, generates sentences, synthesizes speech, and provides speech data. Specifically, it uses the following software components:

[1562] Emotion engine: An open-source library for analyzing user emotional information (e.g., Librosa).

[1563] Generative AI model: GPT-4 for generating natural language sentences.

[1564] Speech synthesis system: Google Text-to-Speech API that converts text into audio data.

[1565] 2. Terminal

[1566] The terminal is equipped with an in-vehicle display, microphone, and camera, and has the ability to collect user input information and communicate with a server.

[1567] In-vehicle display: A display device that allows passengers to input information via touch.

[1568] Microphone: An input device for capturing the user's voice.

[1569] Camera: A device for capturing the user's facial expressions and acquiring emotional information.

[1570] 3. Users

[1571] Users can operate the system to receive the necessary guidance. For example, passengers can request guidance by operating the touch panel or by voice while on board.

[1572] Processing flow

[1573] When the server receives a request from a user, it performs the following process.

[1574] 1. Request Analysis

[1575] The server analyzes the request and extracts the user's emotional information, voice setting information, and guidance conditions.

[1576] 2. Emotional Information Analysis

[1577] An emotion engine is used to analyze emotional information from the user's tone of voice and facial expressions.

[1578] 3. Sentence generation

[1579] Based on the analysis results, prompts are input to a generative AI model (GPT-4) to generate guidance sentences that match the user's emotions. For example, if the user seems anxious, the following prompts are used:

[1580] "The user seems anxious. Please generate gentle, relaxing instructions."

[1581] 4. Speech Synthesis

[1582] The generated text is input into a speech synthesis system (Google Text-to-Speech API) to generate voice data.

[1583] 5. Provision of audio data

[1584] The generated voice data is converted into an appropriate format and provided to the user. Specifically, a URL link is generated and sent to the user's device.

[1585] Specific examples

[1586] For example, if a passenger is stuck in traffic and requests, "Tell me how to have fun in traffic jams," the system will process the request as follows:

[1587] The server analyzes the request and detects anxiety from the user's tone of voice, then feeds the following prompt to the generative AI model:

[1588] "Passengers seem anxious in traffic. Please generate a guide to make them feel happy."

[1589] The text generated by the AI ​​model is input into the Google Text-to-Speech API to generate audio data, and a URL link to that audio data is provided to the device so that passengers can play it back.

[1590] The above is an embodiment of the present invention. The present invention makes it possible to provide flexible and appropriate guidance according to the emotional state of passengers, thereby realizing a comfortable travel experience.

[1591] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1592] Program processing steps

[1593] Step 1:

[1594] The device receives user input information (guidance conditions, voice settings, and emotion information). This input information is collected through touch operations, voice input, facial recognition, etc. After collecting the input information, the device formats it into JSON format and sends it to the server as an HTTP request.

[1595] Step 2:

[1596] The server receives a JSON-formatted HTTP request sent from the device. It analyzes the information contained in the request (guidance conditions, voice settings, and emotion information) and decodes it into an appropriate format. The analyzed data is used in the next emotion analysis step.

[1597] Step 3:

[1598] The server uses an emotion engine to recognize the user's emotion based on the received user emotion information. Specifically, it analyzes the voice and text data to identify the user's emotional state (e.g., anxiety, relief, enjoyment, etc.). The analysis results are used in the next step.

[1599] Step 4:

[1600] The server inputs the emotional information obtained from the emotion engine into a generative AI model (GPT-4) to create a prompt. This prompt specifically indicates the theme and tone of the text to be generated. For example, when providing guidance to an anxious passenger, the prompt might say, "The user seems anxious. Please generate a gentle, relaxing message." This prompt is used to have the generative AI model generate the text.

[1601] Step 5:

[1602] The server receives the text output from the generative AI model and then inputs it into a speech synthesis system. The speech synthesis system (e.g., Google Text-to-Speech API) converts the text data into audio data. This audio data is based on the text generated by the generative AI model and is adjusted according to the voice settings previously set by the user.

[1603] Step 6:

[1604] The server receives the audio data generated by the speech synthesis system and converts it into a format that can be played on the user's device (e.g., MP3 format). The converted audio data is then generated as a download link or streaming URL.

[1605] Step 7:

[1606] The server sends an HTTP response to the device, which includes a download link or streaming URL for the generated audio data. The device receives this response and provides the user with an option to play the audio data.

[1607] Step 8:

[1608] The terminal provides an interface for the user to play the audio data, allowing the user to listen to the guidance, and displays the interface again so that the user can make further requests as necessary.

[1609] The above steps make it possible to provide optimal voice guidance based on the user's emotional information. As a specific example of operation, if a user inputs "Please tell me how to have fun even in traffic jams," this information is sent to the server, and appropriate guidance is provided using the generative AI model and speech synthesis system.

[1610] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1611] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1612] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1613] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1614] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1615] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1616] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1617] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1618] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1619] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1620] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1621] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1622] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1623] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1624] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1625] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1626] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1627] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1628] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1629] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1630] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1631] The following is further disclosed regarding the above embodiment.

[1632] (Claim 1)

[1633] means for analyzing the generated text;

[1634] means for generating predetermined voice data based on the analyzed sentence;

[1635] means for converting the generated audio data into an outputtable format;

[1636] means for providing the generated voice data to a user;

[1637] A system including:

[1638] (Claim 2)

[1639] 2. The system according to claim 1, wherein the means for generating the audio data operates based on audio setting information selected in advance by a user.

[1640] (Claim 3)

[1641] 2. The system of claim 1, wherein the means for receiving a request from a user operates through a user interface.

[1642] "Example 1"

[1643] (Claim 1)

[1644] means for receiving a request including a condition from a user;

[1645] means for generating a sentence using a generative model based on the received conditions;

[1646] a synthesis means for converting the generated sentence into speech data;

[1647] means for converting the generated audio data into an outputtable format;

[1648] means for providing the generated voice data to a user;

[1649] A system including:

[1650] (Claim 2)

[1651] 2. The system according to claim 1, wherein the means for generating the audio data operates based on audio setting information selected in advance by a user.

[1652] (Claim 3)

[1653] 2. The system according to claim 1, further comprising means for inputting a user's requested condition via an interface and transmitting the condition to a server.

[1654] "Application Example 1"

[1655] (Claim 1)

[1656] means for receiving a request from a user;

[1657] A means for generating guidance text using a generative AI model based on the received request;

[1658] A means for analyzing the generated guidance text;

[1659] means for generating predetermined voice data based on the analyzed sentence;

[1660] means for converting the generated audio data into an outputtable format;

[1661] means for providing the generated voice data to a user;

[1662] A system including:

[1663] (Claim 2)

[1664] 2. The system according to claim 1, wherein the means for generating the audio data operates based on audio setting information selected in advance by a user.

[1665] (Claim 3)

[1666] 2. The system of claim 1, wherein the means for receiving a request from a user operates through a user interface.

[1667] "Example 2: Combining Emotion Engines"

[1668] (Claim 1)

[1669] means for analyzing the generated text;

[1670] means for generating predetermined voice data based on the analyzed sentence;

[1671] means for converting the generated audio data into an outputtable format;

[1672] means for providing the generated voice data to a user;

[1673] means for analyzing emotion information received from a user;

[1674] means for adjusting the tone and content of the generated text based on the analyzed emotional information;

[1675] a user interface means for receiving and formatting user input;

[1676] A system including:

[1677] (Claim 2)

[1678] 2. The system according to claim 1, wherein the means for generating the audio data operates based on audio setting information selected in advance by a user.

[1679] (Claim 3)

[1680] 2. The system of claim 1, wherein the means for receiving a request from a user operates through a user interface.

[1681] "Application example 2 when combining emotion engines"

[1682] (Claim 1)

[1683] means for analyzing the generated text;

[1684] means for generating predetermined voice data based on the analyzed sentence;

[1685] means for converting the generated audio data into an outputtable format;

[1686] means for providing the generated voice data to a user;

[1687] A means for analyzing the user's emotional information and generating sentences according to the emotional information using a generation AI model based on the analysis results;

[1688] A system including:

[1689] (Claim 2)

[1690] 2. The system according to claim 1, wherein the means for generating the audio data operates based on audio setting information selected in advance by a user.

[1691] (Claim 3)

[1692] 2. The system of claim 1, wherein the means for receiving a request from a user operates through a user interface.

[1693] (Claim 4)

[1694] 2. The system according to claim 1, wherein the means for collecting the user's emotional information uses an in-vehicle display, a camera, and a microphone.

[1695] (Claim 5)

[1696] The system described in claim 1, characterized in that the generative AI model inputs prompt sentences according to the passenger's emotional state and generates guidance sentences that match the emotion. [Explanation of symbols]

[1697] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for analyzing the generated text; means for generating predetermined voice data based on the analyzed sentence; means for converting the generated audio data into an outputtable format; means for providing the generated voice data to a user; A system including:

2. 2. The system according to claim 1, wherein the means for generating the audio data operates based on audio setting information previously selected by a user.

3. 2. The system of claim 1, wherein the means for receiving a request from a user operates through a user interface.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A