System

The system allows users to personalize vehicle interactions by setting personality and voice characteristics, addressing the limitations of conventional interfaces to enhance user engagement and experience.

JP2026016203APending Publication Date: 2026-02-03SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117293
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Conventional vehicle interactive interfaces provide limited information and instructions, resulting in one-way interaction and lack of personalization, making it difficult for users to develop attachment and familiarity with their vehicles.

Method used

A system that allows users to set the vehicle's personality and voice characteristics, using input means, analysis means, voice generation means, transmission means, conversion means, uploading means, and interaction means to generate personalized and interactive driving experiences.

Benefits of technology

Enables users to freely customize the vehicle's personality and voice, enhancing the driving experience through interactive communication and personalized dialogue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016203000001_ABST
    Figure 2026016203000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: input means for a user to set a personality of a vehicle and a feature of a voice; analysis means for analyzing information set by the input means; voice generation means for generating voice data based on the information analyzed by the analysis means; and transmission means for transmitting the generated voice data to a terminal of the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional vehicle interactive interfaces generally provide limited information and instructions from the vehicle, resulting in one-way interaction between the user and the vehicle, making it difficult to provide a personalized driving experience. Furthermore, there are also limited ways for users to increase their attachment to and familiarity with their vehicle. Therefore, to make the driving experience richer and more enjoyable, a system is needed that allows users to set the vehicle's personality and voice characteristics and then have a dialogue based on these. [Means for solving the problem]

[0005] The present invention provides a system including an input means for a user to set the vehicle's personality and voice characteristics, an analysis means for analyzing the information set by the input means, a voice generation means for generating voice data based on the information analyzed by the analysis means, and a transmission means for transmitting the generated voice data to a user's terminal. The system further includes a conversion means for analyzing the voice data transmitted by the transmission means and converting it into a format usable by an in-vehicle system, an upload means for uploading the converted voice data to the in-vehicle system, and an interaction means for performing an interaction function using the voice data uploaded by the upload means. This system allows a user to freely customize the vehicle's personality and voice and enjoy a personalized interactive driving experience.

[0006] "User" or "Consumer" means the owner or driver of a vehicle who utilizes the system to configure the personality and voice characteristics of the vehicle.

[0007] "Vehicle" refers to a land vehicle equipped with an engine or electric motor used as a means of transportation, such as an automobile.

[0008] "Personality" is one of the characteristics that classify the human feelings and behaviors that the vehicle imitates, and refers to types such as "friendly" and "cool."

[0009] "Voice characteristics" refer to various characteristics of a voice, such as the gender of the voice (male, female) or the age group of the voice (youthful, deep).

[0010] "Input means" refers to devices or interfaces through which a user configures the vehicle's personality and voice characteristics, including, for example, smartphone apps and computer forms.

[0011] The "analysis means" refers to a component having the function of receiving information set by the input means, analyzing it, and determining voice generation parameters suitable for the vehicle.

[0012] "Speech generation means" refers to a component that has the function of generating realistic and natural voice data based on the voice generation parameters determined by the analysis means.

[0013] "Transmission means" refers to a component that has the function of transmitting the generated voice data to the user's terminal.

[0014] "Terminal" refers to a smartphone, tablet, computer, or in-vehicle system used by a user.

[0015] The "conversion means" refers to a component that has the function of analyzing the voice data transmitted by the transmission means and converting it into a format that can be used by the in-vehicle system.

[0016] The "uploading means" refers to a component that has the function of uploading the audio data converted by the conversion means to the in-vehicle system.

[0017] The "interaction means" refers to a component having a function of using the voice data uploaded by the upload means to execute a dialogue between the vehicle and the user. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] MODE FOR CARRYING OUT THE INVENTION

[0040] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. To implement this system, the following program processing is performed.

[0041] User settings processing

[0042] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[0043] The terminal collects the input configuration information and converts it into a data packet, which is then sent to the server.

[0044] Data analysis and speech generation

[0045] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. This includes selecting an appropriate template from existing voice samples.

[0046] A speech synthesis engine in the server then generates realistic, natural-sounding speech data based on the determined speech generation parameters, for example, using neural network-based techniques.

[0047] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[0048] Providing voice data and implementing dialogue functions

[0049] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[0050] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[0051] Finally, while actually driving, users can enjoy interacting with the vehicle using the voice generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?" In this way, conversational communication is realized.

[0052] Specific examples

[0053] As an example, consider the following scenario: A user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0054] When the user starts the engine while driving, the vehicle will speak to the user in a friendly voice, asking, "Good morning! Where are you going today?" In this way, the present invention enables interactive communication with the vehicle, enriching the user's driving experience.

[0055] The processing flow will be explained below.

[0056] Step 1:

[0057] The user launches a dedicated app on their smartphone or computer, and then enters the vehicle's personality and voice characteristics on the app's settings screen. Specifically, the user selects options such as "friendly personality," "male voice," or "youthful voice."

[0058] Step 2:

[0059] The device collects the configuration information entered by the user and converts it into a data packet, which contains details such as the vehicle's personality, the gender of the voice, and the tone of the voice.

[0060] Step 3:

[0061] The device generates data packets and sends them to the server over Wi-Fi or mobile data networks.

[0062] Step 4:

[0063] The server receives the setting information sent from the terminal and stores the received data in a database within the server.

[0064] Step 5:

[0065] The server analyzes the received data. Specifically, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0066] Step 6:

[0067] The server's internal voice synthesis engine generates realistic, natural voice data based on the determined voice generation parameters. This process uses voice synthesis technologies such as neural networks.

[0068] Step 7:

[0069] The server stores the generated voice data and generates new data packets to send to the user's terminal.

[0070] Step 8:

[0071] The server then sends the generated data packets to the user's device, again over Wi-Fi or a mobile data network.

[0072] Step 9:

[0073] The terminal acquires the voice data transmitted from the server.

[0074] Step 10:

[0075] The device analyzes the received voice data and converts it into a format that can be used by the in-vehicle system, specifically into common audio file formats such as WAV and MP3.

[0076] Step 11:

[0077] The device then uploads the converted audio data to the vehicle's system via Bluetooth, USB connection, or other means.

[0078] Step 12:

[0079] The in-car system plays back the audio data while the user is driving. For example, when the engine is started, the vehicle may say, "Good morning. Where are you going today?"

[0080] Through the above steps, users can freely customize the vehicle's personality and voice characteristics, enjoying a personalized interactive driving experience.

[0081] Example 1

[0082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0083] Conventional in-vehicle systems have made it difficult for users to enjoy a personalized experience through interaction with the vehicle. Furthermore, they lacked a means to freely set the vehicle's personality and voice characteristics and generate natural-sounding voice data based on those settings. As a result, users could only receive uniform, standardized voice guidance, limiting their driving experience. To solve this issue, a system was needed that could flexibly generate voice data based on parameters set by the user and provide a personalized experience through interaction with the vehicle.

[0084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0085] In this invention, the server includes input means for a user to set the vehicle's personality and voice characteristics, transmission means for converting the information set by the input means into data packets and transmitting the data to the server, analysis means for analyzing the information transmitted by the transmission means and determining voice generation parameters based on the personality and voice characteristics specified by the user, voice generation means for generating natural voice data using a voice synthesis engine based on the voice generation parameters, and transmission means for transmitting the generated voice data to the user's terminal, thereby enabling the user to realize interactive communication based on personalized voice data.

[0086] "Input means" refers to the interface that the user utilizes to configure the vehicle's personality and voice characteristics.

[0087] The "transmission means" is a device, software, or protocol that has the function of converting the information set by the input means into a data packet and transmitting it to the server.

[0088] The "analysis means" is a device or software that has the function of analyzing the data packets received by the server and determining voice generation parameters based on the personality and voice characteristics specified by the user.

[0089] The "voice generation means" is a device or software that has the function of generating natural voice data using a voice synthesis engine based on the voice generation parameters determined by the analysis means.

[0090] The "conversion means" is a device or software that has the function of analyzing the voice data transmitted by the transmission means and converting it into a format that can be used by the in-vehicle system.

[0091] The "uploading means" is a device or software that has the function of uploading the voice data generated by the conversion means to the in-vehicle system.

[0092] The "interaction means" is a device or software that has the function of executing an interaction function between a user and a vehicle using the voice data uploaded to the in-vehicle system by the upload means.

[0093] In the system of the present invention, the user sets the vehicle's personality and voice characteristics, and the server generates voice data based on that information, realizing interactive communication between the user and the vehicle. To implement this system, the following program processing is required.

[0094] User settings input method

[0095] Users install the dedicated application "VehicleVoiceCustomizer" on their smartphone or computer. The application provides a user interface where users can set the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). Once the settings are complete, users press the "Settings" button to proceed to the next step.

[0096] Sending configuration information

[0097] The device converts the user-entered configuration information into data packets containing information about the selected personality and voice characteristics, which are then sent to the server using Internet communication protocols (e.g., HTTP or HTTPS).

[0098] Data analysis and speech generation methods

[0099] The server receives the data packets sent from the device. It analyzes the received information and determines voice generation parameters based on the personality and voice characteristics specified by the user. This analysis is performed using a voice synthesis engine, such as the Google Cloud Text-to-Speech API. The voice synthesis engine uses neural network-based technology (e.g., WaveNet) to generate natural-sounding voice data. The generated voice data is stored on the server and prepared for the next step.

[0100] Voice data transmission method

[0101] The server converts the generated voice data into data packets and transmits the packets to the terminal, again using the Internet communication protocol.

[0102] A means of converting voice data and uploading it to an in-vehicle system

[0103] The device receives the audio data sent from the server. It uses tools such as FFmpeg to analyze the received audio data and convert it into a format that can be used by the in-vehicle system (e.g., WAV or MP3). The converted audio data is then uploaded to the in-vehicle system using a connection method such as Bluetooth, Wi-Fi, or USB.

[0104] Interaction methods

[0105] When a user gets into the vehicle and starts the engine, the in-vehicle system will begin playing the uploaded voice data. For example, the vehicle may ask in a friendly voice, "Good morning! Where are you going today?" This dialogue function allows users to enjoy interactive communication with the vehicle.

[0106] Specific examples

[0107] Consider the following example: A user opens the app "VehicleVoiceCustomizer" and sets a "friendly personality," a "male voice," and a "young voice in his 30s." The app then converts this configuration information into a data packet and sends it to the server. The server analyzes the information and determines appropriate voice generation parameters. It then generates voice data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech API, WaveNet). This generated voice data is then sent from the server to the device. The device receives the voice data, uses "FFmpeg" to convert the audio file into a format usable by the in-vehicle system, and uploads it to the in-vehicle system via Bluetooth. When the user starts the engine, the vehicle speaks in a friendly voice, asking, "Good morning! Where are you going today?"

[0108] An example of a prompt for a generative AI model is:

[0109] "Generate a friendly vehicle speaking to the user in the youthful voice of a man in his 30s."

[0110] This allows the system to provide interactive communication according to the user's wishes, resulting in richer interaction with the vehicle.

[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0112] Step 1: Enter your user settings

[0113] The user launches the dedicated app "Vehicle Voice Customizer" on their smartphone or computer. Through the app's graphical interface, they can set the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). The information set by the user is temporarily stored within the app. The input here is the personality and voice characteristics selected by the user, and the output is a data packet containing the setting information.

[0114] Step 2: Sending configuration information

[0115] The device converts the user-entered setting information into a data packet. Specifically, the selected personality and voice characteristics are packetized as text data. The converted data packet is sent to the server using the HTTP / HTTPS protocol. The input here is the user-selected setting information, and the output is the data packet sent to the server.

[0116] Step 3: Data analysis

[0117] The server analyzes data packets received from the device. During the analysis, it extracts speech generation parameters based on the personality and voice characteristics specified by the user. For example, it uses the Google Cloud Text-to-Speech API to select a speech template. The input of this process is the received data packets, and the output is the speech generation parameters.

[0118] Step 4: Speech generation

[0119] The speech synthesis engine in the server generates speech data using the speech generation parameters extracted earlier. Here, neural network-based technology (e.g., WaveNet) is used to generate natural and realistic speech data. The input of this step is the speech generation parameters, and the output is the generated speech data.

[0120] Step 5: Sending audio data

[0121] The server reconverts the generated voice data into data packets and sends them to the terminal again using the HTTP / HTTPS protocol. The input of this process is the generated voice data, and the output is the data packets sent from the server.

[0122] Step 6: Receiving and converting audio data

[0123] The device receives the audio data sent from the server. Because the received audio data is compressed and encrypted, it is first analyzed and decrypted. Then, a tool such as FFmpeg is used to convert it into a format that can be used by the in-vehicle system (e.g., WAV or MP3). The input of this step is the received data packet, and the output is audio data in a format suitable for the in-vehicle system.

[0124] Step 7: Upload your audio data

[0125] The terminal uploads the converted voice data to the in-vehicle system via Bluetooth, Wi-Fi, USB, etc. The input of this step is the converted voice data, and the output is the voice data uploaded to the in-vehicle system.

[0126] Step 8: Implementing Interactivity

[0127] When the user starts the vehicle, the in-vehicle system starts playing the uploaded voice data. For example, the vehicle may start a dialogue with the user by saying, "Good morning! Where are you going today?" The input of this step is the voice data uploaded to the in-vehicle system, and the output is the interactive dialogue between the user and the vehicle.

[0128] (Application example 1)

[0129] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0130] Conventional autonomous vehicles have a problem in that the driving experience is mechanical due to the limited information provided by the vehicle and one-way communication with the user. Furthermore, there is a lack of a way to realize a dialogue system with the personality and voice characteristics desired by the user. Furthermore, there is a lack of a means to provide navigation information, traffic information, weather information, maintenance information, etc. in a unified and interactive manner.

[0131] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0132] In this invention, the server includes an input means for a user to set the personality and voice characteristics of a vehicle, an analysis means for analyzing the information set by the input means, a voice generation means for generating voice data based on the information analyzed by the analysis means, a transmission means for transmitting the generated voice data to the user's terminal, a navigation means for providing navigation information for the vehicle, a traffic information providing means for providing traffic information based on information from the navigation means, and an information providing means for providing weather and maintenance information. This allows the user to interact with an autonomous vehicle with a personality and voice characteristics tailored to their preferences, enriching the driving experience. Furthermore, the provision of comprehensive, interactive information improves the user's comfort and safety.

[0133] Understood. Now, I will create definitions for the important words included in the rewritten claims.

[0134] "Mobile object" refers to a vehicle or other conveyance, including those with automatic driving capabilities.

[0135] "Personality" refers to the personified characteristics of the mobile object set by the user, and includes characteristics such as friendly, cool, sporty, etc.

[0136] "Voice characteristics" refers to the characteristics of a voice set by a user, and includes attributes such as gender, age, tone of voice, and pitch.

[0137] "Input means" refers to an interface that allows a user to set the vehicle's personality and voice characteristics, such as a smartphone app or an in-car console.

[0138] The "analysis means" refers to a computer system for analyzing the information set by the input means.

[0139] "Speech generation means" refers to a system for generating speech data based on the information analyzed by the analysis means.

[0140] "Transmission means" refers to a function for transmitting the generated voice data to the user's terminal.

[0141] "Navigation means" refers to a system for guiding a route to a destination of a mobile object.

[0142] "Traffic information providing means" refers to a function that provides the user with traffic conditions based on information from the navigation means.

[0143] "Information provision means" refers to a function for providing weather and maintenance information to users.

[0144] "Conversion means" refers to a computer program for converting transmitted voice data into a format usable by the mobile system.

[0145] "Uploading means" refers to a function for uploading converted audio data to a mobile system.

[0146] "Interaction means" refers to a system that allows a user to interact with a mobile object using uploaded voice data.

[0147] "Entertainment means" refers to the capability to provide audio and visual entertainment over a mobile system.

[0148] In the system of the present invention, the user sets the personality and voice characteristics of the mobile object, and AI generates the "voice" of the mobile object based on that information, realizing interactive communication between the user and the mobile object. To implement this system, the following program processing is performed.

[0149] User settings processing

[0150] First, the user sets the mobile device's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the mobile device's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). The device collects the entered setting information and converts it into a data packet. Then, the data packet is sent to the server.

[0151] Data analysis and speech generation

[0152] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. This process involves selecting an appropriate template from existing voice samples. Next, the server's speech synthesis engine generates natural-sounding voice data based on the determined voice generation parameters. This voice data is generated using, for example, neural network-based technology (e.g., Google Cloud Text-to-Speech API or Amazon Polly). The generated voice data is stored on the server, and a data packet is generated to be sent to the device. This packet is then sent back to the device.

[0153] Providing voice data and implementing dialogue functions

[0154] The receiving device receives the voice data sent from the server, analyzes it, and converts it into a format that can be used by the mobile system. For example, it may be converted into a common audio file format (WAV, MP3, etc.). The device then uploads the converted voice data to the mobile system, where it can be used in the vehicle's navigation and audio systems. Finally, during actual autonomous driving, the user can enjoy interacting with the vehicle using the voice generated by the system. For example, the vehicle may ask, "Hello! Where are you going today?" Furthermore, the navigation and traffic information means may suggest appropriate routes and report traffic conditions, while the entertainment means may play music or provide simple conversations.

[0155] Usage example

[0156] As an example of usage, consider the following scenario: A user opens the app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to the server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device.

[0157] The terminal converts the voice data and uploads it to the vehicle system. During autonomous driving, when the user starts the engine, the vehicle will speak to them in a friendly voice, asking, "Hello! Where are you going today?"

[0158] For example, the Google Maps API is used for navigation, providing route guidance to destinations, and real-time traffic information is provided as a means of providing traffic information. For entertainment, music streaming services such as Spotify are integrated, allowing users to play music according to their preferences.

[0159] Prompt Sentence Examples

[0160] assistant = AIDriverAssistant("Friendly", "Young male voice")

[0161] print(assistant.generate_response("greeting"))

[0162] In this way, the present invention allows interactive communication with the vehicle, enriching the user's driving experience.

[0163] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0164] Step 1:

[0165] The user uses a dedicated app on a smartphone or computer to input data to set the mobile device's personality and voice characteristics. Specifically, the user selects personality traits such as "friendly," "cool," or "sporty" and voice characteristics such as "male voice," "female voice," or "youthful voice" on the app. This is input data, and the collected information is sent to the device as a data packet.

[0166] Step 2:

[0167] The terminal sends a data packet containing user-defined information to the server. This packet contains information about personality and voice characteristics. This information is sent via a data transfer protocol and received by the server.

[0168] Step 3:

[0169] The server receives the data packet sent from the terminal and analyzes its contents. The analysis means analyzes the contents of the data packet (personality, voice characteristics, etc.) and determines appropriate voice generation parameters. For example, based on information such as "friendly" or "youthful male voice," it generates the parameters required for the voice generation engine.

[0170] Step 4:

[0171] The server generates natural-sounding voice data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech API or Amazon Polly) based on the voice generation parameters determined by the analysis means. This generated voice data is converted into an audio file format (WAV, MP3, etc.) through a series of data calculations.

[0172] Step 5:

[0173] The server converts the generated voice data into data packets and transmits the data packets to the terminal, where the voice data may be compressed. The terminal receives the data packets transmitted from the server.

[0174] Step 6:

[0175] The terminal analyzes the voice data sent from the server and converts it into a format that can be used by the mobile system (for example, WAV or MP3 format). The voice data conversion is performed using a format conversion algorithm.

[0176] Step 7:

[0177] The device then uploads the converted voice data to the in-vehicle system, where it can be used in the navigation and audio systems. Specifically, the device stores the voice data in the in-vehicle system's memory or storage.

[0178] Step 8:

[0179] The user enjoys interacting with the vehicle during autonomous driving. The vehicle uses the generated voice to provide interactive navigation, traffic information, weather and maintenance information, and entertainment functions (e.g., music playback and chat). The interactive navigation uses a voice recognition system to understand user input and respond appropriately.

[0180] Example: When the user starts the engine, the vehicle will speak to them in a friendly voice, asking, "Hello! Where are you going today?" It is possible to use the Google Maps API as a navigation tool and provide real-time traffic information.

[0181] In this way, the present invention realizes interactive communication with the user and enriches the experience of using a mobile device.

[0182] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0183] MODE FOR CARRYING OUT THE INVENTION

[0184] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized dialogue.

[0185] User settings processing

[0186] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[0187] The device collects the setting information entered by the user and converts it into a data packet, which is then sent to the server.

[0188] Data analysis and speech generation

[0189] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0190] Next, the server's speech synthesis engine generates realistic, natural-sounding voice data based on the determined voice generation parameters. This process uses speech synthesis technologies such as neural networks.

[0191] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[0192] Providing voice data and implementing dialogue functions

[0193] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[0194] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[0195] While driving, the user can interact with the vehicle using voices generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?"

[0196] Emotion engine integration

[0197] Furthermore, an emotion engine is integrated into the system, which analyzes the user's voice and facial expressions to recognize their emotions, for example, determining whether they are feeling stressed or having fun.

[0198] The server adjusts the voice generation parameters and dialogue content based on the emotional information recognized by the emotion engine, enabling dialogue appropriate to the user's emotional state. For example, if the vehicle recognizes that the user is tired, it might suggest, "Why don't you take a short break today?"

[0199] Specific examples

[0200] For example, a user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0201] The emotion engine analyzes the user's emotions and recognizes when the user is feeling stressed while driving. Based on this information, the system can say, "I'll play some relaxing music," and play music accordingly, making the user's driving experience richer and more personalized.

[0202] The processing flow will be explained below.

[0203] Step 1:

[0204] The user launches a dedicated app on their smartphone or computer, and then enters the vehicle's personality and voice characteristics on the app's settings screen. Specifically, the user selects options such as "friendly personality," "male voice," or "young voice in his 30s."

[0205] Step 2:

[0206] The device collects the configuration information entered by the user and converts it into a data packet, which contains details such as the vehicle's personality, the gender of the voice, and the tone of the voice.

[0207] Step 3:

[0208] The device generates data packets and sends them to the server over Wi-Fi or mobile data networks.

[0209] Step 4:

[0210] The server receives the setting information sent from the terminal and stores the received data in a database within the server.

[0211] Step 5:

[0212] The server analyzes the received data. Specifically, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0213] Step 6:

[0214] The server's internal voice synthesis engine generates realistic, natural voice data based on the determined voice generation parameters. This process uses voice synthesis technologies such as neural networks.

[0215] Step 7:

[0216] The server stores the generated voice data and generates new data packets to send to the user's terminal.

[0217] Step 8:

[0218] The server then sends the generated data packets to the user's device, again over Wi-Fi or a mobile data network.

[0219] Step 9:

[0220] The terminal acquires the voice data transmitted from the server.

[0221] Step 10:

[0222] The device analyzes the received voice data and converts it into a format that can be used by the in-vehicle system, specifically into common audio file formats such as WAV and MP3.

[0223] Step 11:

[0224] The device then uploads the converted audio data to the vehicle's system via Bluetooth, USB connection, or other means.

[0225] Step 12:

[0226] The in-car system plays back the audio data while the user is driving. For example, when the engine is started, the vehicle may say, "Good morning. Where are you going today?"

[0227] Emotion Engine Processing Flow

[0228] Step 13:

[0229] The emotion engine installed in the device captures the user's voice and facial expressions using the device's microphone and camera.

[0230] Step 14:

[0231] The device's emotion engine analyzes the voice and facial expression data it acquires to recognize the user's emotions. For example, it can determine whether the user is feeling "stressed" based on the tone of their voice and facial expression.

[0232] Step 15:

[0233] The device sends the recognized emotion information to the server, which includes information that the user is currently feeling "stressed."

[0234] Step 16:

[0235] The server receives and analyzes the emotional information. Based on this information, it adjusts the generated voice data and dialogue content. For example, it generates new voice data such as "Shall I play some calming music?" to help the user relax.

[0236] Step 17:

[0237] The server sends the newly generated voice data to the terminal.

[0238] Step 18:

[0239] The device then uploads the new voice data it receives back to the in-vehicle system.

[0240] Step 19:

[0241] New audio data will be played while the user is driving: the vehicle will say, "I'm going to play some music to help you relax," and then play appropriate music.

[0242] Through the above steps, the user can enjoy an interactive driving experience that is appropriate to their emotions.

[0243] Example 2

[0244] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0245] In recent years, interactive communication systems have become increasingly important for improving in-vehicle user experiences. However, existing systems are unable to provide personalized dialogue based on the user's individual preferences and emotional state. Furthermore, it is difficult for users to easily configure the vehicle's personality and voice characteristics and quickly provide personalized voice dialogue based on those settings. This has led to problems that reduce user satisfaction and convenience.

[0246] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0247] In this invention, the server includes an input means for a user to set the vehicle's personality and voice characteristics, a transmission means for converting the information set by the input means into a data packet and transmitting it to the server, an analysis means for analyzing the data packet transmitted by the transmission means and determining voice generation parameters based on the user's settings, a voice generation means for generating voice data using the voice generation parameters determined by the analysis means, and a transmission means for transmitting the generated voice data to the user's terminal. This allows the user to easily set the vehicle's personality and voice characteristics and quickly provide personalized voice dialogue based on them. Furthermore, it is possible to provide appropriate dialogue according to the user's emotional state, thereby improving the user experience.

[0248] A "user" is a person who uses this system to set the vehicle's personality and voice characteristics and engage in interactive communication.

[0249] "Vehicle personality" refers to the interaction attitude and atmosphere of the in-vehicle system set by the user, and includes characteristics such as "friendly," "cool," and "sporty."

[0250] "Voice characteristics" refers to the attributes of the voice emitted by the vehicle, and includes parameters such as "male voice," "female voice," "youthful voice," and "deep voice."

[0251] "Input means" refers to the device or interface through which a user configures the vehicle's personality and voice characteristics.

[0252] "Data packet" refers to a series of data containing information set by the user, and is used to send to the server.

[0253] "Transmission means" refers to a communication means for transmitting data packets from a terminal to a server, such as the Internet or a dedicated network.

[0254] The "analysis means" refers to a process or technology for analyzing the transmitted data packets to extract user setting information and determining voice generation parameters based on the extracted information.

[0255] "Voice generation parameters" refer to specific settings and templates for generating voice data based on the vehicle's personality and voice characteristics set by the user.

[0256] "Speech generation means" refers to a technology or engine for generating realistic and natural voice data using the voice generation parameters determined by the analysis means.

[0257] "Conversion means" refers to the process or equipment that converts the received audio data into a format that can be used by the in-vehicle system (e.g., WAV, MP3, etc.).

[0258] "Uploading means" refers to a means for uploading converted audio data to an in-vehicle system.

[0259] "Dialogue means" refers to the technology and functions for dialogue with the user using voice data generated by the in-vehicle system.

[0260] "Emotion analysis means" refers to technology or engines that analyze the user's voice and facial expressions to recognize emotions.

[0261] The "adjustment means" refers to a technique or process for adjusting the voice generation parameters and the dialogue content based on the emotional information recognized by the emotion analysis means.

[0262] MODE FOR CARRYING OUT THE INVENTION

[0263] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized dialogue.

[0264] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[0265] The device collects the setting information entered by the user and converts it into a data packet, which is then sent to the server.

[0266] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0267] Next, the server's speech synthesis engine generates realistic, natural-sounding voice data based on the determined voice generation parameters. This process uses speech synthesis technologies such as neural networks.

[0268] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[0269] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[0270] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[0271] While driving, the user can interact with the vehicle using voices generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?"

[0272] Furthermore, an emotion engine is integrated into the system, which analyzes the user's voice and facial expressions to recognize their emotions, for example, determining whether they are feeling stressed or having fun.

[0273] The server adjusts the voice generation parameters and dialogue content based on the emotional information recognized by the emotion engine, enabling dialogue appropriate to the user's emotional state. For example, if the vehicle recognizes that the user is tired, it might suggest, "Why don't you take a short break today?"

[0274] Specific examples

[0275] For example, a user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0276] The emotion engine analyzes the user's emotions and recognizes when the user is feeling stressed while driving. Based on this information, the system can say, "I'll play some relaxing music," and play music accordingly, making the user's driving experience richer and more personalized.

[0277] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0278] Step 1:

[0279] The user launches the dedicated app.

[0280] Specific action: The user taps or clicks to launch a dedicated application installed on their smartphone or computer.

[0281] Input: User Action

[0282] Output: App home screen display

[0283] Step 2:

[0284] The user selects the vehicle's personality and voice characteristics.

[0285] How it works: Select a personality type such as "friendly," "cool," or "sporty" and voice characteristics such as "male voice," "female voice," or "youthful voice" from the app's graphical interface.

[0286] Input: User-selected personality and voice characteristics

[0287] Output: Selected personality and vocal trait data

[0288] Step 3:

[0289] The terminal converts the setting information into a data packet and transmits it to the server.

[0290] What it does: The application packages the user's selected settings into a data packet and sends it over the Internet to a server.

[0291] Input: User-selected setting information

[0292] Output: Data packets sent to the server

[0293] Step 4:

[0294] The server receives the data packets and performs the analysis.

[0295] Specific operation: The server analyzes the data packets received from the terminal and extracts the user's setting information.

[0296] Input: Data packets from the terminal

[0297] Output: Parsed configuration information

[0298] Step 5:

[0299] The server determines the voice generation parameters.

[0300] Specific operation: Based on the analyzed configuration information, the server selects an appropriate template from the voice sample library and determines voice generation parameters.

[0301] Input: Parsed configuration information

[0302] Output: Speech generation parameters

[0303] Step 6:

[0304] The server generates voice data based on the voice generation parameters.

[0305] How it works: The speech synthesis engine uses speech generation parameters to generate realistic, natural-sounding speech data, using technologies such as neural networks.

[0306] Input: Speech generation parameters

[0307] Output: Generated audio data

[0308] Step 7:

[0309] The server transmits the generated voice data to the terminal.

[0310] Specific operation: The generated voice data is converted into data packets and sent to the terminal.

[0311] Input: Generated audio data

[0312] Output: Data packets to the terminal

[0313] Step 8:

[0314] The device receives the voice data, analyzes it, and converts it into a format that can be used by the in-vehicle system.

[0315] Specific operation: The device analyzes the audio data it receives and converts it into a format that can be used by the in-car system, such as WAV or MP3.

[0316] Input: Audio data sent to the device

[0317] Output: Converted audio data

[0318] Step 9:

[0319] The terminal uploads the audio data to the in-vehicle system.

[0320] What it does: Uploads converted audio data to the vehicle's navigation and audio systems.

[0321] Input: Converted audio data

[0322] Output: Audio data uploaded to the in-car system

[0323] Step 10:

[0324] The user interacts using the generated voice.

[0325] Specific behavior: When the engine is started, the vehicle will speak to you in a preset personality and voice, saying, "Good morning. What music would you like to listen to today?"

[0326] Input: Audio data uploaded to the in-vehicle system

[0327] Output: Interaction with the vehicle

[0328] Step 11:

[0329] The emotion engine analyzes the user's emotions.

[0330] Specific operation: The emotion engine analyzes voice and facial expressions to recognize emotions such as stress and joy.

[0331] Input: User's voice and facial expression data

[0332] Output: Recognized emotion information

[0333] Step 12:

[0334] The server adjusts the voice generation parameters and dialogue content based on the emotion information.

[0335] Specific operation: Based on the emotional information recognized by the emotion engine, the server adjusts the voice generation parameters and dialogue content, for example, to say, "Would you like to take a short break today?"

[0336] Input: Recognized emotion information

[0337] Output: Adjusted voice generation parameters and dialogue content

[0338] (Application example 2)

[0339] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0340] Autonomous vehicles are expected to provide a personalized travel experience through dialogue between passengers and the vehicle. However, current systems are unable to recognize passenger emotions in real time and flexibly adapt the dialogue content with the vehicle based on those emotions. This makes it difficult to achieve interactive communication that truly satisfies passengers.

[0341] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes input means for the user to set the vehicle's personality and voice characteristics, analysis means for analyzing the information set by the input means, voice generation means for generating voice data based on the information analyzed by the analysis means, transmission means for transmitting the generated voice data to the user's terminal, emotion recognition means for analyzing the user's emotions, and adjustment means for adjusting voice generation parameters in accordance with the emotion information obtained by the emotion recognition means. This enables personalized dialogue in real time according to the user's emotions.

[0342] "User" means any person or entity that uses the vehicle or system.

[0343] "Vehicle personality" is a concept that refers to the emotional and behavioral characteristics that a vehicle is supposed to have, which are set by the user.

[0344] "Voice characteristics" are attributes that refer to voice characteristics such as the quality, tone, and timbre of the voice emitted from the vehicle.

[0345] The term "input means" refers to a device or interface that allows a user to input setting information.

[0346] "Analysis means" refers to a system that includes software and hardware for analyzing input information and understanding its meaning and content.

[0347] "Speech generation means" refers to a function that includes a speech synthesis engine and related technologies for generating speech data based on analyzed information.

[0348] "Transmission means" refers to the infrastructure, which refers to the communication technologies and protocols used to transmit the generated audio data to other devices or systems.

[0349] "Emotion recognition means" refers to a system that uses technology or algorithms to analyze and identify emotions from a user's facial expressions, voice, or other input.

[0350] The "adjustment means" is a function that refers to a method or technique for dynamically changing voice generation parameters based on information obtained by the emotion recognition means.

[0351] "Conversion means" is a function that refers to technology or devices for converting transmitted audio data into a format that can be used by the in-vehicle system.

[0352] "Uploading means" refers to the infrastructure that refers to the methods and protocols used to transfer and store the converted audio data in the in-vehicle system.

[0353] "Dialogue means" refers to the technology and functions that enable the in-vehicle system to communicate with the user via voice.

[0354] The "display means" is an apparatus that refers to a device or technology for displaying the dialogue content based on the voice generation parameters adjusted by the adjustment means.

[0355] MODE FOR CARRYING OUT THE INVENTION

[0356] The following describes an embodiment of the present invention: This system has a structure that allows real-time dialogue within a vehicle based on the vehicle's personality and voice characteristics designed by the user.

[0357] System Overview

[0358] The system consists of the following main components:

[0359] 1. Input Method

[0360] Users can configure their vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app's interface is designed to make configuration easy for users. The configured information is collected by the app and converted into data packets.

[0361] 2. Analysis method

[0362] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. During this process, it uses a voice sample library to select a voice template that matches the settings.

[0363] 3. Voice Generation Method

[0364] A speech synthesis engine (such as Google Cloud TTS or Amazon Polly) in the server generates realistic and natural voice data based on the analyzed data. The generated voice data is stored in the server and a data packet is generated for processing.

[0365] 4. Transmission Method

[0366] The audio data is sent from the server to the device using a common data transfer protocol (e.g., HTTP or HTTPS).

[0367] 5. Emotion recognition means

[0368] An emotion engine (for example, Microsoft Azure's Emotion API) analyzes the user's voice and facial expressions to recognize their emotions. This information is collected in real time and sent to a server.

[0369] 6. Adjustment means

[0370] The server uses the information obtained from the emotion recognition means to adjust the voice and dialogue played by the in-vehicle system, using predefined prompts.

[0371] Specific examples

[0372] The user opens the app and selects settings such as "friendly personality," "male voice," and "young voice in his 30s." The app then sends this setting information to the server. The server analyzes the information and determines the appropriate voice generation parameters. The voice synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0373] Prompt Sentence Examples

[0374] When passengers board the vehicle:

[0375] "Hello, how's it going today?"

[0376] If the user says they are tired:

[0377] "Thank you for your hard work. I'll put on some relaxing music."

[0378] In this way, the system can combine emotion recognition and speech generation techniques to provide personalized interactions according to the user's emotions.

[0379] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0380] Step 1:

[0381] The user opens a dedicated app on their smartphone or computer and sets the vehicle's personality and voice characteristics. The input information can include, for example, "friendly personality," "male voice," or "youthful voice in his 30s." The app collects this information as a data packet and sends it to a server.

[0382] Input: User setting information

[0383] Data processing: Converting user setting information into data packets

[0384] Output: Data packets

[0385] Step 2:

[0386] The server receives the data packet sent from the device and starts the analysis process. Based on the configuration information, it references the voice sample library to determine voice generation parameters, and selects a template such as a "friendly male voice."

[0387] Input: Data packet

[0388] Data calculation: Determining voice generation parameters

[0389] Output: Speech generation parameters

[0390] Step 3:

[0391] The server's speech synthesis engine generates voice data based on the determined voice generation parameters. During this process, realistic and natural voices are generated using, for example, Google Cloud TTS or Amazon Polly. The generated voice data is stored on the server.

[0392] Input: Speech generation parameters

[0393] Data Computation: Generation of voice data using a TTS engine

[0394] Output: Audio data

[0395] Step 4:

[0396] The server converts the generated voice data into data packets and sends them to the terminal, using common data transfer protocols (HTTP or HTTPS).

[0397] Input: Audio data

[0398] Data processing: Converting voice data into data packets

[0399] Output: Data packets

[0400] Step 5:

[0401] The device receives the data packets sent from the server, analyzes the audio data, and converts it into a common audio file format (e.g., WAV or MP3).

[0402] Input: Data packet

[0403] Data processing: Convert data packets into audio file format

[0404] Output: Audio file

[0405] Step 6:

[0406] The device then uploads the converted audio file to the in-car system via communication methods such as Bluetooth or Wi-Fi.

[0407] Input: Audio file

[0408] Data Calculation: Uploading Audio Files

[0409] Output: Uploaded audio file

[0410] Step 7:

[0411] The user plays generated audio through the in-car system, initiating a dialogue with the vehicle, for example, the vehicle saying, "Hello, how's your day?"

[0412] Input: Uploaded audio file

[0413] Data calculation: Audio playback by in-car systems

[0414] Output: Voice dialogue

[0415] Step 8:

[0416] The emotion engine analyzes the user's voice and facial expressions to recognize emotions, for example, detecting the user's stress level using Microsoft Azure's Emotion API.

[0417] Input: User's voice and facial expression data

[0418] Data Computing: Emotion Analysis and Recognition

[0419] Output: Emotional information

[0420] Step 9:

[0421] The server adjusts the voice generation parameters using the adjustment means based on the emotion information obtained by the emotion recognition means. For example, if the server recognizes that the user is tired, it suggests playing relaxing music.

[0422] Input: Emotion information

[0423] Data calculation: Adjustment of voice generation parameters

[0424] Output: Adjusted speech generation parameters

[0425] Step 10:

[0426] The in-vehicle system then adjusts the dialogue content based on the adjusted voice generation parameters to provide a dialogue tailored to the user. For example, the vehicle might say, "Thank you for your hard work. I'll play some relaxing music."

[0427] Input: Adjusted speech generation parameters

[0428] Data calculation: Change of dialogue content

[0429] Output: Personalized dialogue

[0430] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0431] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0432] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0433] [Second embodiment]

[0434] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0435] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0436] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0437] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0438] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0439] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0440] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0441] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0442] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0443] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0444] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0445] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0446] MODE FOR CARRYING OUT THE INVENTION

[0447] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. To implement this system, the following program processing is performed.

[0448] User settings processing

[0449] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[0450] The terminal collects the input configuration information and converts it into a data packet, which is then sent to the server.

[0451] Data analysis and speech generation

[0452] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. This includes selecting an appropriate template from existing voice samples.

[0453] A speech synthesis engine in the server then generates realistic, natural-sounding speech data based on the determined speech generation parameters, for example, using neural network-based techniques.

[0454] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[0455] Providing voice data and implementing dialogue functions

[0456] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[0457] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[0458] Finally, while actually driving, users can enjoy interacting with the vehicle using the voice generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?" In this way, conversational communication is realized.

[0459] Specific examples

[0460] As an example, consider the following scenario: A user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0461] When the user starts the engine while driving, the vehicle will speak to the user in a friendly voice, asking, "Good morning! Where are you going today?" In this way, the present invention enables interactive communication with the vehicle, enriching the user's driving experience.

[0462] The processing flow will be explained below.

[0463] Step 1:

[0464] The user launches a dedicated app on their smartphone or computer, and then enters the vehicle's personality and voice characteristics on the app's settings screen. Specifically, the user selects options such as "friendly personality," "male voice," or "youthful voice."

[0465] Step 2:

[0466] The device collects the configuration information entered by the user and converts it into a data packet, which contains details such as the vehicle's personality, the gender of the voice, and the tone of the voice.

[0467] Step 3:

[0468] The device generates data packets and sends them to the server over Wi-Fi or mobile data networks.

[0469] Step 4:

[0470] The server receives the setting information sent from the terminal and stores the received data in a database within the server.

[0471] Step 5:

[0472] The server analyzes the received data. Specifically, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0473] Step 6:

[0474] The server's internal voice synthesis engine generates realistic, natural voice data based on the determined voice generation parameters. This process uses voice synthesis technologies such as neural networks.

[0475] Step 7:

[0476] The server stores the generated voice data and generates new data packets to send to the user's terminal.

[0477] Step 8:

[0478] The server then sends the generated data packets to the user's device, again over Wi-Fi or a mobile data network.

[0479] Step 9:

[0480] The terminal acquires the voice data transmitted from the server.

[0481] Step 10:

[0482] The device analyzes the received voice data and converts it into a format that can be used by the in-vehicle system, specifically into common audio file formats such as WAV and MP3.

[0483] Step 11:

[0484] The device then uploads the converted audio data to the vehicle's system via Bluetooth, USB connection, or other means.

[0485] Step 12:

[0486] The in-car system plays back the audio data while the user is driving. For example, when the engine is started, the vehicle may say, "Good morning. Where are you going today?"

[0487] Through the above steps, users can freely customize the vehicle's personality and voice characteristics, enjoying a personalized interactive driving experience.

[0488] Example 1

[0489] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0490] Conventional in-vehicle systems have made it difficult for users to enjoy a personalized experience through interaction with the vehicle. Furthermore, they lacked a means to freely set the vehicle's personality and voice characteristics and generate natural-sounding voice data based on those settings. As a result, users could only receive uniform, standardized voice guidance, limiting their driving experience. To solve this issue, a system was needed that could flexibly generate voice data based on parameters set by the user and provide a personalized experience through interaction with the vehicle.

[0491] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0492] In this invention, the server includes input means for a user to set the vehicle's personality and voice characteristics, transmission means for converting the information set by the input means into data packets and transmitting the data to the server, analysis means for analyzing the information transmitted by the transmission means and determining voice generation parameters based on the personality and voice characteristics specified by the user, voice generation means for generating natural voice data using a voice synthesis engine based on the voice generation parameters, and transmission means for transmitting the generated voice data to the user's terminal, thereby enabling the user to realize interactive communication based on personalized voice data.

[0493] "Input means" refers to the interface that the user utilizes to configure the vehicle's personality and voice characteristics.

[0494] The "transmission means" is a device, software, or protocol that has the function of converting the information set by the input means into a data packet and transmitting it to the server.

[0495] The "analysis means" is a device or software that has the function of analyzing the data packets received by the server and determining voice generation parameters based on the personality and voice characteristics specified by the user.

[0496] The "voice generation means" is a device or software that has the function of generating natural voice data using a voice synthesis engine based on the voice generation parameters determined by the analysis means.

[0497] The "conversion means" is a device or software that has the function of analyzing the voice data transmitted by the transmission means and converting it into a format that can be used by the in-vehicle system.

[0498] The "uploading means" is a device or software that has the function of uploading the voice data generated by the conversion means to the in-vehicle system.

[0499] The "interaction means" is a device or software that has the function of executing an interaction function between a user and a vehicle using the voice data uploaded to the in-vehicle system by the upload means.

[0500] In the system of the present invention, the user sets the vehicle's personality and voice characteristics, and the server generates voice data based on that information, realizing interactive communication between the user and the vehicle. To implement this system, the following program processing is required.

[0501] User settings input method

[0502] Users install the dedicated application "VehicleVoiceCustomizer" on their smartphone or computer. The application provides a user interface where users can set the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). Once the settings are complete, users press the "Settings" button to proceed to the next step.

[0503] Sending configuration information

[0504] The device converts the user-entered configuration information into data packets containing information about the selected personality and voice characteristics, which are then sent to the server using Internet communication protocols (e.g., HTTP or HTTPS).

[0505] Data analysis and speech generation methods

[0506] The server receives the data packets sent from the device. It analyzes the received information and determines voice generation parameters based on the personality and voice characteristics specified by the user. This analysis is performed using a voice synthesis engine, such as the Google Cloud Text-to-Speech API. The voice synthesis engine uses neural network-based technology (e.g., WaveNet) to generate natural-sounding voice data. The generated voice data is stored on the server and prepared for the next step.

[0507] Voice data transmission method

[0508] The server converts the generated voice data into data packets and transmits the packets to the terminal, again using the Internet communication protocol.

[0509] A means of converting voice data and uploading it to an in-vehicle system

[0510] The device receives the audio data sent from the server. It uses tools such as FFmpeg to analyze the received audio data and convert it into a format that can be used by the in-vehicle system (e.g., WAV or MP3). The converted audio data is then uploaded to the in-vehicle system using a connection method such as Bluetooth, Wi-Fi, or USB.

[0511] Interaction methods

[0512] When a user gets into the vehicle and starts the engine, the in-vehicle system will begin playing the uploaded voice data. For example, the vehicle may ask in a friendly voice, "Good morning! Where are you going today?" This dialogue function allows users to enjoy interactive communication with the vehicle.

[0513] Specific examples

[0514] Consider the following example: A user opens the app "VehicleVoiceCustomizer" and sets a "friendly personality," a "male voice," and a "young voice in his 30s." The app then converts this configuration information into a data packet and sends it to the server. The server analyzes the information and determines appropriate voice generation parameters. It then generates voice data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech API, WaveNet). This generated voice data is then sent from the server to the device. The device receives the voice data, uses "FFmpeg" to convert the audio file into a format usable by the in-vehicle system, and uploads it to the in-vehicle system via Bluetooth. When the user starts the engine, the vehicle speaks in a friendly voice, asking, "Good morning! Where are you going today?"

[0515] An example of a prompt for a generative AI model is:

[0516] "Generate a friendly vehicle speaking to the user in the youthful voice of a man in his 30s."

[0517] This allows the system to provide interactive communication according to the user's wishes, resulting in richer interaction with the vehicle.

[0518] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0519] Step 1: Enter your user settings

[0520] The user launches the dedicated app "Vehicle Voice Customizer" on their smartphone or computer. Through the app's graphical interface, they can set the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). The information set by the user is temporarily stored within the app. The input here is the personality and voice characteristics selected by the user, and the output is a data packet containing the setting information.

[0521] Step 2: Sending configuration information

[0522] The device converts the user-entered setting information into a data packet. Specifically, the selected personality and voice characteristics are packetized as text data. The converted data packet is sent to the server using the HTTP / HTTPS protocol. The input here is the user-selected setting information, and the output is the data packet sent to the server.

[0523] Step 3: Data analysis

[0524] The server analyzes data packets received from the device. During the analysis, it extracts speech generation parameters based on the personality and voice characteristics specified by the user. For example, it uses the Google Cloud Text-to-Speech API to select a speech template. The input of this process is the received data packets, and the output is the speech generation parameters.

[0525] Step 4: Speech generation

[0526] The speech synthesis engine in the server generates speech data using the speech generation parameters extracted earlier. Here, neural network-based technology (e.g., WaveNet) is used to generate natural and realistic speech data. The input of this step is the speech generation parameters, and the output is the generated speech data.

[0527] Step 5: Sending audio data

[0528] The server reconverts the generated voice data into data packets and sends them to the terminal again using the HTTP / HTTPS protocol. The input of this process is the generated voice data, and the output is the data packets sent from the server.

[0529] Step 6: Receiving and converting audio data

[0530] The device receives the audio data sent from the server. Because the received audio data is compressed and encrypted, it is first analyzed and decrypted. Then, a tool such as FFmpeg is used to convert it into a format that can be used by the in-vehicle system (e.g., WAV or MP3). The input of this step is the received data packet, and the output is audio data in a format suitable for the in-vehicle system.

[0531] Step 7: Upload your audio data

[0532] The terminal uploads the converted voice data to the in-vehicle system via Bluetooth, Wi-Fi, USB, etc. The input of this step is the converted voice data, and the output is the voice data uploaded to the in-vehicle system.

[0533] Step 8: Implementing Interactivity

[0534] When the user starts the vehicle, the in-vehicle system starts playing the uploaded voice data. For example, the vehicle may start a dialogue with the user by saying, "Good morning! Where are you going today?" The input of this step is the voice data uploaded to the in-vehicle system, and the output is the interactive dialogue between the user and the vehicle.

[0535] (Application example 1)

[0536] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0537] Conventional autonomous vehicles have a problem in that the driving experience is mechanical due to the limited information provided by the vehicle and one-way communication with the user. Furthermore, there is a lack of a way to realize a dialogue system with the personality and voice characteristics desired by the user. Furthermore, there is a lack of a means to provide navigation information, traffic information, weather information, maintenance information, etc. in a unified and interactive manner.

[0538] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0539] In this invention, the server includes an input means for a user to set the personality and voice characteristics of a vehicle, an analysis means for analyzing the information set by the input means, a voice generation means for generating voice data based on the information analyzed by the analysis means, a transmission means for transmitting the generated voice data to the user's terminal, a navigation means for providing navigation information for the vehicle, a traffic information providing means for providing traffic information based on information from the navigation means, and an information providing means for providing weather and maintenance information. This allows the user to interact with an autonomous vehicle with a personality and voice characteristics tailored to their preferences, enriching the driving experience. Furthermore, the provision of comprehensive, interactive information improves the user's comfort and safety.

[0540] Understood. Now, I will create definitions for the important words included in the rewritten claims.

[0541] "Mobile object" refers to a vehicle or other conveyance, including those with automatic driving capabilities.

[0542] "Personality" refers to the personified characteristics of the mobile object set by the user, and includes characteristics such as friendly, cool, sporty, etc.

[0543] "Voice characteristics" refers to the characteristics of a voice set by a user, and includes attributes such as gender, age, tone of voice, and pitch.

[0544] "Input means" refers to an interface that allows a user to set the vehicle's personality and voice characteristics, such as a smartphone app or an in-car console.

[0545] The "analysis means" refers to a computer system for analyzing the information set by the input means.

[0546] "Speech generation means" refers to a system for generating speech data based on the information analyzed by the analysis means.

[0547] "Transmission means" refers to a function for transmitting the generated voice data to the user's terminal.

[0548] "Navigation means" refers to a system for guiding a route to a destination of a mobile object.

[0549] "Traffic information providing means" refers to a function that provides the user with traffic conditions based on information from the navigation means.

[0550] "Information provision means" refers to a function for providing weather and maintenance information to users.

[0551] "Conversion means" refers to a computer program for converting transmitted voice data into a format usable by the mobile system.

[0552] "Uploading means" refers to a function for uploading converted audio data to a mobile system.

[0553] "Interaction means" refers to a system that allows a user to interact with a mobile object using uploaded voice data.

[0554] "Entertainment means" refers to the capability to provide audio and visual entertainment over a mobile system.

[0555] In the system of the present invention, the user sets the personality and voice characteristics of the mobile object, and AI generates the "voice" of the mobile object based on that information, realizing interactive communication between the user and the mobile object. To implement this system, the following program processing is performed.

[0556] User settings processing

[0557] First, the user sets the mobile device's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the mobile device's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). The device collects the entered setting information and converts it into a data packet. Then, the data packet is sent to the server.

[0558] Data analysis and speech generation

[0559] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. This process involves selecting an appropriate template from existing voice samples. Next, the server's speech synthesis engine generates natural-sounding voice data based on the determined voice generation parameters. This voice data is generated using, for example, neural network-based technology (e.g., Google Cloud Text-to-Speech API or Amazon Polly). The generated voice data is stored on the server, and a data packet is generated to be sent to the device. This packet is then sent back to the device.

[0560] Providing voice data and implementing dialogue functions

[0561] The receiving device receives the voice data sent from the server, analyzes it, and converts it into a format that can be used by the mobile system. For example, it may be converted into a common audio file format (WAV, MP3, etc.). The device then uploads the converted voice data to the mobile system, where it can be used in the vehicle's navigation and audio systems. Finally, during actual autonomous driving, the user can enjoy interacting with the vehicle using the voice generated by the system. For example, the vehicle may ask, "Hello! Where are you going today?" Furthermore, the navigation and traffic information means may suggest appropriate routes and report traffic conditions, while the entertainment means may play music or provide simple conversations.

[0562] Usage example

[0563] As an example of usage, consider the following scenario: A user opens the app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to the server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device.

[0564] The terminal converts the voice data and uploads it to the vehicle system. During autonomous driving, when the user starts the engine, the vehicle will speak to them in a friendly voice, asking, "Hello! Where are you going today?"

[0565] For example, the Google Maps API is used for navigation, providing route guidance to destinations, and real-time traffic information is provided as a means of providing traffic information. For entertainment, music streaming services such as Spotify are integrated, allowing users to play music according to their preferences.

[0566] Prompt Sentence Examples

[0567] assistant = AIDriverAssistant("Friendly", "Young male voice")

[0568] print(assistant.generate_response("greeting"))

[0569] In this way, the present invention allows interactive communication with the vehicle, enriching the user's driving experience.

[0570] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0571] Step 1:

[0572] The user uses a dedicated app on a smartphone or computer to input data to set the mobile device's personality and voice characteristics. Specifically, the user selects personality traits such as "friendly," "cool," or "sporty" and voice characteristics such as "male voice," "female voice," or "youthful voice" on the app. This is input data, and the collected information is sent to the device as a data packet.

[0573] Step 2:

[0574] The terminal sends a data packet containing user-defined information to the server. This packet contains information about personality and voice characteristics. This information is sent via a data transfer protocol and received by the server.

[0575] Step 3:

[0576] The server receives the data packet sent from the terminal and analyzes its contents. The analysis means analyzes the contents of the data packet (personality, voice characteristics, etc.) and determines appropriate voice generation parameters. For example, based on information such as "friendly" or "youthful male voice," it generates the parameters required for the voice generation engine.

[0577] Step 4:

[0578] The server generates natural-sounding voice data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech API or Amazon Polly) based on the voice generation parameters determined by the analysis means. This generated voice data is converted into an audio file format (WAV, MP3, etc.) through a series of data calculations.

[0579] Step 5:

[0580] The server converts the generated voice data into data packets and transmits the data packets to the terminal, where the voice data may be compressed. The terminal receives the data packets transmitted from the server.

[0581] Step 6:

[0582] The terminal analyzes the voice data sent from the server and converts it into a format that can be used by the mobile system (for example, WAV or MP3 format). The voice data conversion is performed using a format conversion algorithm.

[0583] Step 7:

[0584] The device then uploads the converted voice data to the in-vehicle system, where it can be used in the navigation and audio systems. Specifically, the device stores the voice data in the in-vehicle system's memory or storage.

[0585] Step 8:

[0586] The user enjoys interacting with the vehicle during autonomous driving. The vehicle uses the generated voice to provide interactive navigation, traffic information, weather and maintenance information, and entertainment functions (e.g., music playback and chat). The interactive navigation uses a voice recognition system to understand user input and respond appropriately.

[0587] Example: When the user starts the engine, the vehicle will speak to them in a friendly voice, asking, "Hello! Where are you going today?" It is possible to use the Google Maps API as a navigation tool and provide real-time traffic information.

[0588] In this way, the present invention realizes interactive communication with the user and enriches the experience of using a mobile device.

[0589] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0590] MODE FOR CARRYING OUT THE INVENTION

[0591] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized dialogue.

[0592] User settings processing

[0593] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[0594] The device collects the setting information entered by the user and converts it into a data packet, which is then sent to the server.

[0595] Data analysis and speech generation

[0596] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0597] Next, the server's speech synthesis engine generates realistic, natural-sounding voice data based on the determined voice generation parameters. This process uses speech synthesis technologies such as neural networks.

[0598] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[0599] Providing voice data and implementing dialogue functions

[0600] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[0601] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[0602] While driving, the user can interact with the vehicle using voices generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?"

[0603] Emotion engine integration

[0604] Furthermore, an emotion engine is integrated into the system, which analyzes the user's voice and facial expressions to recognize their emotions, for example, determining whether they are feeling stressed or having fun.

[0605] The server adjusts the voice generation parameters and dialogue content based on the emotional information recognized by the emotion engine, enabling dialogue appropriate to the user's emotional state. For example, if the vehicle recognizes that the user is tired, it might suggest, "Why don't you take a short break today?"

[0606] Specific examples

[0607] For example, a user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0608] The emotion engine analyzes the user's emotions and recognizes when the user is feeling stressed while driving. Based on this information, the system can say, "I'll play some relaxing music," and play music accordingly, making the user's driving experience richer and more personalized.

[0609] The processing flow will be explained below.

[0610] Step 1:

[0611] The user launches a dedicated app on their smartphone or computer, and then enters the vehicle's personality and voice characteristics on the app's settings screen. Specifically, the user selects options such as "friendly personality," "male voice," or "young voice in his 30s."

[0612] Step 2:

[0613] The device collects the configuration information entered by the user and converts it into a data packet, which contains details such as the vehicle's personality, the gender of the voice, and the tone of the voice.

[0614] Step 3:

[0615] The device generates data packets and sends them to the server over Wi-Fi or mobile data networks.

[0616] Step 4:

[0617] The server receives the setting information sent from the terminal and stores the received data in a database within the server.

[0618] Step 5:

[0619] The server analyzes the received data. Specifically, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0620] Step 6:

[0621] The server's internal voice synthesis engine generates realistic, natural voice data based on the determined voice generation parameters. This process uses voice synthesis technologies such as neural networks.

[0622] Step 7:

[0623] The server stores the generated voice data and generates new data packets to send to the user's terminal.

[0624] Step 8:

[0625] The server then sends the generated data packets to the user's device, again over Wi-Fi or a mobile data network.

[0626] Step 9:

[0627] The terminal acquires the voice data transmitted from the server.

[0628] Step 10:

[0629] The device analyzes the received voice data and converts it into a format that can be used by the in-vehicle system, specifically into common audio file formats such as WAV and MP3.

[0630] Step 11:

[0631] The device then uploads the converted audio data to the vehicle's system via Bluetooth, USB connection, or other means.

[0632] Step 12:

[0633] The in-car system plays back the audio data while the user is driving. For example, when the engine is started, the vehicle may say, "Good morning. Where are you going today?"

[0634] Emotion Engine Processing Flow

[0635] Step 13:

[0636] The emotion engine installed in the device captures the user's voice and facial expressions using the device's microphone and camera.

[0637] Step 14:

[0638] The device's emotion engine analyzes the voice and facial expression data it acquires to recognize the user's emotions. For example, it can determine whether the user is feeling "stressed" based on the tone of their voice and facial expression.

[0639] Step 15:

[0640] The device sends the recognized emotion information to the server, which includes information that the user is currently feeling "stressed."

[0641] Step 16:

[0642] The server receives and analyzes the emotional information. Based on this information, it adjusts the generated voice data and dialogue content. For example, it generates new voice data such as "Shall I play some calming music?" to help the user relax.

[0643] Step 17:

[0644] The server sends the newly generated voice data to the terminal.

[0645] Step 18:

[0646] The device then uploads the new voice data it receives back to the in-vehicle system.

[0647] Step 19:

[0648] New audio data will be played while the user is driving: the vehicle will say, "I'm going to play some music to help you relax," and then play appropriate music.

[0649] Through the above steps, the user can enjoy an interactive driving experience that is appropriate to their emotions.

[0650] Example 2

[0651] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0652] In recent years, interactive communication systems have become increasingly important for improving in-vehicle user experiences. However, existing systems are unable to provide personalized dialogue based on the user's individual preferences and emotional state. Furthermore, it is difficult for users to easily configure the vehicle's personality and voice characteristics and quickly provide personalized voice dialogue based on those settings. This has led to problems that reduce user satisfaction and convenience.

[0653] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0654] In this invention, the server includes an input means for a user to set the vehicle's personality and voice characteristics, a transmission means for converting the information set by the input means into a data packet and transmitting it to the server, an analysis means for analyzing the data packet transmitted by the transmission means and determining voice generation parameters based on the user's settings, a voice generation means for generating voice data using the voice generation parameters determined by the analysis means, and a transmission means for transmitting the generated voice data to the user's terminal. This allows the user to easily set the vehicle's personality and voice characteristics and quickly provide personalized voice dialogue based on them. Furthermore, it is possible to provide appropriate dialogue according to the user's emotional state, thereby improving the user experience.

[0655] A "user" is a person who uses this system to set the vehicle's personality and voice characteristics and engage in interactive communication.

[0656] "Vehicle personality" refers to the interaction attitude and atmosphere of the in-vehicle system set by the user, and includes characteristics such as "friendly," "cool," and "sporty."

[0657] "Voice characteristics" refers to the attributes of the voice emitted by the vehicle, and includes parameters such as "male voice," "female voice," "youthful voice," and "deep voice."

[0658] "Input means" refers to the device or interface through which a user configures the vehicle's personality and voice characteristics.

[0659] "Data packet" refers to a series of data containing information set by the user, and is used to send to the server.

[0660] "Transmission means" refers to a communication means for transmitting data packets from a terminal to a server, such as the Internet or a dedicated network.

[0661] The "analysis means" refers to a process or technology for analyzing the transmitted data packets to extract user setting information and determining voice generation parameters based on the extracted information.

[0662] "Voice generation parameters" refer to specific settings and templates for generating voice data based on the vehicle's personality and voice characteristics set by the user.

[0663] "Speech generation means" refers to a technology or engine for generating realistic and natural voice data using the voice generation parameters determined by the analysis means.

[0664] "Conversion means" refers to the process or equipment that converts the received audio data into a format that can be used by the in-vehicle system (e.g., WAV, MP3, etc.).

[0665] "Uploading means" refers to a means for uploading converted audio data to an in-vehicle system.

[0666] "Dialogue means" refers to the technology and functions for dialogue with the user using voice data generated by the in-vehicle system.

[0667] "Emotion analysis means" refers to technology or engines that analyze the user's voice and facial expressions to recognize emotions.

[0668] The "adjustment means" refers to a technique or process for adjusting the voice generation parameters and the dialogue content based on the emotional information recognized by the emotion analysis means.

[0669] MODE FOR CARRYING OUT THE INVENTION

[0670] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized dialogue.

[0671] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[0672] The device collects the setting information entered by the user and converts it into a data packet, which is then sent to the server.

[0673] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0674] Next, the server's speech synthesis engine generates realistic, natural-sounding voice data based on the determined voice generation parameters. This process uses speech synthesis technologies such as neural networks.

[0675] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[0676] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[0677] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[0678] While driving, the user can interact with the vehicle using voices generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?"

[0679] Furthermore, an emotion engine is integrated into the system, which analyzes the user's voice and facial expressions to recognize their emotions, for example, determining whether they are feeling stressed or having fun.

[0680] The server adjusts the voice generation parameters and dialogue content based on the emotional information recognized by the emotion engine, enabling dialogue appropriate to the user's emotional state. For example, if the vehicle recognizes that the user is tired, it might suggest, "Why don't you take a short break today?"

[0681] Specific examples

[0682] For example, a user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0683] The emotion engine analyzes the user's emotions and recognizes when the user is feeling stressed while driving. Based on this information, the system can say, "I'll play some relaxing music," and play music accordingly, making the user's driving experience richer and more personalized.

[0684] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0685] Step 1:

[0686] The user launches the dedicated app.

[0687] Specific action: The user taps or clicks to launch a dedicated application installed on their smartphone or computer.

[0688] Input: User Action

[0689] Output: App home screen display

[0690] Step 2:

[0691] The user selects the vehicle's personality and voice characteristics.

[0692] How it works: Select a personality type such as "friendly," "cool," or "sporty" and voice characteristics such as "male voice," "female voice," or "youthful voice" from the app's graphical interface.

[0693] Input: User-selected personality and voice characteristics

[0694] Output: Selected personality and vocal trait data

[0695] Step 3:

[0696] The terminal converts the setting information into a data packet and transmits it to the server.

[0697] What it does: The application packages the user's selected settings into a data packet and sends it over the Internet to a server.

[0698] Input: User-selected setting information

[0699] Output: Data packets sent to the server

[0700] Step 4:

[0701] The server receives the data packets and performs the analysis.

[0702] Specific operation: The server analyzes the data packets received from the terminal and extracts the user's setting information.

[0703] Input: Data packets from the terminal

[0704] Output: Parsed configuration information

[0705] Step 5:

[0706] The server determines the voice generation parameters.

[0707] Specific operation: Based on the analyzed configuration information, the server selects an appropriate template from the voice sample library and determines voice generation parameters.

[0708] Input: Parsed configuration information

[0709] Output: Speech generation parameters

[0710] Step 6:

[0711] The server generates voice data based on the voice generation parameters.

[0712] How it works: The speech synthesis engine uses speech generation parameters to generate realistic, natural-sounding speech data, using technologies such as neural networks.

[0713] Input: Speech generation parameters

[0714] Output: Generated audio data

[0715] Step 7:

[0716] The server transmits the generated voice data to the terminal.

[0717] Specific operation: The generated voice data is converted into data packets and sent to the terminal.

[0718] Input: Generated audio data

[0719] Output: Data packets to the terminal

[0720] Step 8:

[0721] The device receives the voice data, analyzes it, and converts it into a format that can be used by the in-vehicle system.

[0722] Specific operation: The device analyzes the audio data it receives and converts it into a format that can be used by the in-car system, such as WAV or MP3.

[0723] Input: Audio data sent to the device

[0724] Output: Converted audio data

[0725] Step 9:

[0726] The terminal uploads the audio data to the in-vehicle system.

[0727] What it does: Uploads converted audio data to the vehicle's navigation and audio systems.

[0728] Input: Converted audio data

[0729] Output: Audio data uploaded to the in-car system

[0730] Step 10:

[0731] The user interacts using the generated voice.

[0732] Specific behavior: When the engine is started, the vehicle will speak to you in a preset personality and voice, saying, "Good morning. What music would you like to listen to today?"

[0733] Input: Audio data uploaded to the in-vehicle system

[0734] Output: Interaction with the vehicle

[0735] Step 11:

[0736] The emotion engine analyzes the user's emotions.

[0737] Specific operation: The emotion engine analyzes voice and facial expressions to recognize emotions such as stress and joy.

[0738] Input: User's voice and facial expression data

[0739] Output: Recognized emotion information

[0740] Step 12:

[0741] The server adjusts the voice generation parameters and dialogue content based on the emotion information.

[0742] Specific operation: Based on the emotional information recognized by the emotion engine, the server adjusts the voice generation parameters and dialogue content, for example, to say, "Would you like to take a short break today?"

[0743] Input: Recognized emotion information

[0744] Output: Adjusted voice generation parameters and dialogue content

[0745] (Application example 2)

[0746] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0747] Autonomous vehicles are expected to provide a personalized travel experience through dialogue between passengers and the vehicle. However, current systems are unable to recognize passenger emotions in real time and flexibly adapt the dialogue content with the vehicle based on those emotions. This makes it difficult to achieve interactive communication that truly satisfies passengers.

[0748] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes input means for the user to set the vehicle's personality and voice characteristics, analysis means for analyzing the information set by the input means, voice generation means for generating voice data based on the information analyzed by the analysis means, transmission means for transmitting the generated voice data to the user's terminal, emotion recognition means for analyzing the user's emotions, and adjustment means for adjusting voice generation parameters in accordance with the emotion information obtained by the emotion recognition means. This enables personalized dialogue in real time according to the user's emotions.

[0749] "User" means any person or entity that uses the vehicle or system.

[0750] "Vehicle personality" is a concept that refers to the emotional and behavioral characteristics that a vehicle is supposed to have, which are set by the user.

[0751] "Voice characteristics" are attributes that refer to voice characteristics such as the quality, tone, and timbre of the voice emitted from the vehicle.

[0752] The term "input means" refers to a device or interface that allows a user to input setting information.

[0753] "Analysis means" refers to a system that includes software and hardware for analyzing input information and understanding its meaning and content.

[0754] "Speech generation means" refers to a function that includes a speech synthesis engine and related technologies for generating speech data based on analyzed information.

[0755] "Transmission means" refers to the infrastructure, which refers to the communication technologies and protocols used to transmit the generated audio data to other devices or systems.

[0756] "Emotion recognition means" refers to a system that uses technology or algorithms to analyze and identify emotions from a user's facial expressions, voice, or other input.

[0757] The "adjustment means" is a function that refers to a method or technique for dynamically changing voice generation parameters based on information obtained by the emotion recognition means.

[0758] "Conversion means" is a function that refers to technology or devices for converting transmitted audio data into a format that can be used by the in-vehicle system.

[0759] "Uploading means" refers to the infrastructure that refers to the methods and protocols used to transfer and store the converted audio data in the in-vehicle system.

[0760] "Dialogue means" refers to the technology and functions that enable the in-vehicle system to communicate with the user via voice.

[0761] The "display means" is an apparatus that refers to a device or technology for displaying the dialogue content based on the voice generation parameters adjusted by the adjustment means.

[0762] MODE FOR CARRYING OUT THE INVENTION

[0763] The following describes an embodiment of the present invention: This system has a structure that allows real-time dialogue within a vehicle based on the vehicle's personality and voice characteristics designed by the user.

[0764] System Overview

[0765] The system consists of the following main components:

[0766] 1. Input Method

[0767] Users can configure their vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app's interface is designed to make configuration easy for users. The configured information is collected by the app and converted into data packets.

[0768] 2. Analysis method

[0769] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. During this process, it uses a voice sample library to select a voice template that matches the settings.

[0770] 3. Voice Generation Method

[0771] A speech synthesis engine (such as Google Cloud TTS or Amazon Polly) in the server generates realistic and natural voice data based on the analyzed data. The generated voice data is stored in the server and a data packet is generated for processing.

[0772] 4. Transmission Method

[0773] The audio data is sent from the server to the device using a common data transfer protocol (e.g., HTTP or HTTPS).

[0774] 5. Emotion recognition means

[0775] An emotion engine (for example, Microsoft Azure's Emotion API) analyzes the user's voice and facial expressions to recognize their emotions. This information is collected in real time and sent to a server.

[0776] 6. Adjustment means

[0777] The server uses the information obtained from the emotion recognition means to adjust the voice and dialogue played by the in-vehicle system, using predefined prompts.

[0778] Specific examples

[0779] The user opens the app and selects settings such as "friendly personality," "male voice," and "young voice in his 30s." The app then sends this setting information to the server. The server analyzes the information and determines the appropriate voice generation parameters. The voice synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0780] Prompt Sentence Examples

[0781] When passengers board the vehicle:

[0782] "Hello, how's it going today?"

[0783] If the user says they are tired:

[0784] "Thank you for your hard work. I'll put on some relaxing music."

[0785] In this way, the system can combine emotion recognition and speech generation techniques to provide personalized interactions according to the user's emotions.

[0786] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0787] Step 1:

[0788] The user opens a dedicated app on their smartphone or computer and sets the vehicle's personality and voice characteristics. The input information can include, for example, "friendly personality," "male voice," or "youthful voice in his 30s." The app collects this information as a data packet and sends it to a server.

[0789] Input: User setting information

[0790] Data processing: Converting user setting information into data packets

[0791] Output: Data packets

[0792] Step 2:

[0793] The server receives the data packet sent from the device and starts the analysis process. Based on the configuration information, it references the voice sample library to determine voice generation parameters, and selects a template such as a "friendly male voice."

[0794] Input: Data packet

[0795] Data calculation: Determining voice generation parameters

[0796] Output: Speech generation parameters

[0797] Step 3:

[0798] The server's speech synthesis engine generates voice data based on the determined voice generation parameters. During this process, realistic and natural voices are generated using, for example, Google Cloud TTS or Amazon Polly. The generated voice data is stored on the server.

[0799] Input: Speech generation parameters

[0800] Data Computation: Generation of voice data using a TTS engine

[0801] Output: Audio data

[0802] Step 4:

[0803] The server converts the generated voice data into data packets and sends them to the terminal, using common data transfer protocols (HTTP or HTTPS).

[0804] Input: Audio data

[0805] Data processing: Converting voice data into data packets

[0806] Output: Data packets

[0807] Step 5:

[0808] The device receives the data packets sent from the server, analyzes the audio data, and converts it into a common audio file format (e.g., WAV or MP3).

[0809] Input: Data packet

[0810] Data processing: Convert data packets into audio file format

[0811] Output: Audio file

[0812] Step 6:

[0813] The device then uploads the converted audio file to the in-car system via communication methods such as Bluetooth or Wi-Fi.

[0814] Input: Audio file

[0815] Data Calculation: Uploading Audio Files

[0816] Output: Uploaded audio file

[0817] Step 7:

[0818] The user plays generated audio through the in-car system, initiating a dialogue with the vehicle, for example, the vehicle saying, "Hello, how's your day?"

[0819] Input: Uploaded audio file

[0820] Data calculation: Audio playback by in-car systems

[0821] Output: Voice dialogue

[0822] Step 8:

[0823] The emotion engine analyzes the user's voice and facial expressions to recognize emotions, for example, detecting the user's stress level using Microsoft Azure's Emotion API.

[0824] Input: User's voice and facial expression data

[0825] Data Computing: Emotion Analysis and Recognition

[0826] Output: Emotional information

[0827] Step 9:

[0828] The server adjusts the voice generation parameters using the adjustment means based on the emotion information obtained by the emotion recognition means. For example, if the server recognizes that the user is tired, it suggests playing relaxing music.

[0829] Input: Emotion information

[0830] Data calculation: Adjustment of voice generation parameters

[0831] Output: Adjusted speech generation parameters

[0832] Step 10:

[0833] The in-vehicle system then adjusts the dialogue content based on the adjusted voice generation parameters to provide a dialogue tailored to the user. For example, the vehicle might say, "Thank you for your hard work. I'll play some relaxing music."

[0834] Input: Adjusted speech generation parameters

[0835] Data calculation: Change of dialogue content

[0836] Output: Personalized dialogue

[0837] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0838] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0839] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0840] [Third embodiment]

[0841] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0842] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0843] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0844] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0845] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0846] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0847] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0848] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0849] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0850] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0851] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0852] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0853] MODE FOR CARRYING OUT THE INVENTION

[0854] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. To implement this system, the following program processing is performed.

[0855] User settings processing

[0856] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[0857] The terminal collects the input configuration information and converts it into a data packet, which is then sent to the server.

[0858] Data analysis and speech generation

[0859] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. This includes selecting an appropriate template from existing voice samples.

[0860] A speech synthesis engine in the server then generates realistic, natural-sounding speech data based on the determined speech generation parameters, for example, using neural network-based techniques.

[0861] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[0862] Providing voice data and implementing dialogue functions

[0863] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[0864] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[0865] Finally, while actually driving, users can enjoy interacting with the vehicle using the voice generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?" In this way, conversational communication is realized.

[0866] Specific examples

[0867] As an example, consider the following scenario: A user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[0868] When the user starts the engine while driving, the vehicle will speak to the user in a friendly voice, asking, "Good morning! Where are you going today?" In this way, the present invention enables interactive communication with the vehicle, enriching the user's driving experience.

[0869] The processing flow will be explained below.

[0870] Step 1:

[0871] The user launches a dedicated app on their smartphone or computer, and then enters the vehicle's personality and voice characteristics on the app's settings screen. Specifically, the user selects options such as "friendly personality," "male voice," or "youthful voice."

[0872] Step 2:

[0873] The device collects the configuration information entered by the user and converts it into a data packet, which contains details such as the vehicle's personality, the gender of the voice, and the tone of the voice.

[0874] Step 3:

[0875] The device generates data packets and sends them to the server over Wi-Fi or mobile data networks.

[0876] Step 4:

[0877] The server receives the setting information sent from the terminal and stores the received data in a database within the server.

[0878] Step 5:

[0879] The server analyzes the received data. Specifically, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[0880] Step 6:

[0881] The server's internal voice synthesis engine generates realistic, natural voice data based on the determined voice generation parameters. This process uses voice synthesis technologies such as neural networks.

[0882] Step 7:

[0883] The server stores the generated voice data and generates new data packets to send to the user's terminal.

[0884] Step 8:

[0885] The server then sends the generated data packets to the user's device, again over Wi-Fi or a mobile data network.

[0886] Step 9:

[0887] The terminal acquires the voice data transmitted from the server.

[0888] Step 10:

[0889] The device analyzes the received voice data and converts it into a format that can be used by the in-vehicle system, specifically into common audio file formats such as WAV and MP3.

[0890] Step 11:

[0891] The device then uploads the converted audio data to the vehicle's system via Bluetooth, USB connection, or other means.

[0892] Step 12:

[0893] The in-car system plays back the audio data while the user is driving. For example, when the engine is started, the vehicle may say, "Good morning. Where are you going today?"

[0894] Through the above steps, users can freely customize the vehicle's personality and voice characteristics, enjoying a personalized interactive driving experience.

[0895] Example 1

[0896] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0897] Conventional in-vehicle systems have made it difficult for users to enjoy a personalized experience through interaction with the vehicle. Furthermore, they lacked a means to freely set the vehicle's personality and voice characteristics and generate natural-sounding voice data based on those settings. As a result, users could only receive uniform, standardized voice guidance, limiting their driving experience. To solve this issue, a system was needed that could flexibly generate voice data based on parameters set by the user and provide a personalized experience through interaction with the vehicle.

[0898] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0899] In this invention, the server includes input means for a user to set the vehicle's personality and voice characteristics, transmission means for converting the information set by the input means into data packets and transmitting the data to the server, analysis means for analyzing the information transmitted by the transmission means and determining voice generation parameters based on the personality and voice characteristics specified by the user, voice generation means for generating natural voice data using a voice synthesis engine based on the voice generation parameters, and transmission means for transmitting the generated voice data to the user's terminal, thereby enabling the user to realize interactive communication based on personalized voice data.

[0900] "Input means" refers to the interface that the user utilizes to configure the vehicle's personality and voice characteristics.

[0901] The "transmission means" is a device, software, or protocol that has the function of converting the information set by the input means into a data packet and transmitting it to the server.

[0902] The "analysis means" is a device or software that has the function of analyzing the data packets received by the server and determining voice generation parameters based on the personality and voice characteristics specified by the user.

[0903] The "voice generation means" is a device or software that has the function of generating natural voice data using a voice synthesis engine based on the voice generation parameters determined by the analysis means.

[0904] The "conversion means" is a device or software that has the function of analyzing the voice data transmitted by the transmission means and converting it into a format that can be used by the in-vehicle system.

[0905] The "uploading means" is a device or software that has the function of uploading the voice data generated by the conversion means to the in-vehicle system.

[0906] The "interaction means" is a device or software that has the function of executing an interaction function between a user and a vehicle using the voice data uploaded to the in-vehicle system by the upload means.

[0907] In the system of the present invention, the user sets the vehicle's personality and voice characteristics, and the server generates voice data based on that information, realizing interactive communication between the user and the vehicle. To implement this system, the following program processing is required.

[0908] User settings input method

[0909] Users install the dedicated application "VehicleVoiceCustomizer" on their smartphone or computer. The application provides a user interface where users can set the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). Once the settings are complete, users press the "Settings" button to proceed to the next step.

[0910] Sending configuration information

[0911] The device converts the user-entered configuration information into data packets containing information about the selected personality and voice characteristics, which are then sent to the server using Internet communication protocols (e.g., HTTP or HTTPS).

[0912] Data analysis and speech generation methods

[0913] The server receives the data packets sent from the device. It analyzes the received information and determines voice generation parameters based on the personality and voice characteristics specified by the user. This analysis is performed using a voice synthesis engine, such as the Google Cloud Text-to-Speech API. The voice synthesis engine uses neural network-based technology (e.g., WaveNet) to generate natural-sounding voice data. The generated voice data is stored on the server and prepared for the next step.

[0914] Voice data transmission method

[0915] The server converts the generated voice data into data packets and transmits the packets to the terminal, again using the Internet communication protocol.

[0916] A means of converting voice data and uploading it to an in-vehicle system

[0917] The device receives the audio data sent from the server. It uses tools such as FFmpeg to analyze the received audio data and convert it into a format that can be used by the in-vehicle system (e.g., WAV or MP3). The converted audio data is then uploaded to the in-vehicle system using a connection method such as Bluetooth, Wi-Fi, or USB.

[0918] Interaction methods

[0919] When a user gets into the vehicle and starts the engine, the in-vehicle system will begin playing the uploaded voice data. For example, the vehicle may ask in a friendly voice, "Good morning! Where are you going today?" This dialogue function allows users to enjoy interactive communication with the vehicle.

[0920] Specific examples

[0921] Consider the following example: A user opens the app "VehicleVoiceCustomizer" and sets a "friendly personality," a "male voice," and a "young voice in his 30s." The app then converts this configuration information into a data packet and sends it to the server. The server analyzes the information and determines appropriate voice generation parameters. It then generates voice data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech API, WaveNet). This generated voice data is then sent from the server to the device. The device receives the voice data, uses "FFmpeg" to convert the audio file into a format usable by the in-vehicle system, and uploads it to the in-vehicle system via Bluetooth. When the user starts the engine, the vehicle speaks in a friendly voice, asking, "Good morning! Where are you going today?"

[0922] An example of a prompt for a generative AI model is:

[0923] "Generate a friendly vehicle speaking to the user in the youthful voice of a man in his 30s."

[0924] This allows the system to provide interactive communication according to the user's wishes, resulting in richer interaction with the vehicle.

[0925] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0926] Step 1: Enter your user settings

[0927] The user launches the dedicated app "Vehicle Voice Customizer" on their smartphone or computer. Through the app's graphical interface, they can set the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). The information set by the user is temporarily stored within the app. The input here is the personality and voice characteristics selected by the user, and the output is a data packet containing the setting information.

[0928] Step 2: Sending configuration information

[0929] The device converts the user-entered setting information into a data packet. Specifically, the selected personality and voice characteristics are packetized as text data. The converted data packet is sent to the server using the HTTP / HTTPS protocol. The input here is the user-selected setting information, and the output is the data packet sent to the server.

[0930] Step 3: Data analysis

[0931] The server analyzes data packets received from the device. During the analysis, it extracts speech generation parameters based on the personality and voice characteristics specified by the user. For example, it uses the Google Cloud Text-to-Speech API to select a speech template. The input of this process is the received data packets, and the output is the speech generation parameters.

[0932] Step 4: Speech generation

[0933] The speech synthesis engine in the server generates speech data using the speech generation parameters extracted earlier. Here, neural network-based technology (e.g., WaveNet) is used to generate natural and realistic speech data. The input of this step is the speech generation parameters, and the output is the generated speech data.

[0934] Step 5: Sending audio data

[0935] The server reconverts the generated voice data into data packets and sends them to the terminal again using the HTTP / HTTPS protocol. The input of this process is the generated voice data, and the output is the data packets sent from the server.

[0936] Step 6: Receiving and converting audio data

[0937] The device receives the audio data sent from the server. Because the received audio data is compressed and encrypted, it is first analyzed and decrypted. Then, a tool such as FFmpeg is used to convert it into a format that can be used by the in-vehicle system (e.g., WAV or MP3). The input of this step is the received data packet, and the output is audio data in a format suitable for the in-vehicle system.

[0938] Step 7: Upload your audio data

[0939] The terminal uploads the converted voice data to the in-vehicle system via Bluetooth, Wi-Fi, USB, etc. The input of this step is the converted voice data, and the output is the voice data uploaded to the in-vehicle system.

[0940] Step 8: Implementing Interactivity

[0941] When the user starts the vehicle, the in-vehicle system starts playing the uploaded voice data. For example, the vehicle may start a dialogue with the user by saying, "Good morning! Where are you going today?" The input of this step is the voice data uploaded to the in-vehicle system, and the output is the interactive dialogue between the user and the vehicle.

[0942] (Application example 1)

[0943] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0944] Conventional autonomous vehicles have a problem in that the driving experience is mechanical due to the limited information provided by the vehicle and one-way communication with the user. Furthermore, there is a lack of a way to realize a dialogue system with the personality and voice characteristics desired by the user. Furthermore, there is a lack of a means to provide navigation information, traffic information, weather information, maintenance information, etc. in a unified and interactive manner.

[0945] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0946] In this invention, the server includes an input means for a user to set the personality and voice characteristics of a vehicle, an analysis means for analyzing the information set by the input means, a voice generation means for generating voice data based on the information analyzed by the analysis means, a transmission means for transmitting the generated voice data to the user's terminal, a navigation means for providing navigation information for the vehicle, a traffic information providing means for providing traffic information based on information from the navigation means, and an information providing means for providing weather and maintenance information. This allows the user to interact with an autonomous vehicle with a personality and voice characteristics tailored to their preferences, enriching the driving experience. Furthermore, the provision of comprehensive, interactive information improves the user's comfort and safety.

[0947] Understood. Now, I will create definitions for the important words included in the rewritten claims.

[0948] "Mobile object" refers to a vehicle or other conveyance, including those with automatic driving capabilities.

[0949] "Personality" refers to the personified characteristics of the mobile object set by the user, and includes characteristics such as friendly, cool, sporty, etc.

[0950] "Voice characteristics" refers to the characteristics of a voice set by a user, and includes attributes such as gender, age, tone of voice, and pitch.

[0951] "Input means" refers to an interface that allows a user to set the vehicle's personality and voice characteristics, such as a smartphone app or an in-car console.

[0952] The "analysis means" refers to a computer system for analyzing the information set by the input means.

[0953] "Speech generation means" refers to a system for generating speech data based on the information analyzed by the analysis means.

[0954] "Transmission means" refers to a function for transmitting the generated voice data to the user's terminal.

[0955] "Navigation means" refers to a system for guiding a route to a destination of a mobile object.

[0956] "Traffic information providing means" refers to a function that provides the user with traffic conditions based on information from the navigation means.

[0957] "Information provision means" refers to a function for providing weather and maintenance information to users.

[0958] "Conversion means" refers to a computer program for converting transmitted voice data into a format usable by the mobile system.

[0959] "Uploading means" refers to a function for uploading converted audio data to a mobile system.

[0960] "Interaction means" refers to a system that allows a user to interact with a mobile object using uploaded voice data.

[0961] "Entertainment means" refers to the capability to provide audio and visual entertainment over a mobile system.

[0962] In the system of the present invention, the user sets the personality and voice characteristics of the mobile object, and AI generates the "voice" of the mobile object based on that information, realizing interactive communication between the user and the mobile object. To implement this system, the following program processing is performed.

[0963] User settings processing

[0964] First, the user sets the mobile device's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the mobile device's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). The device collects the entered setting information and converts it into a data packet. Then, the data packet is sent to the server.

[0965] Data analysis and speech generation

[0966] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. This process involves selecting an appropriate template from existing voice samples. Next, the server's speech synthesis engine generates natural-sounding voice data based on the determined voice generation parameters. This voice data is generated using, for example, neural network-based technology (e.g., Google Cloud Text-to-Speech API or Amazon Polly). The generated voice data is stored on the server, and a data packet is generated to be sent to the device. This packet is then sent back to the device.

[0967] Providing voice data and implementing dialogue functions

[0968] The receiving device receives the voice data sent from the server, analyzes it, and converts it into a format that can be used by the mobile system. For example, it may be converted into a common audio file format (WAV, MP3, etc.). The device then uploads the converted voice data to the mobile system, where it can be used in the vehicle's navigation and audio systems. Finally, during actual autonomous driving, the user can enjoy interacting with the vehicle using the voice generated by the system. For example, the vehicle may ask, "Hello! Where are you going today?" Furthermore, the navigation and traffic information means may suggest appropriate routes and report traffic conditions, while the entertainment means may play music or provide simple conversations.

[0969] Usage example

[0970] As an example of usage, consider the following scenario: A user opens the app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to the server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device.

[0971] The terminal converts the voice data and uploads it to the vehicle system. During autonomous driving, when the user starts the engine, the vehicle will speak to them in a friendly voice, asking, "Hello! Where are you going today?"

[0972] For example, the Google Maps API is used for navigation, providing route guidance to destinations, and real-time traffic information is provided as a means of providing traffic information. For entertainment, music streaming services such as Spotify are integrated, allowing users to play music according to their preferences.

[0973] Prompt Sentence Examples

[0974] assistant = AIDriverAssistant("Friendly", "Young male voice")

[0975] print(assistant.generate_response("greeting"))

[0976] In this way, the present invention allows interactive communication with the vehicle, enriching the user's driving experience.

[0977] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0978] Step 1:

[0979] The user uses a dedicated app on a smartphone or computer to input data to set the mobile device's personality and voice characteristics. Specifically, the user selects personality traits such as "friendly," "cool," or "sporty" and voice characteristics such as "male voice," "female voice," or "youthful voice" on the app. This is input data, and the collected information is sent to the device as a data packet.

[0980] Step 2:

[0981] The terminal sends a data packet containing user-defined information to the server. This packet contains information about personality and voice characteristics. This information is sent via a data transfer protocol and received by the server.

[0982] Step 3:

[0983] The server receives the data packet sent from the terminal and analyzes its contents. The analysis means analyzes the contents of the data packet (personality, voice characteristics, etc.) and determines appropriate voice generation parameters. For example, based on information such as "friendly" or "youthful male voice," it generates the parameters required for the voice generation engine.

[0984] Step 4:

[0985] The server generates natural-sounding voice data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech API or Amazon Polly) based on the voice generation parameters determined by the analysis means. This generated voice data is converted into an audio file format (WAV, MP3, etc.) through a series of data calculations.

[0986] Step 5:

[0987] The server converts the generated voice data into data packets and transmits the data packets to the terminal, where the voice data may be compressed. The terminal receives the data packets transmitted from the server.

[0988] Step 6:

[0989] The terminal analyzes the voice data sent from the server and converts it into a format that can be used by the mobile system (for example, WAV or MP3 format). The voice data conversion is performed using a format conversion algorithm.

[0990] Step 7:

[0991] The device then uploads the converted voice data to the in-vehicle system, where it can be used in the navigation and audio systems. Specifically, the device stores the voice data in the in-vehicle system's memory or storage.

[0992] Step 8:

[0993] The user enjoys interacting with the vehicle during autonomous driving. The vehicle uses the generated voice to provide interactive navigation, traffic information, weather and maintenance information, and entertainment functions (e.g., music playback and chat). The interactive navigation uses a voice recognition system to understand user input and respond appropriately.

[0994] Example: When the user starts the engine, the vehicle will speak to them in a friendly voice, asking, "Hello! Where are you going today?" It is possible to use the Google Maps API as a navigation tool and provide real-time traffic information.

[0995] In this way, the present invention realizes interactive communication with the user and enriches the experience of using a mobile device.

[0996] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0997] MODE FOR CARRYING OUT THE INVENTION

[0998] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized dialogue.

[0999] User settings processing

[1000] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[1001] The device collects the setting information entered by the user and converts it into a data packet, which is then sent to the server.

[1002] Data analysis and speech generation

[1003] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[1004] Next, the server's speech synthesis engine generates realistic, natural-sounding voice data based on the determined voice generation parameters. This process uses speech synthesis technologies such as neural networks.

[1005] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[1006] Providing voice data and implementing dialogue functions

[1007] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[1008] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[1009] While driving, the user can interact with the vehicle using voices generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?"

[1010] Emotion engine integration

[1011] Furthermore, an emotion engine is integrated into the system, which analyzes the user's voice and facial expressions to recognize their emotions, for example, determining whether they are feeling stressed or having fun.

[1012] The server adjusts the voice generation parameters and dialogue content based on the emotional information recognized by the emotion engine, enabling dialogue appropriate to the user's emotional state. For example, if the vehicle recognizes that the user is tired, it might suggest, "Why don't you take a short break today?"

[1013] Specific examples

[1014] For example, a user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[1015] The emotion engine analyzes the user's emotions and recognizes when the user is feeling stressed while driving. Based on this information, the system can say, "I'll play some relaxing music," and play music accordingly, making the user's driving experience richer and more personalized.

[1016] The processing flow will be explained below.

[1017] Step 1:

[1018] The user launches a dedicated app on their smartphone or computer, and then enters the vehicle's personality and voice characteristics on the app's settings screen. Specifically, the user selects options such as "friendly personality," "male voice," or "young voice in his 30s."

[1019] Step 2:

[1020] The device collects the configuration information entered by the user and converts it into a data packet, which contains details such as the vehicle's personality, the gender of the voice, and the tone of the voice.

[1021] Step 3:

[1022] The device generates data packets and sends them to the server over Wi-Fi or mobile data networks.

[1023] Step 4:

[1024] The server receives the setting information sent from the terminal and stores the received data in a database within the server.

[1025] Step 5:

[1026] The server analyzes the received data. Specifically, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[1027] Step 6:

[1028] The server's internal voice synthesis engine generates realistic, natural voice data based on the determined voice generation parameters. This process uses voice synthesis technologies such as neural networks.

[1029] Step 7:

[1030] The server stores the generated voice data and generates new data packets to send to the user's terminal.

[1031] Step 8:

[1032] The server then sends the generated data packets to the user's device, again over Wi-Fi or a mobile data network.

[1033] Step 9:

[1034] The terminal acquires the voice data transmitted from the server.

[1035] Step 10:

[1036] The device analyzes the received voice data and converts it into a format that can be used by the in-vehicle system, specifically into common audio file formats such as WAV and MP3.

[1037] Step 11:

[1038] The device then uploads the converted audio data to the vehicle's system via Bluetooth, USB connection, or other means.

[1039] Step 12:

[1040] The in-car system plays back the audio data while the user is driving. For example, when the engine is started, the vehicle may say, "Good morning. Where are you going today?"

[1041] Emotion Engine Processing Flow

[1042] Step 13:

[1043] The emotion engine installed in the device captures the user's voice and facial expressions using the device's microphone and camera.

[1044] Step 14:

[1045] The device's emotion engine analyzes the voice and facial expression data it acquires to recognize the user's emotions. For example, it can determine whether the user is feeling "stressed" based on the tone of their voice and facial expression.

[1046] Step 15:

[1047] The device sends the recognized emotion information to the server, which includes information that the user is currently feeling "stressed."

[1048] Step 16:

[1049] The server receives and analyzes the emotional information. Based on this information, it adjusts the generated voice data and dialogue content. For example, it generates new voice data such as "Shall I play some calming music?" to help the user relax.

[1050] Step 17:

[1051] The server sends the newly generated voice data to the terminal.

[1052] Step 18:

[1053] The device then uploads the new voice data it receives back to the in-vehicle system.

[1054] Step 19:

[1055] New audio data will be played while the user is driving: the vehicle will say, "I'm going to play some music to help you relax," and then play appropriate music.

[1056] Through the above steps, the user can enjoy an interactive driving experience that is appropriate to their emotions.

[1057] Example 2

[1058] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1059] In recent years, interactive communication systems have become increasingly important for improving in-vehicle user experiences. However, existing systems are unable to provide personalized dialogue based on the user's individual preferences and emotional state. Furthermore, it is difficult for users to easily configure the vehicle's personality and voice characteristics and quickly provide personalized voice dialogue based on those settings. This has led to problems that reduce user satisfaction and convenience.

[1060] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1061] In this invention, the server includes an input means for a user to set the vehicle's personality and voice characteristics, a transmission means for converting the information set by the input means into a data packet and transmitting it to the server, an analysis means for analyzing the data packet transmitted by the transmission means and determining voice generation parameters based on the user's settings, a voice generation means for generating voice data using the voice generation parameters determined by the analysis means, and a transmission means for transmitting the generated voice data to the user's terminal. This allows the user to easily set the vehicle's personality and voice characteristics and quickly provide personalized voice dialogue based on them. Furthermore, it is possible to provide appropriate dialogue according to the user's emotional state, thereby improving the user experience.

[1062] A "user" is a person who uses this system to set the vehicle's personality and voice characteristics and engage in interactive communication.

[1063] "Vehicle personality" refers to the interaction attitude and atmosphere of the in-vehicle system set by the user, and includes characteristics such as "friendly," "cool," and "sporty."

[1064] "Voice characteristics" refers to the attributes of the voice emitted by the vehicle, and includes parameters such as "male voice," "female voice," "youthful voice," and "deep voice."

[1065] "Input means" refers to the device or interface through which a user configures the vehicle's personality and voice characteristics.

[1066] "Data packet" refers to a series of data containing information set by the user, and is used to send to the server.

[1067] "Transmission means" refers to a communication means for transmitting data packets from a terminal to a server, such as the Internet or a dedicated network.

[1068] The "analysis means" refers to a process or technology for analyzing the transmitted data packets to extract user setting information and determining voice generation parameters based on the extracted information.

[1069] "Voice generation parameters" refer to specific settings and templates for generating voice data based on the vehicle's personality and voice characteristics set by the user.

[1070] "Speech generation means" refers to a technology or engine for generating realistic and natural voice data using the voice generation parameters determined by the analysis means.

[1071] "Conversion means" refers to the process or equipment that converts the received audio data into a format that can be used by the in-vehicle system (e.g., WAV, MP3, etc.).

[1072] "Uploading means" refers to a means for uploading converted audio data to an in-vehicle system.

[1073] "Dialogue means" refers to the technology and functions for dialogue with the user using voice data generated by the in-vehicle system.

[1074] "Emotion analysis means" refers to technology or engines that analyze the user's voice and facial expressions to recognize emotions.

[1075] The "adjustment means" refers to a technique or process for adjusting the voice generation parameters and the dialogue content based on the emotional information recognized by the emotion analysis means.

[1076] MODE FOR CARRYING OUT THE INVENTION

[1077] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized dialogue.

[1078] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[1079] The device collects the setting information entered by the user and converts it into a data packet, which is then sent to the server.

[1080] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[1081] Next, the server's speech synthesis engine generates realistic, natural-sounding voice data based on the determined voice generation parameters. This process uses speech synthesis technologies such as neural networks.

[1082] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[1083] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[1084] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[1085] While driving, the user can interact with the vehicle using voices generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?"

[1086] Furthermore, an emotion engine is integrated into the system, which analyzes the user's voice and facial expressions to recognize their emotions, for example, determining whether they are feeling stressed or having fun.

[1087] The server adjusts the voice generation parameters and dialogue content based on the emotional information recognized by the emotion engine, enabling dialogue appropriate to the user's emotional state. For example, if the vehicle recognizes that the user is tired, it might suggest, "Why don't you take a short break today?"

[1088] Specific examples

[1089] For example, a user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[1090] The emotion engine analyzes the user's emotions and recognizes when the user is feeling stressed while driving. Based on this information, the system can say, "I'll play some relaxing music," and play music accordingly, making the user's driving experience richer and more personalized.

[1091] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1092] Step 1:

[1093] The user launches the dedicated app.

[1094] Specific action: The user taps or clicks to launch a dedicated application installed on their smartphone or computer.

[1095] Input: User Action

[1096] Output: App home screen display

[1097] Step 2:

[1098] The user selects the vehicle's personality and voice characteristics.

[1099] How it works: Select a personality type such as "friendly," "cool," or "sporty" and voice characteristics such as "male voice," "female voice," or "youthful voice" from the app's graphical interface.

[1100] Input: User-selected personality and voice characteristics

[1101] Output: Selected personality and vocal trait data

[1102] Step 3:

[1103] The terminal converts the setting information into a data packet and transmits it to the server.

[1104] What it does: The application packages the user's selected settings into a data packet and sends it over the Internet to a server.

[1105] Input: User-selected setting information

[1106] Output: Data packets sent to the server

[1107] Step 4:

[1108] The server receives the data packets and performs the analysis.

[1109] Specific operation: The server analyzes the data packets received from the terminal and extracts the user's setting information.

[1110] Input: Data packets from the terminal

[1111] Output: Parsed configuration information

[1112] Step 5:

[1113] The server determines the voice generation parameters.

[1114] Specific operation: Based on the analyzed configuration information, the server selects an appropriate template from the voice sample library and determines voice generation parameters.

[1115] Input: Parsed configuration information

[1116] Output: Speech generation parameters

[1117] Step 6:

[1118] The server generates voice data based on the voice generation parameters.

[1119] How it works: The speech synthesis engine uses speech generation parameters to generate realistic, natural-sounding speech data, using technologies such as neural networks.

[1120] Input: Speech generation parameters

[1121] Output: Generated audio data

[1122] Step 7:

[1123] The server transmits the generated voice data to the terminal.

[1124] Specific operation: The generated voice data is converted into data packets and sent to the terminal.

[1125] Input: Generated audio data

[1126] Output: Data packets to the terminal

[1127] Step 8:

[1128] The device receives the voice data, analyzes it, and converts it into a format that can be used by the in-vehicle system.

[1129] Specific operation: The device analyzes the audio data it receives and converts it into a format that can be used by the in-car system, such as WAV or MP3.

[1130] Input: Audio data sent to the device

[1131] Output: Converted audio data

[1132] Step 9:

[1133] The terminal uploads the audio data to the in-vehicle system.

[1134] What it does: Uploads converted audio data to the vehicle's navigation and audio systems.

[1135] Input: Converted audio data

[1136] Output: Audio data uploaded to the in-car system

[1137] Step 10:

[1138] The user interacts using the generated voice.

[1139] Specific behavior: When the engine is started, the vehicle will speak to you in a preset personality and voice, saying, "Good morning. What music would you like to listen to today?"

[1140] Input: Audio data uploaded to the in-vehicle system

[1141] Output: Interaction with the vehicle

[1142] Step 11:

[1143] The emotion engine analyzes the user's emotions.

[1144] Specific operation: The emotion engine analyzes voice and facial expressions to recognize emotions such as stress and joy.

[1145] Input: User's voice and facial expression data

[1146] Output: Recognized emotion information

[1147] Step 12:

[1148] The server adjusts the voice generation parameters and dialogue content based on the emotion information.

[1149] Specific operation: Based on the emotional information recognized by the emotion engine, the server adjusts the voice generation parameters and dialogue content, for example, to say, "Would you like to take a short break today?"

[1150] Input: Recognized emotion information

[1151] Output: Adjusted voice generation parameters and dialogue content

[1152] (Application example 2)

[1153] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1154] Autonomous vehicles are expected to provide a personalized travel experience through dialogue between passengers and the vehicle. However, current systems are unable to recognize passenger emotions in real time and flexibly adapt the dialogue content with the vehicle based on those emotions. This makes it difficult to achieve interactive communication that truly satisfies passengers.

[1155] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes input means for the user to set the vehicle's personality and voice characteristics, analysis means for analyzing the information set by the input means, voice generation means for generating voice data based on the information analyzed by the analysis means, transmission means for transmitting the generated voice data to the user's terminal, emotion recognition means for analyzing the user's emotions, and adjustment means for adjusting voice generation parameters in accordance with the emotion information obtained by the emotion recognition means. This enables personalized dialogue in real time according to the user's emotions.

[1156] "User" means any person or entity that uses the vehicle or system.

[1157] "Vehicle personality" is a concept that refers to the emotional and behavioral characteristics that a vehicle is supposed to have, which are set by the user.

[1158] "Voice characteristics" are attributes that refer to voice characteristics such as the quality, tone, and timbre of the voice emitted from the vehicle.

[1159] The term "input means" refers to a device or interface that allows a user to input setting information.

[1160] "Analysis means" refers to a system that includes software and hardware for analyzing input information and understanding its meaning and content.

[1161] "Speech generation means" refers to a function that includes a speech synthesis engine and related technologies for generating speech data based on analyzed information.

[1162] "Transmission means" refers to the infrastructure, which refers to the communication technologies and protocols used to transmit the generated audio data to other devices or systems.

[1163] "Emotion recognition means" refers to a system that uses technology or algorithms to analyze and identify emotions from a user's facial expressions, voice, or other input.

[1164] The "adjustment means" is a function that refers to a method or technique for dynamically changing voice generation parameters based on information obtained by the emotion recognition means.

[1165] "Conversion means" is a function that refers to technology or devices for converting transmitted audio data into a format that can be used by the in-vehicle system.

[1166] "Uploading means" refers to the infrastructure that refers to the methods and protocols used to transfer and store the converted audio data in the in-vehicle system.

[1167] "Dialogue means" refers to the technology and functions that enable the in-vehicle system to communicate with the user via voice.

[1168] The "display means" is an apparatus that refers to a device or technology for displaying the dialogue content based on the voice generation parameters adjusted by the adjustment means.

[1169] MODE FOR CARRYING OUT THE INVENTION

[1170] The following describes an embodiment of the present invention: This system has a structure that allows real-time dialogue within a vehicle based on the vehicle's personality and voice characteristics designed by the user.

[1171] System Overview

[1172] The system consists of the following main components:

[1173] 1. Input Method

[1174] Users can configure their vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app's interface is designed to make configuration easy for users. The configured information is collected by the app and converted into data packets.

[1175] 2. Analysis method

[1176] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. During this process, it uses a voice sample library to select a voice template that matches the settings.

[1177] 3. Voice Generation Method

[1178] A speech synthesis engine (such as Google Cloud TTS or Amazon Polly) in the server generates realistic and natural voice data based on the analyzed data. The generated voice data is stored in the server and a data packet is generated for processing.

[1179] 4. Transmission Method

[1180] The audio data is sent from the server to the device using a common data transfer protocol (e.g., HTTP or HTTPS).

[1181] 5. Emotion recognition means

[1182] An emotion engine (for example, Microsoft Azure's Emotion API) analyzes the user's voice and facial expressions to recognize their emotions. This information is collected in real time and sent to a server.

[1183] 6. Adjustment means

[1184] The server uses the information obtained from the emotion recognition means to adjust the voice and dialogue played by the in-vehicle system, using predefined prompts.

[1185] Specific examples

[1186] The user opens the app and selects settings such as "friendly personality," "male voice," and "young voice in his 30s." The app then sends this setting information to the server. The server analyzes the information and determines the appropriate voice generation parameters. The voice synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[1187] Prompt Sentence Examples

[1188] When passengers board the vehicle:

[1189] "Hello, how's it going today?"

[1190] If the user says they are tired:

[1191] "Thank you for your hard work. I'll put on some relaxing music."

[1192] In this way, the system can combine emotion recognition and speech generation techniques to provide personalized interactions according to the user's emotions.

[1193] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1194] Step 1:

[1195] The user opens a dedicated app on their smartphone or computer and sets the vehicle's personality and voice characteristics. The input information can include, for example, "friendly personality," "male voice," or "youthful voice in his 30s." The app collects this information as a data packet and sends it to a server.

[1196] Input: User setting information

[1197] Data processing: Converting user setting information into data packets

[1198] Output: Data packets

[1199] Step 2:

[1200] The server receives the data packet sent from the device and starts the analysis process. Based on the configuration information, it references the voice sample library to determine voice generation parameters, and selects a template such as a "friendly male voice."

[1201] Input: Data packet

[1202] Data calculation: Determining voice generation parameters

[1203] Output: Speech generation parameters

[1204] Step 3:

[1205] The server's speech synthesis engine generates voice data based on the determined voice generation parameters. During this process, realistic and natural voices are generated using, for example, Google Cloud TTS or Amazon Polly. The generated voice data is stored on the server.

[1206] Input: Speech generation parameters

[1207] Data Computation: Generation of voice data using a TTS engine

[1208] Output: Audio data

[1209] Step 4:

[1210] The server converts the generated voice data into data packets and sends them to the terminal, using common data transfer protocols (HTTP or HTTPS).

[1211] Input: Audio data

[1212] Data processing: Converting voice data into data packets

[1213] Output: Data packets

[1214] Step 5:

[1215] The device receives the data packets sent from the server, analyzes the audio data, and converts it into a common audio file format (e.g., WAV or MP3).

[1216] Input: Data packet

[1217] Data processing: Convert data packets into audio file format

[1218] Output: Audio file

[1219] Step 6:

[1220] The device then uploads the converted audio file to the in-car system via communication methods such as Bluetooth or Wi-Fi.

[1221] Input: Audio file

[1222] Data Calculation: Uploading Audio Files

[1223] Output: Uploaded audio file

[1224] Step 7:

[1225] The user plays generated audio through the in-car system, initiating a dialogue with the vehicle, for example, the vehicle saying, "Hello, how's your day?"

[1226] Input: Uploaded audio file

[1227] Data calculation: Audio playback by in-car systems

[1228] Output: Voice dialogue

[1229] Step 8:

[1230] The emotion engine analyzes the user's voice and facial expressions to recognize emotions, for example, detecting the user's stress level using Microsoft Azure's Emotion API.

[1231] Input: User's voice and facial expression data

[1232] Data Computing: Emotion Analysis and Recognition

[1233] Output: Emotional information

[1234] Step 9:

[1235] The server adjusts the voice generation parameters using the adjustment means based on the emotion information obtained by the emotion recognition means. For example, if the server recognizes that the user is tired, it suggests playing relaxing music.

[1236] Input: Emotion information

[1237] Data calculation: Adjustment of voice generation parameters

[1238] Output: Adjusted speech generation parameters

[1239] Step 10:

[1240] The in-vehicle system then adjusts the dialogue content based on the adjusted voice generation parameters to provide a dialogue tailored to the user. For example, the vehicle might say, "Thank you for your hard work. I'll play some relaxing music."

[1241] Input: Adjusted speech generation parameters

[1242] Data calculation: Change of dialogue content

[1243] Output: Personalized dialogue

[1244] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1245] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1246] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1247] [Fourth embodiment]

[1248] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1249] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1250] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1251] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1252] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1253] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1254] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1255] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1256] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1257] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1258] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1259] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1260] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1261] MODE FOR CARRYING OUT THE INVENTION

[1262] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. To implement this system, the following program processing is performed.

[1263] User settings processing

[1264] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[1265] The terminal collects the input configuration information and converts it into a data packet, which is then sent to the server.

[1266] Data analysis and speech generation

[1267] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. This includes selecting an appropriate template from existing voice samples.

[1268] A speech synthesis engine in the server then generates realistic, natural-sounding speech data based on the determined speech generation parameters, for example, using neural network-based techniques.

[1269] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[1270] Providing voice data and implementing dialogue functions

[1271] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[1272] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[1273] Finally, while actually driving, users can enjoy interacting with the vehicle using the voice generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?" In this way, conversational communication is realized.

[1274] Specific examples

[1275] As an example, consider the following scenario: A user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[1276] When the user starts the engine while driving, the vehicle will speak to the user in a friendly voice, asking, "Good morning! Where are you going today?" In this way, the present invention enables interactive communication with the vehicle, enriching the user's driving experience.

[1277] The processing flow will be explained below.

[1278] Step 1:

[1279] The user launches a dedicated app on their smartphone or computer, and then enters the vehicle's personality and voice characteristics on the app's settings screen. Specifically, the user selects options such as "friendly personality," "male voice," or "youthful voice."

[1280] Step 2:

[1281] The device collects the configuration information entered by the user and converts it into a data packet, which contains details such as the vehicle's personality, the gender of the voice, and the tone of the voice.

[1282] Step 3:

[1283] The device generates data packets and sends them to the server over Wi-Fi or mobile data networks.

[1284] Step 4:

[1285] The server receives the setting information sent from the terminal and stores the received data in a database within the server.

[1286] Step 5:

[1287] The server analyzes the received data. Specifically, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[1288] Step 6:

[1289] The server's internal voice synthesis engine generates realistic, natural voice data based on the determined voice generation parameters. This process uses voice synthesis technologies such as neural networks.

[1290] Step 7:

[1291] The server stores the generated voice data and generates new data packets to send to the user's terminal.

[1292] Step 8:

[1293] The server then sends the generated data packets to the user's device, again over Wi-Fi or a mobile data network.

[1294] Step 9:

[1295] The terminal acquires the voice data transmitted from the server.

[1296] Step 10:

[1297] The device analyzes the received voice data and converts it into a format that can be used by the in-vehicle system, specifically into common audio file formats such as WAV and MP3.

[1298] Step 11:

[1299] The device then uploads the converted audio data to the vehicle's system via Bluetooth, USB connection, or other means.

[1300] Step 12:

[1301] The in-car system plays back the audio data while the user is driving. For example, when the engine is started, the vehicle may say, "Good morning. Where are you going today?"

[1302] Through the above steps, users can freely customize the vehicle's personality and voice characteristics, enjoying a personalized interactive driving experience.

[1303] Example 1

[1304] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1305] Conventional in-vehicle systems have made it difficult for users to enjoy a personalized experience through interaction with the vehicle. Furthermore, they lacked a means to freely set the vehicle's personality and voice characteristics and generate natural-sounding voice data based on those settings. As a result, users could only receive uniform, standardized voice guidance, limiting their driving experience. To solve this issue, a system was needed that could flexibly generate voice data based on parameters set by the user and provide a personalized experience through interaction with the vehicle.

[1306] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1307] In this invention, the server includes input means for a user to set the vehicle's personality and voice characteristics, transmission means for converting the information set by the input means into data packets and transmitting the data to the server, analysis means for analyzing the information transmitted by the transmission means and determining voice generation parameters based on the personality and voice characteristics specified by the user, voice generation means for generating natural voice data using a voice synthesis engine based on the voice generation parameters, and transmission means for transmitting the generated voice data to the user's terminal, thereby enabling the user to realize interactive communication based on personalized voice data.

[1308] "Input means" refers to the interface that the user utilizes to configure the vehicle's personality and voice characteristics.

[1309] The "transmission means" is a device, software, or protocol that has the function of converting the information set by the input means into a data packet and transmitting it to the server.

[1310] The "analysis means" is a device or software that has the function of analyzing the data packets received by the server and determining voice generation parameters based on the personality and voice characteristics specified by the user.

[1311] The "voice generation means" is a device or software that has the function of generating natural voice data using a voice synthesis engine based on the voice generation parameters determined by the analysis means.

[1312] The "conversion means" is a device or software that has the function of analyzing the voice data transmitted by the transmission means and converting it into a format that can be used by the in-vehicle system.

[1313] The "uploading means" is a device or software that has the function of uploading the voice data generated by the conversion means to the in-vehicle system.

[1314] The "interaction means" is a device or software that has the function of executing an interaction function between a user and a vehicle using the voice data uploaded to the in-vehicle system by the upload means.

[1315] In the system of the present invention, the user sets the vehicle's personality and voice characteristics, and the server generates voice data based on that information, realizing interactive communication between the user and the vehicle. To implement this system, the following program processing is required.

[1316] User settings input method

[1317] Users install the dedicated application "VehicleVoiceCustomizer" on their smartphone or computer. The application provides a user interface where users can set the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). Once the settings are complete, users press the "Settings" button to proceed to the next step.

[1318] Sending configuration information

[1319] The device converts the user-entered configuration information into data packets containing information about the selected personality and voice characteristics, which are then sent to the server using Internet communication protocols (e.g., HTTP or HTTPS).

[1320] Data analysis and speech generation methods

[1321] The server receives the data packets sent from the device. It analyzes the received information and determines voice generation parameters based on the personality and voice characteristics specified by the user. This analysis is performed using a voice synthesis engine, such as the Google Cloud Text-to-Speech API. The voice synthesis engine uses neural network-based technology (e.g., WaveNet) to generate natural-sounding voice data. The generated voice data is stored on the server and prepared for the next step.

[1322] Voice data transmission method

[1323] The server converts the generated voice data into data packets and transmits the packets to the terminal, again using the Internet communication protocol.

[1324] A means of converting voice data and uploading it to an in-vehicle system

[1325] The device receives the audio data sent from the server. It uses tools such as FFmpeg to analyze the received audio data and convert it into a format that can be used by the in-vehicle system (e.g., WAV or MP3). The converted audio data is then uploaded to the in-vehicle system using a connection method such as Bluetooth, Wi-Fi, or USB.

[1326] Interaction methods

[1327] When a user gets into the vehicle and starts the engine, the in-vehicle system will begin playing the uploaded voice data. For example, the vehicle may ask in a friendly voice, "Good morning! Where are you going today?" This dialogue function allows users to enjoy interactive communication with the vehicle.

[1328] Specific examples

[1329] Consider the following example: A user opens the app "VehicleVoiceCustomizer" and sets a "friendly personality," a "male voice," and a "young voice in his 30s." The app then converts this configuration information into a data packet and sends it to the server. The server analyzes the information and determines appropriate voice generation parameters. It then generates voice data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech API, WaveNet). This generated voice data is then sent from the server to the device. The device receives the voice data, uses "FFmpeg" to convert the audio file into a format usable by the in-vehicle system, and uploads it to the in-vehicle system via Bluetooth. When the user starts the engine, the vehicle speaks in a friendly voice, asking, "Good morning! Where are you going today?"

[1330] An example of a prompt for a generative AI model is:

[1331] "Generate a friendly vehicle speaking to the user in the youthful voice of a man in his 30s."

[1332] This allows the system to provide interactive communication according to the user's wishes, resulting in richer interaction with the vehicle.

[1333] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1334] Step 1: Enter your user settings

[1335] The user launches the dedicated app "Vehicle Voice Customizer" on their smartphone or computer. Through the app's graphical interface, they can set the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). The information set by the user is temporarily stored within the app. The input here is the personality and voice characteristics selected by the user, and the output is a data packet containing the setting information.

[1336] Step 2: Sending configuration information

[1337] The device converts the user-entered setting information into a data packet. Specifically, the selected personality and voice characteristics are packetized as text data. The converted data packet is sent to the server using the HTTP / HTTPS protocol. The input here is the user-selected setting information, and the output is the data packet sent to the server.

[1338] Step 3: Data analysis

[1339] The server analyzes data packets received from the device. During the analysis, it extracts speech generation parameters based on the personality and voice characteristics specified by the user. For example, it uses the Google Cloud Text-to-Speech API to select a speech template. The input of this process is the received data packets, and the output is the speech generation parameters.

[1340] Step 4: Speech generation

[1341] The speech synthesis engine in the server generates speech data using the speech generation parameters extracted earlier. Here, neural network-based technology (e.g., WaveNet) is used to generate natural and realistic speech data. The input of this step is the speech generation parameters, and the output is the generated speech data.

[1342] Step 5: Sending audio data

[1343] The server reconverts the generated voice data into data packets and sends them to the terminal again using the HTTP / HTTPS protocol. The input of this process is the generated voice data, and the output is the data packets sent from the server.

[1344] Step 6: Receiving and converting audio data

[1345] The device receives the audio data sent from the server. Because the received audio data is compressed and encrypted, it is first analyzed and decrypted. Then, a tool such as FFmpeg is used to convert it into a format that can be used by the in-vehicle system (e.g., WAV or MP3). The input of this step is the received data packet, and the output is audio data in a format suitable for the in-vehicle system.

[1346] Step 7: Upload your audio data

[1347] The terminal uploads the converted voice data to the in-vehicle system via Bluetooth, Wi-Fi, USB, etc. The input of this step is the converted voice data, and the output is the voice data uploaded to the in-vehicle system.

[1348] Step 8: Implementing Interactivity

[1349] When the user starts the vehicle, the in-vehicle system starts playing the uploaded voice data. For example, the vehicle may start a dialogue with the user by saying, "Good morning! Where are you going today?" The input of this step is the voice data uploaded to the in-vehicle system, and the output is the interactive dialogue between the user and the vehicle.

[1350] (Application example 1)

[1351] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1352] Conventional autonomous vehicles have a problem in that the driving experience is mechanical due to the limited information provided by the vehicle and one-way communication with the user. Furthermore, there is a lack of a way to realize a dialogue system with the personality and voice characteristics desired by the user. Furthermore, there is a lack of a means to provide navigation information, traffic information, weather information, maintenance information, etc. in a unified and interactive manner.

[1353] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1354] In this invention, the server includes an input means for a user to set the personality and voice characteristics of a vehicle, an analysis means for analyzing the information set by the input means, a voice generation means for generating voice data based on the information analyzed by the analysis means, a transmission means for transmitting the generated voice data to the user's terminal, a navigation means for providing navigation information for the vehicle, a traffic information providing means for providing traffic information based on information from the navigation means, and an information providing means for providing weather and maintenance information. This allows the user to interact with an autonomous vehicle with a personality and voice characteristics tailored to their preferences, enriching the driving experience. Furthermore, the provision of comprehensive, interactive information improves the user's comfort and safety.

[1355] Understood. Now, I will create definitions for the important words included in the rewritten claims.

[1356] "Mobile object" refers to a vehicle or other conveyance, including those with automatic driving capabilities.

[1357] "Personality" refers to the personified characteristics of the mobile object set by the user, and includes characteristics such as friendly, cool, sporty, etc.

[1358] "Voice characteristics" refers to the characteristics of a voice set by a user, and includes attributes such as gender, age, tone of voice, and pitch.

[1359] "Input means" refers to an interface that allows a user to set the vehicle's personality and voice characteristics, such as a smartphone app or an in-car console.

[1360] The "analysis means" refers to a computer system for analyzing the information set by the input means.

[1361] "Speech generation means" refers to a system for generating speech data based on the information analyzed by the analysis means.

[1362] "Transmission means" refers to a function for transmitting the generated voice data to the user's terminal.

[1363] "Navigation means" refers to a system for guiding a route to a destination of a mobile object.

[1364] "Traffic information providing means" refers to a function that provides the user with traffic conditions based on information from the navigation means.

[1365] "Information provision means" refers to a function for providing weather and maintenance information to users.

[1366] "Conversion means" refers to a computer program for converting transmitted voice data into a format usable by the mobile system.

[1367] "Uploading means" refers to a function for uploading converted audio data to a mobile system.

[1368] "Interaction means" refers to a system that allows a user to interact with a mobile object using uploaded voice data.

[1369] "Entertainment means" refers to the capability to provide audio and visual entertainment over a mobile system.

[1370] In the system of the present invention, the user sets the personality and voice characteristics of the mobile object, and AI generates the "voice" of the mobile object based on that information, realizing interactive communication between the user and the mobile object. To implement this system, the following program processing is performed.

[1371] User settings processing

[1372] First, the user sets the mobile device's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the mobile device's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice"). The device collects the entered setting information and converts it into a data packet. Then, the data packet is sent to the server.

[1373] Data analysis and speech generation

[1374] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. This process involves selecting an appropriate template from existing voice samples. Next, the server's speech synthesis engine generates natural-sounding voice data based on the determined voice generation parameters. This voice data is generated using, for example, neural network-based technology (e.g., Google Cloud Text-to-Speech API or Amazon Polly). The generated voice data is stored on the server, and a data packet is generated to be sent to the device. This packet is then sent back to the device.

[1375] Providing voice data and implementing dialogue functions

[1376] The receiving device receives the voice data sent from the server, analyzes it, and converts it into a format that can be used by the mobile system. For example, it may be converted into a common audio file format (WAV, MP3, etc.). The device then uploads the converted voice data to the mobile system, where it can be used in the vehicle's navigation and audio systems. Finally, during actual autonomous driving, the user can enjoy interacting with the vehicle using the voice generated by the system. For example, the vehicle may ask, "Hello! Where are you going today?" Furthermore, the navigation and traffic information means may suggest appropriate routes and report traffic conditions, while the entertainment means may play music or provide simple conversations.

[1377] Usage example

[1378] As an example of usage, consider the following scenario: A user opens the app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to the server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device.

[1379] The terminal converts the voice data and uploads it to the vehicle system. During autonomous driving, when the user starts the engine, the vehicle will speak to them in a friendly voice, asking, "Hello! Where are you going today?"

[1380] For example, the Google Maps API is used for navigation, providing route guidance to destinations, and real-time traffic information is provided as a means of providing traffic information. For entertainment, music streaming services such as Spotify are integrated, allowing users to play music according to their preferences.

[1381] Prompt Sentence Examples

[1382] assistant = AIDriverAssistant("Friendly", "Young male voice")

[1383] print(assistant.generate_response("greeting"))

[1384] In this way, the present invention allows interactive communication with the vehicle, enriching the user's driving experience.

[1385] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1386] Step 1:

[1387] The user uses a dedicated app on a smartphone or computer to input data to set the mobile device's personality and voice characteristics. Specifically, the user selects personality traits such as "friendly," "cool," or "sporty" and voice characteristics such as "male voice," "female voice," or "youthful voice" on the app. This is input data, and the collected information is sent to the device as a data packet.

[1388] Step 2:

[1389] The terminal sends a data packet containing user-defined information to the server. This packet contains information about personality and voice characteristics. This information is sent via a data transfer protocol and received by the server.

[1390] Step 3:

[1391] The server receives the data packet sent from the terminal and analyzes its contents. The analysis means analyzes the contents of the data packet (personality, voice characteristics, etc.) and determines appropriate voice generation parameters. For example, based on information such as "friendly" or "youthful male voice," it generates the parameters required for the voice generation engine.

[1392] Step 4:

[1393] The server generates natural-sounding voice data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech API or Amazon Polly) based on the voice generation parameters determined by the analysis means. This generated voice data is converted into an audio file format (WAV, MP3, etc.) through a series of data calculations.

[1394] Step 5:

[1395] The server converts the generated voice data into data packets and transmits the data packets to the terminal, where the voice data may be compressed. The terminal receives the data packets transmitted from the server.

[1396] Step 6:

[1397] The terminal analyzes the voice data sent from the server and converts it into a format that can be used by the mobile system (for example, WAV or MP3 format). The voice data conversion is performed using a format conversion algorithm.

[1398] Step 7:

[1399] The device then uploads the converted voice data to the in-vehicle system, where it can be used in the navigation and audio systems. Specifically, the device stores the voice data in the in-vehicle system's memory or storage.

[1400] Step 8:

[1401] The user enjoys interacting with the vehicle during autonomous driving. The vehicle uses the generated voice to provide interactive navigation, traffic information, weather and maintenance information, and entertainment functions (e.g., music playback and chat). The interactive navigation uses a voice recognition system to understand user input and respond appropriately.

[1402] Example: When the user starts the engine, the vehicle will speak to them in a friendly voice, asking, "Hello! Where are you going today?" It is possible to use the Google Maps API as a navigation tool and provide real-time traffic information.

[1403] In this way, the present invention realizes interactive communication with the user and enriches the experience of using a mobile device.

[1404] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1405] MODE FOR CARRYING OUT THE INVENTION

[1406] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized dialogue.

[1407] User settings processing

[1408] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[1409] The device collects the setting information entered by the user and converts it into a data packet, which is then sent to the server.

[1410] Data analysis and speech generation

[1411] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[1412] Next, the server's speech synthesis engine generates realistic, natural-sounding voice data based on the determined voice generation parameters. This process uses speech synthesis technologies such as neural networks.

[1413] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[1414] Providing voice data and implementing dialogue functions

[1415] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[1416] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[1417] While driving, the user can interact with the vehicle using voices generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?"

[1418] Emotion engine integration

[1419] Furthermore, an emotion engine is integrated into the system, which analyzes the user's voice and facial expressions to recognize their emotions, for example, determining whether they are feeling stressed or having fun.

[1420] The server adjusts the voice generation parameters and dialogue content based on the emotional information recognized by the emotion engine, enabling dialogue appropriate to the user's emotional state. For example, if the vehicle recognizes that the user is tired, it might suggest, "Why don't you take a short break today?"

[1421] Specific examples

[1422] For example, a user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[1423] The emotion engine analyzes the user's emotions and recognizes when the user is feeling stressed while driving. Based on this information, the system can say, "I'll play some relaxing music," and play music accordingly, making the user's driving experience richer and more personalized.

[1424] The processing flow will be explained below.

[1425] Step 1:

[1426] The user launches a dedicated app on their smartphone or computer, and then enters the vehicle's personality and voice characteristics on the app's settings screen. Specifically, the user selects options such as "friendly personality," "male voice," or "young voice in his 30s."

[1427] Step 2:

[1428] The device collects the configuration information entered by the user and converts it into a data packet, which contains details such as the vehicle's personality, the gender of the voice, and the tone of the voice.

[1429] Step 3:

[1430] The device generates data packets and sends them to the server over Wi-Fi or mobile data networks.

[1431] Step 4:

[1432] The server receives the setting information sent from the terminal and stores the received data in a database within the server.

[1433] Step 5:

[1434] The server analyzes the received data. Specifically, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[1435] Step 6:

[1436] The server's internal voice synthesis engine generates realistic, natural voice data based on the determined voice generation parameters. This process uses voice synthesis technologies such as neural networks.

[1437] Step 7:

[1438] The server stores the generated voice data and generates new data packets to send to the user's terminal.

[1439] Step 8:

[1440] The server then sends the generated data packets to the user's device, again over Wi-Fi or a mobile data network.

[1441] Step 9:

[1442] The terminal acquires the voice data transmitted from the server.

[1443] Step 10:

[1444] The device analyzes the received voice data and converts it into a format that can be used by the in-vehicle system, specifically into common audio file formats such as WAV and MP3.

[1445] Step 11:

[1446] The device then uploads the converted audio data to the vehicle's system via Bluetooth, USB connection, or other means.

[1447] Step 12:

[1448] The in-car system plays back the audio data while the user is driving. For example, when the engine is started, the vehicle may say, "Good morning. Where are you going today?"

[1449] Emotion Engine Processing Flow

[1450] Step 13:

[1451] The emotion engine installed in the device captures the user's voice and facial expressions using the device's microphone and camera.

[1452] Step 14:

[1453] The device's emotion engine analyzes the voice and facial expression data it acquires to recognize the user's emotions. For example, it can determine whether the user is feeling "stressed" based on the tone of their voice and facial expression.

[1454] Step 15:

[1455] The device sends the recognized emotion information to the server, which includes information that the user is currently feeling "stressed."

[1456] Step 16:

[1457] The server receives and analyzes the emotional information. Based on this information, it adjusts the generated voice data and dialogue content. For example, it generates new voice data such as "Shall I play some calming music?" to help the user relax.

[1458] Step 17:

[1459] The server sends the newly generated voice data to the terminal.

[1460] Step 18:

[1461] The device then uploads the new voice data it receives back to the in-vehicle system.

[1462] Step 19:

[1463] New audio data will be played while the user is driving: the vehicle will say, "I'm going to play some music to help you relax," and then play appropriate music.

[1464] Through the above steps, the user can enjoy an interactive driving experience that is appropriate to their emotions.

[1465] Example 2

[1466] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1467] In recent years, interactive communication systems have become increasingly important for improving in-vehicle user experiences. However, existing systems are unable to provide personalized dialogue based on the user's individual preferences and emotional state. Furthermore, it is difficult for users to easily configure the vehicle's personality and voice characteristics and quickly provide personalized voice dialogue based on those settings. This has led to problems that reduce user satisfaction and convenience.

[1468] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1469] In this invention, the server includes an input means for a user to set the vehicle's personality and voice characteristics, a transmission means for converting the information set by the input means into a data packet and transmitting it to the server, an analysis means for analyzing the data packet transmitted by the transmission means and determining voice generation parameters based on the user's settings, a voice generation means for generating voice data using the voice generation parameters determined by the analysis means, and a transmission means for transmitting the generated voice data to the user's terminal. This allows the user to easily set the vehicle's personality and voice characteristics and quickly provide personalized voice dialogue based on them. Furthermore, it is possible to provide appropriate dialogue according to the user's emotional state, thereby improving the user experience.

[1470] A "user" is a person who uses this system to set the vehicle's personality and voice characteristics and engage in interactive communication.

[1471] "Vehicle personality" refers to the interaction attitude and atmosphere of the in-vehicle system set by the user, and includes characteristics such as "friendly," "cool," and "sporty."

[1472] "Voice characteristics" refers to the attributes of the voice emitted by the vehicle, and includes parameters such as "male voice," "female voice," "youthful voice," and "deep voice."

[1473] "Input means" refers to the device or interface through which a user configures the vehicle's personality and voice characteristics.

[1474] "Data packet" refers to a series of data containing information set by the user, and is used to send to the server.

[1475] "Transmission means" refers to a communication means for transmitting data packets from a terminal to a server, such as the Internet or a dedicated network.

[1476] The "analysis means" refers to a process or technology for analyzing the transmitted data packets to extract user setting information and determining voice generation parameters based on the extracted information.

[1477] "Voice generation parameters" refer to specific settings and templates for generating voice data based on the vehicle's personality and voice characteristics set by the user.

[1478] "Speech generation means" refers to a technology or engine for generating realistic and natural voice data using the voice generation parameters determined by the analysis means.

[1479] "Conversion means" refers to the process or equipment that converts the received audio data into a format that can be used by the in-vehicle system (e.g., WAV, MP3, etc.).

[1480] "Uploading means" refers to a means for uploading converted audio data to an in-vehicle system.

[1481] "Dialogue means" refers to the technology and functions for dialogue with the user using voice data generated by the in-vehicle system.

[1482] "Emotion analysis means" refers to technology or engines that analyze the user's voice and facial expressions to recognize emotions.

[1483] The "adjustment means" refers to a technique or process for adjusting the voice generation parameters and the dialogue content based on the emotional information recognized by the emotion analysis means.

[1484] MODE FOR CARRYING OUT THE INVENTION

[1485] The system of the present invention allows the user to set the vehicle's personality and voice characteristics, and then AI generates the vehicle's "voice" based on that information, realizing interactive communication between the user and the vehicle. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized dialogue.

[1486] First, the user sets the vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app provides a graphical interface that allows the user to easily select the vehicle's personality (e.g., "friendly," "cool," "sporty") and voice characteristics (e.g., "male voice," "female voice," "youthful voice," "deep voice").

[1487] The device collects the setting information entered by the user and converts it into a data packet, which is then sent to the server.

[1488] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. For example, it selects a "friendly male voice" template from the voice sample library.

[1489] Next, the server's speech synthesis engine generates realistic, natural-sounding voice data based on the determined voice generation parameters. This process uses speech synthesis technologies such as neural networks.

[1490] The generated voice data is stored in the server, and a data packet is generated to be sent to the terminal. This packet is then sent back to the terminal.

[1491] The receiving device receives the audio data sent from the server, analyzes it, and converts it into a format that can be used by the in-vehicle system, such as a common audio file format (WAV, MP3, etc.).

[1492] The device then uploads the converted audio data to the vehicle's system, where it can be used by the vehicle's navigation and audio systems.

[1493] While driving, the user can interact with the vehicle using voices generated by the system. For example, when starting the engine, the vehicle will say, "Good morning. What music would you like to listen to today?"

[1494] Furthermore, an emotion engine is integrated into the system, which analyzes the user's voice and facial expressions to recognize their emotions, for example, determining whether they are feeling stressed or having fun.

[1495] The server adjusts the voice generation parameters and dialogue content based on the emotional information recognized by the emotion engine, enabling dialogue appropriate to the user's emotional state. For example, if the vehicle recognizes that the user is tired, it might suggest, "Why don't you take a short break today?"

[1496] Specific examples

[1497] For example, a user opens an app and selects a "friendly personality," a "male voice," and a "youthful voice in his 30s." The app then sends this configuration information to a server. The server analyzes the information and determines appropriate voice generation parameters. The speech synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[1498] The emotion engine analyzes the user's emotions and recognizes when the user is feeling stressed while driving. Based on this information, the system can say, "I'll play some relaxing music," and play music accordingly, making the user's driving experience richer and more personalized.

[1499] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1500] Step 1:

[1501] The user launches the dedicated app.

[1502] Specific action: The user taps or clicks to launch a dedicated application installed on their smartphone or computer.

[1503] Input: User Action

[1504] Output: App home screen display

[1505] Step 2:

[1506] The user selects the vehicle's personality and voice characteristics.

[1507] How it works: Select a personality type such as "friendly," "cool," or "sporty" and voice characteristics such as "male voice," "female voice," or "youthful voice" from the app's graphical interface.

[1508] Input: User-selected personality and voice characteristics

[1509] Output: Selected personality and vocal trait data

[1510] Step 3:

[1511] The terminal converts the setting information into a data packet and transmits it to the server.

[1512] What it does: The application packages the user's selected settings into a data packet and sends it over the Internet to a server.

[1513] Input: User-selected setting information

[1514] Output: Data packets sent to the server

[1515] Step 4:

[1516] The server receives the data packets and performs the analysis.

[1517] Specific operation: The server analyzes the data packets received from the terminal and extracts the user's setting information.

[1518] Input: Data packets from the terminal

[1519] Output: Parsed configuration information

[1520] Step 5:

[1521] The server determines the voice generation parameters.

[1522] Specific operation: Based on the analyzed configuration information, the server selects an appropriate template from the voice sample library and determines voice generation parameters.

[1523] Input: Parsed configuration information

[1524] Output: Speech generation parameters

[1525] Step 6:

[1526] The server generates voice data based on the voice generation parameters.

[1527] How it works: The speech synthesis engine uses speech generation parameters to generate realistic, natural-sounding speech data, using technologies such as neural networks.

[1528] Input: Speech generation parameters

[1529] Output: Generated audio data

[1530] Step 7:

[1531] The server transmits the generated voice data to the terminal.

[1532] Specific operation: The generated voice data is converted into data packets and sent to the terminal.

[1533] Input: Generated audio data

[1534] Output: Data packets to the terminal

[1535] Step 8:

[1536] The device receives the voice data, analyzes it, and converts it into a format that can be used by the in-vehicle system.

[1537] Specific operation: The device analyzes the audio data it receives and converts it into a format that can be used by the in-car system, such as WAV or MP3.

[1538] Input: Audio data sent to the device

[1539] Output: Converted audio data

[1540] Step 9:

[1541] The terminal uploads the audio data to the in-vehicle system.

[1542] What it does: Uploads converted audio data to the vehicle's navigation and audio systems.

[1543] Input: Converted audio data

[1544] Output: Audio data uploaded to the in-car system

[1545] Step 10:

[1546] The user interacts using the generated voice.

[1547] Specific behavior: When the engine is started, the vehicle will speak to you in a preset personality and voice, saying, "Good morning. What music would you like to listen to today?"

[1548] Input: Audio data uploaded to the in-vehicle system

[1549] Output: Interaction with the vehicle

[1550] Step 11:

[1551] The emotion engine analyzes the user's emotions.

[1552] Specific operation: The emotion engine analyzes voice and facial expressions to recognize emotions such as stress and joy.

[1553] Input: User's voice and facial expression data

[1554] Output: Recognized emotion information

[1555] Step 12:

[1556] The server adjusts the voice generation parameters and dialogue content based on the emotion information.

[1557] Specific operation: Based on the emotional information recognized by the emotion engine, the server adjusts the voice generation parameters and dialogue content, for example, to say, "Would you like to take a short break today?"

[1558] Input: Recognized emotion information

[1559] Output: Adjusted voice generation parameters and dialogue content

[1560] (Application example 2)

[1561] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1562] Autonomous vehicles are expected to provide a personalized travel experience through dialogue between passengers and the vehicle. However, current systems are unable to recognize passenger emotions in real time and flexibly adapt the dialogue content with the vehicle based on those emotions. This makes it difficult to achieve interactive communication that truly satisfies passengers.

[1563] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes input means for the user to set the vehicle's personality and voice characteristics, analysis means for analyzing the information set by the input means, voice generation means for generating voice data based on the information analyzed by the analysis means, transmission means for transmitting the generated voice data to the user's terminal, emotion recognition means for analyzing the user's emotions, and adjustment means for adjusting voice generation parameters in accordance with the emotion information obtained by the emotion recognition means. This enables personalized dialogue in real time according to the user's emotions.

[1564] "User" means any person or entity that uses the vehicle or system.

[1565] "Vehicle personality" is a concept that refers to the emotional and behavioral characteristics that a vehicle is supposed to have, which are set by the user.

[1566] "Voice characteristics" are attributes that refer to voice characteristics such as the quality, tone, and timbre of the voice emitted from the vehicle.

[1567] The term "input means" refers to a device or interface that allows a user to input setting information.

[1568] "Analysis means" refers to a system that includes software and hardware for analyzing input information and understanding its meaning and content.

[1569] "Speech generation means" refers to a function that includes a speech synthesis engine and related technologies for generating speech data based on analyzed information.

[1570] "Transmission means" refers to the infrastructure, which refers to the communication technologies and protocols used to transmit the generated audio data to other devices or systems.

[1571] "Emotion recognition means" refers to a system that uses technology or algorithms to analyze and identify emotions from a user's facial expressions, voice, or other input.

[1572] The "adjustment means" is a function that refers to a method or technique for dynamically changing voice generation parameters based on information obtained by the emotion recognition means.

[1573] "Conversion means" is a function that refers to technology or devices for converting transmitted audio data into a format that can be used by the in-vehicle system.

[1574] "Uploading means" refers to the infrastructure that refers to the methods and protocols used to transfer and store the converted audio data in the in-vehicle system.

[1575] "Dialogue means" refers to the technology and functions that enable the in-vehicle system to communicate with the user via voice.

[1576] The "display means" is an apparatus that refers to a device or technology for displaying the dialogue content based on the voice generation parameters adjusted by the adjustment means.

[1577] MODE FOR CARRYING OUT THE INVENTION

[1578] The following describes an embodiment of the present invention: This system has a structure that allows real-time dialogue within a vehicle based on the vehicle's personality and voice characteristics designed by the user.

[1579] System Overview

[1580] The system consists of the following main components:

[1581] 1. Input Method

[1582] Users can configure their vehicle's personality and voice characteristics using a dedicated app on their smartphone or computer. The app's interface is designed to make configuration easy for users. The configured information is collected by the app and converted into data packets.

[1583] 2. Analysis method

[1584] The server receives and analyzes the configuration information sent from the device. Based on the received data, it determines voice generation parameters based on the personality and voice characteristics specified by the user. During this process, it uses a voice sample library to select a voice template that matches the settings.

[1585] 3. Voice Generation Method

[1586] A speech synthesis engine (such as Google Cloud TTS or Amazon Polly) in the server generates realistic and natural voice data based on the analyzed data. The generated voice data is stored in the server and a data packet is generated for processing.

[1587] 4. Transmission Method

[1588] The audio data is sent from the server to the device using a common data transfer protocol (e.g., HTTP or HTTPS).

[1589] 5. Emotion recognition means

[1590] An emotion engine (for example, Microsoft Azure's Emotion API) analyzes the user's voice and facial expressions to recognize their emotions. This information is collected in real time and sent to a server.

[1591] 6. Adjustment means

[1592] The server uses the information obtained from the emotion recognition means to adjust the voice and dialogue played by the in-vehicle system, using predefined prompts.

[1593] Specific examples

[1594] The user opens the app and selects settings such as "friendly personality," "male voice," and "young voice in his 30s." The app then sends this setting information to the server. The server analyzes the information and determines the appropriate voice generation parameters. The voice synthesis engine then generates speech based on these parameters and sends the voice data to the user's device. The device then converts the voice data and uploads it to the in-car system.

[1595] Prompt Sentence Examples

[1596] When passengers board the vehicle:

[1597] "Hello, how's it going today?"

[1598] If the user says they are tired:

[1599] "Thank you for your hard work. I'll put on some relaxing music."

[1600] In this way, the system can combine emotion recognition and speech generation techniques to provide personalized interactions according to the user's emotions.

[1601] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1602] Step 1:

[1603] The user opens a dedicated app on their smartphone or computer and sets the vehicle's personality and voice characteristics. The input information can include, for example, "friendly personality," "male voice," or "youthful voice in his 30s." The app collects this information as a data packet and sends it to a server.

[1604] Input: User setting information

[1605] Data processing: Converting user setting information into data packets

[1606] Output: Data packets

[1607] Step 2:

[1608] The server receives the data packet sent from the device and starts the analysis process. Based on the configuration information, it references the voice sample library to determine voice generation parameters, and selects a template such as a "friendly male voice."

[1609] Input: Data packet

[1610] Data calculation: Determining voice generation parameters

[1611] Output: Speech generation parameters

[1612] Step 3:

[1613] The server's speech synthesis engine generates voice data based on the determined voice generation parameters. During this process, realistic and natural voices are generated using, for example, Google Cloud TTS or Amazon Polly. The generated voice data is stored on the server.

[1614] Input: Speech generation parameters

[1615] Data Computation: Generation of voice data using a TTS engine

[1616] Output: Audio data

[1617] Step 4:

[1618] The server converts the generated voice data into data packets and sends them to the terminal, using common data transfer protocols (HTTP or HTTPS).

[1619] Input: Audio data

[1620] Data processing: Converting voice data into data packets

[1621] Output: Data packets

[1622] Step 5:

[1623] The device receives the data packets sent from the server, analyzes the audio data, and converts it into a common audio file format (e.g., WAV or MP3).

[1624] Input: Data packet

[1625] Data processing: Convert data packets into audio file format

[1626] Output: Audio file

[1627] Step 6:

[1628] The device then uploads the converted audio file to the in-car system via communication methods such as Bluetooth or Wi-Fi.

[1629] Input: Audio file

[1630] Data Calculation: Uploading Audio Files

[1631] Output: Uploaded audio file

[1632] Step 7:

[1633] The user plays generated audio through the in-car system, initiating a dialogue with the vehicle, for example, the vehicle saying, "Hello, how's your day?"

[1634] Input: Uploaded audio file

[1635] Data calculation: Audio playback by in-car systems

[1636] Output: Voice dialogue

[1637] Step 8:

[1638] The emotion engine analyzes the user's voice and facial expressions to recognize emotions, for example, detecting the user's stress level using Microsoft Azure's Emotion API.

[1639] Input: User's voice and facial expression data

[1640] Data Computing: Emotion Analysis and Recognition

[1641] Output: Emotional information

[1642] Step 9:

[1643] The server adjusts the voice generation parameters using the adjustment means based on the emotion information obtained by the emotion recognition means. For example, if the server recognizes that the user is tired, it suggests playing relaxing music.

[1644] Input: Emotion information

[1645] Data calculation: Adjustment of voice generation parameters

[1646] Output: Adjusted speech generation parameters

[1647] Step 10:

[1648] The in-vehicle system then adjusts the dialogue content based on the adjusted voice generation parameters to provide a dialogue tailored to the user. For example, the vehicle might say, "Thank you for your hard work. I'll play some relaxing music."

[1649] Input: Adjusted speech generation parameters

[1650] Data calculation: Change of dialogue content

[1651] Output: Personalized dialogue

[1652] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1653] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1654] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1655] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1656] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1657] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1658] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1659] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1660] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1661] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1662] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1663] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1664] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1665] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1666] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1667] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1668] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1669] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1670] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1671] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1672] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1673] The following is further disclosed regarding the above embodiment.

[1674] (Claim 1)

[1675] input means for a user to set the vehicle's personality and voice characteristics;

[1676] analysis means for analyzing the information set by the input means;

[1677] a voice generating means for generating voice data based on the information analyzed by the analyzing means;

[1678] a transmitting means for transmitting the generated voice data to a user terminal;

[1679] A system including:

[1680] (Claim 2)

[1681] a conversion means for analyzing the voice data transmitted by the transmission means and converting the voice data into a format that can be used by the in-vehicle system;

[1682] uploading means for uploading the converted voice data to an in-vehicle system;

[1683] an interaction means for executing an interaction function using the voice data uploaded by the upload means;

[1684] The system of claim 1 further comprising:

[1685] (Claim 3)

[1686] 2. The system of claim 1, wherein the analyzing means includes classifying means for classifying the vehicle personality and voice characteristics set by the user.

[1687] "Example 1"

[1688] (Claim 1)

[1689] input means for a user to set the vehicle's personality and voice characteristics;

[1690] a transmitting means for converting the information set by the input means into a data packet and transmitting the data packet to a server;

[1691] an analysis means for analyzing the information transmitted by the transmission means and determining voice generation parameters based on the personality and voice characteristics designated by the user;

[1692] a voice generating means for generating natural voice data using a voice synthesis engine based on the voice generation parameters;

[1693] a transmitting means for transmitting the generated voice data to a user terminal;

[1694] A system including:

[1695] (Claim 2)

[1696] a conversion means for analyzing the voice data transmitted by the transmission means and converting the voice data into a format that can be used by the in-vehicle system;

[1697] uploading means for uploading the converted voice data to an in-vehicle system;

[1698] an interaction means for executing an interaction function using the voice data uploaded by the upload means;

[1699] The system of claim 1 further comprising:

[1700] (Claim 3)

[1701] 2. The system of claim 1, wherein the analyzing means includes classifying means for classifying the vehicle personality and voice characteristics set by the user.

[1702] "Application Example 1"

[1703] Okay, so let's rewrite the original patent claim by adding the characteristic features of the application example.

[1704] (Claim 1)

[1705] input means for a user to set the personality and voice characteristics of the mobile unit;

[1706] analysis means for analyzing the information set by the input means;

[1707] a voice generating means for generating voice data based on the information analyzed by the analyzing means;

[1708] a transmitting means for transmitting the generated voice data to a user terminal;

[1709] a navigation means for providing navigation information for the moving body;

[1710] a traffic information providing means for providing traffic information based on information from the navigation means;

[1711] an information providing means for providing weather and maintenance information;

[1712] A system including:

[1713] (Claim 2)

[1714] a conversion means for analyzing the voice data transmitted by the transmission means and converting the voice data into a format that can be used in a mobile system;

[1715] uploading means for uploading the converted voice data to a mobile system;

[1716] an interaction means for executing an interaction function using the voice data uploaded by the upload means;

[1717] entertainment means for providing audio and visual entertainment over said mobile system;

[1718] The system of claim 1 further comprising:

[1719] (Claim 3)

[1720] The analyzing means includes a classifying means for classifying the personality and voice characteristics of the mobile object set by the user.

[1721] 10. The system of claim 1.

[1722] "Example 2: Combining Emotion Engines"

[1723] (Claim 1)

[1724] input means for a user to set the vehicle's personality and voice characteristics;

[1725] a transmitting means for converting the information set by the input means into a data packet and transmitting the data packet to a server;

[1726] analysis means for analyzing the data packets transmitted by said transmission means and determining voice generation parameters based on user settings;

[1727] a voice generating means for generating voice data using the voice generation parameters determined by the analyzing means;

[1728] a transmitting means for transmitting the generated voice data to a user terminal;

[1729] A system including:

[1730] (Claim 2)

[1731] a conversion means for receiving and analyzing the voice data transmitted by the transmission means and converting the voice data into a format that can be used by the in-vehicle system;

[1732] uploading means for uploading the converted voice data to an in-vehicle system;

[1733] an interaction means for executing an interaction function using the voice data uploaded by the upload means;

[1734] emotion analysis means for analyzing the emotions of a user;

[1735] an adjusting means for adjusting voice generation parameters and dialogue contents based on the emotion information recognized by the emotion analyzing means;

[1736] The system of claim 1 further comprising:

[1737] (Claim 3)

[1738] 2. The system of claim 1, wherein the analyzing means includes classifying means for classifying the vehicle personality and voice characteristics set by the user.

[1739] "Application example 2 when combining emotion engines"

[1740] (Claim 1)

[1741] input means for a user to set the vehicle's personality and voice characteristics;

[1742] analysis means for analyzing the information set by the input means;

[1743] a voice generating means for generating voice data based on the information analyzed by the analyzing means;

[1744] a transmitting means for transmitting the generated voice data to a user terminal;

[1745] emotion recognition means for analyzing the emotion of a user;

[1746] an adjustment means for adjusting a voice generation parameter in accordance with the emotion information obtained by the emotion recognition means;

[1747] A system including:

[1748] (Claim 2)

[1749] a conversion means for analyzing the voice data transmitted by the transmission means and converting the voice data into a format that can be used by the in-vehicle system;

[1750] uploading means for uploading the converted voice data to an in-vehicle system;

[1751] an interaction means for executing an interaction function using the voice data uploaded by the upload means;

[1752] a display means for displaying a dialogue content based on the speech generation parameters adjusted by the adjustment means;

[1753] The system of claim 1 further comprising:

[1754] (Claim 3)

[1755] 2. The system of claim 1, wherein the analyzing means includes classifying means for classifying the vehicle personality and voice characteristics set by the user. [Explanation of symbols]

[1756] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. input means for a user to set the vehicle's personality and voice characteristics; analysis means for analyzing the information set by the input means; a voice generating means for generating voice data based on the information analyzed by the analyzing means; a transmitting means for transmitting the generated voice data to a user terminal; A system including:

2. a conversion means for analyzing the voice data transmitted by the transmission means and converting the voice data into a format that can be used by the in-vehicle system; uploading means for uploading the converted voice data to an in-vehicle system; an interaction means for executing an interaction function using the voice data uploaded by the upload means; The system of claim 1 further comprising:

3. 2. The system of claim 1, wherein said analyzing means includes classifying means for classifying vehicle personality and voice characteristics set by a user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A