system
The system simplifies voice conversion by using a recording, transmission, and conversion process with generative AI, enabling easy and efficient voice transformation for various applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
The conventional process of converting a user's voice into another voice is complicated and difficult to execute easily.
A system comprising a recording unit, a transmission unit, and a conversion unit that utilizes a generative AI to analyze and convert a user's voice into another voice, with a providing unit to deliver the converted voice to the user, utilizing devices like smartphones and cloud servers for processing.
Enables easy and efficient conversion of a user's voice into another voice, allowing for applications in entertainment, education, and security, with features like noise cancellation, real-time adjustments, and optimal delivery methods.
Smart Images

Figure 2026073032000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that the process of converting a user's voice into another voice is complicated and difficult to execute easily.
[0005] The system according to the embodiment aims to easily convert a user's voice into another voice.
Means for Solving the Problems
[0006] The system according to the embodiment includes a recording unit, a transmission unit, a conversion unit, and a providing unit. The recording unit records a user's voice. The transmission unit transmits the data recorded by the recording unit to a generative AI. The conversion unit analyzes the data transmitted by the transmission unit and converts it into another voice. The providing unit provides the voice converted by the conversion unit to the user. [Effects of the Invention]
[0007] The system according to this embodiment can easily convert a user's voice into another voice. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The voice conversion system according to an embodiment of the present invention is a system that records a user's voice, converts it to another voice using a generation AI, and provides it to the user. The voice conversion system works by having the user record their own voice and send the recording data to the generation AI. The generation AI analyzes the data and converts it to another voice. The converted voice is then provided to the user. This system allows for easy voice conversion and is expected to have applications in fields such as entertainment, education, and security. For example, if a user records their own voice and sends the recording data to the generation AI, the generation AI converts that voice to a different voice and provides it to the user. This mechanism allows for easy voice conversion and can be used for various purposes. As a result, the voice conversion system can automatically convert and provide the user's voice.
[0029] The voice conversion system according to this embodiment comprises a recording unit, a transmission unit, a conversion unit, and a provision unit. The recording unit records the user's voice. The recording unit can record the user's voice, for example, using a smartphone application. The recording unit may also have a noise-canceling function to record the user's voice in high quality. The recording unit starts recording, for example, when the user launches a smartphone application and presses the record button. When recording is finished, the recording unit saves the recorded data. The transmission unit sends the data recorded by the recording unit to a generation AI. The transmission unit can send the data to a generation AI on the cloud, for example. The transmission unit can also dynamically adjust the data compression ratio to optimize the data transmission speed. The transmission unit can upload the recorded data to a cloud server and send it to the generation AI, for example. The conversion unit analyzes the data transmitted by the transmission unit and converts it into a different voice. The conversion unit can convert the voice, for example, by changing specific parameters of the audio data. The conversion unit uses the generation AI to convert the user's voice into a different voice. The conversion unit, for example, modifies the frequency and amplitude of the audio data to adjust the tone and pitch of the voice. The providing unit provides the voice converted by the conversion unit to the user. The providing unit can, for example, provide the converted voice to the user as an audio file. The providing unit may also have an audio filtering function to play the voice with optimal sound quality for the user's device. The providing unit can, for example, allow the user to download the converted voice to their smartphone. Thus, the voice conversion system according to this embodiment can record the user's voice, convert it to another voice using a generation AI, and provide it.
[0030] The recording unit records the user's voice. For example, the recording unit can record the user's voice using a smartphone app. The recording unit may also have noise cancellation capabilities to ensure high-quality recording of the user's voice. Specifically, the recording unit uses active noise cancellation technology to reduce ambient noise. This technology allows for clear audio recording by having the microphone pick up ambient sounds and generating sound waves with the opposite phase to cancel out the noise. The recording unit starts recording, for example, when the user launches a smartphone app and presses the record button. Once recording is complete, the recording unit saves the recording data. The recording data is stored in local or cloud storage and used for subsequent processing. The recording unit also has a function to monitor the audio quality in real time during recording and automatically adjust recording settings as needed. For example, if the audio volume is too low during recording, it can automatically increase the sensitivity to record at an appropriate volume. This allows the recording unit to record the user's voice in high quality and provide data optimized for subsequent processing.
[0031] The transmitting unit sends data recorded by the recording unit to the generating AI. The transmitting unit can, for example, send data to a generating AI on the cloud. The transmitting unit can also dynamically adjust the data compression ratio to optimize the data transmission speed. Specifically, the transmitting unit selects the optimal compression algorithm according to the network bandwidth and communication conditions to maximize data transmission efficiency. For example, if the communication environment is good, it sends high-quality data with a low compression ratio, and if the communication environment is poor, it maintains transmission speed by reducing the amount of data with a high compression ratio. The transmitting unit can, for example, upload the recorded data to a cloud server and send it to the generating AI. The transmitting unit also has a function to automatically retransmit data if an error occurs during transmission. This ensures that data is transmitted reliably and that the generating AI receives accurate data. Furthermore, the transmitting unit has a data encryption function, reducing the risk of data being intercepted by a third party during transmission. This allows the transmitting unit to transmit data to the generating AI safely and efficiently, improving the reliability of the entire system.
[0032] The conversion unit analyzes the data transmitted by the transmission unit and converts it into a different voice. The conversion unit can convert a voice, for example, by changing specific parameters of the audio data. The conversion unit uses generative AI to convert the user's voice into a different voice. Specifically, the generative AI analyzes the audio features of the audio data, such as frequency components, amplitude, and formants, and adjusts the tone and pitch of the voice by changing these parameters. For example, the generative AI can make the voice sound younger by raising the frequency components of the user's voice, or conversely, make it sound calmer by lowering them. Because the conversion unit can analyze and convert audio data in real time, the user can immediately hear the converted voice. Furthermore, the conversion unit also has the function of learning the characteristics of the user's voice and automatically adjusting the conversion parameters to be optimal for each individual user. As a result, the conversion unit can convert the user's voice into a different voice that sounds natural and unnatural, and provide a variety of voice variations to meet the user's needs.
[0033] The provider unit provides the user with the voice converted by the conversion unit. For example, the provider unit can provide the user with the converted voice as an audio file. The provider unit may also have audio filtering functions to play the voice with optimal sound quality on the user's device. Specifically, the provider unit has a function to allow the user to download the audio file to their smartphone or computer. The provider unit applies audio filtering technologies such as noise reduction and echo cancellation to optimize the audio file format for the user's device and improve sound quality during playback. For example, the provider unit provides the converted voice audio file in common formats such as MP3 or WAV, making it easy for the user to play. Furthermore, the provider unit also has a function to allow the user to share the converted voice. For example, the provider unit can generate a link to share the audio file with other users via email or messaging apps. This allows the provider unit to provide the converted voice to the user quickly and conveniently, and to allow the user to use that voice in a variety of ways.
[0034] The recording unit can record the user's voice using a smartphone app. For example, recording begins when the user launches the smartphone app and presses the record button. The recording unit saves the recorded data when recording is complete. The recording unit makes it easy for users to record their voices using a smartphone app. The recording unit may also include features such as noise cancellation to ensure high-quality recording of the user's voice in the smartphone app. The recording unit can also save the recorded data in real time when the smartphone app is recording the user's voice. This makes it easy for users to record their voices using a smartphone app.
[0035] The transmission unit can send data to a generative AI on the cloud. For example, the transmission unit can upload recorded data to a cloud server and send it to the generative AI. The transmission unit improves processing power by sending data to a generative AI on the cloud. For example, the transmission unit can process large amounts of data quickly because the cloud server has high processing power. The transmission unit can also dynamically adjust the data compression ratio when sending data to a generative AI on the cloud. For example, the transmission unit can increase the compression ratio to increase the transmission speed when the network is congested. The transmission unit can reduce the load on the user's device by sending data to a generative AI on the cloud. This improves processing power by sending data to a generative AI on the cloud.
[0036] The conversion unit can transform a voice by changing specific parameters of the audio data. For example, the conversion unit can change the frequency and amplitude of the audio data to adjust the tone and pitch of the voice. The conversion unit can transform a voice by changing specific parameters of the audio data. For example, the conversion unit can change the formants of the audio data to alter the characteristics of the voice. The conversion unit can use a generative AI when changing specific parameters of the audio data. For example, the generative AI analyzes the audio data and changes specific parameters. The conversion unit can transform the user's voice into a different voice by changing specific parameters of the audio data. This makes it possible to transform a voice by changing specific parameters of the audio data.
[0037] The service provider can provide the converted voice to the user. For example, the service provider can provide the converted voice to the user as an audio file. By providing the converted voice to the user, the service provider can enable the user to use the converted voice. For example, the service provider can have the user download the converted voice to their smartphone. When providing the converted voice to the user, the service provider can use an audio filtering function. For example, the service provider can use an audio filtering function to optimize the sound quality of the converted voice. By providing the converted voice to the user, the service provider can enable the user to use the converted voice for various purposes. In this way, by providing the converted voice to the user, the user can use the converted voice.
[0038] The recording unit can adjust the tone and pitch of the user's voice in real time during recording. For example, if the user's voice is too low, the recording unit will raise the pitch in real time. If the user's voice is too high, the recording unit will lower the pitch in real time. If the user's voice tone is inconsistent, the recording unit will adjust the tone to be uniform in real time. The recording unit can use generative AI to adjust the tone and pitch of the user's voice in real time. For example, the generative AI analyzes the user's voice and adjusts the tone and pitch. The recording unit records the adjusted voice in real time. This allows for higher quality recordings by adjusting the tone and pitch of the user's voice in real time.
[0039] The recording unit can have a filtering function added to automatically remove background noise during recording. For example, the recording unit can automatically remove wind noise that occurs during recording. The recording unit can automatically remove car noise that occurs during recording. The recording unit can automatically remove human voices that occur during recording. The recording unit can use generative AI to automatically remove background noise. For example, the generative AI analyzes the recording data and removes the noise. The recording unit records clear audio with the noise removed. This makes clear recording possible by automatically removing background noise.
[0040] The recording unit can suggest the optimal recording environment based on the user's geographical location information during recording. For example, if the user is outdoors, the recording unit will suggest a location with little wind. If the user is indoors, the recording unit will suggest a quiet room. If the user is in a car, the recording unit will suggest a location with little engine noise. The recording unit can use generative AI to suggest the optimal recording environment based on the user's geographical location information. For example, the generative AI will analyze the user's location information and suggest the optimal recording environment. The recording unit will suggest the optimal recording environment based on the user's geographical location information. This allows for a better recording environment to be provided by suggesting the optimal recording environment based on the user's geographical location information.
[0041] The recording unit can automatically apply the optimal recording settings by referring to the user's past recording data during recording. For example, the recording unit can automatically apply the microphone settings the user has used in the past. The recording unit can apply noise cancellation settings by referring to ambient sounds the user has recorded in the past. The recording unit can apply the optimal settings by referring to the tone and pitch of the user's voice recorded in the past. The recording unit can use generative AI to automatically apply the optimal recording settings by referring to the user's past recording data. For example, the recording unit can use generative AI to analyze past recording data and select the optimal settings. The recording unit can then refer to the user's past recording data and automatically apply the optimal recording settings. This allows the optimal recording settings to be automatically applied by referring to the user's past recording data.
[0042] The transmission unit can dynamically adjust the data compression ratio to optimize transmission speed during data transmission. For example, if the network is congested, the transmission unit increases the compression ratio to increase transmission speed. If the network is not congested, the transmission unit decreases the compression ratio to transmit high-quality data. If the user's device battery level is low, the transmission unit increases the compression ratio to shorten transmission time. The transmission unit can use generative AI to dynamically adjust the data compression ratio. For example, the generative AI analyzes the network conditions and selects the optimal compression ratio. The transmission unit optimizes transmission speed by dynamically adjusting the data compression ratio during data transmission. This allows for optimized transmission speed by dynamically adjusting the data compression ratio.
[0043] The transmitting unit can select the optimal transmission route based on the load status of the destination server when transmitting data. For example, if the destination server is congested, the transmitting unit will send the data to another server. If the destination server is not congested, the transmitting unit will send the data to that server. If the destination server is under heavy load, the transmitting unit will temporarily delay transmission. The transmitting unit can use generative AI to select the optimal transmission route based on the load status of the destination server. For example, the generative AI will analyze the server load status and select the optimal transmission route. The transmitting unit selects the optimal transmission route based on the load status of the destination server when transmitting data. This enables efficient data transmission by selecting the optimal transmission route based on the load status of the destination server.
[0044] The transmission unit can monitor the user's network connection status in real time during data transmission and select the optimal transmission method. For example, if the user is connected to Wi-Fi, the transmission unit will select a high-speed transmission method. If the user is connected to mobile data, the transmission unit will select a transmission method that minimizes data usage. If the user's network connection is unstable, the transmission unit will temporarily delay transmission. The transmission unit can use generative AI to monitor the user's network connection status in real time and select the optimal transmission method. For example, the generative AI will analyze the network connection status and select the optimal transmission method. The transmission unit monitors the user's network connection status in real time during data transmission and selects the optimal transmission method. This allows the transmission unit to select the optimal transmission method by monitoring the user's network connection status in real time.
[0045] The transmitting unit can adjust the transmission method based on the battery level of the user's device when transmitting data. For example, if the user's device battery level is low, the transmitting unit will increase the compression ratio of the transmitted data to shorten the transmission time. If the user's device battery level is sufficient, the transmitting unit will transmit high-quality data. If the user's device battery level is moderate, the transmitting unit will select a balanced transmission method. The transmitting unit can use generative AI to adjust the transmission method based on the battery level of the user's device. For example, the generative AI will analyze the battery level and select the optimal transmission method. The transmitting unit adjusts the transmission method based on the battery level of the user's device when transmitting data. This enables efficient data transmission by adjusting the transmission method based on the battery level of the user's device.
[0046] The conversion unit can apply algorithms that convert the user's voice to a different voice while preserving its unique characteristics. For example, the conversion unit can change only the tone of the user's voice without altering its pitch. The conversion unit can change the voice quality while preserving the rhythm of the user's voice. The conversion unit can change the timbre while preserving the accent of the user's voice. The conversion unit can use a generative AI to convert the user's voice to a different voice while preserving its unique characteristics. For example, the generative AI analyzes the user's voice characteristics and applies an appropriate algorithm. The conversion unit applies an algorithm that converts the user's voice to a different voice while preserving its unique characteristics. This enables natural voice conversion by preserving the user's voice characteristics while converting it to a different voice.
[0047] The conversion unit can be enhanced with a function to convert the voice to one that corresponds to a specific language or dialect during the conversion process. For example, the conversion unit can convert the user's voice from English to Japanese. The conversion unit can convert the user's voice from standard Japanese to Kansai dialect. The conversion unit can convert the user's voice from French to German. The conversion unit can use a generative AI to convert the voice to one that corresponds to a specific language or dialect. For example, the conversion unit can use a generative AI to convert the user's voice to the appropriate language or dialect using a language model. The conversion unit can be enhanced with a function to convert the voice to one that corresponds to a specific language or dialect during the conversion process. This enables multilingual voice conversion by converting the voice to one that corresponds to a specific language or dialect.
[0048] The conversion unit can automatically apply the optimal conversion parameters by referring to the user's past conversion history during conversion. For example, the conversion unit automatically applies conversion parameters that the user has used in the past. The conversion unit selects the optimal voice quality from the user's past conversion history. The conversion unit analyzes the user's past conversion history and applies the most preferred conversion parameters. The conversion unit can use a generation AI to automatically apply the optimal conversion parameters by referring to the user's past conversion history. For example, the generation AI analyzes the past conversion history and selects the optimal parameters. The conversion unit automatically applies the optimal conversion parameters by referring to the user's past conversion history during conversion. This allows the optimal conversion parameters to be automatically applied by referring to the user's past conversion history.
[0049] The conversion unit can select a conversion algorithm according to the intended use of the user's voice during conversion. For example, for entertainment purposes, the conversion unit converts to a humorous voice. For educational purposes, the conversion unit converts to a clear and easy-to-understand voice. For security purposes, the conversion unit converts to a highly reliable voice. The conversion unit can use a generation AI to select a conversion algorithm according to the intended use of the user's voice. For example, the generation AI analyzes the intended use and selects an appropriate conversion algorithm. The conversion unit selects a conversion algorithm according to the intended use of the user's voice during conversion. This allows for more appropriate voice conversion by selecting a conversion algorithm according to the intended use of the user's voice.
[0050] The service provider can add audio filtering functionality at the time of delivery to play audio at the optimal sound quality for the user's device. For example, when playing on a smartphone, the service provider adjusts the sound quality to match the device's speaker characteristics. When playing on headphones, the service provider optimizes sound localization and balance. When playing on an in-car audio system, the service provider removes echoes and noise to provide clear sound quality. The service provider can use generative AI to add audio filtering functionality to play audio at the optimal sound quality for the user's device. For example, the service provider can use generative AI to analyze the device's characteristics and provide optimal sound quality. The service provider adds audio filtering functionality at the time of delivery to play audio at the optimal sound quality for the user's device. This enables higher quality playback by adding audio filtering functionality to play audio at the optimal sound quality for the user's device.
[0051] The service provider can select the optimal service method by referring to the user's past usage history at the time of service provision. For example, the service provider can automatically apply the playback method that the user has previously preferred. The service provider can select the optimal volume setting from the user's past usage history. The service provider can analyze the user's past usage history and apply the most preferred playback method. The service provider can use generative AI to select the optimal service method by referring to the user's past usage history. For example, the service provider can use generative AI to analyze past usage history and select the optimal service method. The service provider selects the optimal service method by referring to the user's past usage history at the time of service provision. This allows the service provider to select the optimal service method by referring to the user's past usage history.
[0052] The service provider can select the optimal service delivery method according to the type of user's device at the time of delivery. For example, when playing on a smartphone, the service provider provides a display method that matches the screen size. When playing on a tablet, the service provider provides a display method optimized for a large screen. When playing on a PC, the service provider provides a high-resolution display method. The service provider can use a generation AI to select the optimal service delivery method according to the type of user's device. For example, the generation AI analyzes the characteristics of the device and selects the optimal service delivery method. The service provider selects the optimal service delivery method according to the type of user's device at the time of delivery. This allows for more appropriate playback by selecting the optimal service delivery method according to the type of user's device.
[0053] The service provider can select the optimal service delivery method based on the user's network connection status at the time of delivery. For example, if the user is connected to Wi-Fi, the service provider will select a high-speed service delivery method. If the user is connected to mobile data, the service provider will select a service delivery method that minimizes data usage. If the user's network connection is unstable, the service provider will temporarily delay the delivery. The service provider can use generative AI to select the optimal service delivery method based on the user's network connection status. For example, the generative AI will analyze the network connection status and select the optimal service delivery method. The service provider selects the optimal service delivery method based on the user's network connection status at the time of delivery. This enables efficient playback by selecting the optimal service delivery method based on the user's network connection status.
[0054] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0055] The recording unit can analyze the characteristics of the user's voice in real time while recording it and automatically apply the optimal recording settings. For example, if the user's voice is biased towards the low frequency range, the recording unit will automatically apply settings to emphasize the low frequencies. If the user's voice is biased towards the high frequency range, it will apply settings to emphasize the high frequencies. Furthermore, if the volume of the user's voice is not constant, the recording unit can automatically apply settings to adjust the volume to be uniform. This enables optimal recording tailored to the characteristics of the user's voice.
[0056] The voice conversion unit can select a conversion algorithm based on the intended use of the user's voice. For example, for entertainment purposes, it converts the voice to a humorous one. For educational purposes, it converts the voice to a clear and easy-to-understand one. For security purposes, it converts the voice to a highly reliable one. The voice conversion unit uses a generation AI to analyze the intended use of the user's voice and selects the appropriate conversion algorithm. This enables optimal voice conversion tailored to the intended use of the user's voice.
[0057] The recording unit can be enhanced with a filtering function that automatically removes background noise when recording the user's voice. For example, it can automatically remove wind noise, car noise, and human speech during recording. The recording unit uses a generation AI to analyze the recording data and remove noise. This enables clear recordings by automatically removing background noise.
[0058] The service provider can add audio filtering functionality at the time of delivery to ensure optimal sound quality for the user's device. For example, when playing on a smartphone, the sound quality is adjusted to match the device's speaker characteristics. When playing with headphones, the sound localization and balance are optimized. When playing on an in-car audio system, echoes and noise are removed to provide clear sound quality. The service provider uses generative AI to analyze the device's characteristics and provide optimal sound quality. By adding audio filtering functionality to ensure optimal sound quality for the user's device, higher quality playback becomes possible.
[0059] The conversion unit can be enhanced with a function to convert the voice to one that corresponds to a specific language or dialect during the conversion process. For example, it can convert the user's voice from English to Japanese, from standard Japanese to Kansai dialect, or from French to German. The conversion unit uses a generative AI to analyze a language model and convert the user's voice to the appropriate language or dialect. This enables multilingual voice conversion by converting the voice to one that corresponds to a specific language or dialect.
[0060] The following briefly describes the processing flow for example form 1.
[0061] Step 1: The recording unit records the user's voice. The recording unit can record the user's voice using, for example, a smartphone app. The recording unit may also have a noise-canceling function to record the user's voice in high quality. The recording unit starts recording when, for example, the user launches the smartphone app and presses the record button. When the recording is finished, the recording unit saves the recorded data. Step 2: The transmitting unit sends the data recorded by the recording unit to the generating AI. The transmitting unit can send the data to, for example, a generating AI on the cloud. The transmitting unit can also dynamically adjust the data compression ratio to optimize the data transmission speed. For example, the transmitting unit uploads the recorded data to a cloud server and sends it to the generating AI. Step 3: The conversion unit analyzes the data transmitted by the transmission unit and converts it into a different voice. The conversion unit can convert the voice by, for example, changing specific parameters of the audio data. The conversion unit uses generative AI to convert the user's voice into a different voice. The conversion unit can adjust the tone and pitch of the voice by, for example, changing the frequency and amplitude of the audio data. Step 4: The provider unit provides the user with the voice converted by the conversion unit. The provider unit can, for example, provide the converted voice to the user as an audio file. The provider unit may also have an audio filtering function to play the audio at the optimal sound quality for the user's device. The provider unit can, for example, have the user download the converted voice to their smartphone.
[0062] (Example of form 2) The voice conversion system according to an embodiment of the present invention is a system that records a user's voice, converts it to another voice using a generation AI, and provides it to the user. The voice conversion system works by having the user record their own voice and send the recording data to the generation AI. The generation AI analyzes the data and converts it to another voice. The converted voice is then provided to the user. This system allows for easy voice conversion and is expected to have applications in fields such as entertainment, education, and security. For example, if a user records their own voice and sends the recording data to the generation AI, the generation AI converts that voice to a different voice and provides it to the user. This mechanism allows for easy voice conversion and can be used for various purposes. As a result, the voice conversion system can automatically convert and provide the user's voice.
[0063] The voice conversion system according to this embodiment comprises a recording unit, a transmission unit, a conversion unit, and a provision unit. The recording unit records the user's voice. The recording unit can record the user's voice, for example, using a smartphone application. The recording unit may also have a noise-canceling function to record the user's voice in high quality. The recording unit starts recording, for example, when the user launches a smartphone application and presses the record button. When recording is finished, the recording unit saves the recorded data. The transmission unit sends the data recorded by the recording unit to a generation AI. The transmission unit can send the data to a generation AI on the cloud, for example. The transmission unit can also dynamically adjust the data compression ratio to optimize the data transmission speed. The transmission unit can upload the recorded data to a cloud server and send it to the generation AI, for example. The conversion unit analyzes the data transmitted by the transmission unit and converts it into a different voice. The conversion unit can convert the voice, for example, by changing specific parameters of the audio data. The conversion unit uses the generation AI to convert the user's voice into a different voice. The conversion unit, for example, modifies the frequency and amplitude of the audio data to adjust the tone and pitch of the voice. The providing unit provides the voice converted by the conversion unit to the user. The providing unit can, for example, provide the converted voice to the user as an audio file. The providing unit may also have an audio filtering function to play the voice with optimal sound quality for the user's device. The providing unit can, for example, allow the user to download the converted voice to their smartphone. Thus, the voice conversion system according to this embodiment can record the user's voice, convert it to another voice using a generation AI, and provide it.
[0064] The recording unit records the user's voice. For example, the recording unit can record the user's voice using a smartphone app. The recording unit may also have noise cancellation capabilities to ensure high-quality recording of the user's voice. Specifically, the recording unit uses active noise cancellation technology to reduce ambient noise. This technology allows for clear audio recording by having the microphone pick up ambient sounds and generating sound waves with the opposite phase to cancel out the noise. The recording unit starts recording, for example, when the user launches a smartphone app and presses the record button. Once recording is complete, the recording unit saves the recording data. The recording data is stored in local or cloud storage and used for subsequent processing. The recording unit also has a function to monitor the audio quality in real time during recording and automatically adjust recording settings as needed. For example, if the audio volume is too low during recording, it can automatically increase the sensitivity to record at an appropriate volume. This allows the recording unit to record the user's voice in high quality and provide data optimized for subsequent processing.
[0065] The transmitting unit sends data recorded by the recording unit to the generating AI. The transmitting unit can, for example, send data to a generating AI on the cloud. The transmitting unit can also dynamically adjust the data compression ratio to optimize the data transmission speed. Specifically, the transmitting unit selects the optimal compression algorithm according to the network bandwidth and communication conditions to maximize data transmission efficiency. For example, if the communication environment is good, it sends high-quality data with a low compression ratio, and if the communication environment is poor, it maintains transmission speed by reducing the amount of data with a high compression ratio. The transmitting unit can, for example, upload the recorded data to a cloud server and send it to the generating AI. The transmitting unit also has a function to automatically retransmit data if an error occurs during transmission. This ensures that data is transmitted reliably and that the generating AI receives accurate data. Furthermore, the transmitting unit has a data encryption function, reducing the risk of data being intercepted by a third party during transmission. This allows the transmitting unit to transmit data to the generating AI safely and efficiently, improving the reliability of the entire system.
[0066] The conversion unit analyzes the data transmitted by the transmission unit and converts it into a different voice. The conversion unit can convert a voice, for example, by changing specific parameters of the audio data. The conversion unit uses generative AI to convert the user's voice into a different voice. Specifically, the generative AI analyzes the audio features of the audio data, such as frequency components, amplitude, and formants, and adjusts the tone and pitch of the voice by changing these parameters. For example, the generative AI can make the voice sound younger by raising the frequency components of the user's voice, or conversely, make it sound calmer by lowering them. Because the conversion unit can analyze and convert audio data in real time, the user can immediately hear the converted voice. Furthermore, the conversion unit also has the function of learning the characteristics of the user's voice and automatically adjusting the conversion parameters to be optimal for each individual user. As a result, the conversion unit can convert the user's voice into a different voice that sounds natural and unnatural, and provide a variety of voice variations to meet the user's needs.
[0067] The provider unit provides the user with the voice converted by the conversion unit. For example, the provider unit can provide the user with the converted voice as an audio file. The provider unit may also have audio filtering functions to play the voice with optimal sound quality on the user's device. Specifically, the provider unit has a function to allow the user to download the audio file to their smartphone or computer. The provider unit applies audio filtering technologies such as noise reduction and echo cancellation to optimize the audio file format for the user's device and improve sound quality during playback. For example, the provider unit provides the converted voice audio file in common formats such as MP3 or WAV, making it easy for the user to play. Furthermore, the provider unit also has a function to allow the user to share the converted voice. For example, the provider unit can generate a link to share the audio file with other users via email or messaging apps. This allows the provider unit to provide the converted voice to the user quickly and conveniently, and to allow the user to use that voice in a variety of ways.
[0068] The recording unit can record the user's voice using a smartphone app. For example, recording begins when the user launches the smartphone app and presses the record button. The recording unit saves the recorded data when recording is complete. The recording unit makes it easy for users to record their voices using a smartphone app. The recording unit may also include features such as noise cancellation to ensure high-quality recording of the user's voice in the smartphone app. The recording unit can also save the recorded data in real time when the smartphone app is recording the user's voice. This makes it easy for users to record their voices using a smartphone app.
[0069] The transmission unit can send data to a generative AI on the cloud. For example, the transmission unit can upload recorded data to a cloud server and send it to the generative AI. The transmission unit improves processing power by sending data to a generative AI on the cloud. For example, the transmission unit can process large amounts of data quickly because the cloud server has high processing power. The transmission unit can also dynamically adjust the data compression ratio when sending data to a generative AI on the cloud. For example, the transmission unit can increase the compression ratio to increase the transmission speed when the network is congested. The transmission unit can reduce the load on the user's device by sending data to a generative AI on the cloud. This improves processing power by sending data to a generative AI on the cloud.
[0070] The conversion unit can transform a voice by changing specific parameters of the audio data. For example, the conversion unit can change the frequency and amplitude of the audio data to adjust the tone and pitch of the voice. The conversion unit can transform a voice by changing specific parameters of the audio data. For example, the conversion unit can change the formants of the audio data to alter the characteristics of the voice. The conversion unit can use a generative AI when changing specific parameters of the audio data. For example, the generative AI analyzes the audio data and changes specific parameters. The conversion unit can transform the user's voice into a different voice by changing specific parameters of the audio data. This makes it possible to transform a voice by changing specific parameters of the audio data.
[0071] The service provider can provide the converted voice to the user. For example, the service provider can provide the converted voice to the user as an audio file. By providing the converted voice to the user, the service provider can enable the user to use the converted voice. For example, the service provider can have the user download the converted voice to their smartphone. When providing the converted voice to the user, the service provider can use an audio filtering function. For example, the service provider can use an audio filtering function to optimize the sound quality of the converted voice. By providing the converted voice to the user, the service provider can enable the user to use the converted voice for various purposes. In this way, by providing the converted voice to the user, the user can use the converted voice.
[0072] The recording unit can estimate the user's emotions and adjust the recording start time based on the estimated emotions. For example, if the user is nervous, the recording unit will wait until the user is relaxed before starting recording. If the user is excited, the recording unit will start recording immediately to capture that emotion. If the user is calm, the recording unit will start recording at a natural timing. The recording unit can use emotion estimation functions, such as an emotion engine or generative AI, to estimate the user's emotions. For example, the emotion engine analyzes the user's facial expressions and voice to estimate emotions. The recording unit adjusts the recording start time based on the estimated emotions. This allows for more natural recordings by adjusting the recording start time based on the user's emotions.
[0073] The recording unit can adjust the tone and pitch of the user's voice in real time during recording. For example, if the user's voice is too low, the recording unit will raise the pitch in real time. If the user's voice is too high, the recording unit will lower the pitch in real time. If the user's voice tone is inconsistent, the recording unit will adjust the tone to be uniform in real time. The recording unit can use generative AI to adjust the tone and pitch of the user's voice in real time. For example, the generative AI analyzes the user's voice and adjusts the tone and pitch. The recording unit records the adjusted voice in real time. This allows for higher quality recordings by adjusting the tone and pitch of the user's voice in real time.
[0074] The recording unit can have a filtering function added to automatically remove background noise during recording. For example, the recording unit can automatically remove wind noise that occurs during recording. The recording unit can automatically remove car noise that occurs during recording. The recording unit can automatically remove human voices that occur during recording. The recording unit can use generative AI to automatically remove background noise. For example, the generative AI analyzes the recording data and removes the noise. The recording unit records clear audio with the noise removed. This makes clear recording possible by automatically removing background noise.
[0075] The recording unit can estimate the user's emotions and adjust the recording length based on the estimated emotions. For example, if the user is relaxed, the recording unit will record for a longer duration. If the user is in a hurry, the recording unit will record for a shorter duration. If the user is excited, the recording unit will continue recording until the emotion subsides. The recording unit can use emotion estimation capabilities, such as an emotion engine or generative AI, to estimate the user's emotions. For example, the emotion engine analyzes the user's facial expressions and voice to estimate their emotions. The recording unit then adjusts the recording length based on the estimated emotions. This allows for more appropriate recordings by adjusting the recording length based on the user's emotions.
[0076] The recording unit can suggest the optimal recording environment based on the user's geographical location information during recording. For example, if the user is outdoors, the recording unit will suggest a location with little wind. If the user is indoors, the recording unit will suggest a quiet room. If the user is in a car, the recording unit will suggest a location with little engine noise. The recording unit can use generative AI to suggest the optimal recording environment based on the user's geographical location information. For example, the generative AI will analyze the user's location information and suggest the optimal recording environment. The recording unit will suggest the optimal recording environment based on the user's geographical location information. This allows for a better recording environment to be provided by suggesting the optimal recording environment based on the user's geographical location information.
[0077] The recording unit can automatically apply the optimal recording settings by referring to the user's past recording data during recording. For example, the recording unit can automatically apply the microphone settings the user has used in the past. The recording unit can apply noise cancellation settings by referring to ambient sounds the user has recorded in the past. The recording unit can apply the optimal settings by referring to the tone and pitch of the user's voice recorded in the past. The recording unit can use generative AI to automatically apply the optimal recording settings by referring to the user's past recording data. For example, the recording unit can use generative AI to analyze past recording data and select the optimal settings. The recording unit can then refer to the user's past recording data and automatically apply the optimal recording settings. This allows the optimal recording settings to be automatically applied by referring to the user's past recording data.
[0078] The transmitting unit can estimate the user's emotions and adjust the timing of data transmission based on the estimated emotions. For example, if the user is relaxed, the transmitting unit will transmit data immediately. If the user is tense, the transmitting unit will wait until the user is relaxed before transmitting data. If the user is in a hurry, the transmitting unit will transmit data instantly. The transmitting unit can use emotion estimation capabilities, such as an emotion engine or generative AI, to estimate the user's emotions. For example, the emotion engine can analyze the user's facial expressions and voice to estimate their emotions. The transmitting unit adjusts the timing of data transmission based on the estimated emotions. This allows for data transmission at a more appropriate time by adjusting the timing of data transmission based on the user's emotions.
[0079] The transmission unit can dynamically adjust the data compression ratio to optimize transmission speed during data transmission. For example, if the network is congested, the transmission unit increases the compression ratio to increase transmission speed. If the network is not congested, the transmission unit decreases the compression ratio to transmit high-quality data. If the user's device battery level is low, the transmission unit increases the compression ratio to shorten transmission time. The transmission unit can use generative AI to dynamically adjust the data compression ratio. For example, the generative AI analyzes the network conditions and selects the optimal compression ratio. The transmission unit optimizes transmission speed by dynamically adjusting the data compression ratio during data transmission. This allows for optimized transmission speed by dynamically adjusting the data compression ratio.
[0080] The transmitting unit can select the optimal transmission route based on the load status of the destination server when transmitting data. For example, if the destination server is congested, the transmitting unit will send the data to another server. If the destination server is not congested, the transmitting unit will send the data to that server. If the destination server is under heavy load, the transmitting unit will temporarily delay transmission. The transmitting unit can use generative AI to select the optimal transmission route based on the load status of the destination server. For example, the generative AI will analyze the server load status and select the optimal transmission route. The transmitting unit selects the optimal transmission route based on the load status of the destination server when transmitting data. This enables efficient data transmission by selecting the optimal transmission route based on the load status of the destination server.
[0081] The transmitting unit can estimate the user's emotions and prioritize the data to be transmitted based on the estimated emotions. For example, if the user is in a hurry, the transmitting unit will prioritize transmitting important data. If the user is relaxed, the transmitting unit will transmit all data equally. If the user is excited, the transmitting unit will prioritize transmitting data related to those emotions. The transmitting unit can use emotion estimation capabilities, such as an emotion engine or generative AI, to estimate the user's emotions. For example, the emotion engine can analyze the user's facial expressions and voice to estimate their emotions. The transmitting unit then prioritizes the data to be transmitted based on the estimated emotions. This allows for the priority transmission of important data based on the user's emotions.
[0082] The transmission unit can monitor the user's network connection status in real time during data transmission and select the optimal transmission method. For example, if the user is connected to Wi-Fi, the transmission unit will select a high-speed transmission method. If the user is connected to mobile data, the transmission unit will select a transmission method that minimizes data usage. If the user's network connection is unstable, the transmission unit will temporarily delay transmission. The transmission unit can use generative AI to monitor the user's network connection status in real time and select the optimal transmission method. For example, the generative AI will analyze the network connection status and select the optimal transmission method. The transmission unit monitors the user's network connection status in real time during data transmission and selects the optimal transmission method. This allows the transmission unit to select the optimal transmission method by monitoring the user's network connection status in real time.
[0083] The transmitting unit can adjust the transmission method based on the battery level of the user's device when transmitting data. For example, if the user's device battery level is low, the transmitting unit will increase the compression ratio of the transmitted data to shorten the transmission time. If the user's device battery level is sufficient, the transmitting unit will transmit high-quality data. If the user's device battery level is moderate, the transmitting unit will select a balanced transmission method. The transmitting unit can use generative AI to adjust the transmission method based on the battery level of the user's device. For example, the generative AI will analyze the battery level and select the optimal transmission method. The transmitting unit adjusts the transmission method based on the battery level of the user's device when transmitting data. This enables efficient data transmission by adjusting the transmission method based on the battery level of the user's device.
[0084] The voice conversion unit can estimate the user's emotions and adjust the tone and pitch of the converted voice based on the estimated emotions. For example, if the user is relaxed, the unit will convert to a calm tone of voice. If the user is excited, the unit will convert to an energetic tone of voice. If the user is sad, the unit will convert to a calm tone of voice. The voice conversion unit can use emotion estimation functions, such as an emotion engine or generative AI, to estimate the user's emotions. For example, the emotion engine analyzes the user's facial expressions and voice to estimate emotions. The voice conversion unit adjusts the tone and pitch of the converted voice based on the estimated emotions. This allows for a more natural voice conversion by adjusting the tone and pitch of the converted voice based on the user's emotions.
[0085] The conversion unit can apply algorithms that convert the user's voice to a different voice while preserving its unique characteristics. For example, the conversion unit can change only the tone of the user's voice without altering its pitch. The conversion unit can change the voice quality while preserving the rhythm of the user's voice. The conversion unit can change the timbre while preserving the accent of the user's voice. The conversion unit can use a generative AI to convert the user's voice to a different voice while preserving its unique characteristics. For example, the generative AI analyzes the user's voice characteristics and applies an appropriate algorithm. The conversion unit applies an algorithm that converts the user's voice to a different voice while preserving its unique characteristics. This enables natural voice conversion by preserving the user's voice characteristics while converting it to a different voice.
[0086] The conversion unit can be enhanced with a function to convert the voice to one that corresponds to a specific language or dialect during the conversion process. For example, the conversion unit can convert the user's voice from English to Japanese. The conversion unit can convert the user's voice from standard Japanese to Kansai dialect. The conversion unit can convert the user's voice from French to German. The conversion unit can use a generative AI to convert the voice to one that corresponds to a specific language or dialect. For example, the conversion unit can use a generative AI to convert the user's voice to the appropriate language or dialect using a language model. The conversion unit can be enhanced with a function to convert the voice to one that corresponds to a specific language or dialect during the conversion process. This enables multilingual voice conversion by converting the voice to one that corresponds to a specific language or dialect.
[0087] The voice conversion unit can estimate the user's emotions and adjust the length of the converted voice based on the estimated emotions. For example, if the user is in a hurry, the unit converts to a short, concise voice. If the user is relaxed, the unit converts to a longer voice that includes detailed explanations. If the user is excited, the unit converts to a voice of a length that emphasizes the emotion. The voice conversion unit can use emotion estimation functions, such as an emotion engine or generative AI, to estimate the user's emotions. For example, the emotion engine analyzes the user's facial expressions and voice to estimate emotions. The unit then adjusts the length of the converted voice based on the estimated emotions. This allows for more appropriate voice conversion by adjusting the length of the converted voice based on the user's emotions.
[0088] The conversion unit can automatically apply the optimal conversion parameters by referring to the user's past conversion history during conversion. For example, the conversion unit automatically applies conversion parameters that the user has used in the past. The conversion unit selects the optimal voice quality from the user's past conversion history. The conversion unit analyzes the user's past conversion history and applies the most preferred conversion parameters. The conversion unit can use a generation AI to automatically apply the optimal conversion parameters by referring to the user's past conversion history. For example, the generation AI analyzes the past conversion history and selects the optimal parameters. The conversion unit automatically applies the optimal conversion parameters by referring to the user's past conversion history during conversion. This allows the optimal conversion parameters to be automatically applied by referring to the user's past conversion history.
[0089] The conversion unit can select a conversion algorithm according to the intended use of the user's voice during conversion. For example, for entertainment purposes, the conversion unit converts to a humorous voice. For educational purposes, the conversion unit converts to a clear and easy-to-understand voice. For security purposes, the conversion unit converts to a highly reliable voice. The conversion unit can use a generation AI to select a conversion algorithm according to the intended use of the user's voice. For example, the generation AI analyzes the intended use and selects an appropriate conversion algorithm. The conversion unit selects a conversion algorithm according to the intended use of the user's voice during conversion. This allows for more appropriate voice conversion by selecting a conversion algorithm according to the intended use of the user's voice.
[0090] The voice delivery unit can estimate the user's emotions and adjust the voice playback method based on the estimated emotions. For example, if the user is relaxed, the voice delivery unit will play at a gentle volume. If the user is excited, the voice delivery unit will play at a loud volume. If the user is sad, the voice delivery unit will play at a calm volume. The voice delivery unit can use an emotion estimation function, such as an emotion engine or generative AI, to estimate the user's emotions. For example, the emotion engine will analyze the user's facial expressions and voice to estimate the emotions. The voice delivery unit will adjust the voice playback method based on the estimated emotions. This allows for more appropriate playback by adjusting the voice playback method based on the user's emotions.
[0091] The service provider can add audio filtering functionality at the time of delivery to play audio at the optimal sound quality for the user's device. For example, when playing on a smartphone, the service provider adjusts the sound quality to match the device's speaker characteristics. When playing on headphones, the service provider optimizes sound localization and balance. When playing on an in-car audio system, the service provider removes echoes and noise to provide clear sound quality. The service provider can use generative AI to add audio filtering functionality to play audio at the optimal sound quality for the user's device. For example, the service provider can use generative AI to analyze the device's characteristics and provide optimal sound quality. The service provider adds audio filtering functionality at the time of delivery to play audio at the optimal sound quality for the user's device. This enables higher quality playback by adding audio filtering functionality to play audio at the optimal sound quality for the user's device.
[0092] The service provider can select the optimal service method by referring to the user's past usage history at the time of service provision. For example, the service provider can automatically apply the playback method that the user has previously preferred. The service provider can select the optimal volume setting from the user's past usage history. The service provider can analyze the user's past usage history and apply the most preferred playback method. The service provider can use generative AI to select the optimal service method by referring to the user's past usage history. For example, the service provider can use generative AI to analyze past usage history and select the optimal service method. The service provider selects the optimal service method by referring to the user's past usage history at the time of service provision. This allows the service provider to select the optimal service method by referring to the user's past usage history.
[0093] The service provider can estimate the user's emotions and determine the priority of the voices it provides based on those emotions. For example, if the user is in a hurry, the service provider will prioritize providing important voices. If the user is relaxed, the service provider will provide all voices equally. If the user is excited, the service provider will prioritize providing voices related to those emotions. The service provider can use emotion estimation functionality, such as an emotion engine or generative AI, to estimate the user's emotions. For example, the emotion engine can analyze the user's facial expressions and voice to estimate their emotions. Based on the estimated emotions, the service provider will determine the priority of the voices it provides. This allows the service provider to prioritize important voices by determining the priority of voices based on the user's emotions.
[0094] The service provider can select the optimal service delivery method according to the type of user's device at the time of delivery. For example, when playing on a smartphone, the service provider provides a display method that matches the screen size. When playing on a tablet, the service provider provides a display method optimized for a large screen. When playing on a PC, the service provider provides a high-resolution display method. The service provider can use a generation AI to select the optimal service delivery method according to the type of user's device. For example, the generation AI analyzes the characteristics of the device and selects the optimal service delivery method. The service provider selects the optimal service delivery method according to the type of user's device at the time of delivery. This allows for more appropriate playback by selecting the optimal service delivery method according to the type of user's device.
[0095] The service provider can select the optimal service delivery method based on the user's network connection status at the time of delivery. For example, if the user is connected to Wi-Fi, the service provider will select a high-speed service delivery method. If the user is connected to mobile data, the service provider will select a service delivery method that minimizes data usage. If the user's network connection is unstable, the service provider will temporarily delay the delivery. The service provider can use generative AI to select the optimal service delivery method based on the user's network connection status. For example, the generative AI will analyze the network connection status and select the optimal service delivery method. The service provider selects the optimal service delivery method based on the user's network connection status at the time of delivery. This enables efficient playback by selecting the optimal service delivery method based on the user's network connection status.
[0096] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0097] The recording unit can analyze the characteristics of the user's voice in real time while recording it and automatically apply the optimal recording settings. For example, if the user's voice is biased towards the low frequency range, the recording unit will automatically apply settings to emphasize the low frequencies. If the user's voice is biased towards the high frequency range, it will apply settings to emphasize the high frequencies. Furthermore, if the volume of the user's voice is not constant, the recording unit can automatically apply settings to adjust the volume to be uniform. This enables optimal recording tailored to the characteristics of the user's voice.
[0098] The transmission unit can estimate the user's emotions and determine the priority of data transmission based on those emotions. For example, if the user is stressed, important data will be sent first. If the user is relaxed, all data will be sent evenly. If the user is in a hurry, data will be sent immediately. The transmission unit estimates the user's emotions using an emotion engine or generative AI and determines the priority of data transmission based on those emotions. This enables appropriate data transmission according to the user's emotions.
[0099] The voice conversion unit can select a conversion algorithm based on the intended use of the user's voice. For example, for entertainment purposes, it converts the voice to a humorous one. For educational purposes, it converts the voice to a clear and easy-to-understand one. For security purposes, it converts the voice to a highly reliable one. The voice conversion unit uses a generation AI to analyze the intended use of the user's voice and selects the appropriate conversion algorithm. This enables optimal voice conversion tailored to the intended use of the user's voice.
[0100] The service provider can estimate the user's emotions and adjust the voice playback method based on the estimated emotions. For example, if the user is relaxed, the voice will be played at a gentle volume. If the user is excited, the voice will be played at a loud volume. If the user is sad, the voice will be played at a calming volume. The service provider uses an emotion engine or generative AI to estimate the user's emotions and adjusts the voice playback method based on those emotions. This enables appropriate playback according to the user's emotions.
[0101] The recording unit can be enhanced with a filtering function that automatically removes background noise when recording the user's voice. For example, it can automatically remove wind noise, car noise, and human speech during recording. The recording unit uses a generation AI to analyze the recording data and remove noise. This enables clear recordings by automatically removing background noise.
[0102] The voice conversion unit can estimate the user's emotions and adjust the tone and pitch of the converted voice based on those emotions. For example, if the user is relaxed, the voice will be converted to a calm tone. If the user is excited, the voice will be converted to an energetic tone. If the user is sad, the voice will be converted to a calm tone. The voice conversion unit estimates the user's emotions using an emotion engine or generative AI and adjusts the tone and pitch of the converted voice based on those emotions. This enables natural voice conversion based on the user's emotions.
[0103] The service provider can add audio filtering functionality at the time of delivery to ensure optimal sound quality for the user's device. For example, when playing on a smartphone, the sound quality is adjusted to match the device's speaker characteristics. When playing with headphones, the sound localization and balance are optimized. When playing on an in-car audio system, echoes and noise are removed to provide clear sound quality. The service provider uses generative AI to analyze the device's characteristics and provide optimal sound quality. By adding audio filtering functionality to ensure optimal sound quality for the user's device, higher quality playback becomes possible.
[0104] The recording unit can estimate the user's emotions and adjust the recording length based on that estimation. For example, if the user is relaxed, it will record for a longer duration. If the user is in a hurry, it will record for a shorter duration. If the user is excited, it will continue recording until their emotions subside. The recording unit uses an emotion engine or generative AI to estimate the user's emotions and adjusts the recording length based on those emotions. This enables appropriate recordings based on the user's emotions.
[0105] The conversion unit can be enhanced with a function to convert the voice to one that corresponds to a specific language or dialect during the conversion process. For example, it can convert the user's voice from English to Japanese, from standard Japanese to Kansai dialect, or from French to German. The conversion unit uses a generative AI to analyze a language model and convert the user's voice to the appropriate language or dialect. This enables multilingual voice conversion by converting the voice to one that corresponds to a specific language or dialect.
[0106] The delivery unit can estimate the user's emotions and determine the priority of the voices to deliver based on those emotions. For example, if the user is in a hurry, important voices will be prioritized. If the user is relaxed, all voices will be delivered equally. If the user is excited, voices related to that emotion will be prioritized. The delivery unit estimates the user's emotions using an emotion engine or generative AI and determines the priority of the voices to deliver based on those emotions. This makes it possible to deliver appropriate voices based on the user's emotions.
[0107] The following briefly describes the processing flow for example form 2.
[0108] Step 1: The recording unit records the user's voice. The recording unit can record the user's voice using, for example, a smartphone app. The recording unit may also have a noise-canceling function to record the user's voice in high quality. The recording unit starts recording when, for example, the user launches the smartphone app and presses the record button. When the recording is finished, the recording unit saves the recorded data. Step 2: The transmitting unit sends the data recorded by the recording unit to the generating AI. The transmitting unit can send the data to, for example, a generating AI on the cloud. The transmitting unit can also dynamically adjust the data compression ratio to optimize the data transmission speed. For example, the transmitting unit uploads the recorded data to a cloud server and sends it to the generating AI. Step 3: The conversion unit analyzes the data transmitted by the transmission unit and converts it into a different voice. The conversion unit can convert the voice by, for example, changing specific parameters of the audio data. The conversion unit uses generative AI to convert the user's voice into a different voice. The conversion unit can adjust the tone and pitch of the voice by, for example, changing the frequency and amplitude of the audio data. Step 4: The provider unit provides the user with the voice converted by the conversion unit. The provider unit can, for example, provide the converted voice to the user as an audio file. The provider unit may also have an audio filtering function to play the audio at the optimal sound quality for the user's device. The provider unit can, for example, have the user download the converted voice to their smartphone.
[0109] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0110] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0111] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0112] Each of the multiple elements described above, including the recording unit, transmission unit, conversion unit, and providing unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the recording unit records the user's voice using the microphone 38B of the smart device 14. The transmission unit transmits the recorded data to the data processing unit 12 via the communication I / F 44 of the smart device 14. The conversion unit is implemented by the specific processing unit 290 of the data processing unit 12 and converts the user's voice into another voice using a generation AI. The providing unit provides the converted voice to the user through the output device 40 of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0113] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0114] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0115] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0116] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0117] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0118] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0119] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0120] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0121] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0122] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0123] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0124] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0125] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0126] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0127] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0128] Each of the multiple elements described above, including the recording unit, transmission unit, conversion unit, and providing unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the recording unit records the user's voice using the microphone 238 of the smart glasses 214. The transmission unit transmits the recorded data to the data processing unit 12 via the communication I / F 44 of the smart glasses 214. The conversion unit is implemented by the specific processing unit 290 of the data processing unit 12 and converts the user's voice into another voice using generation AI. The providing unit provides the converted voice to the user through the speaker 240 of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0129] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0130] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0131] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0132] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0133] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0134] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0135] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0136] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0137] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0138] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0139] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0140] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0141] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0142] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0143] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0144] Each of the multiple elements described above, including the recording unit, transmission unit, conversion unit, and providing unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the recording unit records the user's voice using the microphone 238 of the headset terminal 314. The transmission unit transmits the recorded data to the data processing unit 12 via the communication I / F 44 of the headset terminal 314. The conversion unit is implemented by the specific processing unit 290 of the data processing unit 12 and converts the user's voice into another voice using a generation AI. The providing unit provides the converted voice to the user through the speaker 240 of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0145] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0146] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0147] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0148] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0149] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0150] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0151] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0152] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0153] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0154] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0155] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0156] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0157] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0158] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0159] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0160] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0161] Each of the multiple elements described above, including the recording unit, transmission unit, conversion unit, and providing unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the recording unit records the user's voice using the microphone 238 of the robot 414. The transmission unit transmits the recorded data to the data processing unit 12 via the communication I / F 44 of the robot 414. The conversion unit is implemented by the specific processing unit 290 of the data processing unit 12 and converts the user's voice into another voice using a generation AI. The providing unit provides the converted voice to the user through the speaker 240 of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0162] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0163] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0164] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0165] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0166] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0167] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0168] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0169] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0170] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0171] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0172] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0173] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0174] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0175] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0176] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0177] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0178] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0179] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0180] (Note 1) A recording unit that records user voices, A transmission unit that transmits the data recorded by the recording unit to a generating AI, A conversion unit analyzes the data transmitted by the aforementioned transmission unit and converts it into a different voice, The system includes a providing unit that provides the voice converted by the conversion unit to the user. A system characterized by the following features. (Note 2) The aforementioned recording unit is Record the user's voice using a smartphone app. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned transmitting unit Send data to a generative AI in the cloud. The system described in Appendix 1, characterized by the features described herein. (Note 4) The conversion unit is The voice is transformed by changing specific parameters in the audio data. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned supply unit is, Provide the converted voice to the user. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned recording unit is It estimates the user's emotions and adjusts the recording start time based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned recording unit is During recording, the tone and pitch of the user's voice are adjusted in real time. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned recording unit is Add a filtering function that automatically removes background noise during recording. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned recording unit is It estimates the user's emotions and adjusts the length of the recording based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned recording unit is During recording, the system suggests the optimal recording environment based on the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned recording unit is During recording, the system automatically applies optimal recording settings by referencing the user's past recording data. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned transmitting unit It estimates the user's emotions and adjusts the timing of data transmission based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned transmitting unit During data transmission, the data compression ratio is dynamically adjusted to optimize transmission speed. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned transmitting unit When sending data, the system selects the optimal transmission route based on the load status of the destination server. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned transmitting unit It estimates the user's emotions and determines the priority of transmitted data based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned transmitting unit During data transmission, the system monitors the user's network connection status in real time and selects the optimal transmission method. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned transmitting unit When transmitting data, the transmission method is adjusted based on the battery level of the user's device. The system described in Appendix 1, characterized by the features described herein. (Note 18) The conversion unit is It estimates the user's emotions and adjusts the tone and pitch of the converted voice based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The conversion unit is During the conversion process, an algorithm is applied that preserves the characteristics of the user's voice while converting it to a different voice. The system described in Appendix 1, characterized by the features described herein. (Note 20) The conversion unit is Add a feature to convert the voice to one that corresponds to a specific language or dialect during the conversion process. The system described in Appendix 1, characterized by the features described herein. (Note 21) The conversion unit is It estimates the user's emotions and adjusts the length of the converted voice based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The conversion unit is During conversion, the system automatically applies the optimal conversion parameters by referring to the user's past conversion history. The system described in Appendix 1, characterized by the features described herein. (Note 23) The conversion unit is During the conversion process, the conversion algorithm is selected according to the intended use of the user's voice. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned supply unit is, It estimates the user's emotions and adjusts the voice playback method based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned supply unit is, When the product is released, an audio filtering function will be added to ensure optimal sound quality for the user's device. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned supply unit is, When providing the service, the optimal delivery method is selected by referring to the user's past usage history. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned supply unit is, It estimates the user's emotions and determines the priority of the voice messages to deliver based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned supply unit is, At the time of delivery, the optimal delivery method will be selected according to the type of device the user is using. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned supply unit is, At the time of delivery, the optimal delivery method will be selected based on the user's network connectivity status. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0181] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A recording unit that records user voices, A transmission unit that transmits the data recorded by the recording unit to the generating AI, A conversion unit analyzes the data transmitted by the aforementioned transmission unit and converts it into a different voice, The system includes a providing unit that provides the voice converted by the conversion unit to the user. A system characterized by the following features.
2. The aforementioned recording unit is Record the user's voice using a smartphone app. The system according to feature 1.
3. The aforementioned transmitting unit Send data to a generating AI in the cloud. The system according to feature 1.
4. The conversion unit is The voice is transformed by changing specific parameters in the audio data. The system according to feature 1.
5. The aforementioned supply unit is, Provide the converted voice to the user. The system according to feature 1.
6. The aforementioned recording unit is It estimates the user's emotions and adjusts the recording start time based on the estimated emotions. The system according to feature 1.
7. The aforementioned recording unit is During recording, the tone and pitch of the user's voice are adjusted in real time. The system according to feature 1.
8. The aforementioned recording unit is Add a filtering function that automatically removes background noise during recording. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A