system
The system addresses voice interruptions and quality degradation by using AI to generate and integrate complementary audio with original audio in real-time, ensuring high-quality and uninterrupted calls with privacy protection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Voice interruptions and quality degradation in wireless communication terminals due to base station failures or environmental changes cause discomfort and stress, especially in business and emergency communications, leading to difficulties in conveying call content effectively.
A data processing system that utilizes AI-based audio analysis to detect interruptions and quality degradation in real-time, generates complementary audio to fill gaps, and combines it with original audio using edge computing for low-latency processing, ensuring high-quality and uninterrupted calls while protecting user privacy.
The system provides seamless and high-quality voice communication by filling audio gaps and minimizing latency, enhancing user experience and ensuring privacy through real-time data encryption and deletion.
Smart Images

Figure 2026070249000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a call using a wireless communication terminal such as a mobile phone, voice interruption or noise caused by a base station failure or a change in the reception environment leads to a deterioration in call quality, causing problems of discomfort and stress to users. Also, since this makes it difficult to convey the call content, it is regarded as a serious problem especially in business scenes and emergency communications. It is an object to solve such instability of voice quality and provide a high-quality and uninterrupted call experience.
Means for Solving the Problems
[0005] This invention provides a data processing means that analyzes audio data received during a call in real time and detects interruptions or degradation of audio quality. This data processing means includes an audio generation means that uses AI-based audio analysis technology to analyze the context of the received audio and the characteristics of the speaker, thereby generating audio that complements the interrupted or degraded parts. Furthermore, an audio mixing means is provided that combines the generated complementary audio with the original audio data and outputs it in a natural form, providing high-quality audio that does not cause discomfort to the user. In addition, by making full use of edge computing technology, real-time processing of audio data is achieved with low latency, creating a call environment that does not cause users to feel any delay. All communication data is encrypted and deleted immediately after processing is completed, so user privacy is also ensured. As a result, this invention proposes a system that simultaneously reduces discomfort during calls and provides clear communication.
[0006] "Voice data" refers to human voices collected from communication devices such as mobile phones, converted into a digital format, and used as data for the content of phone calls.
[0007] "Data processing means" refers to components of a system or device that have the function of receiving and analyzing audio data and performing supplementary processing as needed.
[0008] "Speech generation means" refers to the process or technology of generating new speech from analyzed speech data to supplement any missing or degraded parts.
[0009] "Audio mixing means" refers to technologies and devices that combine generated complementary audio with the original audio data to produce a consistent and natural output.
[0010] "Network communication methods" refer to a set of technologies and protocols for exchanging data between communication-capable terminals, with the aim of minimizing latency.
[0011] "Data protection measures" refer to technical measures taken to protect audio data from unauthorized external access and tampering, such as data encryption and deletion after processing.
[0012] Edge computing technology is a technique that processes data not on a central server, but on a distributed network edge closer to where the data is generated. This reduces processing delays and improves real-time performance. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a system for improving the quality of voice communication and is implemented as follows.
[0035] Terminal operation
[0036] The user initiates a call without requiring any special actions. The device captures the user's voice using the microphone and generates audio data. The generated audio data is then transmitted to the server via the communication line.
[0037] Server operation
[0038] The server receives audio data transmitted from the terminal. This audio data is analyzed in real time by data processing equipment. During the analysis process, AI technology is used to evaluate the context of the conversation and the quality of the audio, and to determine whether there are any interruptions or noise.
[0039] If necessary, the server uses a speech generation mechanism to generate new audio data to fill in the gaps in the audio. After the supplementary audio is generated, a speech mixing mechanism combines it with the original audio data to create a consistent audio stream.
[0040] Final output and user experience
[0041] This audio stream is retransmitted to the device with low latency and high quality. The sound is played through the device's speaker, allowing the user to enjoy a seamless and smooth conversation experience.
[0042] Data Protection
[0043] All data handled during a call is encrypted, so it cannot be illegally obtained during communication. The server protects user privacy by promptly deleting all processed data after processing the voice data is complete.
[0044] Specific example
[0045] For example, when a user makes a call in the basement of a building, the signal may be unstable and the audio may be interrupted. The user's device sends the captured audio data to a server. The server uses AI to fill in the audio segments lost due to signal interruptions and generates consistent audio data. This provides the other party with a conversation that does not feel interrupted. The user can experience high-quality calls even in places with poor communication environments, such as basements. This systematic functionality provides added value to the user in this embodiment of the invention.
[0046] The following describes the processing flow.
[0047] Step 1:
[0048] The device captures the user's spoken voice using a microphone and converts it into digital audio data. It then divides the converted audio data into packets in real time and prepares them for transmission to the server.
[0049] Step 2:
[0050] The server receives voice data packets sent from the terminal. It stores the received data in a processing buffer and prepares it for analysis.
[0051] Step 3:
[0052] The server passes the audio data stored in the buffer to the AI engine for audio quality analysis. The analysis evaluates the spectral information of the audio to detect interruptions and noise.
[0053] Step 4:
[0054] When the AI engine detects a gap, the server analyzes the context of the preceding audio data. It then infers the audio data that needs to be supplemented and synthesizes the supplementary audio using an audio generation method.
[0055] Step 5:
[0056] The server passes the synthesized supplemental audio and the original audio data to the audio mixing device to generate a single, consistent audio stream.
[0057] Step 6:
[0058] The server compresses the generated audio stream and sends it to the terminal using edge computing technology, minimizing latency.
[0059] Step 7:
[0060] The device receives the audio stream sent from the server and plays it back to the user through the speaker. This allows the user to experience a natural, continuous conversation.
[0061] Step 8:
[0062] The server immediately deletes all processed audio data once the call ends, protecting the user's communication content.
[0063] (Example 1)
[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0065] Conventional voice communication systems often experienced interruptions in poor communication environments, making it difficult to maintain smooth conversations. Furthermore, security concerns included the risk of unauthorized data acquisition during communication, resulting in insufficient protection of user privacy. Technologies are needed to address these issues of voice interruptions and quality degradation.
[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0067] In this invention, the server includes information processing means for receiving and analyzing voice information in real time, voice generation means for detecting interruptions and quality degradation in the analyzed voice and generating supplementary voice, and voice integration means for combining the generated supplementary voice with the original voice information and outputting it. This makes it possible to provide consistent voice communication without interruption even when the communication environment deteriorates, and to improve the user's conversation experience.
[0068] "Voice information" refers to data obtained by capturing the user's speech with a microphone and converting it into digital data.
[0069] "Real-time" means processing audio data instantly, resulting in data transmission and processing at a level where users do not perceive any delay.
[0070] "Information processing means" refers to system components that analyze received audio information and detect interruptions and noise.
[0071] "Speech generation means" refers to a technology that has the function of generating and supplementing lost audio portions based on analyzed data.
[0072] "Audio integration means" refers to a process for combining the generated supplementary audio with the original audio data to output a continuous audio stream.
[0073] "Communication method" refers to network technology that has the function of transmitting encrypted voice data with low latency.
[0074] "Information protection measures" refer to a system that protects privacy by encrypting audio data and deleting it immediately after processing is complete.
[0075] "Acquisition means" refers to a part of a system that captures the user's voice using a microphone and generates audio data.
[0076] This invention is a system for improving the quality of voice communication, and is particularly aimed at providing high-quality voice even in unstable communication environments. When a user initiates a call, the terminal uses its built-in microphone to acquire voice and converts it into data as digital voice information. This voice information is then encoded and sent to a server.
[0077] The server processes the received audio information in real time. AI technology is used as the information processing method to detect audio interruptions and noise. Using a generative AI model, generated audio is created that fills in the interrupted audio portions based on the analysis results. Furthermore, the server uses an audio integration method to integrate the supplemented audio with the original audio information, generating a seamless audio stream.
[0078] The generated audio stream is retransmitted to the terminal using a low-latency, secure communication method. The terminal decodes this audio stream and plays it back through the speaker, allowing the user to experience uninterrupted, high-quality audio. All audio information is encrypted and immediately deleted by the server once the call ends, ensuring privacy.
[0079] For example, even if the user is underground, this system allows for a call without any noticeable audio interruption. An example of a prompt from the generated AI model might be, "Please explain the steps to analyze the audio data from the underground call and fill in any lost segments."
[0080] In this way, this invention can provide users with a better voice communication environment.
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The user initiates a call using the terminal. The terminal uses its built-in microphone to capture the user's voice and converts the analog audio signal into digital data. At this time, the terminal performs sampling to prepare the audio data as packets. The input is the user's analog voice, and the output is digital audio data.
[0084] Step 2:
[0085] The terminal encrypts the generated digital audio data and sends it to the server using a secure communication protocol. By utilizing Wi-Fi or mobile data communication for data transmission, low-latency data transmission is achieved. The input is unencrypted digital audio data, and the output is encrypted audio data.
[0086] Step 3:
[0087] The server receives audio data transmitted from the terminal. At this stage, the server uses information processing tools to decrypt the audio data, filter out noise, and perform speech recognition. Through information analysis, it identifies any interruptions or degradation in audio quality. The input is encrypted audio data, and the output is the analyzed audio data.
[0088] Step 4:
[0089] Based on the analysis results, the server uses a generative AI model to generate supplementary audio for the missing sections. The generative AI model performs natural-sounding supplementation by considering past conversational context and speech characteristics. The input is the analyzed audio data and its contextual information, and the output is the supplemented audio data.
[0090] Step 5:
[0091] The server mixes the supplemented audio data with the original audio data. A consistent audio stream is constructed using an audio integration mechanism. This process optimizes the temporal consistency and intelligibility of the audio. The input is the supplemented and original audio data, and the output is the integrated audio stream.
[0092] Step 6:
[0093] The server encrypts the integrated audio stream and sends it to the terminal. For data protection, communication is again conducted through a secure protocol. The input is the integrated audio stream, and the output is the encrypted audio stream.
[0094] Step 7:
[0095] The device decodes the received audio stream and plays it through the speaker. The device performs digital signal processing to play the audio at optimal volume and quality. This allows the user to enjoy a high-quality, uninterrupted audio experience. The input is an encrypted audio stream, and the output is the analog audio being played back.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] In communication between customers and delivery personnel, there is a need to achieve smooth communication with high-quality voice even in environments with unstable radio wave conditions. In particular, there is a lack of technology to compensate for voice interruptions and quality degradation and provide a clear voice experience in places where communication environments tend to deteriorate, such as underground parking lots and recessed areas of buildings.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes information processing means for receiving and analyzing voice data in real time, voice generation means for detecting interruptions and quality degradation in the analyzed voice and generating supplementary voice, and communication improvement means for improving voice quality in an environment specifically designed for communication between customers and delivery personnel. This enables high-quality voice calls with minimal voice interruptions and noise, even in unstable communication environments.
[0101] "Audio data" refers to audio signals converted into digital data, and is a data format that includes the content of the audio.
[0102] "Information processing means" refers to a device or program for receiving and analyzing audio data in real time, primarily for analyzing audio interruptions and quality.
[0103] "Speech generation means" refers to a technology or device that detects interruptions in analyzed speech data and generates new speech data to fill in the missing parts.
[0104] "Speech synthesis means" refers to a device or technology that combines generated supplementary speech with the original speech data to output consistent speech.
[0105] "Network communication means" refers to communication technology or equipment for transmitting voice data with low latency and ensuring real-time performance.
[0106] "Communication improvement measures" refer to technologies or methods for improving the quality of voice communication between customers and delivery personnel.
[0107] "Data protection measures" refer to technologies or devices that ensure privacy by encrypting voice data and deleting it immediately after processing.
[0108] Edge computing technology is a computing technique that performs data processing on devices at the network edge rather than on the server side, promoting real-time performance and low latency.
[0109] This invention is a system for improving the quality of voice communication, aiming to maintain high quality voice communication between customers and delivery personnel. The system primarily includes multiple means for implementing technologies that receive, analyze, and supplement voice data.
[0110] First, the device captures the user's voice using a microphone and sends that audio data to the server in real time. A smartphone is used as the device, and a software library such as FFmpeg is used for audio capture.
[0111] The server analyzes the received audio data using an AI analysis engine such as TENSORFLOW®. This analysis detects audio interruptions and quality degradation while considering the audio context and communication environment. If necessary, the server generates new audio data via an audio generation system to ensure consistent audio.
[0112] The generated supplementary audio is combined with the original audio by a speech synthesis system and reconstructed as high-quality audio data. During this process, the audio data is securely protected using encryption technologies such as OpenSSL.
[0113] The synthesized speech data is retransmitted to the terminal in real time via network communication, providing users with a smooth voice experience. This system enables seamless communication between customers and delivery personnel, even in environments with unstable radio waves, such as underground parking lots.
[0114] For example, when a delivery person makes a phone call to a customer from an underground parking lot, the audio may normally cut out. However, by using this technology, the server can fill in the gaps and deliver a clear audio message to the customer.
[0115] An example of a prompt message is: "The delivery person is calling the customer to confirm the order. However, the signal is poor and the audio is frequently interrupted. How can AI be used to improve the audio quality?"
[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0117] Step 1:
[0118] The user's device captures their voice using a microphone. This voice input is converted into a digital signal and sent to the server in real time. FFmpeg is used for voice capture.
[0119] Step 2:
[0120] The server begins processing digital audio data received from the user's terminal. The server analyzes this data using an AI analysis engine such as TensorFlow to detect the context of the audio, noise due to the communication environment, and audio interruptions. The output of the analysis includes the locations of the interruptions and an evaluation of the audio quality.
[0121] Step 3:
[0122] If the analysis detects a gap in the audio, the server uses AI technology to fill in the missing audio segments. This process utilizes a generative AI model to generate new audio data that doesn't sound interrupted. This generated audio data becomes the output.
[0123] Step 4:
[0124] Next, the server uses speech synthesis to combine the generated supplementary speech with the original speech data. This mixing produces a consistent speech stream. The output is reconstructed speech data that sounds smooth.
[0125] Step 5:
[0126] The server retransmits the reconstructed audio stream to the end user's device with low latency. During this process, encryption technologies such as OpenSSL are used to securely protect the data during transmission. The user's device receives this audio stream and plays it through its speakers, providing a clear audio experience.
[0127] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0128] This invention is a system for voice communication, such as telephone calls, that compensates for interruptions in voice data while recognizing the user's emotions to make the conversation more natural.
[0129] Terminal operation
[0130] The user initiates a call in the usual manner. The device captures the user's voice using the microphone and generates audio data in digital format. This audio data is transmitted to the server in real time.
[0131] Server operation
[0132] The server receives audio data transmitted from the terminal, stores it in a buffer, and begins analyzing the data. In the initial stages of analysis, the server uses an AI engine to detect audio interruptions and noise. It also uses an emotion engine to analyze and recognize the user's emotions in real time.
[0133] The server generates supplementary audio as needed. During this process, the analysis results of the emotion engine are incorporated to create supplementary audio with a tone and intonation that reflects the user's emotions. The audio generation method performs appropriate supplementation while considering the context of the conversation and the user's emotions. The generated supplementary audio is then combined with the original audio data by an audio mixing method.
[0134] Final output and user experience
[0135] The audio stream, including supplementary voice, sent from the server is received by the device and played back to the user through the speaker. This allows the user to experience a natural conversation that is seamless and appropriately reflects their emotions.
[0136] Data Protection
[0137] The server encrypts all data during transmission and stores it securely, and immediately deletes the voice data after processing, thus protecting user privacy.
[0138] Specific example
[0139] For example, when a user calls a friend during an emotional moment, their voice may be transmitted with a tremor. The emotion engine recognizes this emotional state from the user's voice data and adjusts the tone and intonation of the voice. Based on this information, the server fills in the gaps to match the emotion, allowing the recipient of the call to understand the user's emotions more clearly and enabling more effective communication. This invention results in high-quality voice calls that take the user's emotions into account, improving the user experience.
[0140] The following describes the processing flow.
[0141] Step 1:
[0142] The device captures the user's spoken voice using a microphone and converts it into digital audio data. This audio data is then divided into packets in real time and prepared for transmission to the server over the network.
[0143] Step 2:
[0144] The server receives voice data packets sent from the terminal. The received data is temporarily stored in a buffer and organized for analysis.
[0145] Step 3:
[0146] The server passes the audio data stored in the buffer to the AI engine. The AI engine analyzes the spectral information of the audio signal to detect interruptions and noise in the audio.
[0147] Step 4:
[0148] The server inputs voice data into the emotion engine. The emotion engine identifies emotional states such as stress, joy, and anger from the user's voice and generates numerical emotion parameters as a result of the analysis.
[0149] Step 5:
[0150] The server integrates the analysis results from the AI engine and the emotion engine, and uses a speech generation method to fill in the gaps in the audio. The supplementary audio is synthesized as speech with tone and intonation adjusted considering the context and emotion parameters.
[0151] Step 6:
[0152] The server inputs the supplemented audio data into the audio mixing system and processes it to seamlessly blend it with the original audio data. This results in a consistent audio stream.
[0153] Step 7:
[0154] The server transmits the completed audio stream to the terminal via edge computing with minimal latency. All communication is encrypted during this process.
[0155] Step 8:
[0156] The device receives the audio stream sent from the server and plays it back to the user through the speaker. This allows the user to have a seamless and natural conversation experience.
[0157] Step 9:
[0158] The server immediately deletes all audio data upon the end of a call, protecting the privacy of the user's communication.
[0159] (Example 2)
[0160] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0161] In voice communication, poor communication environments and data interruptions can degrade voice quality, leading to problems with accurately conveying the speaker's intentions and emotions. Furthermore, the inability to properly reflect emotions makes natural conversation difficult. In addition, the risk of personal information leakage necessitates the protection of privacy, which is a crucial issue.
[0162] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0163] In this invention, the server includes information processing means for receiving and analyzing voice information in real time, voice generation means for detecting interruptions and quality degradation of voice and generating supplementary voice, and emotion analysis means for analyzing the speaker's emotions and adjusting the tone and intonation of the supplementary voice based on the analysis results. This makes it possible to achieve natural voice communication that appropriately reflects the speaker's emotions without interruptions in voice information. Furthermore, since the voice information is encrypted, privacy is also protected.
[0164] "Audio information" refers to audio data, including human voices, acquired through input devices such as microphones.
[0165] "Information processing means" are components for analyzing and transforming input data, and are responsible for processing audio information in real time.
[0166] A "speech generation means" is a technical element that has the function of generating speech to compensate for interrupted or degraded speech and to enable natural communication.
[0167] A "speech synthesis means" is a component that combines the generated supplementary speech with the original speech information to produce an output.
[0168] "Communication methods" refer to network technologies and protocols used to transmit data with low latency.
[0169] "Data protection measures" refer to technical measures taken to prevent the leakage of personal information and protect privacy, such as encrypting voice information and deleting data after processing.
[0170] "Emotion analysis means" refers to a technological element that identifies the speaker's emotions from audio information and adjusts the tone and intonation of supplementary audio based on the results.
[0171] "Edge computing technology" refers to a technology that performs processing in a distributed computing environment to achieve low latency, and enables real-time data processing on edge devices.
[0172] This invention provides a system that offers natural conversation by reinforcing interruptions in voice information during voice communication while recognizing the user's emotions in real time. The following describes the configuration for implementing this system.
[0173] Terminal operation
[0174] The user initiates a call using the terminal. The terminal uses its built-in microphone to capture the user's voice information. This voice information is converted into a digital format and transmitted to the server in real time. In this process, the terminal uses standard communication protocols, such as TCP / IP, to transmit data.
[0175] Server operation
[0176] The server receives audio information transmitted from the terminal. The server is equipped with information processing capabilities, and the audio information is stored in a buffer and analyzed. During analysis, the server uses an AI engine to detect audio interruptions and noise. This includes a process that effectively filters out noise using a generative AI model. In addition, emotion analysis capabilities are used to identify the user's emotions from the audio. The server estimates whether the speaker is experiencing emotions such as "joy," "sadness," or "anger," and based on the results, it uses speech generation capabilities to generate supplementary audio. This supplementary audio is composed of a tone and intonation appropriate to the user's emotions and is combined with the original audio information by speech synthesis capabilities.
[0177] Specific example
[0178] For example, consider a scenario where a user shares their thoughts about essential items with a friend. If the user's voice sounds shaky, the emotion analysis system can recognize this emotion and generate a reassuring, calming tone of voice to complement it. This allows the recipient to understand the user's emotions more deeply, enriching the communication.
[0179] Privacy protection
[0180] The server has data protection measures in place to encrypt voice information and store it securely for temporary use. After processing is complete, the voice information is immediately deleted, ensuring user privacy.
[0181] Example of a prompt
[0182] "Please introduce a system that analyzes a user's emotions in real time during a phone call and fills in gaps in the conversation in a way that reflects those emotions."
[0183] This invention is expected to make voice communication more natural and enriching, thereby improving the user experience.
[0184] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0185] Step 1:
[0186] The user initiates a call using the terminal. The terminal acquires audio information through the microphone and converts the analog audio into a digital format. For example, a sampling rate of 48kHz is used for this conversion. The digitized audio information is transmitted to the server in real time via the TCP / IP protocol. The input is analog audio, and the output is digital audio data.
[0187] Step 2:
[0188] The server receives digital audio data transmitted from the terminal. The server stores this audio data in a buffer and begins analysis using information processing equipment. Here, the input is digital audio data, and the output is audio information prepared for analysis.
[0189] Step 3:
[0190] The server uses an AI engine to detect interruptions and noise in the audio data. During this process, a generative AI model is utilized to perform filtering and audio correction. The input is refined audio information, and the output is audio data with reduced noise.
[0191] Step 4:
[0192] The server uses emotion analysis tools to analyze the user's emotions in real time from the audio data. Feature quantities (e.g., pitch and speed) are extracted and input into an emotion recognition model. The model's output is an emotion label such as "joy," "sadness," or "anger."
[0193] Step 5:
[0194] The server generates supplementary speech based on the results of emotion analysis. The speech generation method generates speech with tone and intonation that corresponds to the user's emotions. A generation AI model is used in this process. The input is emotion labels, and the output is supplementary speech.
[0195] Step 6:
[0196] The server mixes the generated supplementary audio with the original audio data using a speech synthesis system. This creates an audio stream in which the missing parts are naturally filled in. The input is the supplementary audio and the original audio data, and the output is the mixed audio stream.
[0197] Step 7:
[0198] The server sends the mixed audio stream to the terminal in real time. The terminal reconstructs the received audio stream and plays it through the speaker. This allows the user to enjoy a seamless, natural conversation. The input is the mixed audio stream, and the output is the audio played from the speaker.
[0199] (Application Example 2)
[0200] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0201] In food delivery and other forms of communication, emotions may not be fully conveyed during voice calls. While users want to sense the emotions of delivery drivers and operators, current voice communication systems often make this difficult due to interruptions and quality degradation. This challenge hinders improvements in user experience and reliability.
[0202] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0203] In this invention, the server includes information processing means for receiving and analyzing audio data in real time, audio generation means for detecting interruptions and quality degradation in the analyzed audio and generating supplementary audio, and emotion recognition means for recognizing the speaker's emotional state using emotion analysis technology and adjusting the tone and intonation of the supplementary audio based on that. As a result, the user can experience natural conversations that are free from interruptions and that appropriately reflect emotions.
[0204] "Audio data" refers to information that represents sound in digital format and is processed in real time.
[0205] "Information processing means" refers to the capabilities of a system that receives and analyzes audio data, and includes devices and programs that play a role in detecting interruptions or degradation of audio quality.
[0206] A "speech generation means" is a mechanism for synthesizing and generating new speech to compensate for detected interruptions or quality degradation.
[0207] "Audio mixing means" refers to a processing function that appropriately combines the original audio data with the generated supplementary audio and outputs it.
[0208] "Communication methods" refer to network technologies and equipment for transmitting and receiving voice data and its analysis results with low latency.
[0209] "Data protection measures" refer to technologies and methods that encrypt voice data and immediately delete it after transmission or processing to ensure privacy.
[0210] "Emotion identification means" refers to a technology that analyzes the speaker's emotional state from audio data and uses that information to add tone and intonation to the supplementary audio.
[0211] In this invention, first, the terminal captures the user's voice in real time. The voice is converted into a digital format and transmitted to a server via network communication means. The server receives and analyzes the voice data using information processing means to detect interruptions and quality degradation.
[0212] The emotion recognition means allows the server to recognize the emotional state within the audio data, and the audio generation means generates supplementary audio with adjusted tone and intonation based on the analysis results. This supplementary audio is then combined with the original audio data by the audio mixing means.
[0213] The communication method achieves particularly low latency data transmission, allowing users to enjoy a natural voice experience without experiencing interruptions in conversation. Data protection measures encrypt voice data to safeguard user privacy and delete it immediately after processing.
[0214] One concrete example of implementation is in food delivery services, where, when a delivery person converses with a user, supplementary voices are generated that convey the delivery person's enthusiasm, allowing the user to experience a more friendly and approachable communication.
[0215] An example of a prompt message would be: "Analyze the emotions during this call and generate audio that will make the user feel more excited about the delivery person and improve their satisfaction."
[0216] In terms of specific hardware and software, high-sensitivity microphones are used for voice capture, high-speed network technology for data transmission, and AI and emotion engines for voice analysis and generation. The configuration may include cloud services as servers or edge devices, enabling real-time processing and data protection.
[0217] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0218] Step 1:
[0219] The terminal captures the user's voice using a high-sensitivity microphone. It converts the input analog audio into a digital format and prepares it for transmission to a server over the network. This process outputs digitized audio data.
[0220] Step 2:
[0221] The server receives digital audio data from terminals in real time using network communication means. The received data is passed to an information processing system and stored in a buffer. The output at this stage is audio data ready for analysis.
[0222] Step 3:
[0223] The server analyzes the stored audio data using information processing tools. An AI engine is used to detect interruptions and quality degradation within the audio data. The output generated by this analysis includes information about audio interruptions.
[0224] Step 4:
[0225] Using emotion recognition technology, the server recognizes the speaker's emotional state based on the analyzed audio data. This process uses an emotion analysis model to perform data transformations for detecting tone and intonation. The output is an analysis result that possesses emotional characteristics.
[0226] Step 5:
[0227] The server uses speech generation technology to generate supplementary speech based on the analysis results. In this process, a generation AI model is used to incorporate tones and intonations that blend naturally with the original conversation and to fill in any gaps. The generated supplementary speech is then output.
[0228] Step 6:
[0229] The server's audio mixing mechanism combines the original audio data with the supplementary audio. The mixing process generates a natural audio stream with interruptions corrected. The output of this process is the integrated audio data.
[0230] Step 7:
[0231] The server sends the generated audio stream back to the terminal. Low latency is crucial for the data transmitted through the communication method. The output of this process is an audio stream ready for playback.
[0232] Step 8:
[0233] The device plays the received audio stream to the user through its speaker. This allows the user to experience a more natural and enhanced conversation, increasing engagement. The output here is the audio that is played back.
[0234] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0235] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0236] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0237] [Second Embodiment]
[0238] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0239] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0240] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0241] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0242] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0243] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0244] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0245] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0246] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0247] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0248] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0249] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0250] This invention is a system for improving the quality of voice communication and is implemented as follows.
[0251] Terminal operation
[0252] The user initiates a call without requiring any special actions. The device captures the user's voice using the microphone and generates audio data. The generated audio data is then transmitted to the server via the communication line.
[0253] Server operation
[0254] The server receives audio data transmitted from the terminal. This audio data is analyzed in real time by data processing equipment. During the analysis process, AI technology is used to evaluate the context of the conversation and the quality of the audio, and to determine whether there are any interruptions or noise.
[0255] If necessary, the server uses a speech generation mechanism to generate new audio data to fill in the gaps in the audio. After the supplementary audio is generated, a speech mixing mechanism combines it with the original audio data to create a consistent audio stream.
[0256] Final output and user experience
[0257] This audio stream is retransmitted to the device with low latency and high quality. The sound is played through the device's speaker, allowing the user to enjoy a seamless and smooth conversation experience.
[0258] Data Protection
[0259] All data handled during a call is encrypted, so it cannot be illegally obtained during communication. The server protects user privacy by promptly deleting all processed data after processing the voice data is complete.
[0260] Specific example
[0261] For example, when a user makes a call in the basement of a building, the signal may be unstable and the audio may be interrupted. The user's device sends the captured audio data to a server. The server uses AI to fill in the audio segments lost due to signal interruptions and generates consistent audio data. This provides the other party with a conversation that does not feel interrupted. The user can experience high-quality calls even in places with poor communication environments, such as basements. This systematic functionality provides added value to the user in this embodiment of the invention.
[0262] The following describes the processing flow.
[0263] Step 1:
[0264] The device captures the user's spoken voice using a microphone and converts it into digital audio data. It then divides the converted audio data into packets in real time and prepares them for transmission to the server.
[0265] Step 2:
[0266] The server receives voice data packets sent from the terminal. It stores the received data in a processing buffer and prepares it for analysis.
[0267] Step 3:
[0268] The server passes the audio data stored in the buffer to the AI engine for audio quality analysis. The analysis evaluates the spectral information of the audio to detect interruptions and noise.
[0269] Step 4:
[0270] When the AI engine detects a gap, the server analyzes the context of the preceding audio data. It then infers the audio data where completion is needed and synthesizes the completion audio using the audio generation method.
[0271] Step 5:
[0272] The server passes the synthesized supplemental audio and the original audio data to the audio mixing device to generate a single, consistent audio stream.
[0273] Step 6:
[0274] The server compresses the generated audio stream and sends it to the terminal using edge computing technology, minimizing latency.
[0275] Step 7:
[0276] The terminal receives the audio stream transmitted from the server and plays it back to the user through the speaker. As a result, the user can experience a natural and continuous conversation.
[0277] Step 8:
[0278] When the call ends, the server immediately deletes all the processed audio data to protect the user's communication content.
[0279] (Example 1)
[0280] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0281] In a conventional voice communication system, the voice may be interrupted when the communication environment deteriorates, making it difficult to maintain a smooth conversation. Also, as a security issue, there is a risk that data may be illegally obtained during communication, and the user's privacy is not sufficiently protected. Technologies for dealing with voice interruptions and quality degradation are required.
[0282] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0283] In this invention, the server includes information processing means for receiving and analyzing voice information in real time, voice generation means for detecting interruptions and quality degradation of the analyzed voice and generating complementary voice, and voice integration means for combining the generated complementary voice with the original voice information and outputting it. As a result, even when the communication environment deteriorates, consistent voice communication can be provided without interruption, and the user's conversation experience can be improved.
[0284] [[ID=三十一]] "Voice information" is what is obtained by a microphone from the user's speech and converted into digital data.
[0285] "Real-time" means to process voice data immediately, enabling data transmission and processing at a level where users do not feel latency.
[0286] "Information processing means" is a system component for analyzing received voice information and detecting breaks and noise.
[0287] "Voice generation means" is a technology with the function of generating and complementing lost voice parts based on the analyzed data.
[0288] "Voice integration means" is a process for combining the generated complementary voice with the original voice data and outputting a continuous voice stream.
[0289] "Communication means" is a network technology with the function of transferring encrypted voice data with low latency.
[0290] "Information protection means" is a mechanism for protecting privacy by encrypting voice data and immediately deleting it when the processing is completed.
[0291] "Acquisition means" is a part of the system that captures the user's voice using a microphone and generates voice data.
[0292] This invention is a system for improving the quality of voice communication, aiming to provide high-quality voice especially in unstable communication environments. When the user starts a call, the terminal uses the built-in microphone to acquire the voice and converts it into data as digital voice information. This voice information is encoded and transmitted to the server.
[0293] The server processes the received audio information in real time. AI technology is used as the information processing method to detect audio interruptions and noise. Using a generative AI model, generated audio is created that fills in the interrupted audio portions based on the analysis results. Furthermore, the server uses an audio integration method to integrate the supplemented audio with the original audio information, generating a seamless audio stream.
[0294] The generated audio stream is retransmitted to the terminal using a low-latency, secure communication method. The terminal decodes this audio stream and plays it back through the speaker, allowing the user to experience uninterrupted, high-quality audio. All audio information is encrypted and immediately deleted by the server once the call ends, ensuring privacy.
[0295] For example, even if the user is underground, this system allows for a call without any noticeable audio interruption. An example of a prompt from the generated AI model might be, "Please explain the steps to analyze the audio data from the underground call and fill in any lost segments."
[0296] In this way, this invention can provide users with a better voice communication environment.
[0297] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0298] Step 1:
[0299] The user initiates a call using the terminal. The terminal uses its built-in microphone to capture the user's voice and converts the analog audio signal into digital data. At this time, the terminal performs sampling to prepare the audio data as packets. The input is the user's analog voice, and the output is digital audio data.
[0300] Step 2:
[0301] The terminal encrypts the generated digital audio data and transmits it to the server using a secure communication protocol. When sending the data, by leveraging Wi-Fi or mobile data communication, data transmission with low latency is achieved. The input is the digital audio data before encryption, and the output is the encrypted audio data.
[0302] Step 3:
[0303] The server receives the audio data sent from the terminal. At this stage, the server utilizes information processing means to decrypt the audio data, perform noise filtering and speech recognition. Through information analysis, interruptions and quality degradation of the speech are discriminated. The input is the encrypted audio data, and the output is the analyzed audio data.
[0304] Step 4:
[0305] Based on the analysis results, the server utilizes a generation AI model to generate complementary audio for the interrupted parts. The generation AI model performs natural complementation considering past conversation contexts and speech characteristics. The input is the analyzed audio data and its context information, and the output is the complemented audio data.
[0306] Step 5:
[0307] The server mixes the complemented audio data with the original audio data. Using audio integration means, a consistent audio stream is constructed. In this process, the temporal consistency and audibility of the audio are optimized. The input is the complemented audio data and the original audio data, and the output is the integrated audio stream.
[0308] Step 6:
[0309] The server encrypts the integrated audio stream and transmits it to the terminal. For data protection, the communication is again carried out through a secure protocol. The input is the integrated audio stream, and the output is the encrypted audio stream.
[0310] Step 7:
[0311] The device decodes the received audio stream and plays it through the speaker. The device performs digital signal processing to play the audio at optimal volume and quality. This allows the user to enjoy a high-quality, uninterrupted audio experience. The input is an encrypted audio stream, and the output is the analog audio being played back.
[0312] (Application Example 1)
[0313] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0314] In communication between customers and delivery personnel, there is a need to achieve smooth communication with high-quality voice even in environments with unstable radio wave conditions. In particular, there is a lack of technology to compensate for voice interruptions and quality degradation and provide a clear voice experience in places where communication environments tend to deteriorate, such as underground parking lots and recessed areas of buildings.
[0315] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0316] In this invention, the server includes information processing means for receiving and analyzing voice data in real time, voice generation means for detecting interruptions and quality degradation in the analyzed voice and generating supplementary voice, and communication improvement means for improving voice quality in an environment specifically designed for communication between customers and delivery personnel. This enables high-quality voice calls with minimal voice interruptions and noise, even in unstable communication environments.
[0317] "Audio data" refers to audio signals converted into digital data, and is a data format that includes the content of the audio.
[0318] "Information processing means" refers to a device or program for receiving and analyzing audio data in real time, primarily for analyzing audio interruptions and quality.
[0319] "Speech generation means" refers to a technology or device that detects interruptions in analyzed speech data and generates new speech data to fill in the missing parts.
[0320] "Speech synthesis means" refers to a device or technology that combines generated supplementary speech with the original speech data to output consistent speech.
[0321] "Network communication means" refers to communication technology or equipment for transmitting voice data with low latency and ensuring real-time performance.
[0322] "Communication improvement measures" refer to technologies or methods for improving the quality of voice communication between customers and delivery personnel.
[0323] "Data protection measures" refer to technologies or devices that ensure privacy by encrypting voice data and deleting it immediately after processing.
[0324] Edge computing technology is a computing technique that performs data processing on devices at the network edge rather than on the server side, promoting real-time performance and low latency.
[0325] This invention is a system for improving the quality of voice communication, aiming to maintain high quality voice communication between customers and delivery personnel. The system primarily includes multiple means for implementing technologies that receive, analyze, and supplement voice data.
[0326] First, the device captures the user's voice using a microphone and sends that audio data to the server in real time. A smartphone is used as the device, and a software library such as FFmpeg is used for audio capture.
[0327] The server analyzes the received audio data using an AI analysis engine such as TensorFlow. This analysis detects audio interruptions and quality degradation while considering the audio context and communication environment. If necessary, the server generates new audio data via an audio generation system to ensure consistent audio.
[0328] The generated supplementary audio is combined with the original audio by a speech synthesis system and reconstructed as high-quality audio data. During this process, the audio data is securely protected using encryption technologies such as OpenSSL.
[0329] The synthesized speech data is retransmitted to the terminal in real time via network communication, providing users with a smooth voice experience. This system enables seamless communication between customers and delivery personnel, even in environments with unstable radio waves, such as underground parking lots.
[0330] For example, when a delivery person makes a phone call to a customer from an underground parking lot, the audio may normally cut out. However, by using this technology, the server can fill in the gaps and deliver a clear audio message to the customer.
[0331] An example of a prompt message is: "The delivery person is calling the customer to confirm the order. However, the signal is poor and the audio is frequently interrupted. How can AI be used to improve the audio quality?"
[0332] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0333] Step 1:
[0334] The user's device captures their voice using a microphone. This voice input is converted into a digital signal and sent to the server in real time. FFmpeg is used for voice capture.
[0335] Step 2:
[0336] The server begins processing digital audio data received from the user's terminal. The server analyzes this data using an AI analysis engine such as TensorFlow to detect the context of the audio, noise due to the communication environment, and audio interruptions. The output of the analysis includes the locations of the interruptions and an evaluation of the audio quality.
[0337] Step 3:
[0338] If the analysis detects a gap in the audio, the server uses AI technology to fill in the missing audio segments. This process utilizes a generative AI model to generate new audio data that doesn't sound interrupted. This generated audio data becomes the output.
[0339] Step 4:
[0340] Next, the server uses speech synthesis to combine the generated supplementary speech with the original speech data. This mixing produces a consistent speech stream. The output is reconstructed speech data that sounds smooth.
[0341] Step 5:
[0342] The server retransmits the reconstructed audio stream to the end user's device with low latency. During this process, encryption technologies such as OpenSSL are used to securely protect the data during transmission. The user's device receives this audio stream and plays it through its speakers, providing a clear audio experience.
[0343] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0344] This invention is a system for voice communication, such as telephone calls, that compensates for interruptions in voice data while recognizing the user's emotions to make the conversation more natural.
[0345] Terminal operation
[0346] The user initiates a call in the usual manner. The device captures the user's voice using the microphone and generates audio data in digital format. This audio data is transmitted to the server in real time.
[0347] Server operation
[0348] The server receives audio data transmitted from the terminal, stores it in a buffer, and begins analyzing the data. In the initial stages of analysis, the server uses an AI engine to detect audio interruptions and noise. It also uses an emotion engine to analyze and recognize the user's emotions in real time.
[0349] The server generates supplementary audio as needed. During this process, the analysis results of the emotion engine are incorporated to create supplementary audio with a tone and intonation that reflects the user's emotions. The audio generation method performs appropriate supplementation while considering the context of the conversation and the user's emotions. The generated supplementary audio is then combined with the original audio data by an audio mixing method.
[0350] Final output and user experience
[0351] The audio stream, including supplementary voice, sent from the server is received by the device and played back to the user through the speaker. This allows the user to experience a natural conversation that is seamless and appropriately reflects their emotions.
[0352] Data Protection
[0353] The server encrypts all data during transmission and stores it securely, and immediately deletes the voice data after processing, thus protecting user privacy.
[0354] Specific example
[0355] For example, when a user calls a friend during an emotional moment, their voice may be transmitted with a tremor. The emotion engine recognizes this emotional state from the user's voice data and adjusts the tone and intonation of the voice. Based on this information, the server fills in the gaps to match the emotion, allowing the recipient of the call to understand the user's emotions more clearly and enabling more effective communication. This invention results in high-quality voice calls that take the user's emotions into account, improving the user experience.
[0356] The following describes the processing flow.
[0357] Step 1:
[0358] The device captures the user's spoken voice using a microphone and converts it into digital audio data. This audio data is then divided into packets in real time and prepared for transmission to the server over the network.
[0359] Step 2:
[0360] The server receives voice data packets sent from the terminal. The received data is temporarily stored in a buffer and organized for analysis.
[0361] Step 3:
[0362] The server passes the audio data stored in the buffer to the AI engine. The AI engine analyzes the spectral information of the audio signal to detect interruptions and noise in the audio.
[0363] Step 4:
[0364] The server inputs voice data into the emotion engine. The emotion engine identifies emotional states such as stress, joy, and anger from the user's voice and generates numerical emotion parameters as a result of the analysis.
[0365] Step 5:
[0366] The server integrates the analysis results from the AI engine and the emotion engine, and uses a speech generation method to fill in the gaps in the audio. The supplementary audio is synthesized as speech with tone and intonation adjusted considering the context and emotion parameters.
[0367] Step 6:
[0368] The server inputs the supplemented audio data into the audio mixing system and processes it to seamlessly blend it with the original audio data. This results in a consistent audio stream.
[0369] Step 7:
[0370] The server transmits the completed audio stream to the terminal via edge computing with minimal latency. All communication is encrypted during this process.
[0371] Step 8:
[0372] The device receives the audio stream sent from the server and plays it back to the user through the speaker. This allows the user to have a seamless and natural conversation experience.
[0373] Step 9:
[0374] The server immediately deletes all audio data upon the end of a call, protecting the privacy of the user's communication.
[0375] (Example 2)
[0376] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0377] In voice communication, poor communication environments and data interruptions can degrade voice quality, leading to problems with accurately conveying the speaker's intentions and emotions. Furthermore, the inability to properly reflect emotions makes natural conversation difficult. In addition, the risk of personal information leakage necessitates the protection of privacy, which is a crucial issue.
[0378] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0379] In this invention, the server includes information processing means for receiving and analyzing voice information in real time, voice generation means for detecting interruptions and quality degradation of voice and generating supplementary voice, and emotion analysis means for analyzing the speaker's emotions and adjusting the tone and intonation of the supplementary voice based on the analysis results. This makes it possible to achieve natural voice communication that appropriately reflects the speaker's emotions without interruptions in voice information. Furthermore, since the voice information is encrypted, privacy is also protected.
[0380] "Audio information" refers to audio data, including human voices, acquired through input devices such as microphones.
[0381] "Information processing means" are components for analyzing and transforming input data, and are responsible for processing audio information in real time.
[0382] A "speech generation means" is a technical element that has the function of generating speech to compensate for interrupted or degraded speech and to enable natural communication.
[0383] A "speech synthesis means" is a component that combines the generated supplementary speech with the original speech information to produce an output.
[0384] "Communication methods" refer to network technologies and protocols used to transmit data with low latency.
[0385] "Data protection measures" refer to technical measures taken to prevent the leakage of personal information and protect privacy, such as encrypting voice information and deleting data after processing.
[0386] "Emotion analysis means" refers to a technological element that identifies the speaker's emotions from audio information and adjusts the tone and intonation of supplementary audio based on the results.
[0387] "Edge computing technology" refers to a technology that performs processing in a distributed computing environment to achieve low latency, and enables real-time data processing on edge devices.
[0388] This invention provides a system that offers natural conversation by reinforcing interruptions in voice information during voice communication while recognizing the user's emotions in real time. The following describes the configuration for implementing this system.
[0389] Terminal operation
[0390] The user initiates a call using the terminal. The terminal uses its built-in microphone to capture the user's voice information. This voice information is converted into a digital format and transmitted to the server in real time. In this process, the terminal uses standard communication protocols, such as TCP / IP, to transmit data.
[0391] Server operation
[0392] The server receives audio information transmitted from the terminal. The server is equipped with information processing capabilities, and the audio information is stored in a buffer and analyzed. During analysis, the server uses an AI engine to detect audio interruptions and noise. This includes a process that effectively filters out noise using a generative AI model. In addition, emotion analysis capabilities are used to identify the user's emotions from the audio. The server estimates whether the speaker is experiencing emotions such as "joy," "sadness," or "anger," and based on the results, it uses speech generation capabilities to generate supplementary audio. This supplementary audio is composed of a tone and intonation appropriate to the user's emotions and is combined with the original audio information by speech synthesis capabilities.
[0393] Specific example
[0394] For example, consider a scenario where a user shares their thoughts about essential items with a friend. If the user's voice sounds shaky, the emotion analysis system can recognize this emotion and generate a reassuring, calming tone of voice to complement it. This allows the recipient to understand the user's emotions more deeply, enriching the communication.
[0395] Privacy protection
[0396] The server has data protection measures in place to encrypt voice information and store it securely for temporary use. After processing is complete, the voice information is immediately deleted, ensuring user privacy.
[0397] Example of a prompt
[0398] "Please introduce a system that analyzes a user's emotions in real time during a phone call and fills in gaps in the conversation in a way that reflects those emotions."
[0399] This invention is expected to make voice communication more natural and enriching, thereby improving the user experience.
[0400] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0401] Step 1:
[0402] The user initiates a call using the terminal. The terminal acquires audio information through the microphone and converts the analog audio into a digital format. For example, a sampling rate of 48kHz is used for this conversion. The digitized audio information is transmitted to the server in real time via the TCP / IP protocol. The input is analog audio, and the output is digital audio data.
[0403] Step 2:
[0404] The server receives digital audio data transmitted from the terminal. The server stores this audio data in a buffer and begins analysis using information processing equipment. Here, the input is digital audio data, and the output is audio information prepared for analysis.
[0405] Step 3:
[0406] The server uses an AI engine to detect interruptions and noise in the audio data. During this process, a generative AI model is utilized to perform filtering and audio correction. The input is refined audio information, and the output is audio data with reduced noise.
[0407] Step 4:
[0408] The server uses emotion analysis tools to analyze the user's emotions in real time from the audio data. Feature quantities (e.g., pitch and speed) are extracted and input into an emotion recognition model. The model's output is an emotion label such as "joy," "sadness," or "anger."
[0409] Step 5:
[0410] The server generates supplementary speech based on the results of emotion analysis. The speech generation method generates speech with tone and intonation that corresponds to the user's emotions. A generation AI model is used in this process. The input is emotion labels, and the output is supplementary speech.
[0411] Step 6:
[0412] The server mixes the generated supplementary audio with the original audio data using a speech synthesis system. This creates an audio stream in which the missing parts are naturally filled in. The input is the supplementary audio and the original audio data, and the output is the mixed audio stream.
[0413] Step 7:
[0414] The server sends the mixed audio stream to the terminal in real time. The terminal reconstructs the received audio stream and plays it through the speaker. This allows the user to enjoy a seamless, natural conversation. The input is the mixed audio stream, and the output is the audio played from the speaker.
[0415] (Application Example 2)
[0416] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0417] In food delivery and other forms of communication, emotions may not be fully conveyed during voice calls. While users want to sense the emotions of delivery drivers and operators, current voice communication systems often make this difficult due to interruptions and quality degradation. This challenge hinders improvements in user experience and reliability.
[0418] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0419] In this invention, the server includes information processing means for receiving and analyzing audio data in real time, audio generation means for detecting interruptions and quality degradation in the analyzed audio and generating supplementary audio, and emotion recognition means for recognizing the speaker's emotional state using emotion analysis technology and adjusting the tone and intonation of the supplementary audio based on that. As a result, the user can experience natural conversations that are free from interruptions and that appropriately reflect emotions.
[0420] "Audio data" refers to information that represents sound in digital format and is processed in real time.
[0421] "Information processing means" refers to the capabilities of a system that receives and analyzes audio data, and includes devices and programs that play a role in detecting interruptions or degradation of audio quality.
[0422] A "speech generation means" is a mechanism for synthesizing and generating new speech to compensate for detected interruptions or quality degradation.
[0423] "Audio mixing means" refers to a processing function that appropriately combines the original audio data with the generated supplementary audio and outputs it.
[0424] "Communication methods" refer to network technologies and equipment for transmitting and receiving voice data and its analysis results with low latency.
[0425] "Data protection measures" refer to technologies and methods that encrypt voice data and immediately delete it after transmission or processing to ensure privacy.
[0426] "Emotion identification means" refers to a technology that analyzes the speaker's emotional state from audio data and uses that information to add tone and intonation to the supplementary audio.
[0427] In this invention, first, the terminal captures the user's voice in real time. The voice is converted into a digital format and transmitted to a server via network communication means. The server receives and analyzes the voice data using information processing means to detect interruptions and quality degradation.
[0428] The emotion recognition means allows the server to recognize the emotional state within the audio data, and the audio generation means generates supplementary audio with adjusted tone and intonation based on the analysis results. This supplementary audio is then combined with the original audio data by the audio mixing means.
[0429] The communication method achieves particularly low latency data transmission, allowing users to enjoy a natural voice experience without experiencing interruptions in conversation. Data protection measures encrypt voice data to safeguard user privacy and delete it immediately after processing.
[0430] One concrete example of implementation is in food delivery services, where, when a delivery person converses with a user, supplementary voices are generated that convey the delivery person's enthusiasm, allowing the user to experience a more friendly and approachable communication.
[0431] An example of a prompt message would be: "Analyze the emotions during this call and generate audio that will make the user feel more excited about the delivery person and improve their satisfaction."
[0432] In terms of specific hardware and software, high-sensitivity microphones are used for voice capture, high-speed network technology for data transmission, and AI and emotion engines for voice analysis and generation. The configuration may include cloud services as servers or edge devices, enabling real-time processing and data protection.
[0433] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0434] Step 1:
[0435] The terminal captures the user's voice using a high-sensitivity microphone. It converts the input analog audio into a digital format and prepares it for transmission to a server over the network. This process outputs digitized audio data.
[0436] Step 2:
[0437] The server receives digital audio data from terminals in real time using network communication means. The received data is passed to an information processing system and stored in a buffer. The output at this stage is audio data ready for analysis.
[0438] Step 3:
[0439] The server analyzes the stored audio data using information processing tools. An AI engine is used to detect interruptions and quality degradation within the audio data. The output generated by this analysis includes information about audio interruptions.
[0440] Step 4:
[0441] Using emotion recognition technology, the server recognizes the speaker's emotional state based on the analyzed audio data. This process uses an emotion analysis model to perform data transformations for detecting tone and intonation. The output is an analysis result that possesses emotional characteristics.
[0442] Step 5:
[0443] The server uses speech generation technology to generate supplementary speech based on the analysis results. In this process, a generation AI model is used to incorporate tones and intonations that blend naturally with the original conversation and to fill in any gaps. The generated supplementary speech is then output.
[0444] Step 6:
[0445] The server's audio mixing mechanism combines the original audio data with the supplementary audio. The mixing process generates a natural audio stream with interruptions corrected. The output of this process is the integrated audio data.
[0446] Step 7:
[0447] The server sends the generated audio stream back to the terminal. Low latency is crucial for the data transmitted through the communication method. The output of this process is an audio stream ready for playback.
[0448] Step 8:
[0449] The device plays the received audio stream to the user through its speaker. This allows the user to experience a more natural and enhanced conversation, increasing engagement. The output here is the audio that is played back.
[0450] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0451] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0452] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0453] [Third Embodiment]
[0454] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0455] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0456] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0457] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0458] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0459] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0460] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0461] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0462] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0463] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0464] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0465] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0466] This invention is a system for improving the quality of voice communication and is implemented as follows.
[0467] Terminal operation
[0468] The user initiates a call without requiring any special actions. The device captures the user's voice using the microphone and generates audio data. The generated audio data is then transmitted to the server via the communication line.
[0469] Server operation
[0470] The server receives audio data transmitted from the terminal. This audio data is analyzed in real time by data processing equipment. During the analysis process, AI technology is used to evaluate the context of the conversation and the quality of the audio, and to determine whether there are any interruptions or noise.
[0471] If necessary, the server uses a speech generation mechanism to generate new audio data to fill in the gaps in the audio. After the supplementary audio is generated, a speech mixing mechanism combines it with the original audio data to create a consistent audio stream.
[0472] Final output and user experience
[0473] This audio stream is retransmitted to the device with low latency and high quality. The sound is played through the device's speaker, allowing the user to enjoy a seamless and smooth conversation experience.
[0474] Data Protection
[0475] All data handled during a call is encrypted, so it cannot be illegally obtained during communication. The server protects user privacy by promptly deleting all processed data after processing the voice data is complete.
[0476] Specific example
[0477] For example, when a user makes a call in the basement of a building, the signal may be unstable and the audio may be interrupted. The user's device sends the captured audio data to a server. The server uses AI to fill in the audio segments lost due to signal interruptions and generates consistent audio data. This provides the other party with a conversation that does not feel interrupted. The user can experience high-quality calls even in places with poor communication environments, such as basements. This systematic functionality provides added value to the user in this embodiment of the invention.
[0478] The following describes the processing flow.
[0479] Step 1:
[0480] The device captures the user's spoken voice using a microphone and converts it into digital audio data. It then divides the converted audio data into packets in real time and prepares them for transmission to the server.
[0481] Step 2:
[0482] The server receives voice data packets sent from the terminal. It stores the received data in a processing buffer and prepares it for analysis.
[0483] Step 3:
[0484] The server passes the audio data stored in the buffer to the AI engine for audio quality analysis. The analysis evaluates the spectral information of the audio to detect interruptions and noise.
[0485] Step 4:
[0486] When the AI engine detects a gap, the server analyzes the context of the preceding audio data. It then infers the audio data where completion is needed and synthesizes the completion audio using the audio generation method.
[0487] Step 5:
[0488] The server passes the synthesized supplemental audio and the original audio data to the audio mixing device to generate a single, consistent audio stream.
[0489] Step 6:
[0490] The server compresses the generated audio stream and sends it to the terminal using edge computing technology, minimizing latency.
[0491] Step 7:
[0492] The device receives the audio stream sent from the server and plays it back to the user through the speaker. This allows the user to experience a natural, continuous conversation.
[0493] Step 8:
[0494] The server immediately deletes all processed audio data once the call ends, protecting the user's communication content.
[0495] (Example 1)
[0496] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0497] Conventional voice communication systems often experienced interruptions in poor communication environments, making it difficult to maintain smooth conversations. Furthermore, security concerns included the risk of unauthorized data acquisition during communication, resulting in insufficient protection of user privacy. Technologies are needed to address these issues of voice interruptions and quality degradation.
[0498] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0499] In this invention, the server includes information processing means for receiving and analyzing voice information in real time, voice generation means for detecting interruptions and quality degradation in the analyzed voice and generating supplementary voice, and voice integration means for combining the generated supplementary voice with the original voice information and outputting it. This makes it possible to provide consistent voice communication without interruption even when the communication environment deteriorates, and to improve the user's conversation experience.
[0500] "Voice information" refers to data obtained by capturing the user's speech with a microphone and converting it into digital data.
[0501] "Real-time" means processing audio data instantly, resulting in data transmission and processing at a level where users do not perceive any delay.
[0502] "Information processing means" refers to system components that analyze received audio information and detect interruptions and noise.
[0503] "Speech generation means" refers to a technology that has the function of generating and supplementing lost audio portions based on analyzed data.
[0504] "Audio integration means" refers to a process for combining the generated supplementary audio with the original audio data to output a continuous audio stream.
[0505] "Communication method" refers to network technology that has the function of transmitting encrypted voice data with low latency.
[0506] "Information protection measures" refer to a system that protects privacy by encrypting audio data and deleting it immediately after processing is complete.
[0507] "Acquisition means" refers to a part of a system that captures the user's voice using a microphone and generates audio data.
[0508] This invention is a system for improving the quality of voice communication, and is particularly aimed at providing high-quality voice even in unstable communication environments. When a user initiates a call, the terminal uses its built-in microphone to acquire voice and converts it into data as digital voice information. This voice information is then encoded and sent to a server.
[0509] The server processes the received audio information in real time. AI technology is used as the information processing method to detect audio interruptions and noise. Using a generative AI model, generated audio is created that fills in the gaps based on the analysis results. Furthermore, the server uses an audio integration method to integrate the supplemented audio with the original audio information, generating a seamless audio stream.
[0510] The generated audio stream is retransmitted to the terminal using a low-latency, secure communication method. The terminal decodes this audio stream and plays it back through the speaker, allowing the user to experience uninterrupted, high-quality audio. All audio information is encrypted and immediately deleted by the server once the call ends, ensuring privacy.
[0511] For example, even if the user is underground, this system allows for a call without any noticeable audio interruption. An example of a prompt from the generated AI model might be, "Please explain the steps to analyze the audio data from the underground call and fill in any lost segments."
[0512] In this way, this invention can provide users with a better voice communication environment.
[0513] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0514] Step 1:
[0515] The user initiates a call using the terminal. The terminal uses its built-in microphone to capture the user's voice and converts the analog audio signal into digital data. At this time, the terminal performs sampling to prepare the audio data as packets. The input is the user's analog voice, and the output is digital audio data.
[0516] Step 2:
[0517] The terminal encrypts the generated digital audio data and sends it to the server using a secure communication protocol. By utilizing Wi-Fi or mobile data communication for data transmission, low-latency data transmission is achieved. The input is unencrypted digital audio data, and the output is encrypted audio data.
[0518] Step 3:
[0519] The server receives audio data transmitted from the terminal. At this stage, the server uses information processing tools to decrypt the audio data, filter out noise, and perform speech recognition. Through information analysis, it identifies any interruptions or degradation in audio quality. The input is encrypted audio data, and the output is the analyzed audio data.
[0520] Step 4:
[0521] Based on the analysis results, the server uses a generative AI model to generate supplementary audio for the missing sections. The generative AI model performs natural-sounding supplementation by considering past conversational context and speech characteristics. The input is the analyzed audio data and its contextual information, and the output is the supplemented audio data.
[0522] Step 5:
[0523] The server mixes the supplemented audio data with the original audio data. A consistent audio stream is constructed using an audio integration mechanism. This process optimizes the temporal consistency and intelligibility of the audio. The input is the supplemented and original audio data, and the output is the integrated audio stream.
[0524] Step 6:
[0525] The server encrypts the integrated audio stream and sends it to the terminal. For data protection, communication is again conducted through a secure protocol. The input is the integrated audio stream, and the output is the encrypted audio stream.
[0526] Step 7:
[0527] The device decodes the received audio stream and plays it through the speaker. The device performs digital signal processing to play the audio at optimal volume and quality. This allows the user to enjoy a high-quality, uninterrupted audio experience. The input is an encrypted audio stream, and the output is the analog audio being played back.
[0528] (Application Example 1)
[0529] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0530] In communication between customers and delivery personnel, there is a need to achieve smooth communication with high-quality voice even in environments with unstable radio wave conditions. In particular, there is a lack of technology to compensate for voice interruptions and quality degradation and provide a clear voice experience in places where communication environments tend to deteriorate, such as underground parking lots and recessed areas of buildings.
[0531] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0532] In this invention, the server includes information processing means for receiving and analyzing voice data in real time, voice generation means for detecting interruptions and quality degradation in the analyzed voice and generating supplementary voice, and communication improvement means for improving voice quality in an environment specifically designed for communication between customers and delivery personnel. This enables high-quality voice calls with minimal voice interruptions and noise, even in unstable communication environments.
[0533] "Audio data" refers to audio signals converted into digital data, and is a data format that includes the content of the audio.
[0534] "Information processing means" refers to a device or program for receiving and analyzing audio data in real time, primarily for analyzing audio interruptions and quality.
[0535] "Speech generation means" refers to a technology or device that detects interruptions in analyzed speech data and generates new speech data to fill in the missing parts.
[0536] "Speech synthesis means" refers to a device or technology that combines generated supplementary speech with the original speech data to output consistent speech.
[0537] "Network communication means" refers to communication technology or equipment for transmitting voice data with low latency and ensuring real-time performance.
[0538] "Communication improvement measures" refer to technologies or methods for improving the quality of voice communication between customers and delivery personnel.
[0539] "Data protection measures" refer to technologies or devices that ensure privacy by encrypting voice data and deleting it immediately after processing.
[0540] Edge computing technology is a computing technique that performs data processing on devices at the network edge rather than on the server side, promoting real-time performance and low latency.
[0541] This invention is a system for improving the quality of voice communication, aiming to maintain high quality voice communication between customers and delivery personnel. The system primarily includes multiple means for implementing technologies that receive, analyze, and supplement voice data.
[0542] First, the device captures the user's voice using a microphone and sends that audio data to the server in real time. A smartphone is used as the device, and a software library such as FFmpeg is used for audio capture.
[0543] The server analyzes the received audio data using an AI analysis engine such as TensorFlow. This analysis detects audio interruptions and quality degradation while considering the audio context and communication environment. If necessary, the server generates new audio data via an audio generation system to ensure consistent audio.
[0544] The generated supplementary audio is combined with the original audio by a speech synthesis system and reconstructed as high-quality audio data. During this process, the audio data is securely protected using encryption technologies such as OpenSSL.
[0545] The synthesized speech data is retransmitted to the terminal in real time via network communication, providing users with a smooth voice experience. This system enables seamless communication between customers and delivery personnel, even in environments with unstable radio waves, such as underground parking lots.
[0546] For example, when a delivery person makes a phone call to a customer from an underground parking lot, the audio may normally cut out. However, by using this technology, the server can fill in the gaps and deliver a clear audio message to the customer.
[0547] An example of a prompt message is: "The delivery person is calling the customer to confirm the order. However, the signal is poor and the audio is frequently interrupted. How can AI be used to improve the audio quality?"
[0548] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0549] Step 1:
[0550] The user's device captures their voice using a microphone. This voice input is converted into a digital signal and sent to the server in real time. FFmpeg is used for voice capture.
[0551] Step 2:
[0552] The server begins processing digital audio data received from the user's terminal. The server analyzes this data using an AI analysis engine such as TensorFlow to detect the context of the audio, noise due to the communication environment, and audio interruptions. The output of the analysis includes the locations of the interruptions and an evaluation of the audio quality.
[0553] Step 3:
[0554] If the analysis detects a gap in the audio, the server uses AI technology to fill in the missing audio segments. This process utilizes a generative AI model to generate new audio data that doesn't sound interrupted. This generated audio data becomes the output.
[0555] Step 4:
[0556] Next, the server uses speech synthesis to combine the generated supplementary speech with the original speech data. This mixing produces a consistent speech stream. The output is reconstructed speech data that sounds smooth.
[0557] Step 5:
[0558] The server retransmits the reconstructed audio stream to the end user's device with low latency. During this process, encryption technologies such as OpenSSL are used to securely protect the data during transmission. The user's device receives this audio stream and plays it through its speakers, providing a clear audio experience.
[0559] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0560] This invention is a system for voice communication, such as telephone calls, that compensates for interruptions in voice data while recognizing the user's emotions to make the conversation more natural.
[0561] Terminal operation
[0562] The user initiates a call in the usual manner. The device captures the user's voice using the microphone and generates audio data in digital format. This audio data is transmitted to the server in real time.
[0563] Server operation
[0564] The server receives audio data transmitted from the terminal, stores it in a buffer, and begins analyzing the data. In the initial stages of analysis, the server uses an AI engine to detect audio interruptions and noise. It also uses an emotion engine to analyze and recognize the user's emotions in real time.
[0565] The server generates supplementary audio as needed. During this process, the analysis results of the emotion engine are incorporated to create supplementary audio with a tone and intonation that reflects the user's emotions. The audio generation method considers the context of the conversation and the user's emotions to perform appropriate supplementation. The generated supplementary audio is then combined with the original audio data by an audio mixing method.
[0566] Final output and user experience
[0567] The audio stream, including supplementary voice, sent from the server is received by the device and played back to the user through the speaker. This allows the user to experience a natural conversation that is seamless and appropriately reflects their emotions.
[0568] Data Protection
[0569] The server encrypts all data during transmission and stores it securely, and immediately deletes the voice data after processing, thus protecting user privacy.
[0570] Specific example
[0571] For example, if a user calls a friend while in an emotional state, their voice may be transmitted with a tremor. The emotion engine recognizes this emotional state from the user's voice data and adjusts the tone and intonation of the voice. Based on this information, the server fills in the gaps to match the emotion, allowing the recipient of the call to understand the user's emotions more clearly and enabling more effective communication. This invention realizes high-quality voice calls that take the user's emotions into account, improving the user experience.
[0572] The following describes the processing flow.
[0573] Step 1:
[0574] The device captures the user's spoken voice using a microphone and converts it into digital audio data. This audio data is then divided into packets in real time and prepared for transmission to the server over the network.
[0575] Step 2:
[0576] The server receives voice data packets sent from the terminal. The received data is temporarily stored in a buffer and organized for analysis.
[0577] Step 3:
[0578] The server passes the audio data stored in the buffer to the AI engine. The AI engine analyzes the spectral information of the audio signal to detect interruptions and noise in the audio.
[0579] Step 4:
[0580] The server inputs voice data into the emotion engine. The emotion engine identifies emotional states such as stress, joy, and anger from the user's voice and generates numerical emotion parameters as a result of the analysis.
[0581] Step 5:
[0582] The server integrates the analysis results from the AI engine and the emotion engine, and uses a speech generation method to fill in the gaps in the audio. The supplementary audio is synthesized as speech with tone and intonation adjusted considering the context and emotion parameters.
[0583] Step 6:
[0584] The server inputs the supplemented audio data into the audio mixing system and processes it to seamlessly blend it with the original audio data. This results in a consistent audio stream.
[0585] Step 7:
[0586] The server transmits the completed audio stream to the terminal via edge computing with minimal latency. All communication is encrypted during this process.
[0587] Step 8:
[0588] The device receives the audio stream sent from the server and plays it back to the user through the speaker. This allows the user to have a seamless and natural conversation experience.
[0589] Step 9:
[0590] The server immediately deletes all audio data upon the end of a call, protecting the privacy of the user's communication.
[0591] (Example 2)
[0592] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0593] In voice communication, poor communication environments and data interruptions can degrade voice quality, leading to problems with accurately conveying the speaker's intentions and emotions. Furthermore, the inability to properly reflect emotions makes natural conversation difficult. In addition, the risk of personal information leakage necessitates the protection of privacy, which is a crucial issue.
[0594] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0595] In this invention, the server includes information processing means for receiving and analyzing voice information in real time, voice generation means for detecting interruptions and quality degradation of voice and generating supplementary voice, and emotion analysis means for analyzing the speaker's emotions and adjusting the tone and intonation of the supplementary voice based on the analysis results. This makes it possible to achieve natural voice communication that appropriately reflects the speaker's emotions without interruptions in voice information. Furthermore, since the voice information is encrypted, privacy is also protected.
[0596] "Audio information" refers to audio data, including human voices, acquired through input devices such as microphones.
[0597] "Information processing means" are components for analyzing and transforming input data, and are responsible for processing audio information in real time.
[0598] A "speech generation means" is a technical element that has the function of generating speech to compensate for interrupted or degraded speech and to enable natural communication.
[0599] A "speech synthesis means" is a component that combines the generated supplementary speech with the original speech information to produce an output.
[0600] "Communication methods" refer to network technologies and protocols used to transmit data with low latency.
[0601] "Data protection measures" refer to technical measures taken to prevent the leakage of personal information and protect privacy, such as encrypting voice information and deleting data after processing.
[0602] "Emotion analysis means" refers to a technological element that identifies the speaker's emotions from audio information and adjusts the tone and intonation of supplementary audio based on the results.
[0603] "Edge computing technology" refers to a technology that performs processing in a distributed computing environment to achieve low latency, and enables real-time data processing on edge devices.
[0604] This invention provides a system that offers natural conversation by reinforcing interruptions in voice information during voice communication while recognizing the user's emotions in real time. The following describes the configuration for implementing this system.
[0605] Terminal operation
[0606] The user initiates a call using the terminal. The terminal uses its built-in microphone to capture the user's voice information. This voice information is converted into a digital format and transmitted to the server in real time. In this process, the terminal uses standard communication protocols, such as TCP / IP, to transmit data.
[0607] Server operation
[0608] The server receives audio information transmitted from the terminal. The server is equipped with information processing capabilities, and the audio information is stored in a buffer and analyzed. During analysis, the server uses an AI engine to detect audio interruptions and noise. This includes a process that effectively filters out noise using a generative AI model. In addition, emotion analysis capabilities are used to identify the user's emotions from the audio. The server estimates whether the speaker is experiencing emotions such as "joy," "sadness," or "anger," and based on the results, it uses speech generation capabilities to generate supplementary audio. This supplementary audio is composed of a tone and intonation appropriate to the user's emotions and is combined with the original audio information by speech synthesis capabilities.
[0609] Specific example
[0610] For example, consider a scenario where a user shares their thoughts about essential items with a friend. If the user's voice sounds shaky, the emotion analysis system can recognize this emotion and generate a reassuring, calming tone of voice to complement it. This allows the recipient to understand the user's emotions more deeply, enriching the communication.
[0611] Privacy protection
[0612] The server has data protection measures in place to encrypt voice information and store it securely for temporary use. After processing is complete, the voice information is immediately deleted, ensuring user privacy.
[0613] Example of a prompt
[0614] "Please introduce a system that analyzes a user's emotions in real time during a phone call and fills in gaps in the conversation in a way that reflects those emotions."
[0615] This invention is expected to make voice communication more natural and enriching, thereby improving the user experience.
[0616] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0617] Step 1:
[0618] The user initiates a call using the terminal. The terminal acquires audio information through the microphone and converts the analog audio into a digital format. For example, a sampling rate of 48kHz is used for this conversion. The digitized audio information is transmitted to the server in real time via the TCP / IP protocol. The input is analog audio, and the output is digital audio data.
[0619] Step 2:
[0620] The server receives digital audio data transmitted from the terminal. The server stores this audio data in a buffer and begins analysis using information processing equipment. Here, the input is digital audio data, and the output is audio information prepared for analysis.
[0621] Step 3:
[0622] The server uses an AI engine to detect interruptions and noise in the audio data. During this process, a generative AI model is utilized to perform filtering and audio correction. The input is refined audio information, and the output is audio data with reduced noise.
[0623] Step 4:
[0624] The server uses emotion analysis tools to analyze the user's emotions in real time from the audio data. Feature quantities (e.g., pitch and speed) are extracted and input into an emotion recognition model. The model's output is an emotion label such as "joy," "sadness," or "anger."
[0625] Step 5:
[0626] The server generates supplementary speech based on the results of emotion analysis. The speech generation method generates speech with tone and intonation that corresponds to the user's emotions. A generation AI model is used in this process. The input is emotion labels, and the output is supplementary speech.
[0627] Step 6:
[0628] The server mixes the generated supplementary audio with the original audio data using a speech synthesis system. This creates an audio stream in which the missing parts are naturally filled in. The input is the supplementary audio and the original audio data, and the output is the mixed audio stream.
[0629] Step 7:
[0630] The server sends the mixed audio stream to the terminal in real time. The terminal reconstructs the received audio stream and plays it through the speaker. This allows the user to enjoy a seamless, natural conversation. The input is the mixed audio stream, and the output is the audio played from the speaker.
[0631] (Application Example 2)
[0632] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0633] In food delivery and other forms of communication, emotions may not be fully conveyed during voice calls. While users want to sense the emotions of delivery drivers and operators, current voice communication systems often make this difficult due to interruptions and quality degradation. This challenge hinders improvements in user experience and reliability.
[0634] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0635] In this invention, the server includes information processing means for receiving and analyzing audio data in real time, audio generation means for detecting interruptions and quality degradation in the analyzed audio and generating supplementary audio, and emotion recognition means for recognizing the speaker's emotional state using emotion analysis technology and adjusting the tone and intonation of the supplementary audio based on that. As a result, the user can experience natural conversations that are free from interruptions and that appropriately reflect emotions.
[0636] "Audio data" refers to information that represents sound in digital format and is processed in real time.
[0637] "Information processing means" refers to the capabilities of a system that receives and analyzes audio data, and includes devices and programs that play a role in detecting interruptions or degradation of audio quality.
[0638] A "speech generation means" is a mechanism for synthesizing and generating new speech to compensate for detected interruptions or quality degradation.
[0639] "Audio mixing means" refers to a processing function that appropriately combines the original audio data with the generated supplementary audio and outputs it.
[0640] "Communication methods" refer to network technologies and equipment for transmitting and receiving voice data and its analysis results with low latency.
[0641] "Data protection measures" refer to technologies and methods that encrypt voice data and immediately delete it after transmission or processing to ensure privacy.
[0642] "Emotion identification means" refers to a technology that analyzes the speaker's emotional state from audio data and uses that information to add tone and intonation to the supplementary audio.
[0643] In this invention, first, the terminal captures the user's voice in real time. The voice is converted into a digital format and transmitted to a server via network communication means. The server receives and analyzes the voice data using information processing means to detect interruptions and quality degradation.
[0644] The emotion recognition means allows the server to recognize the emotional state within the audio data, and the audio generation means generates supplementary audio with adjusted tone and intonation based on the analysis results. This supplementary audio is then combined with the original audio data by the audio mixing means.
[0645] The communication method achieves particularly low latency data transmission, allowing users to enjoy a natural voice experience without experiencing interruptions in conversation. Data protection measures encrypt voice data to safeguard user privacy and delete it immediately after processing.
[0646] One concrete example of implementation is in food delivery services, where, when a delivery person converses with a user, supplementary voices are generated that convey the delivery person's enthusiasm, allowing the user to experience a more friendly and approachable communication.
[0647] An example of a prompt message would be: "Analyze the emotions during this call and generate audio that will make the user feel more excited about the delivery person and improve their satisfaction."
[0648] In terms of specific hardware and software, high-sensitivity microphones are used for voice capture, high-speed network technology for data transmission, and AI and emotion engines for voice analysis and generation. The configuration may include cloud services as servers or edge devices, enabling real-time processing and data protection.
[0649] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0650] Step 1:
[0651] The terminal captures the user's voice using a high-sensitivity microphone. It converts the input analog audio into a digital format and prepares it for transmission to a server over the network. This process outputs digitized audio data.
[0652] Step 2:
[0653] The server receives digital audio data from terminals in real time using network communication means. The received data is passed to an information processing system and stored in a buffer. The output at this stage is audio data ready for analysis.
[0654] Step 3:
[0655] The server analyzes the stored audio data using information processing tools. An AI engine is used to detect interruptions and quality degradation within the audio data. The output generated by this analysis includes information about audio interruptions.
[0656] Step 4:
[0657] Using emotion recognition technology, the server recognizes the speaker's emotional state based on the analyzed audio data. This process uses an emotion analysis model to perform data transformations for detecting tone and intonation. The output is an analysis result that possesses emotional characteristics.
[0658] Step 5:
[0659] The server uses speech generation technology to generate supplementary speech based on the analysis results. In this process, a generation AI model is used to incorporate tones and intonations that blend naturally with the original conversation and to fill in any gaps. The generated supplementary speech is then output.
[0660] Step 6:
[0661] The server's audio mixing mechanism combines the original audio data with the supplementary audio. The mixing process generates a natural audio stream with interruptions corrected. The output of this process is the integrated audio data.
[0662] Step 7:
[0663] The server sends the generated audio stream back to the terminal. Low latency is crucial for the data transmitted through the communication method. The output of this process is an audio stream ready for playback.
[0664] Step 8:
[0665] The device plays the received audio stream to the user through its speaker. This allows the user to experience a more natural and enhanced conversation, increasing engagement. The output here is the audio that is played back.
[0666] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0667] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0668] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0669] [Fourth Embodiment]
[0670] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0671] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0672] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0673] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0674] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0675] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0676] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0677] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0678] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0679] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0680] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0681] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0682] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0683] This invention is a system for improving the quality of voice communication and is implemented as follows.
[0684] Terminal operation
[0685] The user initiates a call without requiring any special actions. The device captures the user's voice using the microphone and generates audio data. The generated audio data is then transmitted to the server via the communication line.
[0686] Server operation
[0687] The server receives audio data transmitted from the terminal. This audio data is analyzed in real time by data processing equipment. During the analysis process, AI technology is used to evaluate the context of the conversation and the quality of the audio, and to determine whether there are any interruptions or noise.
[0688] If necessary, the server uses a speech generation mechanism to generate new audio data to fill in the gaps in the audio. After the supplementary audio is generated, a speech mixing mechanism combines it with the original audio data to create a consistent audio stream.
[0689] Final output and user experience
[0690] This audio stream is retransmitted to the device with low latency and high quality. The sound is played through the device's speaker, allowing the user to enjoy a seamless and smooth conversation experience.
[0691] Data Protection
[0692] All data handled during a call is encrypted, so it cannot be illegally obtained during communication. The server protects user privacy by promptly deleting all processed data after processing the voice data is complete.
[0693] Specific example
[0694] For example, when a user makes a call in the basement of a building, the signal may be unstable and the audio may be interrupted. The user's device sends the captured audio data to a server. The server uses AI to fill in the audio segments lost due to signal interruptions and generates consistent audio data. This provides the other party with a conversation that does not feel interrupted. The user can experience high-quality calls even in places with poor communication environments, such as basements. This systematic functionality provides added value to the user in this embodiment of the invention.
[0695] The following describes the processing flow.
[0696] Step 1:
[0697] The device captures the user's spoken voice using a microphone and converts it into digital audio data. It then divides the converted audio data into packets in real time and prepares them for transmission to the server.
[0698] Step 2:
[0699] The server receives voice data packets sent from the terminal. It stores the received data in a processing buffer and prepares it for analysis.
[0700] Step 3:
[0701] The server passes the audio data stored in the buffer to the AI engine for audio quality analysis. The analysis evaluates the spectral information of the audio to detect interruptions and noise.
[0702] Step 4:
[0703] When the AI engine detects a gap, the server analyzes the context of the preceding audio data. It then infers the audio data where completion is needed and synthesizes the completion audio using the audio generation method.
[0704] Step 5:
[0705] The server passes the synthesized supplemental audio and the original audio data to the audio mixing device to generate a single, consistent audio stream.
[0706] Step 6:
[0707] The server compresses the generated audio stream and sends it to the terminal using edge computing technology, minimizing latency.
[0708] Step 7:
[0709] The device receives the audio stream sent from the server and plays it back to the user through the speaker. This allows the user to experience a natural, continuous conversation.
[0710] Step 8:
[0711] The server immediately deletes all processed audio data once the call ends, protecting the user's communication content.
[0712] (Example 1)
[0713] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0714] Conventional voice communication systems often experienced interruptions in poor communication environments, making it difficult to maintain smooth conversations. Furthermore, security concerns included the risk of unauthorized data acquisition during communication, resulting in insufficient protection of user privacy. Technologies are needed to address these issues of voice interruptions and quality degradation.
[0715] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0716] In this invention, the server includes information processing means for receiving and analyzing voice information in real time, voice generation means for detecting interruptions and quality degradation in the analyzed voice and generating supplementary voice, and voice integration means for combining the generated supplementary voice with the original voice information and outputting it. This makes it possible to provide consistent voice communication without interruption even when the communication environment deteriorates, and to improve the user's conversation experience.
[0717] "Voice information" refers to data obtained by capturing the user's speech with a microphone and converting it into digital data.
[0718] "Real-time" means processing audio data instantly, resulting in data transmission and processing at a level where users do not perceive any delay.
[0719] "Information processing means" refers to system components that analyze received audio information and detect interruptions and noise.
[0720] "Speech generation means" refers to a technology that has the function of generating and supplementing lost audio portions based on analyzed data.
[0721] "Audio integration means" refers to a process for combining the generated supplementary audio with the original audio data to output a continuous audio stream.
[0722] "Communication method" refers to network technology that has the function of transmitting encrypted voice data with low latency.
[0723] "Information protection measures" refer to a system that protects privacy by encrypting audio data and deleting it immediately after processing is complete.
[0724] "Acquisition means" refers to a part of a system that captures the user's voice using a microphone and generates audio data.
[0725] This invention is a system for improving the quality of voice communication, and is particularly aimed at providing high-quality voice even in unstable communication environments. When a user initiates a call, the terminal uses its built-in microphone to acquire voice and converts it into data as digital voice information. This voice information is then encoded and sent to a server.
[0726] The server processes the received audio information in real time. AI technology is used as the information processing method to detect audio interruptions and noise. Using a generative AI model, generated audio is created that fills in the gaps based on the analysis results. Furthermore, the server uses an audio integration method to integrate the supplemented audio with the original audio information, generating a seamless audio stream.
[0727] The generated audio stream is retransmitted to the terminal using a low-latency, secure communication method. The terminal decodes this audio stream and plays it back through the speaker, allowing the user to experience uninterrupted, high-quality audio. All audio information is encrypted and immediately deleted by the server once the call ends, ensuring privacy.
[0728] For example, even if the user is underground, this system allows for a call without any noticeable audio interruption. An example of a prompt from the generated AI model might be, "Please explain the steps to analyze the audio data from the underground call and fill in any lost segments."
[0729] In this way, this invention can provide users with a better voice communication environment.
[0730] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0731] Step 1:
[0732] The user initiates a call using the terminal. The terminal uses its built-in microphone to capture the user's voice and converts the analog audio signal into digital data. At this time, the terminal performs sampling to prepare the audio data as packets. The input is the user's analog voice, and the output is digital audio data.
[0733] Step 2:
[0734] The terminal encrypts the generated digital audio data and sends it to the server using a secure communication protocol. By utilizing Wi-Fi or mobile data communication for data transmission, low-latency data transmission is achieved. The input is unencrypted digital audio data, and the output is encrypted audio data.
[0735] Step 3:
[0736] The server receives audio data transmitted from the terminal. At this stage, the server uses information processing tools to decrypt the audio data, filter out noise, and perform speech recognition. Through information analysis, it identifies any interruptions or degradation in audio quality. The input is encrypted audio data, and the output is the analyzed audio data.
[0737] Step 4:
[0738] Based on the analysis results, the server uses a generative AI model to generate supplementary audio for the missing sections. The generative AI model performs natural-sounding supplementation by considering past conversational context and speech characteristics. The input is the analyzed audio data and its contextual information, and the output is the supplemented audio data.
[0739] Step 5:
[0740] The server mixes the supplemented audio data with the original audio data. A consistent audio stream is constructed using an audio integration mechanism. This process optimizes the temporal consistency and intelligibility of the audio. The input is the supplemented and original audio data, and the output is the integrated audio stream.
[0741] Step 6:
[0742] The server encrypts the integrated audio stream and sends it to the terminal. For data protection, communication is again conducted through a secure protocol. The input is the integrated audio stream, and the output is the encrypted audio stream.
[0743] Step 7:
[0744] The device decodes the received audio stream and plays it through the speaker. The device performs digital signal processing to play the audio at optimal volume and quality. This allows the user to enjoy a high-quality, uninterrupted audio experience. The input is an encrypted audio stream, and the output is the analog audio being played back.
[0745] (Application Example 1)
[0746] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0747] In communication between customers and delivery personnel, there is a need to achieve smooth communication with high-quality voice even in environments with unstable radio wave conditions. In particular, there is a lack of technology to compensate for voice interruptions and quality degradation and provide a clear voice experience in places where communication environments tend to deteriorate, such as underground parking lots and recessed areas of buildings.
[0748] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0749] In this invention, the server includes information processing means for receiving and analyzing voice data in real time, voice generation means for detecting interruptions and quality degradation in the analyzed voice and generating supplementary voice, and communication improvement means for improving voice quality in an environment specifically designed for communication between customers and delivery personnel. This enables high-quality voice calls with minimal voice interruptions and noise, even in unstable communication environments.
[0750] "Audio data" refers to audio signals converted into digital data, and is a data format that includes the content of the audio.
[0751] "Information processing means" refers to a device or program for receiving and analyzing audio data in real time, primarily for analyzing audio interruptions and quality.
[0752] "Speech generation means" refers to a technology or device that detects interruptions in analyzed speech data and generates new speech data to fill in the missing parts.
[0753] "Speech synthesis means" refers to a device or technology that combines generated supplementary speech with the original speech data to output consistent speech.
[0754] "Network communication means" refers to communication technology or equipment for transmitting voice data with low latency and ensuring real-time performance.
[0755] "Communication improvement measures" refer to technologies or methods for improving the quality of voice communication between customers and delivery personnel.
[0756] "Data protection measures" refer to technologies or devices that ensure privacy by encrypting voice data and deleting it immediately after processing.
[0757] Edge computing technology is a computing technique that performs data processing on devices at the network edge rather than on the server side, promoting real-time performance and low latency.
[0758] This invention is a system for improving the quality of voice communication, aiming to maintain high quality voice communication between customers and delivery personnel. The system primarily includes multiple means for implementing technologies that receive, analyze, and supplement voice data.
[0759] First, the device captures the user's voice using a microphone and sends that audio data to the server in real time. A smartphone is used as the device, and a software library such as FFmpeg is used for audio capture.
[0760] The server analyzes the received audio data using an AI analysis engine such as TensorFlow. This analysis detects audio interruptions and quality degradation while considering the audio context and communication environment. If necessary, the server generates new audio data via an audio generation system to ensure consistent audio.
[0761] The generated supplementary audio is combined with the original audio by a speech synthesis system and reconstructed as high-quality audio data. During this process, the audio data is securely protected using encryption technologies such as OpenSSL.
[0762] The synthesized speech data is retransmitted to the terminal in real time via network communication, providing users with a smooth voice experience. This system enables seamless communication between customers and delivery personnel, even in environments with unstable radio waves, such as underground parking lots.
[0763] For example, when a delivery person makes a phone call to a customer from an underground parking lot, the audio may normally cut out. However, by using this technology, the server can fill in the gaps and deliver a clear audio message to the customer.
[0764] An example of a prompt message is: "The delivery person is calling the customer to confirm the order. However, the signal is poor and the audio is frequently interrupted. How can AI be used to improve the audio quality?"
[0765] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0766] Step 1:
[0767] The user's device captures their voice using a microphone. This voice input is converted into a digital signal and sent to the server in real time. FFmpeg is used for voice capture.
[0768] Step 2:
[0769] The server begins processing digital audio data received from the user's terminal. The server analyzes this data using an AI analysis engine such as TensorFlow to detect the context of the audio, noise due to the communication environment, and audio interruptions. The output of the analysis includes the locations of the interruptions and an evaluation of the audio quality.
[0770] Step 3:
[0771] If the analysis detects a gap in the audio, the server uses AI technology to fill in the missing audio segments. This process utilizes a generative AI model to generate new audio data that doesn't sound interrupted. This generated audio data becomes the output.
[0772] Step 4:
[0773] Next, the server uses speech synthesis to combine the generated supplementary speech with the original speech data. This mixing produces a consistent speech stream. The output is reconstructed speech data that sounds smooth.
[0774] Step 5:
[0775] The server retransmits the reconstructed audio stream to the end user's device with low latency. During this process, encryption technologies such as OpenSSL are used to securely protect the data during transmission. The user's device receives this audio stream and plays it through its speakers, providing a clear audio experience.
[0776] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0777] This invention is a system for voice communication, such as telephone calls, that compensates for interruptions in voice data while recognizing the user's emotions to make the conversation more natural.
[0778] Terminal operation
[0779] The user initiates a call in the usual manner. The device captures the user's voice using the microphone and generates audio data in digital format. This audio data is transmitted to the server in real time.
[0780] Server operation
[0781] The server receives audio data transmitted from the terminal, stores it in a buffer, and begins analyzing the data. In the initial stages of analysis, the server uses an AI engine to detect audio interruptions and noise. It also uses an emotion engine to analyze and recognize the user's emotions in real time.
[0782] The server generates supplementary audio as needed. During this process, the analysis results of the emotion engine are incorporated to create supplementary audio with a tone and intonation that reflects the user's emotions. The audio generation method considers the context of the conversation and the user's emotions to perform appropriate supplementation. The generated supplementary audio is then combined with the original audio data by an audio mixing method.
[0783] Final output and user experience
[0784] The audio stream, including supplementary voice, sent from the server is received by the device and played back to the user through the speaker. This allows the user to experience a natural conversation that is seamless and appropriately reflects their emotions.
[0785] Data Protection
[0786] The server encrypts all data during transmission and stores it securely, and immediately deletes the voice data after processing, thus protecting user privacy.
[0787] Specific example
[0788] For example, if a user calls a friend while in an emotional state, their voice may be transmitted with a tremor. The emotion engine recognizes this emotional state from the user's voice data and adjusts the tone and intonation of the voice. Based on this information, the server fills in the gaps to match the emotion, allowing the recipient of the call to understand the user's emotions more clearly and enabling more effective communication. This invention realizes high-quality voice calls that take the user's emotions into account, improving the user experience.
[0789] The following describes the processing flow.
[0790] Step 1:
[0791] The device captures the user's spoken voice using a microphone and converts it into digital audio data. This audio data is then divided into packets in real time and prepared for transmission to the server over the network.
[0792] Step 2:
[0793] The server receives voice data packets sent from the terminal. The received data is temporarily stored in a buffer and organized for analysis.
[0794] Step 3:
[0795] The server passes the audio data stored in the buffer to the AI engine. The AI engine analyzes the spectral information of the audio signal to detect interruptions and noise in the audio.
[0796] Step 4:
[0797] The server inputs voice data into the emotion engine. The emotion engine identifies emotional states such as stress, joy, and anger from the user's voice and generates numerical emotion parameters as a result of the analysis.
[0798] Step 5:
[0799] The server integrates the analysis results from the AI engine and the emotion engine, and uses a speech generation method to fill in the gaps in the audio. The supplementary audio is synthesized as speech with tone and intonation adjusted considering the context and emotion parameters.
[0800] Step 6:
[0801] The server inputs the supplemented audio data into the audio mixing system and processes it to seamlessly blend it with the original audio data. This results in a consistent audio stream.
[0802] Step 7:
[0803] The server transmits the completed audio stream to the terminal via edge computing with minimal latency. All communication is encrypted during this process.
[0804] Step 8:
[0805] The device receives the audio stream sent from the server and plays it back to the user through the speaker. This allows the user to have a seamless and natural conversation experience.
[0806] Step 9:
[0807] The server immediately deletes all audio data upon the end of a call, protecting the privacy of the user's communication.
[0808] (Example 2)
[0809] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0810] In voice communication, poor communication environments and data interruptions can degrade voice quality, leading to problems with accurately conveying the speaker's intentions and emotions. Furthermore, the inability to properly reflect emotions makes natural conversation difficult. In addition, the risk of personal information leakage necessitates the protection of privacy, which is a crucial issue.
[0811] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0812] In this invention, the server includes information processing means for receiving and analyzing voice information in real time, voice generation means for detecting interruptions and quality degradation of voice and generating supplementary voice, and emotion analysis means for analyzing the speaker's emotions and adjusting the tone and intonation of the supplementary voice based on the analysis results. This makes it possible to achieve natural voice communication that appropriately reflects the speaker's emotions without interruptions in voice information. Furthermore, since the voice information is encrypted, privacy is also protected.
[0813] "Audio information" refers to audio data, including human voices, acquired through input devices such as microphones.
[0814] "Information processing means" are components for analyzing and transforming input data, and are responsible for processing audio information in real time.
[0815] A "speech generation means" is a technical element that has the function of generating speech to compensate for interrupted or degraded speech and to enable natural communication.
[0816] A "speech synthesis means" is a component that combines the generated supplementary speech with the original speech information to produce an output.
[0817] "Communication methods" refer to network technologies and protocols used to transmit data with low latency.
[0818] "Data protection measures" refer to technical measures taken to prevent the leakage of personal information and protect privacy, such as encrypting voice information and deleting data after processing.
[0819] "Emotion analysis means" refers to a technological element that identifies the speaker's emotions from audio information and adjusts the tone and intonation of supplementary audio based on the results.
[0820] "Edge computing technology" refers to a technology that performs processing in a distributed computing environment to achieve low latency, and enables real-time data processing on edge devices.
[0821] This invention provides a system that offers natural conversation by reinforcing interruptions in voice information during voice communication while recognizing the user's emotions in real time. The following describes the configuration for implementing this system.
[0822] Terminal operation
[0823] The user initiates a call using the terminal. The terminal uses its built-in microphone to capture the user's voice information. This voice information is converted into a digital format and transmitted to the server in real time. In this process, the terminal uses standard communication protocols, such as TCP / IP, to transmit data.
[0824] Server operation
[0825] The server receives audio information transmitted from the terminal. The server is equipped with information processing capabilities, and the audio information is stored in a buffer and analyzed. During analysis, the server uses an AI engine to detect audio interruptions and noise. This includes a process that effectively filters out noise using a generative AI model. In addition, emotion analysis capabilities are used to identify the user's emotions from the audio. The server estimates whether the speaker is experiencing emotions such as "joy," "sadness," or "anger," and based on the results, it uses speech generation capabilities to generate supplementary audio. This supplementary audio is composed of a tone and intonation appropriate to the user's emotions and is combined with the original audio information by speech synthesis capabilities.
[0826] Specific example
[0827] For example, consider a scenario where a user shares their thoughts about essential items with a friend. If the user's voice sounds shaky, the emotion analysis system can recognize this emotion and generate a reassuring, calming tone of voice to complement it. This allows the recipient to understand the user's emotions more deeply, enriching the communication.
[0828] Privacy protection
[0829] The server has data protection measures in place to encrypt voice information and store it securely for temporary use. After processing is complete, the voice information is immediately deleted, ensuring user privacy.
[0830] Example of a prompt
[0831] "Please introduce a system that analyzes a user's emotions in real time during a phone call and fills in gaps in the conversation in a way that reflects those emotions."
[0832] This invention is expected to make voice communication more natural and enriching, thereby improving the user experience.
[0833] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0834] Step 1:
[0835] The user initiates a call using the terminal. The terminal acquires audio information through the microphone and converts the analog audio into a digital format. For example, a sampling rate of 48kHz is used for this conversion. The digitized audio information is transmitted to the server in real time via the TCP / IP protocol. The input is analog audio, and the output is digital audio data.
[0836] Step 2:
[0837] The server receives digital audio data transmitted from the terminal. The server stores this audio data in a buffer and begins analysis using information processing equipment. Here, the input is digital audio data, and the output is audio information prepared for analysis.
[0838] Step 3:
[0839] The server uses an AI engine to detect interruptions and noise in the audio data. During this process, a generative AI model is utilized to perform filtering and audio correction. The input is refined audio information, and the output is audio data with reduced noise.
[0840] Step 4:
[0841] The server uses emotion analysis tools to analyze the user's emotions in real time from the audio data. Feature quantities (e.g., pitch and speed) are extracted and input into an emotion recognition model. The model's output is an emotion label such as "joy," "sadness," or "anger."
[0842] Step 5:
[0843] The server generates supplementary speech based on the results of emotion analysis. The speech generation method generates speech with tone and intonation that corresponds to the user's emotions. A generation AI model is used in this process. The input is emotion labels, and the output is supplementary speech.
[0844] Step 6:
[0845] The server mixes the generated supplementary audio with the original audio data using a speech synthesis system. This creates an audio stream in which the missing parts are naturally filled in. The input is the supplementary audio and the original audio data, and the output is the mixed audio stream.
[0846] Step 7:
[0847] The server sends the mixed audio stream to the terminal in real time. The terminal reconstructs the received audio stream and plays it through the speaker. This allows the user to enjoy a seamless, natural conversation. The input is the mixed audio stream, and the output is the audio played from the speaker.
[0848] (Application Example 2)
[0849] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0850] In food delivery and other forms of communication, emotions may not be fully conveyed during voice calls. While users want to sense the emotions of delivery drivers and operators, current voice communication systems often make this difficult due to interruptions and quality degradation. This challenge hinders improvements in user experience and reliability.
[0851] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0852] In this invention, the server includes information processing means for receiving and analyzing audio data in real time, audio generation means for detecting interruptions and quality degradation in the analyzed audio and generating supplementary audio, and emotion recognition means for recognizing the speaker's emotional state using emotion analysis technology and adjusting the tone and intonation of the supplementary audio based on that. As a result, the user can experience natural conversations that are free from interruptions and that appropriately reflect emotions.
[0853] "Audio data" refers to information that represents sound in digital format and is processed in real time.
[0854] "Information processing means" refers to the capabilities of a system that receives and analyzes audio data, and includes devices and programs that play a role in detecting interruptions or degradation of audio quality.
[0855] A "speech generation means" is a mechanism for synthesizing and generating new speech to compensate for detected interruptions or quality degradation.
[0856] "Audio mixing means" refers to a processing function that appropriately combines the original audio data with the generated supplementary audio and outputs it.
[0857] "Communication methods" refer to network technologies and equipment for transmitting and receiving voice data and its analysis results with low latency.
[0858] "Data protection measures" refer to technologies and methods that encrypt voice data and immediately delete it after transmission or processing to ensure privacy.
[0859] "Emotion identification means" refers to a technology that analyzes the speaker's emotional state from audio data and uses that information to add tone and intonation to the supplementary audio.
[0860] In this invention, first, the terminal captures the user's voice in real time. The voice is converted into a digital format and transmitted to a server via network communication means. The server receives and analyzes the voice data using information processing means to detect interruptions and quality degradation.
[0861] The emotion recognition means allows the server to recognize the emotional state within the audio data, and the audio generation means generates supplementary audio with adjusted tone and intonation based on the analysis results. This supplementary audio is then combined with the original audio data by the audio mixing means.
[0862] The communication method achieves particularly low latency data transmission, allowing users to enjoy a natural voice experience without experiencing interruptions in conversation. Data protection measures encrypt voice data to safeguard user privacy and delete it immediately after processing.
[0863] One concrete example of implementation is in food delivery services, where, when a delivery person converses with a user, supplementary voices are generated that convey the delivery person's enthusiasm, allowing the user to experience a more friendly and approachable communication.
[0864] An example of a prompt message would be: "Analyze the emotions during this call and generate audio that will make the user feel more excited about the delivery person and improve their satisfaction."
[0865] In terms of specific hardware and software, high-sensitivity microphones are used for voice capture, high-speed network technology for data transmission, and AI and emotion engines for voice analysis and generation. The configuration may include cloud services as servers or edge devices, enabling real-time processing and data protection.
[0866] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0867] Step 1:
[0868] The terminal captures the user's voice using a high-sensitivity microphone. It converts the input analog audio into a digital format and prepares it for transmission to a server over the network. This process outputs digitized audio data.
[0869] Step 2:
[0870] The server receives digital audio data from terminals in real time using network communication means. The received data is passed to an information processing system and stored in a buffer. The output at this stage is audio data ready for analysis.
[0871] Step 3:
[0872] The server analyzes the stored audio data using information processing tools. An AI engine is used to detect interruptions and quality degradation within the audio data. The output generated by this analysis includes information about audio interruptions.
[0873] Step 4:
[0874] Using emotion recognition technology, the server recognizes the speaker's emotional state based on the analyzed audio data. This process uses an emotion analysis model to perform data transformations for detecting tone and intonation. The output is an analysis result that possesses emotional characteristics.
[0875] Step 5:
[0876] The server uses speech generation technology to generate supplementary speech based on the analysis results. In this process, a generation AI model is used to incorporate tones and intonations that blend naturally with the original conversation and to fill in any gaps. The generated supplementary speech is then output.
[0877] Step 6:
[0878] The server's audio mixing mechanism combines the original audio data with the supplementary audio. The mixing process generates a natural audio stream with interruptions corrected. The output of this process is the integrated audio data.
[0879] Step 7:
[0880] The server sends the generated audio stream back to the terminal. Low latency is crucial for the data transmitted through the communication method. The output of this process is an audio stream ready for playback.
[0881] Step 8:
[0882] The device plays the received audio stream to the user through its speaker. This allows the user to experience a more natural and enhanced conversation, increasing engagement. The output here is the audio that is played back.
[0883] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0884] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0885] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0886] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0887] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0888] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0889] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0890] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0891] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0892] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0893] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0894] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0895] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0896] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0897] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0898] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0899] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0900] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0901] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0902] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0903] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0904] The following is further disclosed regarding the embodiments described above.
[0905] (Claim 1)
[0906] A data processing means for receiving and analyzing audio data in real time,
[0907] A voice generation means that detects interruptions and quality degradation in the analyzed audio and generates supplementary audio,
[0908] A voice mixing means that outputs the generated supplementary audio combined with the original audio data,
[0909] A network communication means that achieves low latency in the transmission of voice data,
[0910] Data protection measures that encrypt voice data to protect privacy and delete it immediately after processing,
[0911] A system that includes this.
[0912] (Claim 2)
[0913] The system according to claim 1, which generates supplementary speech by analyzing the context of a conversation and the characteristics of the speakers in the analysis of speech data.
[0914] (Claim 3)
[0915] The system according to claim 1, which performs real-time processing of audio data using edge computing technology to minimize processing delay.
[0916] "Example 1"
[0917] (Claim 1)
[0918] An information processing means for receiving and analyzing audio information in real time,
[0919] A voice generation means that detects interruptions and quality degradation in the analyzed audio and generates supplementary audio,
[0920] A voice integration means that outputs the generated supplementary audio in combination with the original audio information,
[0921] A communication method that achieves low latency in the transmission of voice information,
[0922] To protect privacy, the audio information is encrypted and immediately deleted after processing is complete as an information protection measure.
[0923] An acquisition means that acquires the speaker's voice using a microphone inside the device and generates the acquired voice as data,
[0924] A transmission means for sending captured audio information via secure communication,
[0925] A signal processing means for reconstructing audio information generated based on analysis using signal processing technology,
[0926] A system that includes this.
[0927] (Claim 2)
[0928] The system according to claim 1, which generates supplementary speech by analyzing the context of a conversation and the characteristics of the speaker in the analysis of speech information.
[0929] (Claim 3)
[0930] The system according to claim 1, which performs real-time processing of audio information on a device and minimizes processing delay.
[0931] "Application Example 1"
[0932] (Claim 1)
[0933] An information processing means for receiving and analyzing audio data in real time,
[0934] A voice generation means that detects interruptions and quality degradation in the analyzed audio and generates supplementary audio,
[0935] A speech synthesis means that outputs the generated supplementary audio combined with the original audio data,
[0936] A network communication means that achieves low latency in the transmission of voice data,
[0937] Communication improvement methods that enhance voice quality in an environment specifically designed for communication between customers and delivery personnel,
[0938] Data protection measures that encrypt voice data to protect privacy and delete it immediately after processing,
[0939] A system that includes this.
[0940] (Claim 2)
[0941] The system according to claim 1, which, in analyzing voice data, takes into account the context of the conversation and fluctuations in the communication environment, and generates supplementary voice to maintain clarity of voice between the customer and the delivery person.
[0942] (Claim 3)
[0943] The system according to claim 1, which performs real-time processing of voice data using edge computing technology, minimizes processing delays, and can quickly respond to changes in the communication environment.
[0944] "Example 2 of combining an emotion engine"
[0945] (Claim 1)
[0946] An information processing means for receiving and analyzing audio information in real time,
[0947] A voice generation means that detects interruptions and quality degradation in the analyzed audio and generates supplementary audio,
[0948] A speech synthesis means that outputs the generated supplementary audio combined with the original audio information,
[0949] A communication method that achieves low latency in the transmission of voice information,
[0950] Data protection measures that encrypt voice information to protect privacy and delete it immediately after processing,
[0951] An emotion analysis method that analyzes the speaker's emotions and adjusts the tone and intonation of the supplementary voice based on the analysis results,
[0952] A system that includes this.
[0953] (Claim 2)
[0954] The system according to claim 1, which generates supplementary speech by analyzing the context of the conversation and the emotions of the speaker in the analysis of speech information.
[0955] (Claim 3)
[0956] The system according to claim 1, which performs real-time processing of audio information using terminal computing technology to minimize processing delay.
[0957] "Application example 2 when combining with an emotional engine"
[0958] (Claim 1)
[0959] An information processing means for receiving and analyzing audio data in real time,
[0960] A voice generation means that detects interruptions and quality degradation in the analyzed audio and generates supplementary audio,
[0961] A voice mixing means that outputs the generated supplementary audio combined with the original audio data,
[0962] A communication method that achieves low latency in the transmission of voice data,
[0963] Data protection measures that encrypt voice data to protect privacy and delete it immediately after processing,
[0964] An emotion recognition means that recognizes the speaker's emotional state using emotion analysis technology and adjusts the tone and intonation of the supplementary voice based on that recognition,
[0965] A system that includes this.
[0966] (Claim 2)
[0967] The system according to claim 1, which generates supplementary speech by analyzing the context of the conversation, the characteristics of the speaker, and the emotional state of the speaker in the analysis of speech data.
[0968] (Claim 3)
[0969] The system according to claim 1, which performs real-time processing of voice data using distributed computing technology and minimizes processing delay. [Explanation of symbols]
[0970] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A data processing means for receiving and analyzing audio data in real time, A voice generation means that detects interruptions and quality degradation in the analyzed audio and generates supplementary audio, A voice mixing means that outputs the generated supplementary audio combined with the original audio data, A network communication means that achieves low latency in the transmission of voice data, Data protection measures that encrypt voice data to protect privacy and delete it immediately after processing, A system that includes this.
2. The system according to claim 1, which generates supplementary speech by analyzing the context of the conversation and the characteristics of the speakers in the analysis of speech data.
3. The system according to claim 1, which performs real-time processing of audio data using edge computing technology to minimize processing delay.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A