system
A system with AI-powered speech recognition and synthesis in a wearable device effectively converts spoken words of hearing-impaired individuals into clear, natural-sounding speech, enhancing communication.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2026-03-24
AI Technical Summary
Conventional technologies have not adequately addressed the conversion of words uttered by hearing-impaired individuals into ordinary speech.
A system comprising a collection unit, analysis unit, and output unit, utilizing AI for speech recognition and synthesis, integrated into a wearable device like a bow tie, to convert spoken words into normal speech and output it clearly.
Enables smooth communication by accurately converting spoken words of hearing-impaired individuals into natural-sounding speech, facilitating interaction in daily life scenarios.
Smart Images

Figure 0007834821000001 
Figure 0007834821000002 
Figure 0007834821000003
Abstract
Description
Technical Field
[0006] , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving user speech, adding the user speech to a prompt including an instruction sentence related to the description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate chatbot speech in response to the user speech.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, the conversion of the words uttered by a hearing-impaired person into ordinary speech has not been sufficiently carried out, and there is room for improvement.
[0005] The system according to the embodiment aims to convert the words uttered by a hearing-impaired person into ordinary speech.
Means for Solving the Problems
[0006] The system according to the embodiment includes a collection unit, an analysis unit, a conversion unit, and an output unit. The collection unit collects audio. The analysis unit analyzes the audio data collected by the collection unit. The conversion unit converts the audio data analyzed by the analysis unit into ordinary speech. The output unit outputs the audio converted by the conversion unit. [Effects of the Invention]
[0007] The system according to this embodiment can convert the words spoken by a person with a hearing impairment into normal speech. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The voice conversion system according to an embodiment of the present invention is a system that converts words spoken by a person with a hearing impairment into normal speech. This voice conversion system collects the words spoken by the person with a hearing impairment using a microphone, and the collected voice data is analyzed by an AI and converted into normal speech. The converted speech is output through a speaker. This system is shaped like a bow tie, is easy to wear and inconspicuous, and is easy to use in daily life. For example, the system collects the words spoken by a person with a hearing impairment using a microphone. The microphone is built into the bow tie and collects the spoken voice with high accuracy. For example, if a person with a hearing impairment says "hello," that voice is collected by the microphone. Next, the AI analyzes the collected voice data. The AI analyzes the collected voice data and identifies the spoken word. For example, it identifies the word "hello" from the collected voice data. Then, the AI converts the identified word into normal speech. For example, it converts the word "hello" into normal speech. This conversion is performed based on voice data that the AI has learned in advance. Finally, the converted speech is output through a speaker. The speaker is also built into the bow tie and outputs the converted speech clearly. For example, a normal voice saying "hello" is output from the speaker. This device is shaped like a bow tie, is easy to wear and inconspicuous, making it easy to use in daily life. For example, by wearing it during meetings or presentations, it converts the words spoken by hearing-impaired individuals into normal voice, enabling smooth communication. Thus, the voice conversion system can convert the words spoken by hearing-impaired individuals into normal voice, enabling smooth communication.
[0029] The voice conversion system according to this embodiment comprises a collection unit, an analysis unit, a conversion unit, and an output unit. The collection unit collects words spoken by a person with a hearing impairment. The collection unit, for example, incorporates a microphone to collect spoken voice with high accuracy. For example, if a person with a hearing impairment says "hello," the collection unit can collect that voice with its microphone. The analysis unit analyzes the voice data collected by the collection unit. The analysis unit, for example, uses AI to analyze the collected voice data and identify the spoken words. For example, the analysis unit can identify the word "hello" from the collected voice data. The conversion unit converts the voice data analyzed by the analysis unit into normal voice. The conversion unit converts the word identified using AI into normal voice. For example, the conversion unit can convert the word "hello" into normal voice. The output unit outputs the voice converted by the conversion unit. The output unit, for example, incorporates a speaker to output the converted voice clearly. For example, the output unit can output the normal voice "hello" from the speaker. As a result, the speech conversion system according to this embodiment can convert the words spoken by a person with a hearing impairment into normal speech, thereby enabling smooth communication.
[0030] The data collection unit collects speech uttered by hearing-impaired individuals. For example, the unit incorporates a microphone to capture spoken audio with high precision. Specifically, the unit uses a high-sensitivity microphone and incorporates noise-canceling technology to reduce ambient noise. This allows for accurate collection of even subtle sounds uttered by hearing-impaired individuals. Furthermore, the unit can capture not only audio but also the speaker's mouth movements and facial expressions with a camera. This allows for more accurate analysis by combining audio and video data. For example, if a hearing-impaired person says "hello," capturing not only the audio with a microphone but also their mouth movements and facial expressions with a camera allows for a more accurate understanding of their intent and emotions. The data collection unit transmits this data to a central database in real time, enabling the analysis unit to access it quickly. Additionally, by coordinating multiple microphones and cameras, the unit can identify the speaker's location and direction, providing an optimal collection environment. This allows the unit to collect speech uttered by hearing-impaired individuals with high precision and from multiple angles, improving the overall system performance.
[0031] The analysis unit analyzes the audio data collected by the collection unit. For example, the analysis unit uses AI to analyze the collected audio data and identify the spoken words. Specifically, the analysis unit uses speech recognition technology to convert the collected audio data into text data. The AI analyzes the waveform of the audio data, identifies phonemes and syllables, and combines them to recognize words. For example, if a person with a hearing impairment says "hello," the analysis unit analyzes the audio data, identifies the phonemes "ko," "n," "ni," "chi," and "ha," and combines them to recognize the word "hello." Furthermore, the analysis unit can also analyze the collected video data and improve the accuracy of speech recognition based on the mouth movements and facial expressions of the speaker. For example, if the speaker's mouth movements match "ko," "n," "ni," "chi," and "ha," the speech recognition result can be reinforced. The analysis unit is required to process this data in real time and output results quickly. In addition, the analysis unit can learn from past audio data and speech patterns to build speech recognition models tailored to individual speakers. This allows the analysis unit to analyze the collected audio data with high accuracy and quickly and accurately identify the spoken words.
[0032] The conversion unit converts the audio data analyzed by the analysis unit into normal speech. For example, the conversion unit converts words identified using AI into normal speech. Specifically, the conversion unit uses speech synthesis technology to convert text data into natural-sounding speech. The AI uses a speech synthesis model to reproduce the speaker's voice quality and intonation, generating natural pronunciation. For example, if the analysis unit identifies the word "hello," the conversion unit generates natural-sounding speech based on that text data. Furthermore, the conversion unit can also generate speech that reflects the speaker's emotions and intentions. For example, if the speaker says "hello" with a smile, the conversion unit generates a bright speech that reflects that emotion. The conversion unit generates this audio data in real time and transmits it to the output unit. The conversion unit can also support multiple languages, converting words spoken in different languages into the appropriate language. This allows the conversion unit to convert the audio data analyzed by the analysis unit into high-quality normal speech, enabling smooth communication.
[0033] The output unit outputs the audio converted by the conversion unit. The output unit, for example, incorporates a speaker to clearly output the converted audio. Specifically, the output unit uses a high-quality speaker to reproduce the converted audio with clear and natural sound quality. Furthermore, the output unit has a function to adjust the volume and tone of the audio, providing optimal audio output depending on the environment and situation. For example, it can output audio at a low volume in a quiet room and at a high volume in a noisy environment. The output unit can also achieve wide-range audio output by coordinating multiple speakers. This allows the output unit to output the converted audio clearly and at an appropriate volume, supporting smooth communication between people with hearing impairments and those around them. In addition to audio output, the output unit can also provide other feedback methods such as text display and vibration notifications. For example, it can display the converted words on a text display device simultaneously with the audio output, providing visual feedback. This allows the output unit to provide diverse communication methods between people with hearing impairments and those around them, enabling smooth information transmission.
[0034] The collection unit incorporates a microphone to improve the accuracy of spoken audio. For example, the collection unit incorporates a microphone to collect spoken audio with high accuracy. For example, the collection unit can use noise cancellation technology to remove ambient noise and improve the accuracy of spoken audio. The collection unit can also improve the accuracy of spoken audio by adjusting the sensitivity of the microphone. For example, by increasing the sensitivity of the microphone, the collection unit can collect even quiet voices. This improves the accuracy of analysis by collecting spoken audio with high accuracy. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can input the audio data collected by the microphone into a generating AI and have the generating AI perform an improvement on the accuracy of the audio data.
[0035] The analysis unit can analyze the collected audio data and identify the spoken words. For example, the analysis unit can use AI to analyze the collected audio data and identify the spoken words. For example, the analysis unit can use a speech recognition algorithm to identify the word "hello" from the collected audio data. The analysis unit can also use a dictionary database to identify the spoken words. For example, the analysis unit can compare the collected audio data with a dictionary database to identify the spoken words. This improves conversion accuracy by identifying the spoken words. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the collected audio data into a generating AI and have the generating AI perform the identification of the spoken words.
[0036] The conversion unit can convert identified words into normal speech. The conversion unit can convert identified words into normal speech using, for example, AI. For example, the conversion unit can convert the word "hello" into normal speech. The conversion unit converts identified words into normal speech based on speech data that the AI has previously learned. For example, the conversion unit can convert the word "hello" into normal speech using speech data that the AI has learned. As a result, by converting identified words into normal speech, the speech of a hearing-impaired person is output as normal speech. Some or all of the above processing in the conversion unit may be performed using, for example, AI, or without AI. For example, the conversion unit can input identified words into a generating AI and have the generating AI perform the conversion to normal speech.
[0037] The output unit can output the converted audio clearly. The output unit, for example, has a built-in speaker to output the converted audio clearly. For example, the output unit can output a normal voice saying "hello" from the speaker. The output unit can also output the converted audio clearly using noise cancellation technology. For example, the output unit can use noise cancellation technology to remove ambient noise and output the converted audio clearly. This ensures that the speech of a person with hearing impairment is clearly conveyed by outputting the converted audio clearly. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the converted audio data into a generating AI and have the generating AI perform clarification of the audio output.
[0038] The voice conversion system is shaped like a bow tie, making it easy to wear and inconspicuous. Its bow tie shape makes it easy to wear and use in everyday life. For example, wearing it during meetings or presentations converts the speech of a hearing-impaired person into normal speech, facilitating smooth communication. The bow tie's visibility can be reduced by adjusting its color, shape, and size. For instance, matching the bow tie's color to the wearer's clothing makes it less noticeable. Simplifying the bow tie's shape also reduces its visibility. This makes the bow tie design easy to use in everyday life.
[0039] The sound collection unit can analyze ambient sounds in real time and collect audio while performing noise cancellation. For example, if the ambient noise is loud, the sound collection unit can enhance noise cancellation to collect audio. For example, if the ambient noise is loud, the sound collection unit can enhance noise cancellation to collect clear audio. The sound collection unit can also minimize noise cancellation in quiet environments to collect natural audio. For example, if the sound collection unit minimizes noise cancellation in quiet environments to collect natural audio. The sound collection unit can also cancel out sudden noises in real time and collect audio. For example, if a sudden noise occurs, the sound collection unit can cancel it in real time to collect clear audio. This enables clear audio collection through noise cancellation. Some or all of the above processing in the sound collection unit may be performed using AI, for example, or without AI. For example, the sound collection unit can input ambient sound data into a generating AI and have the generating AI adjust the noise cancellation.
[0040] The collection unit can learn the user's speech patterns and automatically adjust the collection timing. For example, if the user speaks slowly, the collection unit can adjust the collection timing to match that pace. For example, if the user speaks slowly, the collection unit can adjust the collection timing to match that pace, enabling optimal audio collection. The collection unit can also adjust the collection timing to match the user's fast speech. For example, if the user speaks quickly, the collection unit can adjust the collection timing to match that pace, enabling optimal audio collection. The collection unit can also adjust the collection timing to account for pauses in the user's speech. For example, if the user speaks with pauses, the collection unit can adjust the collection timing to account for those pauses, enabling optimal audio collection. This improves the accuracy of audio collection by adjusting the collection timing according to the user's speech patterns. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can input the user's speech pattern data into a generating AI and have the generating AI perform the adjustment of the collection timing.
[0041] The collection unit can prioritize the collection of highly relevant audio by considering the user's geographical location information during audio collection. For example, if the user is in a specific location, the collection unit can prioritize the collection of audio related to that location. For example, if the collection unit is in a specific location, the collection unit can perform optimal audio collection by prioritizing the collection of audio related to that location. The collection unit can also prioritize the collection of audio related to the user's destination if the user is on the move. For example, if the collection unit is on the move, the collection unit can perform optimal audio collection by prioritizing the collection of audio related to the user's destination. The collection unit can also prioritize the collection of audio related to an event if the user is participating in that event. For example, if the collection unit is participating in an event, the collection unit can perform optimal audio collection by prioritizing the collection of audio related to that event. This allows for the priority collection of highly relevant audio through audio collection based on geographical location information. Some or all of the above processing in the collection unit may be performed using AI, for example, or without using AI. For example, the data collection unit can input the user's geographic location data into the generating AI, allowing the generating AI to determine the priority of voice messages.
[0042] The collection unit can analyze the user's social media activity and collect relevant audio during audio collection. For example, the collection unit can collect audio related to topics the user is discussing on social media. For example, by collecting audio related to topics the user is discussing on social media, the collection unit can perform optimal audio collection. The collection unit can also collect audio related to statements made by people the user follows on social media. For example, by collecting audio related to statements made by people the user follows on social media, the collection unit can perform optimal audio collection. The collection unit can also collect audio related to topics of groups the user participates in on social media. For example, by collecting audio related to topics of groups the user participates in on social media, the collection unit can perform optimal audio collection. This allows for the collection of highly relevant audio based on social media activity. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can input the user's social media activity data into a generating AI and have the generating AI determine the priority of the audio.
[0043] The analysis unit can remove background noise from audio data during analysis, thereby improving the accuracy of spoken words. For example, the analysis unit can improve the accuracy of audio data by removing ambient noise during analysis. The analysis unit can also improve the accuracy of spoken words by removing wind noise during analysis. For example, the analysis unit can improve the accuracy of spoken words by removing wind noise during analysis. The analysis unit can also improve the accuracy of audio data by removing echoes during analysis. For example, the analysis unit can improve the accuracy of audio data by removing echoes during analysis. As a result, the accuracy of spoken words is improved by removing background noise. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input background noise from the audio data into a generating AI and have the generating AI perform noise reduction.
[0044] The analysis unit can improve analysis accuracy by considering the user's speech patterns during analysis. For example, if the user speaks slowly, the analysis unit can improve analysis accuracy to match that pace. For example, if the user speaks slowly, the analysis unit can improve analysis accuracy to match that pace, thereby performing optimal analysis. The analysis unit can also improve analysis accuracy to match the user's fast speech. For example, if the user speaks quickly, the analysis unit can improve analysis accuracy to match that pace, thereby performing optimal analysis. The analysis unit can also improve analysis accuracy by considering pauses in the user's speech. For example, if the user speaks with pauses, the analysis unit can improve analysis accuracy by considering those pauses, thereby performing optimal analysis. As a result, analysis accuracy is improved based on the user's speech patterns. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's speech pattern data into a generating AI and have the generating AI perform the improvement of analysis accuracy.
[0045] The analysis unit can determine the priority of analysis based on the collection date of the audio data during analysis. For example, if the collected audio data is recent, the analysis unit will prioritize its analysis. For example, by prioritizing the analysis of recent audio data, the analysis unit can perform optimal analysis. The analysis unit can also postpone the analysis of older audio data. For example, by postponing the analysis of older audio data, the analysis unit can perform optimal analysis. The analysis unit can also determine the priority of analysis based on the importance of the collected audio data. For example, by determining the priority of analysis based on the importance of the collected audio data, the analysis unit can perform optimal analysis. This allows important data to be analyzed preferentially by determining the priority of analysis based on the collection date. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without using AI. For example, the analysis unit can input the audio data collection date data into a generating AI and have the generating AI perform the determination of the analysis priority.
[0046] The analysis unit can adjust the order of analysis based on the relevance of the audio data during analysis. For example, if the collected audio data is highly relevant, the analysis unit will prioritize analyzing that data. For example, the analysis unit can perform optimal analysis by prioritizing the analysis of highly relevant collected audio data. The analysis unit can also postpone the analysis of less relevant collected audio data. For example, the analysis unit can perform optimal analysis by postponing the analysis of less relevant collected audio data. The analysis unit can also adjust the order of analysis based on the content of the collected audio data. For example, the analysis unit can perform optimal analysis by adjusting the order of analysis based on the content of the collected audio data. This enables efficient analysis by adjusting the order of analysis based on relevance. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the relevance data of the audio data into a generating AI and have the generating AI perform the adjustment of the order of analysis.
[0047] The conversion unit can learn the user's speech patterns during conversion and perform optimal speech conversion. For example, if the user speaks slowly, the conversion unit can convert the speech to match that pace. The conversion unit can also convert the speech to match the pace if the user speaks quickly. The conversion unit can also convert the speech to match the pace if the user speaks quickly. The conversion unit can also take into account pauses when the user speaks. This improves conversion accuracy through speech conversion based on the user's speech patterns. Some or all of the above processing in the conversion unit may be performed using AI, for example, or without AI. For example, the conversion unit can input the user's speech pattern data into a generating AI and have the generating AI perform the speech conversion.
[0048] The conversion unit can add a multilingual conversion function to support different languages during conversion. For example, if the user speaks English, the conversion unit can convert that speech to Japanese. For example, if the user speaks English, the conversion unit can perform optimal speech conversion by converting that speech to Japanese. The conversion unit can also convert the user's speech to English if they speak French. For example, if the user speaks French, the conversion unit can perform optimal speech conversion by converting that speech to English. The conversion unit can also convert the user's speech to French if they speak Spanish. For example, if the user speaks Spanish, the conversion unit can perform optimal speech conversion by converting that speech to French. This enables speech conversion between different languages through the multilingual conversion function. Some or all of the above processing in the conversion unit may be performed using AI, for example, or without AI. For example, the conversion unit can input speech data in different languages into a generating AI and have the generating AI perform multilingual conversion.
[0049] The conversion unit can adjust the conversion algorithm based on the location where the audio data was collected during conversion. For example, when the conversion unit converts audio collected by a user outdoors, it adjusts the conversion algorithm to take wind noise into consideration. For example, when the conversion unit converts audio collected by a user outdoors, it can perform optimal audio conversion by adjusting the conversion algorithm to take wind noise into consideration. The conversion unit can also adjust the conversion algorithm to take echo into consideration when converting audio collected by a user indoors. For example, when the conversion unit converts audio collected by a user indoors, it can perform optimal audio conversion by adjusting the conversion algorithm to take echo into consideration. The conversion unit can also adjust the conversion algorithm to take engine noise into consideration when converting audio collected by a user inside a car. For example, when the conversion unit converts audio collected by a user inside a car, it can perform optimal audio conversion by adjusting the conversion algorithm to take engine noise into consideration. This makes optimal audio conversion possible by adjusting the conversion algorithm based on the collection location. Some or all of the above processing in the conversion unit may be performed using AI, for example, or without using AI. For example, the conversion unit can input data on the location where the audio data was collected into the generating AI, and have the generating AI adjust the conversion algorithm.
[0050] The conversion unit can analyze the user's social media activity during conversion and perform relevant voice conversions. For example, the conversion unit can convert voices related to topics the user is discussing on social media. For example, by converting voices related to topics the user is discussing on social media, the conversion unit can perform optimal voice conversions. The conversion unit can also convert voices related to statements made by people the user follows on social media. For example, by converting voices related to statements made by people the user follows on social media, the conversion unit can perform optimal voice conversions. The conversion unit can also convert voices related to topics in groups the user participates in on social media. For example, by converting voices related to topics in groups the user participates in on social media, the conversion unit can perform optimal voice conversions. This enables highly relevant voice conversions based on social media activity. Some or all of the above processing in the conversion unit may be performed using AI, for example, or without AI. For example, the conversion unit can input the user's social media activity data into a generating AI and have the generating AI perform the voice conversion.
[0051] The output unit can analyze ambient sounds in real time during output and adjust the audio output accordingly. For example, if the surroundings are noisy, the output unit can increase the volume of the audio output. For example, if the surroundings are noisy, the output unit can achieve optimal audio output by increasing the volume of the audio output. The output unit can also decrease the volume of the audio output if the surroundings are quiet. For example, if the surroundings are quiet, the output unit can achieve optimal audio output by decreasing the volume of the audio output. Furthermore, if a sudden sound occurs, the output unit can analyze that sound in real time and adjust the audio output accordingly. For example, if a sudden sound occurs, the output unit can achieve optimal audio output by analyzing that sound in real time. This enables clear audio output based on ambient sounds. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input ambient sound data into a generating AI and have the generating AI perform the audio output adjustment.
[0052] The output unit can improve the accuracy of voice output by considering the user's speech pattern during output. For example, if the user speaks slowly, the output unit can adjust the voice output to match that pace. For example, if the user speaks slowly, the output unit can adjust the voice output to match that pace, thereby achieving optimal voice output. The output unit can also adjust the voice output to match the user's fast speech, for example, by adjusting the voice output to match that pace, thereby achieving optimal voice output. Furthermore, if the user speaks with pauses, the output unit can adjust the voice output to take those pauses into account. For example, if the user speaks with pauses, the output unit can adjust the voice output to take those pauses into account, thereby achieving optimal voice output. This improves output accuracy through voice output based on the user's speech pattern. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the user's speech pattern data into a generating AI and have the generating AI perform the adjustment of the voice output.
[0053] The output unit can perform optimal audio output by considering the user's geographical location information when outputting audio. For example, if the user is in a specific location, the output unit can prioritize outputting audio related to that location. The output unit can also prioritize outputting audio related to the user's destination if the user is on the move. The output unit can also prioritize outputting audio related to the user's destination if the user is participating in a specific event. The output unit can also prioritize outputting audio related to that event if the user is participating in a specific event. This allows for the prioritization of highly relevant audio output based on geographical location information. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the user's geographical location information data into a generating AI and have the generating AI adjust the audio output.
[0054] The output unit can analyze the user's social media activity and output relevant audio when outputting audio. For example, the output unit can output audio related to topics the user is discussing on social media. For example, by outputting audio related to topics the user is discussing on social media, the output unit can achieve optimal audio output. The output unit can also output audio related to statements made by people the user follows on social media. For example, by outputting audio related to statements made by people the user follows on social media, the output unit can achieve optimal audio output. The output unit can also output audio related to topics the user is participating in on social media. For example, by outputting audio related to topics the user is participating in on social media, the output unit can achieve optimal audio output. This allows for the output of highly relevant audio based on social media activity. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the user's social media activity data into a generating AI and have the generating AI adjust the audio output.
[0055] The shape of a bow tie can be improved by changing the material it is made from. For example, the shape of a bow tie can be improved by using a soft material. Alternatively, the shape of a bow tie can be made to withstand long-term use by using a highly durable material. Furthermore, the shape of a bow tie can be made to provide a comfortable fit by using a breathable material. Thus, changing the material improves both the fit and durability. Some or all of the above processes regarding the shape of a bow tie may be performed using AI, or not. For example, material selection data can be input into a generating AI, and the AI can then perform the material change.
[0056] The shape of the bow tie can incorporate additional functions, such as a solar panel to extend battery life. The shape of the bow tie can also incorporate a battery pack to enable extended use. Furthermore, the shape of the bow tie can add a wireless charging function to eliminate the need for manual charging. This allows for extended battery life and longer use through these additional functions. Some or all of the above processing in the shape of the bow tie may be performed using AI, or not. For example, the design data for the additional functions can be input into a generating AI, and the generating AI can be made to incorporate the additional functions.
[0057] The shape of a bow tie can be applied to other accessories. For example, it can be applied to brooches and necklaces. For instance, applying the shape of a bow tie to a brooch expands the range of ways it can be worn. Also, applying the shape of a bow tie to a necklace can enhance its fashion appeal. For example, applying the shape of a bow tie to a necklace can enhance its fashion appeal. Furthermore, applying the shape of a bow tie to a hair accessory allows for a variety of styles. For example, applying the shape of a bow tie to a hair accessory allows for a variety of styles. This expands the range of ways it can be worn by applying it to other accessories. Some or all of the above processing regarding the shape of a bow tie may be performed using AI, or not. For example, the shape of a bow tie can be used to input application data for other accessories into a generating AI, which can then execute the application.
[0058] The shape of the bow tie can be customized, allowing users to change it to their liking. For example, the shape of the bow tie can allow users to select a design, providing a bow tie tailored to their individual preferences. The shape of the bow tie can also allow users to select a color, providing a bow tie tailored to their individual style. For example, the shape of the bow tie can allow users to select a color, providing a bow tie tailored to their individual style. The shape of the bow tie can also allow users to select a material, providing a bow tie tailored to their individual fit. For example, the shape of the bow tie can allow users to select a material, providing a bow tie tailored to their individual fit. This allows for customization of the design, enabling users to wear the bow tie according to their preferences. Some or all of the above processing in the shape of the bow tie may be performed using AI, or not. For example, the shape of the bow tie can input customization data into a generating AI, and have the generating AI perform the design customization.
[0059] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0060] The voice conversion system can also be equipped with a function to learn the user's speech patterns and automatically adjust the sensitivity of the collection unit. For example, if the user speaks slowly, the collection unit can adjust its sensitivity to match the user's pace for optimal voice collection. It can also adjust its sensitivity to match the user's pace if the user speaks quickly. Furthermore, if the user pauses while speaking, the system can adjust its sensitivity to account for those pauses. This improves the accuracy of voice collection by adjusting the sensitivity according to the user's speech patterns.
[0061] The analysis unit can also be equipped with a function to learn the user's speech patterns and automatically adjust the analysis algorithm. For example, if the user speaks slowly, the analysis unit can adjust the analysis algorithm to match that pace and perform optimal analysis. It can also adjust the analysis algorithm to match the user's speaking pace if the user speaks quickly. Furthermore, if the user pauses while speaking, the analysis algorithm can be adjusted to take those pauses into account. As a result, the accuracy of the analysis is improved by adjusting the analysis algorithm according to the user's speech patterns.
[0062] The conversion unit can also be equipped with a multilingual conversion function to support even more languages. For example, if a user speaks English, their voice can be converted to Japanese. If a user speaks French, their voice can be converted to English. Furthermore, if a user speaks Spanish, their voice can be converted to French. This multilingual conversion function enables voice conversion between different languages, facilitating smoother international communication.
[0063] The output unit can also be equipped with a function to analyze ambient noise in real time and adjust the audio output. For example, if the surroundings are noisy, the audio output volume can be increased. Conversely, if the surroundings are quiet, the audio output volume can be decreased. Furthermore, if a sudden sound occurs, it can be analyzed in real time and the audio output can be adjusted accordingly. This enables clear audio output based on ambient noise.
[0064] The voice conversion system can also be equipped with functions to collect and analyze voice data while considering the user's geographical location. For example, the collection unit can prioritize collecting voice data related to a specific location if the user is in that location. Similarly, the analysis unit can prioritize analyzing collected voice data if it is related to a specific location. Furthermore, the conversion unit can adjust its conversion algorithm based on the collection location. This allows for the priority processing of highly relevant voice data through voice collection and analysis based on geographical location information.
[0065] The following briefly describes the processing flow for example form 1.
[0066] Step 1: The collection unit collects the words spoken by the hearing impaired person. The collection unit, for example, has a built-in microphone to collect the spoken voice with high precision. For example, if the hearing impaired person says "hello," the collection unit can collect that voice with its microphone. Step 2: The analysis unit analyzes the audio data collected by the collection unit. The analysis unit analyzes the collected audio data using, for example, AI, and identifies the spoken words. For example, the analysis unit can identify the word "hello" from the collected audio data. Step 3: The conversion unit converts the audio data analyzed by the analysis unit into normal speech. The conversion unit converts words identified using AI into normal speech. For example, the conversion unit can convert the word "hello" into normal speech. Step 4: The output unit outputs the audio converted by the conversion unit. The output unit, for example, has a built-in speaker to output the converted audio clearly. For example, the output unit can output a normal voice saying "hello" through the speaker.
[0067] (Example of form 2) The voice conversion system according to an embodiment of the present invention is a system that converts words spoken by a person with a hearing impairment into normal speech. This voice conversion system collects the words spoken by the person with a hearing impairment using a microphone, and the collected voice data is analyzed by an AI and converted into normal speech. The converted speech is output through a speaker. This system is shaped like a bow tie, is easy to wear and inconspicuous, and is easy to use in daily life. For example, the system collects the words spoken by a person with a hearing impairment using a microphone. The microphone is built into the bow tie and collects the spoken voice with high accuracy. For example, if a person with a hearing impairment says "hello," that voice is collected by the microphone. Next, the AI analyzes the collected voice data. The AI analyzes the collected voice data and identifies the spoken word. For example, it identifies the word "hello" from the collected voice data. Then, the AI converts the identified word into normal speech. For example, it converts the word "hello" into normal speech. This conversion is performed based on voice data that the AI has learned in advance. Finally, the converted speech is output through a speaker. The speaker is also built into the bow tie and outputs the converted speech clearly. For example, a normal voice saying "hello" is output from the speaker. This device is shaped like a bow tie, is easy to wear and inconspicuous, making it easy to use in daily life. For example, by wearing it during meetings or presentations, it converts the words spoken by hearing-impaired individuals into normal voice, enabling smooth communication. Thus, the voice conversion system can convert the words spoken by hearing-impaired individuals into normal voice, enabling smooth communication.
[0068] The voice conversion system according to this embodiment comprises a collection unit, an analysis unit, a conversion unit, and an output unit. The collection unit collects words spoken by a person with a hearing impairment. The collection unit, for example, incorporates a microphone to collect spoken voice with high accuracy. For example, if a person with a hearing impairment says "hello," the collection unit can collect that voice with its microphone. The analysis unit analyzes the voice data collected by the collection unit. The analysis unit, for example, uses AI to analyze the collected voice data and identify the spoken words. For example, the analysis unit can identify the word "hello" from the collected voice data. The conversion unit converts the voice data analyzed by the analysis unit into normal voice. The conversion unit converts the word identified using AI into normal voice. For example, the conversion unit can convert the word "hello" into normal voice. The output unit outputs the voice converted by the conversion unit. The output unit, for example, incorporates a speaker to output the converted voice clearly. For example, the output unit can output the normal voice "hello" from the speaker. As a result, the speech conversion system according to this embodiment can convert the words spoken by a person with a hearing impairment into normal speech, thereby enabling smooth communication.
[0069] The data collection unit collects speech uttered by hearing-impaired individuals. For example, the unit incorporates a microphone to capture spoken audio with high precision. Specifically, the unit uses a high-sensitivity microphone and incorporates noise-canceling technology to reduce ambient noise. This allows for accurate collection of even subtle sounds uttered by hearing-impaired individuals. Furthermore, the unit can capture not only audio but also the speaker's mouth movements and facial expressions with a camera. This allows for more accurate analysis by combining audio and video data. For example, if a hearing-impaired person says "hello," capturing not only the audio with a microphone but also their mouth movements and facial expressions with a camera allows for a more accurate understanding of their intent and emotions. The data collection unit transmits this data to a central database in real time, enabling the analysis unit to access it quickly. Additionally, by coordinating multiple microphones and cameras, the unit can identify the speaker's location and direction, providing an optimal collection environment. This allows the unit to collect speech uttered by hearing-impaired individuals with high precision and from multiple angles, improving the overall system performance.
[0070] The analysis unit analyzes the audio data collected by the collection unit. For example, the analysis unit uses AI to analyze the collected audio data and identify the spoken words. Specifically, the analysis unit uses speech recognition technology to convert the collected audio data into text data. The AI analyzes the waveform of the audio data, identifies phonemes and syllables, and combines them to recognize words. For example, if a person with a hearing impairment says "hello," the analysis unit analyzes the audio data, identifies the phonemes "ko," "n," "ni," "chi," and "ha," and combines them to recognize the word "hello." Furthermore, the analysis unit can also analyze the collected video data and improve the accuracy of speech recognition based on the mouth movements and facial expressions of the speaker. For example, if the speaker's mouth movements match "ko," "n," "ni," "chi," and "ha," the speech recognition result can be reinforced. The analysis unit is required to process this data in real time and output results quickly. In addition, the analysis unit can learn from past audio data and speech patterns to build speech recognition models tailored to individual speakers. This allows the analysis unit to analyze the collected audio data with high accuracy and quickly and accurately identify the spoken words.
[0071] The conversion unit converts the audio data analyzed by the analysis unit into normal speech. For example, the conversion unit converts words identified using AI into normal speech. Specifically, the conversion unit uses speech synthesis technology to convert text data into natural-sounding speech. The AI uses a speech synthesis model to reproduce the speaker's voice quality and intonation, generating natural pronunciation. For example, if the analysis unit identifies the word "hello," the conversion unit generates natural-sounding speech based on that text data. Furthermore, the conversion unit can also generate speech that reflects the speaker's emotions and intentions. For example, if the speaker says "hello" with a smile, the conversion unit generates a bright speech that reflects that emotion. The conversion unit generates this audio data in real time and transmits it to the output unit. The conversion unit can also support multiple languages, converting words spoken in different languages into the appropriate language. This allows the conversion unit to convert the audio data analyzed by the analysis unit into high-quality normal speech, enabling smooth communication.
[0072] The output unit outputs the audio converted by the conversion unit. The output unit, for example, incorporates a speaker to clearly output the converted audio. Specifically, the output unit uses a high-quality speaker to reproduce the converted audio with clear and natural sound quality. Furthermore, the output unit has a function to adjust the volume and tone of the audio, providing optimal audio output depending on the environment and situation. For example, it can output audio at a low volume in a quiet room and at a high volume in a noisy environment. The output unit can also achieve wide-range audio output by coordinating multiple speakers. This allows the output unit to output the converted audio clearly and at an appropriate volume, supporting smooth communication between people with hearing impairments and those around them. In addition to audio output, the output unit can also provide other feedback methods such as text display and vibration notifications. For example, it can display the converted words on a text display device simultaneously with the audio output, providing visual feedback. This allows the output unit to provide diverse communication methods between people with hearing impairments and those around them, enabling smooth information transmission.
[0073] The collection unit incorporates a microphone to improve the accuracy of spoken audio. For example, the collection unit incorporates a microphone to collect spoken audio with high accuracy. For example, the collection unit can use noise cancellation technology to remove ambient noise and improve the accuracy of spoken audio. The collection unit can also improve the accuracy of spoken audio by adjusting the sensitivity of the microphone. For example, by increasing the sensitivity of the microphone, the collection unit can collect even quiet voices. This improves the accuracy of analysis by collecting spoken audio with high accuracy. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can input the audio data collected by the microphone into a generating AI and have the generating AI perform an improvement on the accuracy of the audio data.
[0074] The analysis unit can analyze the collected audio data and identify the spoken words. For example, the analysis unit can use AI to analyze the collected audio data and identify the spoken words. For example, the analysis unit can use a speech recognition algorithm to identify the word "hello" from the collected audio data. The analysis unit can also use a dictionary database to identify the spoken words. For example, the analysis unit can compare the collected audio data with a dictionary database to identify the spoken words. This improves conversion accuracy by identifying the spoken words. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the collected audio data into a generating AI and have the generating AI perform the identification of the spoken words.
[0075] The conversion unit can convert identified words into normal speech. The conversion unit can convert identified words into normal speech using, for example, AI. For example, the conversion unit can convert the word "hello" into normal speech. The conversion unit converts identified words into normal speech based on speech data that the AI has previously learned. For example, the conversion unit can convert the word "hello" into normal speech using speech data that the AI has learned. As a result, by converting identified words into normal speech, the speech of a hearing-impaired person is output as normal speech. Some or all of the above processing in the conversion unit may be performed using, for example, AI, or without AI. For example, the conversion unit can input identified words into a generating AI and have the generating AI perform the conversion to normal speech.
[0076] The output unit can output the converted audio clearly. The output unit, for example, has a built-in speaker to output the converted audio clearly. For example, the output unit can output a normal voice saying "hello" from the speaker. The output unit can also output the converted audio clearly using noise cancellation technology. For example, the output unit can use noise cancellation technology to remove ambient noise and output the converted audio clearly. This ensures that the speech of a person with hearing impairment is clearly conveyed by outputting the converted audio clearly. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the converted audio data into a generating AI and have the generating AI perform clarification of the audio output.
[0077] The voice conversion system is shaped like a bow tie, making it easy to wear and inconspicuous. Its bow tie shape makes it easy to wear and use in everyday life. For example, wearing it during meetings or presentations converts the speech of a hearing-impaired person into normal speech, facilitating smooth communication. The bow tie's visibility can be reduced by adjusting its color, shape, and size. For instance, matching the bow tie's color to the wearer's clothing makes it less noticeable. Simplifying the bow tie's shape also reduces its visibility. This makes the bow tie design easy to use in everyday life.
[0078] The collection unit can estimate the user's emotions and adjust the voice collection sensitivity based on the estimated emotions. For example, if the user is nervous, the collection unit can increase the collection sensitivity to capture even soft voices. For example, if the user is nervous, the collection unit can increase the collection sensitivity to capture even soft voices. The collection unit can also return the collection sensitivity to normal when the user is relaxed to capture natural voices. For example, if the user is relaxed, the collection unit can return the collection sensitivity to normal to capture natural voices. The collection unit can also adjust the collection sensitivity to suppress excessive noise when the user is excited. For example, if the user is excited, the collection unit can suppress excessive noise by adjusting the collection sensitivity. This allows for optimal voice collection by adjusting the collection sensitivity according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can input user emotion data into a generating AI and have the generating AI adjust the data collection sensitivity.
[0079] The sound collection unit can analyze ambient sounds in real time and collect audio while performing noise cancellation. For example, if the ambient noise is loud, the sound collection unit can enhance noise cancellation to collect audio. For example, if the ambient noise is loud, the sound collection unit can enhance noise cancellation to collect clear audio. The sound collection unit can also minimize noise cancellation in quiet environments to collect natural audio. For example, if the sound collection unit minimizes noise cancellation in quiet environments to collect natural audio. The sound collection unit can also cancel out sudden noises in real time and collect audio. For example, if a sudden noise occurs, the sound collection unit can cancel it in real time to collect clear audio. This enables clear audio collection through noise cancellation. Some or all of the above processing in the sound collection unit may be performed using AI, for example, or without AI. For example, the sound collection unit can input ambient sound data into a generating AI and have the generating AI adjust the noise cancellation.
[0080] The collection unit can learn the user's speech patterns and automatically adjust the collection timing. For example, if the user speaks slowly, the collection unit can adjust the collection timing to match that pace. For example, if the user speaks slowly, the collection unit can adjust the collection timing to match that pace, enabling optimal audio collection. The collection unit can also adjust the collection timing to match the user's fast speech. For example, if the user speaks quickly, the collection unit can adjust the collection timing to match that pace, enabling optimal audio collection. The collection unit can also adjust the collection timing to account for pauses in the user's speech. For example, if the user speaks with pauses, the collection unit can adjust the collection timing to account for those pauses, enabling optimal audio collection. This improves the accuracy of audio collection by adjusting the collection timing according to the user's speech patterns. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can input the user's speech pattern data into a generating AI and have the generating AI perform the adjustment of the collection timing.
[0081] The audio collection unit can estimate the user's emotions and determine the priority of audio to collect based on the estimated emotions. For example, if the user is nervous, the collection unit can prioritize collecting important audio. For example, if the user is nervous, the collection unit can achieve optimal audio collection by prioritizing the collection of important audio. Also, if the user is relaxed, the collection unit can collect all audio equally. For example, if the user is relaxed, the collection unit can achieve optimal audio collection by collecting all audio equally. Also, if the user is excited, the collection unit can prioritize collecting audio with strong emotions. For example, if the user is excited, the collection unit can achieve optimal audio collection by prioritizing the collection of audio with strong emotions. This allows for the priority collection of important audio by determining the priority of audio according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the processing described above in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can input user emotion data into a generating AI and have the generating AI determine the priority of the voices.
[0082] The collection unit can prioritize the collection of highly relevant audio by considering the user's geographical location information during audio collection. For example, if the user is in a specific location, the collection unit can prioritize the collection of audio related to that location. For example, if the collection unit is in a specific location, the collection unit can perform optimal audio collection by prioritizing the collection of audio related to that location. The collection unit can also prioritize the collection of audio related to the user's destination if the user is on the move. For example, if the collection unit is on the move, the collection unit can perform optimal audio collection by prioritizing the collection of audio related to the user's destination. The collection unit can also prioritize the collection of audio related to an event if the user is participating in that event. For example, if the collection unit is participating in an event, the collection unit can perform optimal audio collection by prioritizing the collection of audio related to that event. This allows for the priority collection of highly relevant audio through audio collection based on geographical location information. Some or all of the above processing in the collection unit may be performed using AI, for example, or without using AI. For example, the data collection unit can input the user's geographic location data into the generating AI, allowing the generating AI to determine the priority of voice messages.
[0083] The collection unit can analyze the user's social media activity and collect relevant audio during audio collection. For example, the collection unit can collect audio related to topics the user is discussing on social media. For example, by collecting audio related to topics the user is discussing on social media, the collection unit can perform optimal audio collection. The collection unit can also collect audio related to statements made by people the user follows on social media. For example, by collecting audio related to statements made by people the user follows on social media, the collection unit can perform optimal audio collection. The collection unit can also collect audio related to topics of groups the user participates in on social media. For example, by collecting audio related to topics of groups the user participates in on social media, the collection unit can perform optimal audio collection. This allows for the collection of highly relevant audio based on social media activity. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can input the user's social media activity data into a generating AI and have the generating AI determine the priority of the audio.
[0084] The analysis unit can estimate the user's emotions and adjust the analysis algorithm based on the estimated emotions. For example, if the user is tense, the analysis unit can adjust the analysis algorithm to prioritize accuracy. For example, if the user is tense, the analysis unit can perform optimal analysis by adjusting the analysis algorithm to prioritize accuracy. The analysis unit can also adjust the analysis algorithm to prioritize speed if the user is relaxed. For example, if the user is relaxed, the analysis unit can perform optimal analysis by adjusting the analysis algorithm to prioritize speed. The analysis unit can also adjust the analysis algorithm to prioritize balance if the user is excited. For example, if the analysis unit is excited, the analysis unit can perform optimal analysis by adjusting the analysis algorithm to prioritize balance. This improves the accuracy of the analysis by adjusting the analysis algorithm according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input user emotion data into the generating AI and have the generating AI adjust the analysis algorithm.
[0085] The analysis unit can remove background noise from audio data during analysis, thereby improving the accuracy of spoken words. For example, the analysis unit can improve the accuracy of audio data by removing ambient noise during analysis. The analysis unit can also improve the accuracy of spoken words by removing wind noise during analysis. For example, the analysis unit can improve the accuracy of spoken words by removing wind noise during analysis. The analysis unit can also improve the accuracy of audio data by removing echoes during analysis. For example, the analysis unit can improve the accuracy of audio data by removing echoes during analysis. As a result, the accuracy of spoken words is improved by removing background noise. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input background noise from the audio data into a generating AI and have the generating AI perform noise reduction.
[0086] The analysis unit can improve analysis accuracy by considering the user's speech patterns during analysis. For example, if the user speaks slowly, the analysis unit can improve analysis accuracy to match that pace. For example, if the user speaks slowly, the analysis unit can improve analysis accuracy to match that pace, thereby performing optimal analysis. The analysis unit can also improve analysis accuracy to match the user's fast speech. For example, if the user speaks quickly, the analysis unit can improve analysis accuracy to match that pace, thereby performing optimal analysis. The analysis unit can also improve analysis accuracy by considering pauses in the user's speech. For example, if the user speaks with pauses, the analysis unit can improve analysis accuracy by considering those pauses, thereby performing optimal analysis. As a result, analysis accuracy is improved based on the user's speech patterns. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's speech pattern data into a generating AI and have the generating AI perform the improvement of analysis accuracy.
[0087] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated user emotions. For example, if the user is tense, the analysis unit can provide a simple and highly visible display method. For example, if the user is tense, the analysis unit can provide a simple and highly visible display method to display the optimal analysis results. The analysis unit can also provide a display method that includes detailed information if the user is relaxed. For example, if the user is relaxed, the analysis unit can provide a display method that includes detailed information to display the optimal analysis results. The analysis unit can also provide a visually stimulating display method if the user is excited. For example, if the user is excited, the analysis unit can provide a visually stimulating display method to display the optimal analysis results. This improves visibility by adjusting the display method according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input user emotion data into a generating AI and have the generating AI adjust the display method.
[0088] The analysis unit can determine the priority of analysis based on the collection date of the audio data during analysis. For example, if the collected audio data is recent, the analysis unit will prioritize its analysis. For example, by prioritizing the analysis of recent audio data, the analysis unit can perform optimal analysis. The analysis unit can also postpone the analysis of older audio data. For example, by postponing the analysis of older audio data, the analysis unit can perform optimal analysis. The analysis unit can also determine the priority of analysis based on the importance of the collected audio data. For example, by determining the priority of analysis based on the importance of the collected audio data, the analysis unit can perform optimal analysis. This allows important data to be analyzed preferentially by determining the priority of analysis based on the collection date. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the audio data collection date data into a generating AI and have the generating AI perform the determination of the analysis priority.
[0089] The analysis unit can adjust the order of analysis based on the relevance of the audio data during analysis. For example, if the collected audio data is highly relevant, the analysis unit will prioritize analyzing that data. For example, the analysis unit can perform optimal analysis by prioritizing the analysis of highly relevant collected audio data. The analysis unit can also postpone the analysis of less relevant collected audio data. For example, the analysis unit can perform optimal analysis by postponing the analysis of less relevant collected audio data. The analysis unit can also adjust the order of analysis based on the content of the collected audio data. For example, the analysis unit can perform optimal analysis by adjusting the order of analysis based on the content of the collected audio data. This enables efficient analysis by adjusting the order of analysis based on relevance. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the relevance data of the audio data into a generating AI and have the generating AI perform the adjustment of the order of analysis.
[0090] The voice conversion unit can estimate the user's emotions and adjust the tone and pitch of the voice conversion based on the estimated emotions. For example, if the user is nervous, the voice conversion unit can convert the voice in a calm tone. For example, if the user is nervous, the voice conversion unit can perform optimal voice conversion by converting the voice in a calm tone. The voice conversion unit can also convert the voice in a natural tone if the user is relaxed. For example, if the user is relaxed, the voice conversion unit can perform optimal voice conversion by converting the voice in a natural tone. The voice conversion unit can also convert the voice in a bright tone if the user is excited. For example, if the user is excited, the voice conversion unit can perform optimal voice conversion by converting the voice in a bright tone. This enables natural voice conversion by adjusting the tone and pitch according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples. Some or all of the above-described processing in the conversion unit may be performed using AI, for example, or without AI. For example, the conversion unit can input user emotion data into a generating AI and have the generating AI perform tone and pitch adjustments.
[0091] The conversion unit can learn the user's speech patterns during conversion and perform optimal speech conversion. For example, if the user speaks slowly, the conversion unit can convert the speech to match that pace. The conversion unit can also convert the speech to match the pace if the user speaks quickly. The conversion unit can also convert the speech to match the pace if the user speaks quickly. The conversion unit can also take into account pauses when the user speaks. This improves conversion accuracy through speech conversion based on the user's speech patterns. Some or all of the above processing in the conversion unit may be performed using AI, for example, or without AI. For example, the conversion unit can input the user's speech pattern data into a generating AI and have the generating AI perform the speech conversion.
[0092] The conversion unit can add a multilingual conversion function to support different languages during conversion. For example, if the user speaks English, the conversion unit can convert that speech to Japanese. For example, if the user speaks English, the conversion unit can perform optimal speech conversion by converting that speech to Japanese. The conversion unit can also convert the user's speech to English if they speak French. For example, if the user speaks French, the conversion unit can perform optimal speech conversion by converting that speech to English. The conversion unit can also convert the user's speech to French if they speak Spanish. For example, if the user speaks Spanish, the conversion unit can perform optimal speech conversion by converting that speech to French. This enables speech conversion between different languages through the multilingual conversion function. Some or all of the above processing in the conversion unit may be performed using AI, for example, or without AI. For example, the conversion unit can input speech data in different languages into a generating AI and have the generating AI perform multilingual conversion.
[0093] The conversion unit can estimate the user's emotions and determine the priority of the converted audio based on the estimated user emotions. For example, if the user is nervous, the conversion unit can prioritize the conversion of important audio. For example, if the user is nervous, the conversion unit can perform optimal audio conversion by prioritizing the conversion of important audio. Also, if the user is relaxed, the conversion unit can convert all audio equally. For example, if the user is relaxed, the conversion unit can perform optimal audio conversion by converting all audio equally. Also, if the user is excited, the conversion unit can prioritize the conversion of emotionally charged audio. For example, if the user is excited, the conversion unit can perform optimal audio conversion by prioritizing the conversion of emotionally charged audio. This allows for the prioritization of important audio by determining the priority of audio according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the conversion unit may be performed using AI, for example, or without AI. For example, the conversion unit can input user emotion data into a generating AI and have the generating AI determine the priority of the voices.
[0094] The conversion unit can adjust the conversion algorithm based on the location where the audio data was collected during conversion. For example, when converting audio collected by a user outdoors, the conversion unit adjusts the conversion algorithm considering wind noise. For example, when converting audio collected by a user outdoors, the conversion unit can perform optimal audio conversion by adjusting the conversion algorithm considering wind noise. The conversion unit can also adjust the conversion algorithm considering echo when converting audio collected by a user indoors. For example, when converting audio collected by a user indoors, the conversion unit can perform optimal audio conversion by adjusting the conversion algorithm considering echo. The conversion unit can also adjust the conversion algorithm considering engine noise when converting audio collected by a user inside a car. For example, when converting audio collected by a user inside a car, the conversion unit can perform optimal audio conversion by adjusting the conversion algorithm considering engine noise. This makes optimal audio conversion possible by adjusting the conversion algorithm based on the collection location. Some or all of the above processing in the conversion unit may be performed using AI, for example, or without using AI. For example, the conversion unit can input data on the location where the audio data was collected into the generating AI, and have the generating AI adjust the conversion algorithm.
[0095] The conversion unit can analyze the user's social media activity during conversion and perform relevant voice conversions. For example, the conversion unit can convert voices related to topics the user is discussing on social media. For example, the conversion unit can perform optimal voice conversions by converting voices related to topics the user is discussing on social media. The conversion unit can also convert voices related to statements made by people the user follows on social media. For example, the conversion unit can perform optimal voice conversions by converting voices related to statements made by people the user follows on social media. The conversion unit can also convert voices related to topics the user is participating in on social media. For example, the conversion unit can perform optimal voice conversions by converting voices related to topics the user is participating in on social media. This enables highly relevant voice conversions based on social media activity. Some or all of the above processing in the conversion unit may be performed using AI, for example, or without AI. For example, the conversion unit can input the user's social media activity data into a generating AI and have the generating AI perform the voice conversion.
[0096] The output unit can estimate the user's emotions and adjust the volume and tone of the voice output based on the estimated emotions. For example, if the user is nervous, the output unit can output voice in a calm tone. For example, if the user is nervous, the output unit can achieve optimal voice output by outputting voice in a calm tone. The output unit can also output voice in a natural tone if the user is relaxed. For example, if the user is relaxed, the output unit can achieve optimal voice output by outputting voice in a natural tone. The output unit can also output voice in a bright tone if the user is excited. For example, if the output unit is excited, the output unit can achieve optimal voice output by outputting voice in a bright tone. This enables optimal voice output by adjusting the volume and tone according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input user emotion data into a generating AI, which can then perform volume and tone adjustments.
[0097] The output unit can analyze ambient sounds in real time during output and adjust the audio output accordingly. For example, if the surroundings are noisy, the output unit can increase the volume of the audio output. For example, if the surroundings are noisy, the output unit can achieve optimal audio output by increasing the volume of the audio output. The output unit can also decrease the volume of the audio output if the surroundings are quiet. For example, if the surroundings are quiet, the output unit can achieve optimal audio output by decreasing the volume of the audio output. Furthermore, if a sudden sound occurs, the output unit can analyze that sound in real time and adjust the audio output accordingly. For example, if a sudden sound occurs, the output unit can achieve optimal audio output by analyzing that sound in real time. This enables clear audio output based on ambient sounds. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input ambient sound data into a generating AI and have the generating AI perform the audio output adjustment.
[0098] The output unit can improve the accuracy of voice output by considering the user's speech pattern during output. For example, if the user speaks slowly, the output unit can adjust the voice output to match that pace. For example, if the user speaks slowly, the output unit can adjust the voice output to match that pace, thereby achieving optimal voice output. The output unit can also adjust the voice output to match the user's fast speech, for example, by adjusting the voice output to match that pace, thereby achieving optimal voice output. Furthermore, if the user speaks with pauses, the output unit can adjust the voice output to take those pauses into account. For example, if the user speaks with pauses, the output unit can adjust the voice output to take those pauses into account, thereby achieving optimal voice output. This improves output accuracy through voice output based on the user's speech pattern. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the user's speech pattern data into a generating AI and have the generating AI perform the adjustment of the voice output.
[0099] The output unit can estimate the user's emotions and determine the priority of audio output based on the estimated emotions. For example, if the user is nervous, the output unit can prioritize outputting important audio. For example, if the user is nervous, the output unit can achieve optimal audio output by prioritizing the output of important audio. The output unit can also output all audio equally if the user is relaxed. For example, if the user is relaxed, the output unit can achieve optimal audio output by outputting all audio equally. The output unit can also prioritize outputting emotionally charged audio if the user is excited. For example, if the output unit is excited, the output unit can achieve optimal audio output by prioritizing the output of emotionally charged audio. This allows for the prioritization of important audio by determining the priority of audio output according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the processing described above in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input user emotion data into a generating AI and have the generating AI determine the priority of voice output.
[0100] The output unit can perform optimal audio output by considering the user's geographical location information when outputting audio. For example, if the user is in a specific location, the output unit can prioritize outputting audio related to that location. The output unit can also prioritize outputting audio related to the user's destination if the user is on the move. The output unit can also prioritize outputting audio related to the user's destination if the user is participating in a specific event. The output unit can also prioritize outputting audio related to that event if the user is participating in a specific event. This allows for the prioritization of highly relevant audio output based on geographical location information. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the user's geographical location information data into a generating AI and have the generating AI adjust the audio output.
[0101] The output unit can analyze the user's social media activity and output relevant audio when outputting audio. For example, the output unit can output audio related to topics the user is discussing on social media. For example, the output unit can achieve optimal audio output by outputting audio related to topics the user is discussing on social media. The output unit can also output audio related to statements made by people the user follows on social media. For example, the output unit can achieve optimal audio output by outputting audio related to statements made by people the user follows on social media. The output unit can also output audio related to topics the user is participating in on social media. For example, the output unit can achieve optimal audio output by outputting audio related to topics the user is participating in on social media. This allows for the output of highly relevant audio based on social media activity. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the user's social media activity data into a generating AI and have the generating AI adjust the audio output.
[0102] The shape of the bow tie can estimate the user's emotions and adjust its design and color based on those emotions. For example, if the user is nervous, the bow tie can provide a design with calming colors. For example, if the bow tie is nervous, providing a design with calming colors can improve visual comfort. The bow tie can also provide a design with bright colors if the user is relaxed. For example, if the bow tie is relaxed, providing a design with bright colors can improve visual comfort. The bow tie can also provide a design that is visually stimulating if the user is excited. For example, if the bow tie is excited, providing a design that is visually stimulating can improve visual comfort. In this way, visual comfort is improved by adjusting the design and color according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the shape of the bow tie may be performed using AI, for example, or without AI. For example, the shape of a bow tie can be determined by inputting user emotional data into a generative AI, which can then perform design and color adjustments.
[0103] The shape of a bow tie can be improved by changing the material it is made from. For example, the shape of a bow tie can be improved by using a soft material. Alternatively, the shape of a bow tie can be made to withstand long-term use by using a highly durable material. Furthermore, the shape of a bow tie can be made to provide a comfortable fit by using a breathable material. Thus, changing the material improves both the fit and durability. Some or all of the above processes regarding the shape of a bow tie may be performed using AI, or not. For example, material selection data can be input into a generating AI, and the AI can then perform the material change.
[0104] The shape of the bow tie can incorporate additional functions, such as a solar panel to extend battery life. The shape of the bow tie can also incorporate a battery pack to enable extended use. Furthermore, the shape of the bow tie can add a wireless charging function to eliminate the need for manual charging. This allows for extended battery life and longer use through these additional functions. Some or all of the above processing in the shape of the bow tie may be performed using AI, or not. For example, the design data for the additional functions can be input into a generating AI, and the generating AI can be made to incorporate the additional functions.
[0105] The shape of the bow tie can estimate the user's emotions and adjust the way the bow tie is worn based on those emotions. For example, if the user is nervous, the shape of the bow tie can provide an easy way to put it on. For example, if the user is nervous, the shape of the bow tie can provide an easy way to put it on, thereby improving comfort. The shape of the bow tie can also provide variations in how it is worn when the user is relaxed. For example, if the user is relaxed, the shape of the bow tie can provide variations in how it is worn, thereby improving comfort. The shape of the bow tie can also simplify the way it is worn, allowing for quick and easy application when the user is excited. For example, if the shape of the bow tie simplifies the way it is worn, allowing for quick and easy application, thereby improving comfort. In this way, comfort is improved by adjusting the way the bow tie is worn according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes for shaping the bow tie may be performed using AI, for example, or without AI. For example, the shape of the bow tie can be determined by inputting user emotion data into a generating AI, which can then be used to adjust the way the bow tie is worn.
[0106] The shape of a bow tie can be applied to other accessories. For example, it can be applied to brooches and necklaces. For instance, applying the shape of a bow tie to a brooch expands the range of ways it can be worn. Also, applying the shape of a bow tie to a necklace can enhance its fashion appeal. For example, applying the shape of a bow tie to a necklace can enhance its fashion appeal. Furthermore, applying the shape of a bow tie to a hair accessory allows for a variety of styles. For example, applying the shape of a bow tie to a hair accessory allows for a variety of styles. This expands the range of ways it can be worn by applying it to other accessories. Some or all of the above processing regarding the shape of a bow tie may be performed using AI, or not. For example, the shape of a bow tie can be used to input application data for other accessories into a generating AI, which can then execute the application.
[0107] The shape of the bow tie can be customized, allowing users to change it to their liking. For example, the shape of the bow tie can allow users to select a design, providing a bow tie tailored to their individual preferences. The shape of the bow tie can also allow users to select a color, providing a bow tie tailored to their individual style. For example, the shape of the bow tie can allow users to select a color, providing a bow tie tailored to their individual style. The shape of the bow tie can also allow users to select a material, providing a bow tie tailored to their individual fit. For example, the shape of the bow tie can allow users to select a material, providing a bow tie tailored to their individual fit. This allows for customization of the design, enabling users to wear the bow tie according to their preferences. Some or all of the above processing in the shape of the bow tie may be performed using AI, or not. For example, the shape of the bow tie can input customization data into a generating AI, and have the generating AI perform the design customization.
[0108] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0109] The voice conversion system can also be equipped with the ability to analyze the tone and pitch of the user's voice and estimate the user's emotions. For example, the collection unit can analyze the tone and pitch of the user's voice and estimate whether the user is tense or relaxed. The analysis unit can adjust the analysis algorithm based on the estimated emotions to perform a more accurate analysis. The conversion unit can adjust the tone and pitch of the voice conversion based on the estimated emotions to achieve a natural voice conversion. The output unit can adjust the volume and tone of the voice output based on the estimated emotions to provide optimal voice output. This enables voice conversion and output that responds to the user's emotions, resulting in more natural communication.
[0110] The voice conversion system can also be equipped with a function to learn the user's speech patterns and automatically adjust the sensitivity of the collection unit. For example, if the user speaks slowly, the collection unit can adjust its sensitivity to match the user's pace for optimal voice collection. It can also adjust its sensitivity to match the user's pace if the user speaks quickly. Furthermore, if the user pauses while speaking, the system can adjust its sensitivity to account for those pauses. This improves the accuracy of voice collection by adjusting the sensitivity according to the user's speech patterns.
[0111] The analysis unit can also be equipped with a function to learn the user's speech patterns and automatically adjust the analysis algorithm. For example, if the user speaks slowly, the analysis unit can adjust the analysis algorithm to match that pace and perform optimal analysis. It can also adjust the analysis algorithm to match the user's speaking pace if the user speaks quickly. Furthermore, if the user pauses while speaking, the analysis algorithm can be adjusted to take those pauses into account. As a result, the accuracy of the analysis is improved by adjusting the analysis algorithm according to the user's speech patterns.
[0112] The conversion unit can also be equipped with a multilingual conversion function to support even more languages. For example, if a user speaks English, their voice can be converted to Japanese. If a user speaks French, their voice can be converted to English. Furthermore, if a user speaks Spanish, their voice can be converted to French. This multilingual conversion function enables voice conversion between different languages, facilitating smoother international communication.
[0113] The output unit can also be equipped with a function to analyze ambient noise in real time and adjust the audio output. For example, if the surroundings are noisy, the audio output volume can be increased. Conversely, if the surroundings are quiet, the audio output volume can be decreased. Furthermore, if a sudden sound occurs, it can be analyzed in real time and the audio output can be adjusted accordingly. This enables clear audio output based on ambient noise.
[0114] The voice conversion system can also be equipped with the ability to estimate the user's emotions and adjust the tone and pitch of the voice conversion based on those emotions. For example, if the user is tense, the conversion unit can convert the voice in a calm tone. If the user is relaxed, it can convert the voice in a natural tone. Furthermore, if the user is excited, it can convert the voice in a bright tone. This allows for natural voice conversion by adjusting the tone and pitch according to the user's emotions.
[0115] The voice conversion system can also be equipped with a function to estimate the user's emotions and adjust the analysis algorithm based on those emotions. For example, if the user is tense, the analysis unit can adjust the analysis algorithm to prioritize accuracy. If the user is relaxed, it can adjust the analysis algorithm to prioritize speed. Furthermore, if the user is excited, it can adjust the analysis algorithm to prioritize balance. This improves the accuracy of the analysis by adjusting the analysis algorithm according to the user's emotions.
[0116] The voice conversion system can also be equipped with a function to estimate the user's emotions and adjust the volume and tone of the voice output based on those emotions. For example, the output unit can output voice in a calm tone if the user is tense. It can also output voice in a natural tone if the user is relaxed. Furthermore, it can output voice in a bright tone if the user is excited. This allows for optimal voice output by adjusting the volume and tone according to the user's emotions.
[0117] The voice conversion system can also be equipped with a function to estimate the user's emotions and determine the priority of voices based on those emotions. For example, the collection unit can prioritize collecting important voices when the user is tense. It can also collect all voices equally when the user is relaxed. Furthermore, it can prioritize collecting voices with strong emotions when the user is excited. This allows for the priority collection of important voices by determining voice priorities according to the user's emotions.
[0118] The voice conversion system can also be equipped with functions to collect and analyze voice data while considering the user's geographical location. For example, the collection unit can prioritize collecting voice data related to a specific location if the user is in that location. Similarly, the analysis unit can prioritize analyzing collected voice data if it is related to a specific location. Furthermore, the conversion unit can adjust its conversion algorithm based on the collection location. This allows for the priority processing of highly relevant voice data through voice collection and analysis based on geographical location information.
[0119] The following briefly describes the processing flow for example form 2.
[0120] Step 1: The collection unit collects the words spoken by the hearing impaired person. The collection unit, for example, has a built-in microphone to collect the spoken voice with high precision. For example, if the hearing impaired person says "hello," the collection unit can collect that voice with its microphone. Step 2: The analysis unit analyzes the audio data collected by the collection unit. The analysis unit analyzes the collected audio data using, for example, AI, and identifies the spoken words. For example, the analysis unit can identify the word "hello" from the collected audio data. Step 3: The conversion unit converts the audio data analyzed by the analysis unit into normal speech. The conversion unit converts words identified using AI into normal speech. For example, the conversion unit can convert the word "hello" into normal speech. Step 4: The output unit outputs the audio converted by the conversion unit. The output unit, for example, has a built-in speaker to output the converted audio clearly. For example, the output unit can output a normal voice saying "hello" through the speaker.
[0121] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0122] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0123] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0124] Each of the multiple elements described above, including the collection unit, analysis unit, conversion unit, and output unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the collection unit collects words spoken by a hearing-impaired person using the microphone of the smart device 14. The analysis unit analyzes the collected audio data by the identification processing unit 290 of the data processing unit 12 and identifies the spoken words. The conversion unit converts the analyzed audio data by the identification processing unit 290 of the data processing unit 12 into ordinary speech. The output unit outputs the converted speech using the speaker of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0125] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0126] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0127] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0128] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0129] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0130] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0131] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0132] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0133] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0134] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0135] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0136] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0137] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0138] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0139] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0140] Each of the multiple elements described above, including the collection unit, analysis unit, conversion unit, and output unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the collection unit collects words spoken by a hearing-impaired person using the microphone of the smart glasses 214. The analysis unit analyzes the collected audio data by the identification processing unit 290 of the data processing unit 12 and identifies the spoken words. The conversion unit converts the analyzed audio data by the identification processing unit 290 of the data processing unit 12 into ordinary speech. The output unit outputs the converted speech using the speaker of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0141] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0142] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0143] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0144] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0145] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0146] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0147] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0148] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0149] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0150] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0151] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0152] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0153] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0154] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0155] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0156] Each of the multiple elements described above, including the collection unit, analysis unit, conversion unit, and output unit, is implemented in, for example, at least one of the headset terminal 314 and the data processing unit 12. For example, the collection unit collects words spoken by a hearing-impaired person using the microphone of the headset terminal 314. The analysis unit analyzes the collected audio data by, for example, the identification processing unit 290 of the data processing unit 12 and identifies the spoken words. The conversion unit converts the audio data analyzed by the identification processing unit 290 of the data processing unit 12 into ordinary speech. The output unit outputs the converted speech using, for example, the speaker of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0157] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0158] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0159] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0160] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0161] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0162] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0163] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0164] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0165] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0166] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0167] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0168] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0169] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0170] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0171] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0172] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0173] Each of the multiple elements described above, including the collection unit, analysis unit, conversion unit, and output unit, is implemented, for example, in at least one of the robot 414 and the data processing unit 12. For example, the collection unit collects words spoken by a hearing-impaired person using the microphone of the robot 414. The analysis unit analyzes the collected audio data by the identification processing unit 290 of the data processing unit 12 and identifies the spoken words. The conversion unit converts the analyzed audio data by the identification processing unit 290 of the data processing unit 12 into ordinary speech. The output unit outputs the converted speech using the speaker of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0174] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0175] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0176] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0177] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0178] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0179] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0180] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0181] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0182] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0183] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0184] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0185] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0186] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0187] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0188] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0189] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0190] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0191] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0192] (Note 1) A collection unit that collects sound, An analysis unit analyzes the audio data collected by the aforementioned collection unit, A conversion unit that converts the audio data analyzed by the aforementioned analysis unit into ordinary audio, The system comprises an output unit that outputs the audio converted by the conversion unit. A system characterized by the following features. (Note 2) The aforementioned collection unit is It has a built-in microphone to improve the accuracy of spoken audio. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned analysis unit, The collected audio data is analyzed to identify the spoken words. The system described in Appendix 1, characterized by the features described herein. (Note 4) The conversion unit is Convert the identified words into normal speech. The system described in Appendix 1, characterized by the features described herein. (Note 5) The output unit is, Clearly output the converted audio. The system described in Appendix 1, characterized by the features described herein. (Note 6) The device is It has the shape of a bow tie, is easy to put on, and has low visibility. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned collection unit is It estimates the user's emotions and adjusts the voice collection sensitivity based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned collection unit is It analyzes ambient sounds in real time and collects audio while performing noise cancellation. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned collection unit is It learns the user's speech patterns and automatically adjusts the timing of data collection. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned collection unit is It estimates the user's emotions and determines the priority of audio to collect based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned collection unit is When collecting audio, the system prioritizes collecting highly relevant audio by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned collection unit is During audio collection, the system analyzes the user's social media activity and collects relevant audio. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, It estimates the user's emotions and adjusts the analysis algorithm based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, background noise in the audio data is removed, improving the accuracy of the spoken words. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, the user's speech patterns are taken into consideration to improve analysis accuracy. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During analysis, the priority of the analysis is determined based on when the audio data was collected. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, the order of analysis is adjusted based on the relevance of the audio data. The system described in Appendix 1, characterized by the features described herein. (Note 19) The conversion unit is It estimates the user's emotions and adjusts the tone and pitch of the voice conversion based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The conversion unit is During conversion, the system learns the user's speech patterns and performs optimal speech conversion. The system described in Appendix 1, characterized by the features described herein. (Note 21) The conversion unit is During conversion, add a multilingual conversion function to support different languages. The system described in Appendix 1, characterized by the features described herein. (Note 22) The conversion unit is It estimates the user's emotions and determines the priority of the converted audio based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The conversion unit is During conversion, the conversion algorithm is adjusted based on the location where the audio data was collected. The system described in Appendix 1, characterized by the features described herein. (Note 24) The conversion unit is During the conversion process, the system analyzes the user's social media activity and performs relevant voice conversions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The output unit is, It estimates the user's emotions and adjusts the volume and tone of the audio output based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The output unit is, During output, the system analyzes ambient noise in real time and adjusts the audio output accordingly. The system described in Appendix 1, characterized by the features described herein. (Note 27) The output unit is, During output, the system improves the accuracy of voice output by taking into account the user's speech patterns. The system described in Appendix 1, characterized by the features described herein. (Note 28) The output unit is, It estimates the user's emotions and determines the priority of voice output based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The output unit is, When outputting audio, the system takes the user's geographical location into consideration to optimize the audio output. The system described in Appendix 1, characterized by the features described herein. (Note 30) The output unit is, When outputting audio, the system analyzes the user's social media activity and outputs relevant audio. The system described in Appendix 1, characterized by the features described herein. (Note 31) The shape of the bow tie is, The system estimates the user's emotions and adjusts the design and color of the bow tie based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The shape of the bow tie is, The material of the bow tie has been changed to improve comfort and durability. The system described in Appendix 1, characterized by the features described herein. (Note 33) The shape of the bow tie is, The bow tie incorporates additional features, such as a solar panel to extend battery life. The system described in Appendix 1, characterized by the features described herein. (Note 34) The shape of the bow tie is, The system estimates the user's emotions and adjusts the way the bow tie is worn based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 35) The shape of the bow tie is, Applying the shape of a bow tie to other accessories. The system described in Appendix 1, characterized by the features described herein. (Note 36) The shape of the bow tie is, The bow tie design is customizable, allowing users to change it to their liking. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]
[0193] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A collection unit that collects sound, An analysis unit analyzes the audio data collected by the aforementioned collection unit, A conversion unit that converts the audio data analyzed by the aforementioned analysis unit into ordinary audio, The system comprises an output unit that outputs the audio converted by the conversion unit, The aforementioned collection unit is The system analyzes the user's social media activity and prioritizes collecting audio related to at least one of the topics the user discusses on social media, the statements of people the user follows on social media, and the topics of groups the user participates in on social media. A system characterized by the following features.
2. The aforementioned collection unit is It has a built-in microphone to improve the accuracy of spoken audio. The system according to feature 1.
3. The aforementioned analysis unit, The collected audio data is analyzed to identify the spoken words. The system according to feature 1.
4. It has the shape of a bow tie. The system according to feature 1.
5. The aforementioned collection unit is The system estimates the user's emotions and adjusts the voice collection sensitivity based on the estimated user emotions. The system according to feature 1.
6. The aforementioned collection unit is It analyzes ambient sounds in real time and collects audio while performing noise cancellation. The system according to feature 1.
Citation Information
Patent Citations
Voice interaction apparatus
JP2008026463A
Information processing apparatus, and voice correction program
JP2010183444A
Speech enhancing device, speech enhancing program
JP2011170261A
Persona chatbot control method and system
JP2022180282A
Voice processing system, voice processing device and voice processing method
JP2023077444A