system
The system addresses unnatural voice conversion in security measures by using AI to convert and manage voice responses, generating voice clones for natural conversation, thereby enhancing security for vulnerable individuals.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional voice conversion technologies for security measures are unnatural and provide limited effectiveness.
A system comprising a conversion unit, response unit, and management unit that utilizes AI to convert and manage voice responses, generating voice clones based on pre-registered voices to enhance security through natural conversation.
Enhances crime prevention by providing natural and effective voice conversion, improving security for vulnerable individuals such as women and the elderly.
Smart Images

Figure 2026072303000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that voice conversion as a security measure is unnatural and the effect is limited.
[0005] The system according to the embodiment aims to enhance the security effect through natural voice conversion.
Means for Solving the Problems
[0006] The system according to the embodiment includes a conversion unit, a response unit, a management unit, and a generation unit. The conversion unit converts voice. The response unit responds based on the voice converted by the conversion unit. The management unit manages a preset response pattern. The generation unit generates a voice clone.
Effects of the Invention
[0007] The system according to this embodiment can enhance crime prevention effectiveness through natural voice conversion. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The AI Intercom Guardian System according to an embodiment of the present invention is a system that enhances security by using AI to convert responses received through the intercom into the voice of an adult male. When a user responds through the intercom, the AI converts the voice in real time into the voice of an adult male. This reduces the risk of attracting the attention of suspicious individuals, and conversely, increases security by prompting the user to answer the intercom. Furthermore, users can pre-set response patterns tailored to their residence. For example, they can easily use the system by setting patterns such as "Leave it at the front door," "Put it in the delivery box," or "Leave it there." Additionally, the AI utilizes voice cloning technology to use pre-registered real voices (such as a father, son, or boyfriend). This allows for natural conversation with the person on the other end of the intercom. Furthermore, the AI utilizes voice cloning technology to use pre-registered real voices (such as those of a father, son, or boyfriend). This allows for natural conversation with the person on the other end of the intercom. As a result, the AI Intercom Guardian system provides a safer environment for women at home, children left alone, and the elderly. In today's world, where intercom use is increasing, natural conversation can enhance security. Therefore, the AI Intercom Guardian system can improve security.
[0029] The AI intercom guardian system according to this embodiment comprises a conversion unit, a response unit, a management unit, and a generation unit. The conversion unit converts speech. The conversion unit converts speech by, for example, changing the tone, pitch, and speed of the speech. The conversion unit can also convert speech using AI. For example, the conversion unit can change the tone of the speech to convert it to a lower voice. The conversion unit can also change the pitch of the speech to convert it to a more masculine voice. The conversion unit can also change the speed of the speech to achieve more natural conversation. The response unit responds based on the speech converted by the conversion unit. The response unit makes a natural response based on the converted speech, for example. The response unit can also generate a response using AI. For example, the response unit generates an appropriate response based on the converted speech. The response unit can also adjust the timing based on the converted speech to maintain a natural pause. The management unit manages pre-set response patterns. The management unit utilizes, for example, response patterns pre-set by the user. The management unit can also manage response patterns using AI. For example, the management unit selects an appropriate response based on the response patterns set by the user. The management unit can also adjust the content of the response based on the response patterns. The generation unit generates a voice clone. The generation unit generates a voice clone using, for example, a real voice that has been registered in advance. The generation unit can also generate a voice clone using AI. For example, the generation unit generates a voice clone based on a voice that has been registered in advance. The generation unit can also achieve natural conversation based on the voice clone. As a result, the AI intercom guardian system according to the embodiment can enhance its security effectiveness.
[0030] The conversion unit converts speech. For example, it converts speech by changing the tone, pitch, and speed. Specifically, by changing the tone, it can convert the voice to a lower or higher pitch. For example, when a visitor speaks through the intercom, converting their voice to a lower, more intimidating tone can enhance security. It can also convert the voice to a more masculine or feminine tone by changing the pitch. This allows, for example, a woman home alone to respond in a male voice, providing reassurance to visitors. Furthermore, changing the speed of the speech can create a more natural conversation. For example, if a visitor speaks too quickly, the speed can be adjusted to a more easily understandable pace. The conversion unit can also use AI to convert speech. The AI analyzes the characteristics of the speech and selects the optimal conversion method. For example, by analyzing the characteristics of a visitor's voice and making the most effective tone and pitch changes, it achieves a more natural and effective speech conversion. This allows the conversion unit to convert visitors' voices in various ways, improving the overall security effectiveness of the system.
[0031] The response unit responds based on the audio converted by the conversion unit. For example, the response unit provides a natural response based on the converted audio. Specifically, when a visitor speaks through the intercom, it generates an appropriate response based on the audio converted by the conversion unit. The response unit can also generate responses using AI. The AI analyzes the visitor's questions and requests and selects the optimal response. For example, if a visitor says, "I've come to deliver a package," the response unit can generate a response such as, "Thank you. Please leave it at the front door." The response unit can also adjust the timing based on the converted audio to maintain a natural pause. This allows for smoother conversations with visitors and enables responses that feel natural. Furthermore, the response unit can refer to past conversation history and provide more appropriate responses to specific requests or questions from visitors. For example, if the same visitor has come before, the response can be adjusted based on that history. This allows the response unit to have natural conversations with visitors and enhance the overall security effect of the system.
[0032] The management department manages pre-configured response patterns. For example, the management department utilizes response patterns pre-set by users. Specifically, it provides appropriate responses to visitors based on response patterns set by users, such as "Please leave your package at the front door" or "We are unable to assist you at this time." The management department can also manage response patterns using AI. The AI analyzes the user's past response history and the characteristics of visitors to select the optimal response pattern. For example, if a particular visitor comes frequently, the AI can select the most appropriate response for that visitor. The management department can also adjust the content of responses based on response patterns. For example, it can fine-tune response patterns according to the visitor's questions and requests to provide more appropriate responses. Furthermore, the management department allows users to add new response patterns or edit existing ones. This allows users to customize the system to their needs. Through managing response patterns, the management department can improve the overall response accuracy and effectiveness of the system.
[0033] The generation unit generates voice clones. For example, the generation unit generates voice clones using pre-registered real voices. Specifically, it can play back voices recorded by the user to visitors. The generation unit can also generate voice clones using AI. The AI analyzes the characteristics of the user's voice and generates voice clones based on those characteristics. For example, it can analyze the tone, pitch, and speed of the user's voice and generate voice clones that reproduce those characteristics. The generation unit can also achieve natural conversations based on the voice clones. For example, when a visitor speaks through the intercom, the generation unit can respond using the user's voice clone, providing a natural conversation to the visitor. Furthermore, the generation unit can continuously improve the quality of the voice clones. For example, each time a user records new voice, the voice clone can be updated based on that voice, providing more natural and high-quality voice clones. This allows the generation unit to provide natural and effective responses to visitors, enhancing the overall security effect of the system.
[0034] The conversion unit can convert speech in real time. For example, when a user responds via an intercom, the conversion unit converts the speech in real time. The conversion unit can also convert speech in real time using AI. For example, the conversion unit can analyze the user's voice in real time and convert it to the voice of an adult male. The conversion unit can also adjust the tone and pitch of the speech in real time to achieve natural conversation. This enables immediate responses by converting speech in real time. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that analyzes the user's voice in real time and converts it to the voice of an adult male.
[0035] The management unit can utilize pre-configured response patterns. For example, the management unit generates responses based on response patterns pre-configured by the user. The management unit can also manage pre-configured response patterns using AI. For example, the management unit selects an appropriate response based on the response patterns configured by the user. The management unit can also adjust the content of the responses based on the response patterns. This makes it easy to use by utilizing pre-configured response patterns. Some or all of the above processes in the management unit may be performed using AI or not. For example, the management unit can generate responses using an AI model that generates responses based on response patterns configured by the user.
[0036] The generation unit can use pre-registered real voices. For example, the generation unit can generate voice clones based on pre-registered real voices such as those of a father, son, or boyfriend. The generation unit can also use AI to use pre-registered real voices. For example, the generation unit can generate voice clones using an AI model that generates voice clones based on pre-registered voices. The generation unit can also achieve natural conversation based on the voice clones. This makes natural conversation possible by using pre-registered real voices. Some or all of the above-described processes in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that generates voice clones based on pre-registered voices.
[0037] The response unit can respond naturally based on the converted speech. For example, the response unit makes a natural response based on the converted speech. The response unit can also respond based on the converted speech using AI. For example, the response unit generates an appropriate response based on the converted speech. The response unit can also adjust the timing based on the converted speech to maintain a natural pause. This prevents unnatural pauses when responding naturally based on the converted speech. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that generates responses based on converted speech.
[0038] The conversion unit can analyze background noise during speech conversion and perform appropriate noise cancellation. For example, if there is car noise in the background, the AI will cancel that noise and convert it to clear speech. The conversion unit can also cancel television noise in the background and convert it to clear speech. The conversion unit can also cancel wind noise in the background and convert it to clear speech. In this way, clear speech conversion is possible by analyzing background noise and performing noise cancellation. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that analyzes background noise and performs noise cancellation.
[0039] The conversion unit can learn the characteristics of the user's voice during speech conversion to achieve a more natural conversion. For example, the conversion unit can learn the pitch of the user's voice and reflect it in the converted voice. The conversion unit can also learn the user's speaking speed and reflect it in the converted voice. The conversion unit can also learn the user's accent and intonation and reflect it in the converted voice. In this way, by learning the characteristics of the user's voice, a more natural speech conversion becomes possible. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that learns the characteristics of the user's voice and reflects them in the converted voice.
[0040] The conversion unit can perform speech conversion while considering the user's geographical accent and dialect. For example, if the user speaks Kansai dialect, the AI will maintain the Kansai accent during conversion. If the user speaks Tohoku dialect, the AI can maintain the Tohoku accent during conversion. If the user speaks Kyushu dialect, the AI can maintain the Kyushu accent during conversion. This allows for more natural speech conversion by considering the user's geographical accent and dialect. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that performs conversion while considering the user's geographical accent and dialect.
[0041] The conversion unit can perform appropriate conversions by referring to the user's past conversation history during speech conversion. For example, the conversion unit can learn phrases the user has used in the past and reflect them in the converted voice. The conversion unit can also learn the tone and pitch the user has spoken in the past and reflect them in the converted voice. The conversion unit can also learn the speed at which the user has spoken in the past and reflect it in the converted voice. This allows for more appropriate speech conversion by referring to the user's past conversation history. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that performs appropriate conversions by referring to the user's past conversation history.
[0042] The response unit can analyze the tone and content of the other party's voice when responding and generate an appropriate response. For example, if the other party is asking a question, the AI will generate an appropriate answer. The response unit can also generate an appropriate response if the other party is giving an instruction. The response unit can also generate an appropriate response if the other party is expressing gratitude. This allows for more appropriate responses by analyzing the tone and content of the other party's voice. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that analyzes the tone and content of the other party's voice and generates an appropriate response.
[0043] The response unit can achieve natural conversation by referring to the user's past response history when responding. For example, the response unit can learn phrases the user has used in the past and generate natural responses. The response unit can also learn the tone and pitch the user has spoken in the past and generate natural responses. The response unit can also learn the speed at which the user has spoken in the past and generate natural responses. This makes more natural conversation possible by referring to the user's past response history. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate responses using an AI model that achieves natural conversation by referring to the user's past response history.
[0044] The response unit can respond while considering the other party's geographical accent and dialect. For example, if the other party speaks Kansai dialect, the AI will respond in Kansai dialect. If the other party speaks Tohoku dialect, the AI can also respond in Tohoku dialect. If the other party speaks Kyushu dialect, the AI can also respond in Kyushu dialect. This allows for more natural responses by considering the other party's geographical accent and dialect. Some or all of the processing described above in the response unit may be performed using AI or not. For example, the response unit can generate responses using an AI model that takes the other party's geographical accent and dialect into consideration.
[0045] The response unit can analyze the background sounds of the other party and generate an appropriate response when responding. For example, if there is the sound of a car in the background, the AI will take that sound into consideration when responding. The response unit can also take the sound of a television in the background into consideration when responding. The response unit can also take the sound of wind in the background into consideration when responding. This allows for a more appropriate response by analyzing the background sounds of the other party. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that analyzes the background sounds of the other party and generates an appropriate response.
[0046] The management department can analyze the user's past response history during management and propose the optimal response pattern. For example, the management department can learn the response patterns the user has used in the past and propose the optimal pattern. The management department can also learn the tone and pitch the user has spoken in the past and propose the optimal pattern. The management department can also learn the speed at which the user has spoken in the past and propose the optimal pattern. In this way, by analyzing the user's past response history, more appropriate response patterns are proposed. Some or all of the above processes in the management department may be performed using AI or not. For example, the management department can propose response patterns using an AI model that analyzes the user's past response history and proposes the optimal response pattern.
[0047] The management department can customize response patterns based on the user's lifestyle and areas of interest during management. For example, if the user is busy, the management department can have the AI suggest a simple response pattern. If the user is relaxed, the management department can also have the AI suggest a more detailed response pattern. If the user has a specific area of interest, the management department can have the AI suggest a response pattern related to that area. This allows for more appropriate responses by customizing response patterns based on the user's lifestyle and areas of interest. Some or all of the above processing in the management department may be performed using AI or not. For example, the management department can suggest response patterns using an AI model that customizes response patterns based on the user's lifestyle and areas of interest.
[0048] The management department can select response patterns while considering the user's geographical information. For example, if the user lives in the Kansai region, the AI can select a response pattern in the Kansai dialect. If the user lives in the Tohoku region, the AI can select a response pattern in the Tohoku dialect. If the user lives in Kyushu, the AI can select a response pattern in the Kyushu dialect. This allows for the selection of a more appropriate response pattern by considering the user's geographical information. Some or all of the above processing in the management department may be performed using AI or not. For example, the management department can select response patterns using an AI model that selects response patterns while considering the user's geographical information.
[0049] The management department can analyze users' social media activity during management and suggest relevant response patterns. For example, the management department can learn phrases that users frequently use on social media and incorporate them into response patterns. The management department can also learn topics that users discuss on social media and incorporate them into response patterns. The management department can also learn information about accounts that users follow on social media and incorporate it into response patterns. This allows for the suggestion of more appropriate response patterns by analyzing users' social media activity. Some or all of the above processes in the management department may be performed using AI or not. For example, the management department can suggest response patterns using an AI model that analyzes users' social media activity and suggests relevant response patterns.
[0050] The generation unit can generate more natural-sounding voice clones by analyzing the user's past voice data during voice clone generation. For example, the generation unit can learn speaking patterns from the user's past voice data to generate natural-sounding voice clones. The generation unit can also learn specific phrases and expressions from the user's past voice data to generate natural-sounding voice clones. The generation unit can also learn how to express emotions from the user's past voice data to generate natural-sounding voice clones. In this way, more natural-sounding voice clones are generated by analyzing the user's past voice data. Some or all of the above-described processes in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that analyzes the user's past voice data and generates more natural-sounding clones.
[0051] The generation unit can learn the characteristics of the user's voice and generate individually optimized clones when generating voice clones. For example, the generation unit can learn the pitch of the user's voice and generate optimized voice clones. The generation unit can also learn the user's speaking speed and generate optimized voice clones. The generation unit can also learn the user's accent and intonation and generate optimized voice clones. In this way, individually optimized voice clones are generated by learning the characteristics of the user's voice. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that learns the characteristics of the user's voice and generates individually optimized clones.
[0052] The generation unit can generate voice clones while considering the user's geographical accent and dialect. For example, if the user speaks Kansai dialect, the AI can generate a voice clone while maintaining the Kansai accent. If the user speaks Tohoku dialect, the AI can generate a voice clone while maintaining the Tohoku accent. If the user speaks Kyushu dialect, the AI can generate a voice clone while maintaining the Kyushu accent. This allows for the generation of more natural-sounding voice clones by considering the user's geographical accent and dialect. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that generates clones while considering the user's geographical accent and dialect.
[0053] The generation unit can analyze the user's social media activity and generate relevant voice clones when creating voice clones. For example, the generation unit can learn phrases that the user frequently uses on social media and reflect them in the voice clones. The generation unit can also learn topics that the user discusses on social media and reflect them in the voice clones. The generation unit can also learn information about accounts that the user follows on social media and reflect it in the voice clones. This allows for the generation of more appropriate voice clones by analyzing the user's social media activity. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that analyzes the user's social media activity and generates relevant voice clones.
[0054] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0055] The AI intercom guardian system may also include a facial recognition unit. The facial recognition unit recognizes the face of a person approaching the intercom and compares it with a registered facial database. For example, the facial recognition unit can pre-register the faces of family members and friends, and automatically recognize these people when they approach the intercom. The facial recognition unit can also temporarily register the faces of delivery personnel or visitors and recognize them upon their return visits. This allows the facial recognition unit to quickly identify people approaching the intercom and generate appropriate responses. Some or all of the above-described processes in the facial recognition unit may be performed using AI or not. For example, the facial recognition unit can perform facial recognition using an AI model and compare it with a registered facial database.
[0056] The AI intercom guardian system may further include a voice recognition unit. The voice recognition unit recognizes the voice of the person speaking through the intercom and compares it with a registered voice database. For example, the voice recognition unit can pre-register the voices of family members and friends and automatically recognize them when they speak through the intercom. The voice recognition unit can also temporarily register the voices of delivery personnel or visitors and recognize them upon their return visit. This allows the voice recognition unit to quickly identify the person speaking through the intercom and generate an appropriate response. Some or all of the above processing in the voice recognition unit may be performed using AI or not. For example, the voice recognition unit can perform voice recognition using an AI model and compare it with a registered voice database.
[0057] The AI intercom guardian system may further include a motion detection unit. The motion detection unit detects movement around the intercom and identifies abnormal movements. For example, the motion detection unit can monitor the movement of a person approaching the intercom in real time and issue a warning if it detects suspicious movement. The motion detection unit can also record movement around the intercom for later review. This allows the motion detection unit to ensure safety around the intercom and quickly detect abnormal movements. Some or all of the above processing in the motion detection unit may be performed using AI or not. For example, the motion detection unit can perform motion detection using an AI model and detect abnormal movements.
[0058] The AI intercom guardian system may also include a temperature detection unit. The temperature detection unit monitors the temperature around the intercom in real time and detects abnormal temperature changes. For example, the temperature detection unit can issue a warning if the temperature around the intercom rises rapidly. The temperature detection unit can also issue a warning if the temperature around the intercom drops rapidly. This allows the temperature detection unit to quickly detect abnormal temperature changes around the intercom and take appropriate action. Some or all of the above processing in the temperature detection unit may be performed using AI or not. For example, the temperature detection unit can monitor temperature changes using an AI model and detect abnormal temperature changes.
[0059] The AI intercom guardian system may also include an illuminance detection unit. The illuminance detection unit monitors the illuminance around the intercom in real time and detects abnormal changes in illuminance. For example, the illuminance detection unit can issue a warning if the area around the intercom suddenly becomes dark. The illuminance detection unit can also issue a warning if the area around the intercom suddenly becomes bright. This allows the illuminance detection unit to quickly detect abnormal changes in illuminance around the intercom and take appropriate action. Some or all of the above processing in the illuminance detection unit may be performed using AI or not. For example, the illuminance detection unit can monitor changes in illuminance using an AI model and detect abnormal changes in illuminance.
[0060] The following briefly describes the processing flow for example form 1.
[0061] Step 1: The conversion unit converts the voice. For example, it converts the voice by changing the tone, pitch, and speed. It can also convert the voice using AI. Specifically, it can change the tone of the voice to make it lower, change the pitch of the voice to make it more masculine, or change the speed of the voice to make the conversation sound more natural. Step 2: The response unit responds based on the audio converted by the conversion unit. For example, it provides a natural response based on the converted audio. AI can also be used to generate responses, producing appropriate responses based on the converted audio and adjusting the timing to maintain a natural pause. Step 3: The management department manages pre-configured response patterns. For example, it is possible to use AI to manage response patterns by utilizing response patterns pre-configured by users. Specifically, it is possible to select an appropriate response based on the response patterns configured by the user, or to adjust the content of the response based on the response patterns. Step 4: The generation unit generates a voice clone. For example, it can generate a voice clone using a pre-registered real voice. It is also possible to generate a voice clone using AI, generating a voice clone based on a pre-registered voice to achieve natural conversation.
[0062] (Example of form 2) The AI Intercom Guardian System according to an embodiment of the present invention is a system that enhances security by using AI to convert responses received through the intercom into the voice of an adult male. When a user responds through the intercom, the AI converts the voice in real time into the voice of an adult male. This reduces the risk of attracting the attention of suspicious individuals, and conversely, increases security by prompting the user to answer the intercom. Furthermore, users can pre-set response patterns tailored to their residence. For example, they can easily use the system by setting patterns such as "Leave it at the front door," "Put it in the delivery box," or "Leave it there." Additionally, the AI utilizes voice cloning technology to use pre-registered real voices (such as a father, son, or boyfriend). This allows for natural conversation with the person on the other end of the intercom. Furthermore, the AI utilizes voice cloning technology to use pre-registered real voices (such as those of a father, son, or boyfriend). This allows for natural conversation with the person on the other end of the intercom. As a result, the AI Intercom Guardian system provides a safer environment for women at home, children left alone, and the elderly. In today's world, where intercom use is increasing, natural conversation can enhance security. Therefore, the AI Intercom Guardian system can improve security.
[0063] The AI intercom guardian system according to this embodiment comprises a conversion unit, a response unit, a management unit, and a generation unit. The conversion unit converts speech. The conversion unit converts speech by, for example, changing the tone, pitch, and speed of the speech. The conversion unit can also convert speech using AI. For example, the conversion unit can change the tone of the speech to convert it to a lower voice. The conversion unit can also change the pitch of the speech to convert it to a more masculine voice. The conversion unit can also change the speed of the speech to achieve more natural conversation. The response unit responds based on the speech converted by the conversion unit. The response unit makes a natural response based on the converted speech, for example. The response unit can also generate a response using AI. For example, the response unit generates an appropriate response based on the converted speech. The response unit can also adjust the timing based on the converted speech to maintain a natural pause. The management unit manages pre-set response patterns. The management unit utilizes, for example, response patterns pre-set by the user. The management unit can also manage response patterns using AI. For example, the management unit selects an appropriate response based on the response patterns set by the user. The management unit can also adjust the content of the response based on the response patterns. The generation unit generates a voice clone. The generation unit generates a voice clone using, for example, a real voice that has been registered in advance. The generation unit can also generate a voice clone using AI. For example, the generation unit generates a voice clone based on a voice that has been registered in advance. The generation unit can also achieve natural conversation based on the voice clone. As a result, the AI intercom guardian system according to the embodiment can enhance its security effectiveness.
[0064] The conversion unit converts speech. For example, it converts speech by changing the tone, pitch, and speed. Specifically, by changing the tone, it can convert the voice to a lower or higher pitch. For example, when a visitor speaks through the intercom, converting their voice to a lower, more intimidating tone can enhance security. It can also convert the voice to a more masculine or feminine tone by changing the pitch. This allows, for example, a woman home alone to respond in a male voice, providing reassurance to visitors. Furthermore, changing the speed of the speech can create a more natural conversation. For example, if a visitor speaks too quickly, the speed can be adjusted to a more easily understandable pace. The conversion unit can also use AI to convert speech. The AI analyzes the characteristics of the speech and selects the optimal conversion method. For example, by analyzing the characteristics of a visitor's voice and making the most effective tone and pitch changes, it achieves a more natural and effective speech conversion. This allows the conversion unit to convert visitors' voices in various ways, improving the overall security effectiveness of the system.
[0065] The response unit responds based on the audio converted by the conversion unit. For example, the response unit provides a natural response based on the converted audio. Specifically, when a visitor speaks through the intercom, it generates an appropriate response based on the audio converted by the conversion unit. The response unit can also generate responses using AI. The AI analyzes the visitor's questions and requests and selects the optimal response. For example, if a visitor says, "I've come to deliver a package," the response unit can generate a response such as, "Thank you. Please leave it at the front door." The response unit can also adjust the timing based on the converted audio to maintain a natural pause. This allows for smoother conversations with visitors and enables responses that feel natural. Furthermore, the response unit can refer to past conversation history and provide more appropriate responses to specific requests or questions from visitors. For example, if the same visitor has come before, the response can be adjusted based on that history. This allows the response unit to have natural conversations with visitors and enhance the overall security effect of the system.
[0066] The management department manages pre-configured response patterns. For example, the management department utilizes response patterns pre-set by users. Specifically, it provides appropriate responses to visitors based on response patterns set by users, such as "Please leave your package at the front door" or "We are unable to assist you at this time." The management department can also manage response patterns using AI. The AI analyzes the user's past response history and the characteristics of visitors to select the optimal response pattern. For example, if a particular visitor comes frequently, the AI can select the most appropriate response for that visitor. The management department can also adjust the content of responses based on response patterns. For example, it can fine-tune response patterns according to the visitor's questions and requests to provide more appropriate responses. Furthermore, the management department allows users to add new response patterns or edit existing ones. This allows users to customize the system to their needs. Through managing response patterns, the management department can improve the overall response accuracy and effectiveness of the system.
[0067] The generation unit generates voice clones. For example, the generation unit generates voice clones using pre-registered real voices. Specifically, it can play back voices recorded by the user to visitors. The generation unit can also generate voice clones using AI. The AI analyzes the characteristics of the user's voice and generates voice clones based on those characteristics. For example, it can analyze the tone, pitch, and speed of the user's voice and generate voice clones that reproduce those characteristics. The generation unit can also achieve natural conversations based on the voice clones. For example, when a visitor speaks through the intercom, the generation unit can respond using the user's voice clone, providing a natural conversation to the visitor. Furthermore, the generation unit can continuously improve the quality of the voice clones. For example, each time a user records new voice, the voice clone can be updated based on that voice, providing more natural and high-quality voice clones. This allows the generation unit to provide natural and effective responses to visitors, enhancing the overall security effect of the system.
[0068] The conversion unit can convert speech in real time. For example, when a user responds via an intercom, the conversion unit converts the speech in real time. The conversion unit can also convert speech in real time using AI. For example, the conversion unit can analyze the user's voice in real time and convert it to the voice of an adult male. The conversion unit can also adjust the tone and pitch of the speech in real time to achieve natural conversation. This enables immediate responses by converting speech in real time. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that analyzes the user's voice in real time and converts it to the voice of an adult male.
[0069] The management unit can utilize pre-configured response patterns. For example, the management unit generates responses based on response patterns pre-configured by the user. The management unit can also manage pre-configured response patterns using AI. For example, the management unit selects an appropriate response based on the response patterns configured by the user. The management unit can also adjust the content of the responses based on the response patterns. This makes it easy to use by utilizing pre-configured response patterns. Some or all of the above processes in the management unit may be performed using AI or not. For example, the management unit can generate responses using an AI model that generates responses based on response patterns configured by the user.
[0070] The generation unit can use pre-registered real voices. For example, the generation unit can generate voice clones based on pre-registered real voices such as those of a father, son, or boyfriend. The generation unit can also use AI to use pre-registered real voices. For example, the generation unit can generate voice clones using an AI model that generates voice clones based on pre-registered voices. The generation unit can also achieve natural conversation based on the voice clones. This makes natural conversation possible by using pre-registered real voices. Some or all of the above-described processes in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that generates voice clones based on pre-registered voices.
[0071] The response unit can respond naturally based on the converted speech. For example, the response unit makes a natural response based on the converted speech. The response unit can also respond based on the converted speech using AI. For example, the response unit generates an appropriate response based on the converted speech. The response unit can also adjust the timing based on the converted speech to maintain a natural pause. This prevents unnatural pauses when responding naturally based on the converted speech. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that generates responses based on converted speech.
[0072] The conversion unit can estimate the user's emotions and adjust the tone and pitch of the speech conversion based on the estimated emotions. For example, if the user is nervous, the conversion unit can use AI to calm the tone and lower the pitch during conversion. If the user is relaxed, the conversion unit can use AI to brighten the tone and maintain a natural pitch. If the user is in a hurry, the conversion unit can use AI to speed up the tone and raise the pitch during conversion. This allows for more natural speech conversion by adjusting the tone and pitch of the speech conversion according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that estimates the user's emotions and adjusts the tone and pitch of the speech conversion based on the estimated emotions.
[0073] The conversion unit can analyze background noise during speech conversion and perform appropriate noise cancellation. For example, if there is car noise in the background, the AI will cancel that noise and convert it to clear speech. The conversion unit can also cancel television noise in the background and convert it to clear speech. The conversion unit can also cancel wind noise in the background and convert it to clear speech. In this way, clear speech conversion is possible by analyzing background noise and performing noise cancellation. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that analyzes background noise and performs noise cancellation.
[0074] The conversion unit can learn the characteristics of the user's voice during speech conversion to achieve a more natural conversion. For example, the conversion unit can learn the pitch of the user's voice and reflect it in the converted voice. The conversion unit can also learn the user's speaking speed and reflect it in the converted voice. The conversion unit can also learn the user's accent and intonation and reflect it in the converted voice. In this way, by learning the characteristics of the user's voice, a more natural speech conversion becomes possible. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that learns the characteristics of the user's voice and reflects them in the converted voice.
[0075] The conversion unit can estimate the user's emotions and adjust the voice expression after conversion based on the estimated emotions. For example, if the user is angry, the AI can convert the voice to a calm voice. If the user is sad, the AI can also convert the voice to an encouraging voice. If the user is happy, the AI can also convert the voice to a cheerful voice. By adjusting the voice expression after conversion according to the user's emotions, more appropriate voice conversion becomes possible. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that estimates the user's emotions and adjusts the voice expression after conversion based on the estimated emotions.
[0076] The conversion unit can perform speech conversion while considering the user's geographical accent and dialect. For example, if the user speaks Kansai dialect, the AI will maintain the Kansai accent during conversion. If the user speaks Tohoku dialect, the AI can maintain the Tohoku accent during conversion. If the user speaks Kyushu dialect, the AI can maintain the Kyushu accent during conversion. This allows for more natural speech conversion by considering the user's geographical accent and dialect. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that performs conversion while considering the user's geographical accent and dialect.
[0077] The conversion unit can perform appropriate conversions by referring to the user's past conversation history during speech conversion. For example, the conversion unit can learn phrases the user has used in the past and reflect them in the converted voice. The conversion unit can also learn the tone and pitch the user has spoken in the past and reflect them in the converted voice. The conversion unit can also learn the speed at which the user has spoken in the past and reflect it in the converted voice. This allows for more appropriate speech conversion by referring to the user's past conversation history. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that performs appropriate conversions by referring to the user's past conversation history.
[0078] The response unit can estimate the user's emotions and adjust the content and tone of the response based on the estimated emotions. For example, if the user is nervous, the AI will respond in a calm tone. If the user is relaxed, the AI can also respond in a bright tone. If the user is in a hurry, the AI can also respond in a rapid tone. This allows for a more natural response by adjusting the content and tone of the response according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that estimates the user's emotions and adjusts the content and tone of the response based on the estimated emotions.
[0079] The response unit can analyze the tone and content of the other party's voice when responding and generate an appropriate response. For example, if the other party is asking a question, the AI will generate an appropriate answer. The response unit can also generate an appropriate response if the other party is giving an instruction. The response unit can also generate an appropriate response if the other party is expressing gratitude. This allows for more appropriate responses by analyzing the tone and content of the other party's voice. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that analyzes the tone and content of the other party's voice and generates an appropriate response.
[0080] The response unit can achieve natural conversation by referring to the user's past response history when responding. For example, the response unit can learn phrases the user has used in the past and generate natural responses. The response unit can also learn the tone and pitch the user has spoken in the past and generate natural responses. The response unit can also learn the speed at which the user has spoken in the past and generate natural responses. This makes more natural conversation possible by referring to the user's past response history. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate responses using an AI model that achieves natural conversation by referring to the user's past response history.
[0081] The response unit can estimate the user's emotions and adjust the speed and timing of its response based on the estimated emotions. For example, if the user is nervous, the AI will respond at a slow pace. If the user is relaxed, the AI can also respond at a natural pace. If the user is in a hurry, the AI can also respond at a rapid pace. This allows for a more appropriate response by adjusting the speed and timing of the response according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. Generative AIs include, but are not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that estimates the user's emotions and adjusts the speed and timing of the response based on the estimated emotions.
[0082] The response unit can respond while considering the other party's geographical accent and dialect. For example, if the other party speaks Kansai dialect, the AI will respond in Kansai dialect. If the other party speaks Tohoku dialect, the AI can also respond in Tohoku dialect. If the other party speaks Kyushu dialect, the AI can also respond in Kyushu dialect. This allows for more natural responses by considering the other party's geographical accent and dialect. Some or all of the processing described above in the response unit may be performed using AI or not. For example, the response unit can generate responses using an AI model that takes the other party's geographical accent and dialect into consideration.
[0083] The response unit can analyze the background sounds of the other party and generate an appropriate response when responding. For example, if there is the sound of a car in the background, the AI will take that sound into consideration when responding. The response unit can also take the sound of a television in the background into consideration when responding. The response unit can also take the sound of wind in the background into consideration when responding. This allows for a more appropriate response by analyzing the background sounds of the other party. Some or all of the above processing in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that analyzes the background sounds of the other party and generates an appropriate response.
[0084] The management unit can estimate the user's emotions and select response patterns based on the estimated emotions. For example, if the user is nervous, the management unit can have the AI select a simple response pattern. If the user is relaxed, the management unit can have the AI select a more detailed response pattern. If the user is in a hurry, the management unit can have the AI select a quick response pattern. This allows for more appropriate responses by selecting response patterns according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the management unit may be performed using AI or not. For example, the management unit can select response patterns using an AI model that estimates the user's emotions and selects response patterns based on the estimated emotions.
[0085] The management department can analyze the user's past response history during management and propose the optimal response pattern. For example, the management department can learn the response patterns the user has used in the past and propose the optimal pattern. The management department can also learn the tone and pitch the user has spoken in the past and propose the optimal pattern. The management department can also learn the speed at which the user has spoken in the past and propose the optimal pattern. In this way, by analyzing the user's past response history, more appropriate response patterns are proposed. Some or all of the above processes in the management department may be performed using AI or not. For example, the management department can propose response patterns using an AI model that analyzes the user's past response history and proposes the optimal response pattern.
[0086] The management department can customize response patterns based on the user's lifestyle and areas of interest during management. For example, if the user is busy, the management department can have the AI suggest a simple response pattern. If the user is relaxed, the management department can also have the AI suggest a more detailed response pattern. If the user has a specific area of interest, the management department can have the AI suggest a response pattern related to that area. This allows for more appropriate responses by customizing response patterns based on the user's lifestyle and areas of interest. Some or all of the above processing in the management department may be performed using AI or not. For example, the management department can suggest response patterns using an AI model that customizes response patterns based on the user's lifestyle and areas of interest.
[0087] The management unit can estimate the user's emotions and prioritize response patterns based on the estimated emotions. For example, if the user is nervous, the management unit may have the AI prioritize simpler response patterns. If the user is relaxed, the management unit may have the AI prioritize more detailed response patterns. If the user is in a hurry, the management unit may have the AI prioritize faster response patterns. This allows for more appropriate responses by prioritizing response patterns according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the management unit may be performed using AI or not. For example, the management unit may select response patterns using an AI model that estimates the user's emotions and prioritizes response patterns based on the estimated emotions.
[0088] The management department can select response patterns while considering the user's geographical information. For example, if the user lives in the Kansai region, the AI can select a response pattern in the Kansai dialect. If the user lives in the Tohoku region, the AI can select a response pattern in the Tohoku dialect. If the user lives in Kyushu, the AI can select a response pattern in the Kyushu dialect. This allows for the selection of a more appropriate response pattern by considering the user's geographical information. Some or all of the above processing in the management department may be performed using AI or not. For example, the management department can select response patterns using an AI model that selects response patterns while considering the user's geographical information.
[0089] The management department can analyze users' social media activity during management and suggest relevant response patterns. For example, the management department can learn phrases that users frequently use on social media and incorporate them into response patterns. The management department can also learn topics that users discuss on social media and incorporate them into response patterns. The management department can also learn information about accounts that users follow on social media and incorporate it into response patterns. This allows for the suggestion of more appropriate response patterns by analyzing users' social media activity. Some or all of the above processes in the management department may be performed using AI or not. For example, the management department can suggest response patterns using an AI model that analyzes users' social media activity and suggests relevant response patterns.
[0090] The generation unit can estimate the user's emotions and adjust the method of generating the voice clone based on the estimated emotions. For example, if the user is relaxed, the generation unit can generate a voice clone that proceeds at a relaxed pace. If the user is in a hurry, the generation unit can also generate a voice clone that proceeds at a fast pace. If the user is excited, the generation unit can also generate a voice clone with visually stimulating effects added. By adjusting the method of generating the voice clone according to the user's emotions, a more natural voice clone is generated. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that estimates the user's emotions and adjusts the method of generating the voice clone based on the estimated emotions.
[0091] The generation unit can generate more natural-sounding voice clones by analyzing the user's past voice data during voice clone generation. For example, the generation unit can learn speaking patterns from the user's past voice data to generate natural-sounding voice clones. The generation unit can also learn specific phrases and expressions from the user's past voice data to generate natural-sounding voice clones. The generation unit can also learn how to express emotions from the user's past voice data to generate natural-sounding voice clones. In this way, more natural-sounding voice clones are generated by analyzing the user's past voice data. Some or all of the above-described processes in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that analyzes the user's past voice data and generates more natural-sounding clones.
[0092] The generation unit can learn the characteristics of the user's voice and generate individually optimized clones when generating voice clones. For example, the generation unit can learn the pitch of the user's voice and generate optimized voice clones. The generation unit can also learn the user's speaking speed and generate optimized voice clones. The generation unit can also learn the user's accent and intonation and generate optimized voice clones. In this way, individually optimized voice clones are generated by learning the characteristics of the user's voice. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that learns the characteristics of the user's voice and generates individually optimized clones.
[0093] The generation unit can estimate the user's emotions and adjust the tone and pitch of the voice clone based on the estimated emotions. For example, if the user is nervous, the generation unit can generate a voice clone with a calmer tone and lower pitch. If the user is relaxed, the generation unit can also generate a voice clone with a brighter tone and a more natural pitch. If the user is in a hurry, the generation unit can also generate a voice clone with a faster tone and higher pitch. This allows for the generation of more natural-sounding voice clones by adjusting the tone and pitch according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that estimates the user's emotions and adjusts the tone and pitch of the voice clone based on the estimated emotions.
[0094] The generation unit can generate voice clones while considering the user's geographical accent and dialect. For example, if the user speaks Kansai dialect, the AI can generate a voice clone while maintaining the Kansai accent. If the user speaks Tohoku dialect, the AI can generate a voice clone while maintaining the Tohoku accent. If the user speaks Kyushu dialect, the AI can generate a voice clone while maintaining the Kyushu accent. This allows for the generation of more natural-sounding voice clones by considering the user's geographical accent and dialect. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that generates clones while considering the user's geographical accent and dialect.
[0095] The generation unit can analyze the user's social media activity and generate relevant voice clones when creating voice clones. For example, the generation unit can learn phrases that the user frequently uses on social media and reflect them in the voice clones. The generation unit can also learn topics that the user discusses on social media and reflect them in the voice clones. The generation unit can also learn information about accounts that the user follows on social media and reflect it in the voice clones. This allows for the generation of more appropriate voice clones by analyzing the user's social media activity. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that analyzes the user's social media activity and generates relevant voice clones.
[0096] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0097] The AI intercom guardian system may also include a facial recognition unit. The facial recognition unit recognizes the face of a person approaching the intercom and compares it with a registered facial database. For example, the facial recognition unit can pre-register the faces of family members and friends, and automatically recognize these people when they approach the intercom. The facial recognition unit can also temporarily register the faces of delivery personnel or visitors and recognize them upon their return visits. This allows the facial recognition unit to quickly identify people approaching the intercom and generate appropriate responses. Some or all of the above-described processes in the facial recognition unit may be performed using AI or not. For example, the facial recognition unit can perform facial recognition using an AI model and compare it with a registered facial database.
[0098] The AI intercom guardian system may further include a voice recognition unit. The voice recognition unit recognizes the voice of the person speaking through the intercom and compares it with a registered voice database. For example, the voice recognition unit can pre-register the voices of family members and friends and automatically recognize them when they speak through the intercom. The voice recognition unit can also temporarily register the voices of delivery personnel or visitors and recognize them upon their return visit. This allows the voice recognition unit to quickly identify the person speaking through the intercom and generate an appropriate response. Some or all of the above processing in the voice recognition unit may be performed using AI or not. For example, the voice recognition unit can perform voice recognition using an AI model and compare it with a registered voice database.
[0099] The AI intercom guardian system may further include a motion detection unit. The motion detection unit detects movement around the intercom and identifies abnormal movements. For example, the motion detection unit can monitor the movement of a person approaching the intercom in real time and issue a warning if it detects suspicious movement. The motion detection unit can also record movement around the intercom for later review. This allows the motion detection unit to ensure safety around the intercom and quickly detect abnormal movements. Some or all of the above processing in the motion detection unit may be performed using AI or not. For example, the motion detection unit can perform motion detection using an AI model and detect abnormal movements.
[0100] The AI intercom guardian system may also include a temperature detection unit. The temperature detection unit monitors the temperature around the intercom in real time and detects abnormal temperature changes. For example, the temperature detection unit can issue a warning if the temperature around the intercom rises rapidly. The temperature detection unit can also issue a warning if the temperature around the intercom drops rapidly. This allows the temperature detection unit to quickly detect abnormal temperature changes around the intercom and take appropriate action. Some or all of the above processing in the temperature detection unit may be performed using AI or not. For example, the temperature detection unit can monitor temperature changes using an AI model and detect abnormal temperature changes.
[0101] The AI intercom guardian system may also include an illuminance detection unit. The illuminance detection unit monitors the illuminance around the intercom in real time and detects abnormal changes in illuminance. For example, the illuminance detection unit can issue a warning if the area around the intercom suddenly becomes dark. The illuminance detection unit can also issue a warning if the area around the intercom suddenly becomes bright. This allows the illuminance detection unit to quickly detect abnormal changes in illuminance around the intercom and take appropriate action. Some or all of the above processing in the illuminance detection unit may be performed using AI or not. For example, the illuminance detection unit can monitor changes in illuminance using an AI model and detect abnormal changes in illuminance.
[0102] The AI intercom guardian system can further estimate the user's emotions and adjust its response based on those emotions. For example, if the user is nervous, the AI will respond in a calm tone. If the user is relaxed, the AI may respond in a cheerful tone. If the user is in a hurry, the AI may respond in a rapid tone. This allows for a more natural response by adjusting the response according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the processing described above in the response unit may be performed using AI or not. For example, the response unit can generate a response using an AI model that estimates the user's emotions and adjusts the response based on those emotions.
[0103] The AI intercom guardian system can further estimate the user's emotions and adjust the tone and pitch of the voice conversion based on the estimated emotions. For example, if the user is nervous, the AI can convert the voice with a calmer tone and a lower pitch. If the user is relaxed, the AI can also brighten the tone and keep the pitch natural. If the user is in a hurry, the AI can also speed up the tone and raise the pitch. This allows for more natural voice conversion by adjusting the tone and pitch of the voice conversion according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert the voice using an AI model that estimates the user's emotions and adjusts the tone and pitch of the voice conversion based on the estimated emotions.
[0104] The AI intercom guardian system can further estimate the user's emotions and adjust the voice clone generation method based on the estimated emotions. For example, if the user is relaxed, the AI can generate a voice clone that proceeds at a leisurely pace. If the user is in a hurry, the AI can also generate a voice clone that proceeds at a fast pace. If the user is excited, the AI can also generate a voice clone with visually stimulating effects. This allows for the generation of more natural voice clones by adjusting the voice clone generation method according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can generate voice clones using an AI model that estimates the user's emotions and adjusts the voice clone generation method based on the estimated emotions.
[0105] The AI intercom guardian system can further estimate the user's emotions and adjust the voice expression after conversion based on the estimated emotions. For example, if the user is angry, the AI can convert to a calm voice. If the user is sad, the AI can convert to an encouraging voice. If the user is happy, the AI can convert to a cheerful voice. This allows for more appropriate voice conversion by adjusting the voice expression after conversion according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the conversion unit may be performed using AI or not. For example, the conversion unit can convert speech using an AI model that estimates the user's emotions and adjusts the voice expression after conversion based on the estimated emotions.
[0106] The AI intercom guardian system can further estimate the user's emotions and select a response pattern based on the estimated emotions. For example, if the user is nervous, the AI can select a simple response pattern. If the user is relaxed, the AI can select a more detailed response pattern. If the user is in a hurry, the AI can select a quick response pattern. This allows for more appropriate responses by selecting a response pattern according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the management department may be performed using AI or not. For example, the management department can select a response pattern using an AI model that estimates the user's emotions and selects a response pattern based on the estimated emotions.
[0107] The following briefly describes the processing flow for example form 2.
[0108] Step 1: The conversion unit converts the voice. For example, it converts the voice by changing the tone, pitch, and speed. It can also convert the voice using AI. Specifically, it can change the tone of the voice to make it lower, change the pitch of the voice to make it more masculine, or change the speed of the voice to make the conversation sound more natural. Step 2: The response unit responds based on the audio converted by the conversion unit. For example, it provides a natural response based on the converted audio. AI can also be used to generate responses, producing appropriate responses based on the converted audio and adjusting the timing to maintain a natural pause. Step 3: The management department manages pre-configured response patterns. For example, it is possible to use AI to manage response patterns by utilizing response patterns pre-configured by users. Specifically, it is possible to select an appropriate response based on the response patterns configured by the user, or to adjust the content of the response based on the response patterns. Step 4: The generation unit generates a voice clone. For example, it can generate a voice clone using a pre-registered real voice. It is also possible to generate a voice clone using AI, generating a voice clone based on a pre-registered voice to achieve natural conversation.
[0109] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0110] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0111] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0112] Each of the multiple elements described above, including the conversion unit, response unit, management unit, and generation unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the conversion unit is implemented by the processor 46 of the smart device 14 changing the tone, pitch, and speed of the voice. The response unit provides a natural response based on the voice converted by the control unit 46A of the smart device 14. The management unit manages the pre-set response patterns by the specific processing unit 290 of the data processing unit 12. The generation unit generates voice clones by the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0113] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0114] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0115] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0116] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0117] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0118] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0119] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0120] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0121] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0122] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0123] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0124] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0125] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0126] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0127] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0128] Each of the multiple elements described above, including the conversion unit, response unit, management unit, and generation unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the conversion unit is implemented by the processor 46 of the smart glasses 214 changing the tone, pitch, and speed of the voice. The response unit provides a natural response based on the voice converted by the control unit 46A of the smart glasses 214. The management unit manages the pre-set response patterns by the specific processing unit 290 of the data processing unit 12. The generation unit generates voice clones by the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0129] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0130] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0131] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0132] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0133] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0134] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0135] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0136] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0137] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0138] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0139] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0140] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0141] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0142] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0143] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0144] Each of the multiple elements described above, including the conversion unit, response unit, management unit, and generation unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the conversion unit is implemented by the processor 46 of the headset terminal 314 changing the tone, pitch, and speed of the voice. The response unit provides a natural response based on the voice converted by the control unit 46A of the headset terminal 314. The management unit manages the pre-set response patterns by the specific processing unit 290 of the data processing unit 12. The generation unit generates voice clones by the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0145] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0146] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0147] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0148] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0149] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0150] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0151] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0152] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0153] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0154] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0155] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0156] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0157] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0158] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0159] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0160] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0161] Each of the multiple elements described above, including the conversion unit, response unit, management unit, and generation unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the conversion unit is implemented by the robot 414's processor 46 changing the tone, pitch, and speed of the voice. The response unit provides a natural response based on the voice converted by the robot 414's control unit 46A. The management unit manages the pre-set response patterns by the specific processing unit 290 of the data processing unit 12. The generation unit generates voice clones by the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0162] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0163] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0164] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0165] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0166] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0167] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0168] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0169] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0170] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0171] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0172] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0173] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0174] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0175] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0176] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0177] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0178] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0179] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0180] (Note 1) A conversion unit that converts audio, A response unit that responds based on the voice converted by the conversion unit, The management department manages pre-set response patterns, It comprises a generation unit that generates voice clones, A system characterized by the following features. (Note 2) The conversion unit is Convert audio in real time The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned management department, Use pre-set response patterns The system described in Appendix 1, characterized by the features described herein. (Note 4) The generating unit is Uses real voices that have been registered in advance. The system described in Appendix 1, characterized by the features described herein. (Note 5) The response unit is Responds naturally based on the converted speech. The system described in Appendix 1, characterized by the features described herein. (Note 6) The conversion unit is It estimates the user's emotions and adjusts the tone and pitch of the voice conversion based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The conversion unit is During speech conversion, background noise is analyzed and appropriate noise cancellation is performed. The system described in Appendix 1, characterized by the features described herein. (Note 8) The conversion unit is During speech conversion, the system learns the characteristics of the user's voice to achieve a more natural conversion. The system described in Appendix 1, characterized by the features described herein. (Note 9) The conversion unit is It estimates the user's emotions and adjusts the voice expression after conversion based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The conversion unit is During speech conversion, the system takes into account the user's geographical accent and dialect. The system described in Appendix 1, characterized by the features described herein. (Note 11) The conversion unit is During speech conversion, the system refers to the user's past conversation history to perform appropriate conversions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The response unit is It estimates the user's emotions and adjusts the content and tone of the response based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The response unit is When responding, the system analyzes the tone and content of the other person's voice to generate an appropriate response. The system described in Appendix 1, characterized by the features described herein. (Note 14) The response unit is When responding, the system references the user's past response history to enable natural conversation. The system described in Appendix 1, characterized by the features described herein. (Note 15) The response unit is It estimates the user's emotions and adjusts the response speed and timing based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The response unit is When responding, take into account the other person's geographical accent and dialect. The system described in Appendix 1, characterized by the features described herein. (Note 17) The response unit is When responding, the system analyzes the background noise of the other party and generates an appropriate response. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned management department, The system estimates the user's emotions and selects response patterns based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned management department, During management, the system analyzes the user's past response history and suggests the optimal response pattern. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned management department, During administration, the response patterns are customized based on the user's lifestyle and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned management department, The system estimates the user's emotions and prioritizes response patterns based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned management department, During management, select response patterns considering the user's geographical information. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned management department, During management, the system analyzes users' social media activity and suggests relevant response patterns. The system described in Appendix 1, characterized by the features described herein. (Note 24) The generating unit is It estimates the user's emotions and adjusts the voice clone generation method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The generating unit is When generating a voice clone, the system analyzes the user's past voice data to create a more natural-sounding clone. The system described in Appendix 1, characterized by the features described herein. (Note 26) The generating unit is During voice cloning, the system learns the characteristics of the user's voice and generates individually optimized clones. The system described in Appendix 1, characterized by the features described herein. (Note 27) The generating unit is It estimates the user's emotions and adjusts the tone and pitch of the voice clone based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The generating unit is When generating voice clones, the system takes into account the user's geographical accent and dialect. The system described in Appendix 1, characterized by the features described herein. (Note 29) The generating unit is When generating voice clones, the system analyzes the user's social media activity and generates relevant voice clones. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0181] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A conversion unit that converts audio, A response unit that responds based on the voice converted by the conversion unit, The management department manages pre-set response patterns, It comprises a generation unit that generates voice clones, A system characterized by the following features.
2. The conversion unit is Convert audio in real time The system according to feature 1.
3. The aforementioned management department, Use pre-set response patterns The system according to feature 1.
4. The generating unit is Uses real voices that have been registered in advance. The system according to feature 1.
5. The response unit is Responds naturally based on the converted speech. The system according to feature 1.
6. The conversion unit is It estimates the user's emotions and adjusts the tone and pitch of the voice conversion based on the estimated emotions. The system according to feature 1.
7. The conversion unit is During speech conversion, background noise is analyzed and appropriate noise cancellation is performed. The system according to feature 1.
8. The conversion unit is During speech conversion, the system learns the characteristics of the user's voice to achieve a more natural conversion. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A