Reproduction device, reproduction method, program and reproduction system

The playback device converts voice quality of played content in real time, addressing the limitations of conventional applications by allowing users to set desired voice qualities, thereby enhancing user engagement.

JP2025114455APending Publication Date: 2025-08-05DOWANGO KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024189829
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Conventional voice conversion applications are limited to converting the quality of one's own voice and do not address the conversion of voice quality in content played on mobile devices or personal computers, missing the opportunity to enhance user engagement by adapting voice quality to personal preferences.

Method used

A playback device with a content player, virtual audio device, and voice conversion unit that processes audio signals to convert the voice quality of played content in real time, allowing users to set desired voice qualities for output.

Benefits of technology

Enables real-time conversion of voice quality for content playback, enhancing user experience by adapting voice characteristics to personal preferences and improving engagement with content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114455000001_ABST
    Figure 2025114455000001_ABST
Patent Text Reader

Abstract

To provide a reproduction device for converting a voice quality of audio content being reproduced.SOLUTION: A reproduction device 10 includes a content player 11 for reproducing content, a virtual audio device 12 for receiving an audio signal output from the content player 11, a voice conversion unit 13 for receiving an audio signal from the virtual audio device 12 and converting a voice quality of a voice included in the audio signal, and an audio device 14 for outputting an audio signal obtained by converting the voice quality of the voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a playback device, a playback method, a program, and a playback system. [Background technology]

[0002] In recent years, voice conversion technology that converts voice quality in real time has been developed. The technology disclosed in Non-Patent Document 1 can convert one's own voice into the voice of another character in real time. The technology disclosed in Patent Document 1 can convert voice quality by reflecting voice characteristics such as whispering, falsetto, and angry voice. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7179216 [Non-patent literature]

[0004] [Non-Patent Document 1] “Voice Changer”, [online], CoeFont Co., Ltd., Internet〈 URL: https: / / vc.coefont.cloud / 〉 Summary of the Invention [Problem to be solved by the invention]

[0005] Conventional voice conversion applications convert the quality of your own voice by inputting your own voice from a microphone into a deep learning model, but are not designed to convert the quality of the voice of content (videos or radio) being played on a mobile device or personal computer. In fact, the challenge of converting the voice quality of content in real time has not been recognized.

[0006] For example, to improve the situation where users are interested in content such as recorded lectures but are not interested in watching it, it is thought that converting the voice quality of the content to suit the user's preferences would be effective.Also, by converting the voice quality of the content, users may be able to discover new interesting aspects of the content.

[0007] The present disclosure has been made in view of the above, and aims to convert the voice quality of the audio of content being played back. [Means for solving the problem]

[0008] A playback device according to one aspect of the present disclosure includes a playback unit that plays content, a virtual audio device that inputs an audio signal output by the playback unit, a conversion unit that inputs the audio signal from the virtual audio device and converts the voice quality of the audio contained in the audio signal, and an audio device that outputs the audio with the converted voice quality. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to convert the voice quality of the audio of content being played back. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a playback device according to the first embodiment. [Figure 2] FIG. 2 is a flowchart showing an example of the flow of processing performed by the playback device of the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of a sound setting screen of the playback device. [Figure 4] FIG. 4 is a diagram showing an example of the configuration of a playback device according to the second embodiment. [Figure 5] FIG. 5 is a flowchart showing an example of the flow of processing performed by the playback device of the second embodiment. [Figure 6] FIG. 6 is a diagram showing an example of the configuration of a playback device according to the third embodiment. [Figure 7] FIG. 7 is a flowchart showing an example of the flow of processing performed by the playback device of the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] [First embodiment] 1 is a diagram showing an example of the configuration of a playback device 10 according to the first embodiment. The playback device 10 shown in the diagram includes a content player 11, a virtual audio device 12, a voice conversion unit 13, and an audio device 14.

[0012] The playback device 10 may be a mobile terminal such as a wearable device, a smartphone, a tablet, or a portable media player, or a terminal such as a personal computer, a game console, a smart home device, or a smart TV.

[0013] The content player 11 plays back content. Examples of content include content containing human speech, such as remote conferences, lectures, voice chats, radio, and audiobooks. The content may be not only audio content but also video containing audio and video. The content player 11 may play back content that has already been downloaded to the terminal, or may play back content while downloading it from the network (streaming playback).

[0014] The content player 11 may be a different application for each type of content. Examples of applications that can be used as the content player 11 include a remote conference application, a chat application, an internet radio application, an audiobook application, and a video playback application.

[0015] The virtual audio device 12 simulates the input and output of audio signals on software without requiring a physical device. When the output destination of an audio signal is changed from a speaker to the virtual audio device 12, the audio signal is input to the virtual audio device 12. When the input source of an audio signal is changed from a microphone to the virtual audio device 12, an audio application that processes the audio input (for example, a voice conversion application, a recording application, a mixer application, etc.) can input and process the audio signal from the virtual audio device 12. By setting the input and output of the audio signal of the playback device 10 to the virtual audio device 12, the audio output of the content player 11 is input to the virtual audio device 12, and the voice conversion unit 13 receives audio input from the virtual audio device 12.

[0016] 1, when the audio signal output is set to audio device 14 (for example, a speaker or earphones), the audio signal from content player 11 is input to audio device 14, indicated by the dashed arrow. When the audio signal output is set to virtual audio device 12, the audio signal from content player 11 is input to virtual audio device 12, indicated by the solid arrow.

[0017] The virtual audio device 12 may be provided as an application. By installing the application in the playback device 10, the playback device 10 or an application running on the playback device 10 can set audio output and audio input to the virtual audio device 12.

[0018] The voice conversion unit 13 converts the voice of the input audio signal (voice of the content) into a desired voice quality in real time, and outputs the converted voice (audio signal) to the audio device 14. For the voice conversion unit 13, the technology of Patent Document 1 and Non-Patent Document 1, for example, can be used. Specifically, the voice conversion unit 13 inputs the voice of the content into a deep-learned neural network (voice conversion AI) to obtain voice with converted voice quality. The voice conversion unit 13 may perform preprocessing such as noise removal and volume adjustment before inputting the voice of the content into the voice conversion AI.

[0019] The voice conversion unit 13 may input information specifying the post-conversion voice quality to the voice conversion AI to obtain speech converted into any voice quality. For example, if the voice conversion unit 13 can convert into the voice qualities of multiple speakers, the voice conversion unit 13 may accept from the user a specification (speaker identifier) of the speaker with the post-conversion voice quality, input the content voice and the speaker specification to the voice conversion AI, and obtain speech converted into the voice quality of the specified speaker. When accepting a specification of voice quality from the user, the voice conversion unit 13 may present several to several tens of speakers to the user and accept the selection of a speaker with a post-conversion voice quality, or may accept from the user characteristics of the post-conversion voice quality (e.g., male, female, child, adult, high, low, etc.).

[0020] The voice conversion unit 13 may use the technology of Patent Document 1 to input information on a vocalization style (for example, a calm voice, a whisper, a falsetto, an angry voice, etc.) into a voice conversion AI to obtain a voice converted into an arbitrary vocalization style. The voice conversion unit 13 may also present vocalization styles to the user so that the user can select an arbitrary vocalization style.

[0021] The voice conversion unit 13 sets the audio input to the virtual audio device 12 and the audio output to the audio device 14. The user may be able to set the audio output of the voice conversion unit 13. For example, the user can set the audio output of the voice conversion unit 13 to a speaker, earphones, an external device, a recording application, a mixer application, or the like. If the audio output of the voice conversion unit 13 is set to an audio application such as a recording application or a mixer application, the converted voice can be processed by another audio application.

[0022] The voice conversion unit 13 may be provided as an application. The voice conversion unit 13 may have the functionality of the virtual audio device 12.

[0023] The audio device 14 is physical hardware such as a speaker or earphone that actually outputs sound. When an audio signal is input to the audio device 14, sound is output from the speaker, earphone, or the like. By setting the output of the voice conversion unit 13 to the audio device 14, the sound converted by the voice conversion unit 13 is output. Since the output of the content player 11 is set to the virtual audio device 12, the sound of the content played by the content player 11 is not output, but the sound converted by the voice conversion unit 13 is output.

[0024] Next, the flow of processing performed by the playback device 10 of the first embodiment will be described with reference to the flowchart of FIG.

[0025] In step S11, the system audio output is set to the virtual audio device 12. The system audio output is a setting for the default output destination of the sound of the playback device 10. When the system audio output is set to speakers or earphones, the sound output by the playback device 10 (for example, the audio of content played by the content player 11) is output from the speakers or earphones. By setting the system audio output to the virtual audio device 12, the sound (audio signal) output by the playback device 10 is input to the virtual audio device 12.

[0026] If audio output is not set for each application, the sound of the application is output according to the system audio output setting. If audio output is set individually for each application, the sound of the application is output according to the application audio output setting. In step S11, the audio output of the content player 11 may be set to the virtual audio device 12.

[0027] In step S12, the system audio input is set to the virtual audio device 12. The system audio input is a setting for the default input source of sound to the playback device 10. By setting the system audio input to the virtual audio device 12, the voice conversion unit 13 converts the voice quality of the voice input to the virtual audio device 12. For example, if the system audio input is set to a microphone, the voice conversion unit 13 converts the voice quality of the voice collected by the microphone. Note that if the voice conversion unit 13 can set its own audio input, the audio input of the voice conversion unit 13 may be set to the virtual audio device 12 without changing the system audio input.

[0028] In step S13, the audio output of the voice conversion unit 13 is set to the audio device 14. Because the system audio output is set to the virtual audio device 12, if the default value is used, the converted voice will be input to the virtual audio device 12. By setting the audio output of the voice conversion unit 13 to the audio device 14, the voice converted by the voice conversion unit 13 is output from the audio device 14, such as a speaker or earphones.

[0029] FIG. 3 shows an example of a sound setting screen 100 of the playback device 10. The sound setting screen 100 shown in the figure shows a setting field 110 for the system and a setting field 120 for the voice conversion unit 13. The audio output and audio input of the system are set using items 111 and 112. The audio output and audio input of the voice conversion unit 13 are set using items 121 and 122. In FIG. 3, both the audio output and audio input of the system are set to the virtual audio device 12. The audio input of the voice conversion unit 13 is set to the virtual audio device 12, and the audio output is set to the speaker.

[0030] In the processing from steps S11 to S13, the flow of audio signals within the playback device 10 is set as shown by the solid arrows in FIG. 1. The processing from steps S11 to S13 may be performed by the user or by the voice conversion unit 13. For example, when the voice conversion unit 13 is started, the voice conversion unit 13 may change the sound settings. The system audio output and audio input are set to the virtual audio device 12, and the audio output of the voice conversion unit 13 is set to the previous system audio output (for example, speakers or earphones). When the function of the voice conversion unit 13 is stopped, the voice conversion unit 13 returns the sound settings to the state before the voice conversion unit 13 was started.

[0031] If the voice conversion unit 13 has the function of the virtual audio device 12, the audio output of the system or content player 11 is set to the voice conversion unit 13, and the audio output of the voice conversion unit 13 is set to the audio device .

[0032] After the sound is set, when the content player 11 starts playing the content, the audio signal output by the content player 11 is input to the voice conversion unit 13 via the virtual audio device 12 in step S14.

[0033] In step S15, the voice conversion unit 13 converts the voice quality of the input audio signal.

[0034] In step S16, the voice conversion unit 13 outputs the converted voice to the speaker.

[0035] The processing from steps S14 to S16 is repeated, and the voice output from the content player 11 has its voice quality converted in real time by the voice conversion unit 13 and is output from the speaker.

[0036] As described above, the playback device 10 of this embodiment includes a content player 11 that plays back content, a virtual audio device 12 that receives an audio signal output by the content player 11, a voice conversion unit 13 that receives an audio signal from the virtual audio device 12 and converts the voice quality of the voice included in the audio signal, and an audio device 14 that outputs the voice quality converted. This makes it possible to convert the voice quality of the voice of content played back by the content player 11 operating within the playback device 10.

[0037] By setting the default audio output and audio input of the playback device 10 to the virtual audio device 12, and setting the audio output of the voice conversion unit 13 to the audio device 14, it is possible to convert the voice output by the playback device 10 to a desired voice quality. Conventional voice conversion applications convert the voice quality of voice input from an external source such as a microphone, but the playback device 10 can convert the voice quality of voice output by other applications running within the playback device 10.

[0038] [Second embodiment] FIG. 4 is a diagram showing an example of the configuration of a playback device 20 of the second embodiment. The playback device 20 shown in the diagram is configured by adding an audio separation unit 21 and a mixer 22 to the playback device 10 of the first embodiment. When content includes human voices and background environmental sounds or music (hereinafter referred to as background sounds), the playback device 20 separates the voices from the background sounds, converts only the voices, and synthesizes the converted voices with the background sounds. Duplicate descriptions of the same configuration as the playback device 10 of the first embodiment will be omitted here.

[0039] The audio separation unit 21 receives an audio signal output from the content player 11 via the virtual audio device 12, and separates the audio signal into human voice and background sound other than human voice. The audio separation unit 21 inputs the audio signal of human voice to the voice conversion unit 13, and inputs the audio signal of background sound to the mixer 22.

[0040] It is not necessary to completely separate the human voice. Even if some background noise is mixed with the human voice, the voice conversion unit 13 performs preprocessing such as noise removal and volume adjustment to convert the voice into the desired voice quality. Furthermore, even if some human voice is mixed with the background noise, the mixer 22 combines the converted voice with the background noise, so it is not noticeable.

[0041] The voice conversion unit 13 receives only the separated human voice and converts the voice into an arbitrary voice quality in real time. The voice conversion unit 13 is the same as that in the first embodiment.

[0042] The mixer 22 inputs and synthesizes the voice converted by the voice conversion unit 13 and the background sound separated by the voice separation unit 21, and outputs the synthesized voice to the audio device 14. If a delay occurs in the processing of the voice conversion unit 13 and the synthesized voice sounds unnatural, the mixer 22 may buffer the background sound and synthesize the converted voice and the background sound in synchronization with each other.

[0043] The mixer 22 may output only the converted voice to the audio device 14 without synthesizing the converted voice with the background sound. In this case, the voice of the content is converted to a desired voice quality, and an output from which the background sound has been removed is obtained.

[0044] The voice separation unit 21, voice conversion unit 13, and mixer 22 may each be configured as separate applications, or the voice conversion unit 13 may have the functions of the voice separation unit 21 and mixer 22, and the voice separation unit 21, voice conversion unit 13, and mixer 22 may be configured as a single voice conversion application.

[0045] The playback device 20 may be a mobile terminal such as a wearable device, a smartphone, a tablet, or a portable media player, or a terminal such as a personal computer, a game console, a smart home device, or a smart TV.

[0046] Next, the flow of processing performed by the playback device 20 of the second embodiment will be described with reference to the flowchart of FIG.

[0047] Before the conversion process of the playback device 20, sound settings, i.e., the flow of audio signals, are set as shown by the solid arrows in Figure 4. For example, the audio output and audio input of the system are set to the virtual audio device 12. The audio input of the voice separation unit 21 is set to the virtual audio device 12, and the audio output of the voice separation unit 21 is set to the voice conversion unit 13 and the mixer 22. Settings are made so that voice is input to the voice conversion unit 13 and background sound is input to the mixer 22. The audio input of the mixer 22 is set to the voice conversion unit 13 and the voice separation unit 21, and the audio output of the mixer 22 is set to the audio device 14. Note that if the voice conversion unit 13 has the functions of the voice separation unit 21 and the mixer 22, the audio output and audio input of the system are set to the virtual audio device 12, and the audio output of the voice conversion unit 13 is set to the audio device 14, as in the first embodiment.

[0048] After the sound is set, when the content player 11 starts playing back the content, the audio signal output by the content player 11 is input to the audio separator 21 in step S21.

[0049] In step S22, the voice separation unit 21 separates the input audio signal into a human voice and background sound. The audio signal of the human voice is input to the voice conversion unit 13, and the audio signal of the background sound is input to the mixer 22.

[0050] In step S23, the voice conversion unit 13 converts the voice quality of the input audio signal. The converted audio signal is input to the mixer 22.

[0051] In step S24, the mixer 22 mixes the converted audio with the background sound.

[0052] In step S25, the mixer 22 outputs the audio obtained by mixing the converted audio with the background sound from the speaker.

[0053] The processing from steps S21 to S25 is repeated, and the voice output from the content player 11 is converted in real time by the voice conversion unit 13 and output from the speaker.

[0054] The playback device 20 of this embodiment includes a voice separation unit 21 that separates the voice contained in the audio signal of the content from background sounds other than the voice, a voice conversion unit 13 that converts the voice quality of the voice, and a mixer 22 that combines the converted voice with the background sound. This makes it possible to improve the quality of voice conversion for content that includes background sounds other than human voices. Furthermore, if the background sound is not mixed in, it is possible to obtain voice that has been converted to the desired voice quality by separating only the voice from the content.

[0055] [Third embodiment] 6 is a diagram showing an example of the configuration of a playback device 30 of the third embodiment. The playback device 30 shown in the figure includes a voice conversion unit 13, which includes a speech recognition unit 31, a text conversion unit 32, and a voice synthesis unit 33, and converts not only the voice quality but also expressions such as the first-person pronunciation, endings, honorifics, and wording. Duplicate descriptions of the same configuration as the playback device 10 of the first embodiment will be omitted here.

[0056] The voice recognition unit 31 receives an audio signal output from the content player 11 via the virtual audio device 12, recognizes the voice, and converts the voice of the content into text. Existing voice recognition technology can be used to convert human voice into text. The voice recognition unit 31 may use an external voice recognition service. For example, the voice recognition unit 31 transfers the input audio signal to a voice recognition service (voice recognition server) and receives text of the voice recognition result from the voice recognition service.

[0057] The text conversion unit 32 converts the text input from the speech recognition unit 31 into a different expression without significantly changing the content of the text. For example, the text conversion unit 32 converts expressions such as the first person, endings, honorifics, or phrases of sentences in the text into desired expressions. Existing Large Language Model (LLM) technology can be used for text conversion. For example, the text conversion unit 32 inputs instructions to the LLM to change the expression of the text from the text before conversion, and obtains the converted text from the LLM. An external service may be used for the LLM. Alternatively, the text conversion unit 32 may simply replace words in the input text with other words.

[0058] The user may specify the expression of the converted text, such as first person speech and endings.

[0059] The voice synthesis unit 33 synthesizes a voice in which the converted text is read aloud in a desired voice. The voice synthesis unit 33 uses, for example, the neural network of Patent Document 1 as a Text-to-Speech (TTS) model to synthesize a voice in which the text is read aloud in a desired voice. As with the first embodiment, the user may be able to select the voice quality and voice generation method after conversion for the voice to be synthesized. The voice synthesis unit 33 may use an external voice synthesis service. For example, the voice synthesis unit 33 transfers the input text to a voice synthesis service (voice synthesis server) and receives an audio signal of the voice reading the text from the voice synthesis service. The voice quality and voice generation method may be input to the voice synthesis service, and voice may be synthesized with the desired voice quality and voice generation method.

[0060] The configuration of the voice conversion unit 13 described above can be applied to the voice conversion unit 13 of either the playback device 10 or 20 of the first or second embodiment.

[0061] The playback device 30 may be a mobile terminal such as a wearable device, a smartphone, a tablet, or a portable media player, or a terminal such as a personal computer, a game console, a smart home device, or a smart TV.

[0062] Next, the flow of processing performed by the playback device 30 of the third embodiment will be described with reference to the flowchart of FIG.

[0063] Before the conversion process of the playback device 30, sound settings, i.e., the flow of audio signals, are set as shown by the solid arrows in Fig. 6. For example, as in the first embodiment, the audio output and audio input of the system are set to the virtual audio device 12, and the audio output of the voice conversion unit 13 is set to the audio device 14. Note that the dashed arrows in Fig. 6 indicate the flow of text information.

[0064] After the sound is set, when the content player 11 starts playing back the content, the audio signal output by the content player 11 is input to the voice recognition unit 31 in step S31.

[0065] In step S32, the voice recognition unit 31 recognizes the input audio signal and converts it into text.

[0066] In step S33, the text conversion unit 32 converts the text into a desired expression.

[0067] In step S34, the voice synthesis unit 33 synthesizes a voice that reads the text aloud.

[0068] In step S35, the voice synthesis unit 33 outputs the synthesized voice to the speaker.

[0069] The processing from steps S31 to S35 is repeated, and the voice output from the content player 11 has its expression changed, is converted in real time by the voice conversion unit 13, and is output from the speaker.

[0070] As described above, the voice conversion unit 13 includes a voice recognition unit 31 that recognizes the voice contained in the audio signal and converts it into text, a text conversion unit 32 that converts the expression of the text, and a voice synthesis unit 33 that synthesizes a voice that reads out the converted text. This makes it possible to obtain voice that has been converted into a desired expression, such as the first person way of speaking or endings, in addition to the voice quality of the voice of the content.

[0071] [Variations] Next, an example of a modification of the playback devices 10, 20, and 30 will be described.

[0072] When content includes voices from multiple people, the voice conversion unit 13 may convert the voices of each speaker into different voice qualities, or may convert only the voice quality of a specific speaker. For example, the voice separation unit 21 separates the voices of the content into those of each speaker, the voice conversion unit 13 converts the voice qualities of each speaker individually, and the mixer 22 mixes the converted voices.

[0073] If the voice of the content is difficult to hear due to poor pronunciation, the voice quality may be converted to improve pronunciation. For example, the voice recognition unit 31 converts the voice of the content into text, the text conversion unit 32 uses LLM to infer and correct conversion errors caused by poor pronunciation, and the voice synthesis unit 33 synthesizes the corrected text and outputs it.

[0074] In the case of audio-only content, a character corresponding to the converted voice quality may be displayed, and the character may lip-sync in accordance with the audio.

[0075] The voice conversion unit 13 of the playback devices 10, 20, and 30 may use an external voice conversion service. For example, a playback system includes the playback devices 10, 20, and 30 and a voice conversion server. The voice conversion server receives input of a voice (audio signal), converts the voice quality of the input voice, and returns the converted voice. The voice conversion unit 13 of the playback devices 10, 20, and 30 transmits the input audio signal to the voice conversion server and receives the converted audio signal from the voice conversion server.

[0076] The voice conversion server may be a server utilizing the service of Non-Patent Document 1 or the technology of Patent Document 1. The voice conversion server may receive a speaker identifier or parameters specifying the post-conversion voice quality and convert the voice into the specified voice quality. The voice conversion server may also input the post-conversion vocalization style and convert the converted voice into the specified vocalization style.

[0077] Any part or all of the functional units described in this disclosure may be implemented by a program. The programs described in this disclosure may be non-temporarily recorded on a computer-readable recording medium and distributed, distributed via a communication line (including wireless communication) such as the Internet, or distributed in a state where they are installed on any terminal. While a person skilled in the art may conceive additional effects and various modifications of the present invention based on the above description, the aspects of this disclosure are not limited to the individual embodiments described above. Various additions, modifications, and partial deletions are possible within the scope of the conceptual idea and spirit of the present invention as derived from the content defined in the claims and their equivalents. For example, what is described in this disclosure as a single device (or component, the same applies hereinafter) (including what is depicted as a single device in the drawings) may be implemented by multiple devices. Conversely, what is described in this disclosure as multiple devices (including what is depicted as multiple devices in the drawings) may be implemented by a single device. Alternatively, some or all of the means and functions included in one device (e.g., a server) may be included in another device (e.g., a user terminal). Furthermore, a "system" may be made up of one device, or may be made up of two or more devices (for example, a server and a user terminal, or multiple user terminals).

[0078] Furthermore, not all of the features described in this disclosure are essential requirements. In particular, features described in this disclosure but not in the claims can be considered optional additional features.

[0079] Please note that the applicant is only aware of the inventions disclosed in the documents listed in the "Prior Art Documents" section of this disclosure, and the present disclosure does not necessarily aim to solve the problems of the disclosed inventions. The problem that the present disclosure aims to solve should be determined by taking into consideration the entire disclosure. For example, if the present disclosure states that a specific configuration achieves a certain effect, it can also be said that the present disclosure solves a problem that is the reverse of the certain effect. However, it is not necessarily intended that such a specific configuration be an essential requirement. [Explanation of symbols]

[0080] 10,20,30 Playback device 11 Content Player 12 Virtual Audio Devices 13 Voice conversion unit 14 Audio Devices 21 Audio separation section 22 Mixer 31 Voice Recognition Unit 32 Text conversion section 33 Voice synthesis section

Claims

1. a playback unit that plays back content; a virtual audio device that receives the audio signal output from the playback unit; a conversion unit that receives the audio signal from the virtual audio device and converts the voice quality of the voice included in the audio signal; An audio device that outputs the voice with the converted voice quality is provided. playback device.

2. 2. The playback device according to claim 1, The conversion unit a separation unit that separates a voice included in the audio signal from a background sound other than the voice; a voice conversion unit that converts the voice quality of the voice; A mixer is provided for mixing the converted voice with the background sound. playback device.

3. 2. The playback device according to claim 1, The conversion unit a speech recognition unit that recognizes speech contained in the audio signal and converts it into text; a text conversion unit for converting the representation of the text; It has a voice synthesis unit that synthesizes a voice that reads the converted text. playback device.

4. 4. A playback device according to claim 1, The default audio output and audio input of the playback device are set to the virtual audio device, and the audio output of the conversion unit is set to the audio device. playback device.

5. A playback method by a playback device, Play the content and input the audio signal to the virtual audio device. inputting the audio signal from the virtual audio device and converting the voice quality of the voice included in the audio signal; Output the voice with the converted voice quality. How to play.

6. On the computer, A process of playing content and inputting an audio signal to a virtual audio device; a process of inputting the audio signal from the virtual audio device and converting the voice quality of the voice included in the audio signal; Execute a process to output the voice with the converted voice quality. program.

7. A playback system including a playback device and a voice conversion server, The playback device a playback unit that plays back content; a virtual audio device that receives the audio signal output from the playback unit; a conversion unit that receives the audio signal from the virtual audio device and converts the voice quality of the voice included in the audio signal; An audio device that outputs the voice with the converted voice quality is provided. Playback system.

8. 8. The playback system of claim 7, The conversion unit transmits the audio signal to the voice conversion server and receives the audio signal whose voice quality has been converted. Playback system.

Citation Information

Patent Citations

  • Voice conversion device, voice conversion method, voice conversion neural network, program, and recording medium

    JP7179216B1