Content voice / sound processing device, content voice / sound reproducing device, content voice / sound processing method, content voice / sound processing program and recording medium with program recorded thereon

The content voice/sound processing device addresses speech recognition errors by adjusting volume and frequency levels of content voice and sound based on genre, ensuring clear content audibility and reducing user dissatisfaction.

JP2025134096APending Publication Date: 2025-09-17SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024031770
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-04
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Existing devices with limited processing power or without echo cancellation capabilities face challenges in accurately recognizing human speech due to interference from content voice and sound, leading to reduced user satisfaction and increased errors in speech recognition.

Method used

A content voice/sound processing device that includes a genre determination unit to identify the content genre, a human voice recognition unit, and a content voice/sound processing unit to adjust the volume and frequency levels of content voice and sound to enhance audibility and reduce misrecognition, while maintaining overall volume levels.

Benefits of technology

The solution effectively suppresses speech recognition errors and maintains user satisfaction by ensuring content voice and sound remain audible, even during human voice recognition, by dynamically adjusting volume and frequency levels based on content genre.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025134096000001_ABST
    Figure 2025134096000001_ABST
Patent Text Reader

Abstract

To suppress the decline in satisfaction of viewing of contents, while suppressing an error in recognition of voice.SOLUTION: A content voice / sound processing device comprises: a human voice recognition unit which recognizes human voice on the basis of a human voice signal in which the human voice is inputted by a human voice input unit; a genre determination unit which determines a genre of contents being reproduced on the basis of a content signal inputted from a content input unit; and a content voice / sound processing unit which controls at least one of content voice outputted by a content voice / sound output unit due to reproduction of the contents and content sound other than the content voice. The content voice / sound processing unit changes the volume size of at least one of the content voice and the content sound, so as to suppress the reduction of audibility by people in one of the content voice or the content sound.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a content voice / sound processing device, a content voice / sound reproduction device, a content voice / sound processing method, a content voice / sound processing program, and a recording medium on which the program is recorded. [Background technology]

[0002] In recent years, technologies have been developed that recognize human voices and operate devices based on the recognized voices. Devices that utilize such human voice recognition technologies include smartphones and content voice / sound playback devices such as television sets.

[0003] If the device to be operated is a device equipped with a speaker that emits some kind of content voice and content sound, the content voice and content sound will become noise when recognizing human voice.

[0004] Typically, a synthesized signal is generated that contains a content voice signal or a content sound signal that can identify the content voice or content sound emitted by such a device, mixed with a human voice signal that can identify human speech. Therefore, echo cancellation technology is used to remove the content voice signal and content sound signal emitted by the device from the synthesized signal, thereby improving the accuracy of human speech recognition.

[0005] However, because echo cancellation requires high-speed calculations, devices with limited processing power may not have sufficient echo cancellation capabilities. Also, some devices may not have an echo cancellation function. In these cases, the accuracy of human speech recognition is low, which can lead to speech recognition errors.

[0006] In such cases, for example, a technology as described in Patent Document 1 is disclosed for television receivers. According to this technology, content voice and content sound are prevented from being input from a microphone together with human voice during the period when human voice is being recognized. Specifically, when human voice recognition starts, the output volume setting value of a content voice / sound playback device such as a television receiver is reduced to a value equal to or lower than a predetermined value until human voice recognition ends. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-181374 Summary of the Invention [Problem to be solved by the invention]

[0008] According to the technology disclosed in Patent Document 1, when processing human voice recognition, the overall volume of the output sound emitted from the speaker is simply reduced. As a result, for example, when watching a news program, the user finds it difficult to hear the anchor's voice, making it difficult to understand the content of the program. Also, for example, when watching a music program, the user finds the temporary reduction in the overall volume of the output sound emitted from the speaker unpleasant. In either of these cases, the user's satisfaction with viewing the content decreases.

[0009] The present disclosure has been made in consideration of the above-mentioned problems. An object of the present disclosure is to provide a content voice / sound processing device and a content voice / sound playback device that can suppress voice recognition errors while suppressing a decrease in satisfaction with content viewing. Another object of the present disclosure is to provide a content voice / sound processing method used for the content voice / sound processing device, a content voice / sound processing program, and a computer-readable recording medium on which the program is recorded. [Means for solving the problem]

[0010] A content voice / sound processing device according to one embodiment of the present disclosure includes a human voice recognition unit that recognizes human voice based on a human voice signal input by a human voice input unit, a genre determination unit that determines the genre of the content being played based on the content signal input from the content input unit, and a content voice / sound processing unit that controls at least one of the content voice and content sound other than the content voice output by a content voice / sound output unit upon playback of the content, and the content voice / sound processing unit changes the volume of at least one of the content voice and the content sound so as to suppress a decrease in the ease with which the content voice or the content sound can be heard by humans.

[0011] A content voice / sound reproducing device according to one aspect of the present disclosure includes the content voice / sound processing device, a speaker as the content voice / sound output unit, and a microphone as the human voice input unit.

[0012] A content voice / sound processing method of one embodiment of the present disclosure comprises the steps of recognizing a human voice based on a human voice signal input by a human voice input unit, determining the genre of the content being played based on the content signal input from the content input unit, and controlling at least one of the content voice and content sound other than the content voice output by a content voice / sound output unit upon playback of the content, wherein in the step of controlling at least one of the content voice and content sound, the volume of at least one of the content voice and content sound is changed so as to suppress a decrease in the ease with which the content voice or content sound can be heard by humans.

[0013] A content voice / sound processing program of one embodiment of the present disclosure is a content voice / sound processing program for causing a computer to operate as a human voice recognition unit that recognizes human voice based on a human voice signal input by a human voice input unit, a genre determination unit that determines the genre of the content being played based on the content signal input from the content input unit, and a content voice / sound processing unit that controls at least one of the content voice and content sound other than the content voice output by a content voice / sound output unit upon playback of the content, wherein the content voice / sound processing unit changes the volume of at least one of the content voice and the content sound so as to suppress a decrease in the ease with which humans can hear the content voice or the content sound.

[0014] A recording medium according to one aspect of the present disclosure is a computer-readable recording medium on which the content voice / sound processing program is recorded. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a functional block diagram of a content voice / sound reproducing device according to a first embodiment. [Figure 2A] 1 is a first graph for explaining the control of ease of hearing of content voice or content sound by the content voice / sound reproducing device of the first embodiment. [Figure 2B] 10 is a second graph for explaining the control of ease of hearing of the content voice or content sound by the content voice / sound reproducing device of the first embodiment. [Figure 3] 10 is a third graph for explaining the control of ease of hearing of the content voice or content sound of the content voice / sound reproducing device of the content voice according to the first embodiment. [Figure 4]4 is a flowchart illustrating processing executed by the content voice / sound processing device of the first embodiment. [Figure 5] FIG. 10 is a functional block diagram of a content voice / sound reproducing device according to a second embodiment. [Figure 6A] FIG. 10 is a first diagram for explaining the control of ease of hearing of the content voice or content sound in the content voice / sound reproducing device of the second embodiment. [Figure 6B] FIG. 2 is a second diagram for explaining the control of the ease of hearing of the content voice or content sound of the content voice / sound reproducing device of the second embodiment. [Figure 7] 10 is a flowchart illustrating processing executed by a content voice / sound processing device according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0016] Hereinafter, a content voice / sound processing device and a content voice / sound playback device according to embodiments of the present disclosure will be described with reference to the drawings. Also, a content voice / sound processing method, a content voice / sound processing program, and a recording medium on which the program is recorded according to embodiments of the present disclosure will be described with reference to the drawings. In the drawings, identical or equivalent elements are designated by the same reference numerals, and redundant descriptions will not be repeated.

[0017] (Embodiment 1) The content voice / sound reproducing device 100 of the first embodiment will be described with reference to Figures 1 to 4. The content voice / sound reproducing device 100 includes a television receiver, as an example. However, the content voice / sound reproducing device 100 may be a device other than a television receiver, such as a smartphone or a personal computer.

[0018] 1 is a functional block diagram of a content voice / sound reproducing device 100 according to this embodiment. The content voice / sound reproducing device 100 further includes a display panel for displaying images and a control unit for controlling the display panel, but these are not shown in the figure.

[0019] 1, the content voice / sound reproducing device 100 includes a content input unit 1, a genre determination unit 2, a human voice recognition unit 3, and a content voice / sound processing unit 4. The content voice / sound reproducing device 100 also includes a content voice / sound output unit 5, an echo cancellation unit 6, and a human voice input unit 7.

[0020] The genre determination unit 2, the human voice recognition unit 3, the content voice / sound processing unit 4, and the echo cancellation unit 6 constitute the content voice / sound processing device 10. The content voice / sound processing device 10 includes a memory as a computer-readable recording medium. The content voice / sound processing device 10 also includes a processor as a computer that executes predetermined control based on a content voice / sound processing program stored in the memory. The processor includes one or more semiconductor chips. The predetermined control causes the computer to operate as the genre determination unit 2, the human voice recognition unit 3, the content voice / sound processing unit 4, and the echo cancellation unit 6, which will be described below. In other words, the genre determination unit 2, the human voice recognition unit 3, the content voice / sound processing unit 4, and the echo cancellation unit 6 are realized by processing executed by software. The content voice / sound processing device 10 is also referred to as a controller or control unit. However, at least one of the genre determination unit 2, the human voice recognition unit 3, the content voice / sound processing unit 4, and the echo cancellation unit 6 may be realized by hardware including dedicated electronic circuits.

[0021] The content voice / sound processing device 10 processes a content signal received via the content input unit 1 and a human voice signal acquired from the human voice input unit 7. The content voice / sound processing device 10 uses the processed content signal to control the content voice / sound output unit 5, causing the content voice / sound output unit 5 to output the content voice and content sound. The content voice / sound processing device 10 recognizes human voice in the human voice recognition unit 3 using the human voice signal processed by the echo cancellation unit 6. Note that in this specification, content voice is defined as the voice uttered by a person appearing in the content while the content is being played. Also, in this specification, content sound is defined as sound other than the voice uttered by a person appearing in the content, such as background noise of a person appearing in the content, that occurs while the content is being played.

[0022] The specific configuration of the content voice / sound reproducing device 100 will be described below.

[0023] The content input unit 1 is a tuner that receives a content signal capable of identifying the content to be played from outside the content voice / sound playback device 100. However, the content input unit 1 may also be an antenna that receives radio waves containing the content signal or a connector connected to electrical wiring that transmits the content signal. The content input unit 1 extracts the content signal from the received radio waves and transmits the extracted content signal to the genre determination unit 2 and the content voice / sound processing unit 4, respectively.

[0024] In this specification, a content signal includes a content voice signal and a content sound signal. A content voice signal is a signal that can identify a content voice. A content sound signal is a signal that can identify a content sound.

[0025] The genre determination unit 2 receives a content signal capable of identifying content via the content input unit 1. The genre determination unit 2 determines the genre of the content being played based on the content signal input from the content input unit 1. The human voice recognition unit 3 recognizes human voice based on a human voice signal input by the human voice input unit 7.

[0026] The content voice / sound processing unit 4 controls at least one of the content voice and content sound output by the content voice / sound output unit 5 when content is played back. While the human voice recognition unit 3 is recognizing human voice, the content voice / sound processing unit 4 changes the volume of at least one of the content voice and content sound based on the genre determined by the genre determination unit 2. When making this change, the content voice / sound processing unit 4 lowers the overall volume of the content voice and content sound, and prevents a decrease in the ease with which humans can hear the content voice or content sound.

[0027] According to the above configuration, by lowering the overall volume of the content voice and content sound, it is possible to suppress misrecognition of human voice by the human voice recognition unit 3. Furthermore, by suppressing a decrease in the ease with which the content voice or content sound can be heard by humans, it is possible to reduce the risk that the viewer will be unable to understand the content voice due to the process of suppressing misrecognition of human voice. Therefore, it is possible to suppress a decrease in satisfaction with viewing the content.

[0028] In this embodiment, the frequency of the content voice and the frequency of the content sound can be distinguished. Furthermore, while the human voice recognition unit is recognizing human voice, the genre of the content determined by the genre determination unit 2 may be a genre in which the importance of the content voice is higher than the importance of the content sound. In this case, the content voice / sound processing unit 4 reduces the overall volume of the content voice and the content sound, and increases the frequency level of the content voice relative to the frequency level of the content sound.

[0029] In this embodiment, the following three methods (1) to (3) can be considered for increasing the frequency level of the content voice relative to the frequency level of the content sound: (1) Lowering the frequency level of the content sound while maintaining the frequency level of the content voice; (2) Lowering the overall volume of the content voice and the content sound, but increasing the frequency level of the content voice and lowering the frequency level of the content sound; (3) Lowering both the frequency level of the content voice and the frequency level of the content sound, but lowering the frequency level of the content sound to a greater extent than the degree of reduction in the frequency level of the content voice.

[0030] On the other hand, while the human voice recognition unit 3 is recognizing human voice, the genre of the content determined by the genre determination unit 2 may be a genre in which the importance of the content sound is higher than the importance of the content voice. In this case, the content voice / sound processing unit 4 reduces the overall volume of the content voice and content sound, and increases the frequency level of the content sound relative to the frequency level of the content voice.

[0031] In this embodiment, the following three methods (4) to (6) are possible for increasing the frequency level of the content sound relative to the frequency level of the content voice: (4) Lowering the frequency level of the content voice while maintaining the frequency level of the content sound; (5) Lowering the overall volume of the content voice and the content sound, but increasing the frequency level of the content sound and lowering the frequency level of the content voice; (6) Lowering both the frequency level of the content voice and the frequency level of the content sound, but lowering the frequency level of the content voice to a greater extent than the degree of reduction in the frequency level of the content sound.

[0032] There may be a case where the period during which the human voice recognition unit 3 recognizes human voice ends. In this case, the content voice / sound processing unit 4 returns the volume of at least one of the content voice and the content sound to the state before the volume of at least one of the content voice and the content sound was changed.

[0033] To achieve the above, the content voice / sound processing unit 4 includes a decoder 41 and a channel-based audio processing unit 42. The decoder 41 decodes a content signal received via the content input unit, and transmits a content voice signal and a content sound signal extracted from the decoded content signal to the channel-based audio processing unit 42.

[0034] The channel-based acoustic processing unit 42 includes a volume control unit 42A and a downmix processing unit 42B. The volume control unit 42A generates a new content voice signal and a new content sound signal based on the content voice signal and the content sound signal received from the decoder 41. The volume control unit 42A transmits the generated new content voice signal and the generated new content sound signal to the downmix processing unit 42B.

[0035] As described above, while the human voice recognition unit 3 is recognizing human voice, the genre determined by the genre determination unit 2 may be one in which the importance of the content voice is higher than the importance of the content sound. In this case, the volume control unit 42A increases the frequency level of the content voice relative to the frequency level of the content sound.

[0036] Furthermore, as described above, there may be cases where the genre determined by the genre determination unit 2 during the period when the human voice recognition unit 3 is recognizing human voice is one in which the importance of the content sound is higher than the importance of the content voice. In this case, the volume control unit 42A increases the frequency level of the content sound relative to the frequency level of the content voice.

[0037] The signal capable of identifying the content voice and the signal capable of identifying the content sound after the above-mentioned frequency level change are called a new content voice signal and a new content sound signal, respectively.

[0038] The downmix processing unit 42B synthesizes the new content voice signal and the new content sound signal, and transmits the synthesized new content voice signal and the new content sound signal to the content voice / sound output unit 5. The content voice / sound output unit 5 outputs the synthesized content voice and content sound based on the synthesized new content voice signal and the new content sound signal received from the downmix processing unit 42B.

[0039] The content voice / sound output unit 5 is a speaker that emits content voice and content sound based on the new content voice signal and new content sound signal. The content voice / sound output unit 5 synthesizes and outputs the content voice and content sound.

[0040] The echo cancellation unit 6 receives a content voice signal and a content sound signal from the content voice / sound processing unit 4. The echo cancellation unit 6 performs echo cancellation using the content voice signal and the content sound signal. In other words, the echo cancellation unit 6 reduces at least one of a voice noise signal corresponding to the content voice mixed in the human voice signal and a sound noise signal corresponding to the content sound mixed in the human voice signal.

[0041] Therefore, the human voice recognition unit 3 recognizes human voice by processing the human voice signal in which at least one of the voice noise and the sound noise has been reduced by the echo cancellation unit 6. This reduces the risk of errors occurring in the recognition of human voice by the human voice recognition unit 3.

[0042] The human voice input unit 7 is a microphone that outputs content voice and content sound. The human voice input unit 7 converts the physical vibrations of sound propagating through the air into an electrical human voice signal. The converted human voice signal is sent to the human voice recognition unit 3 via the echo cancellation unit 6.

[0043] Fig. 2A is a first graph for explaining the control of the audibility of the content voice or content sound by the content voice / sound reproduction device 100 of this embodiment. Fig. 2B is a second graph for explaining the control of the audibility of the content voice or content sound by the content voice / sound reproduction device 100 of this embodiment.

[0044] 2A, 2B and 3, the function of controlling the ease of hearing of the content voice of the content voice / sound reproducing device 100 of this embodiment will be specifically described.

[0045] For example, when suppressing a decrease in the audibility of the content voice, the content voice / sound processing unit 4 analyzes the characteristics of the content voice signal, applies a pattern recognition system to detect the presence of dialogue from moment to moment, and performs the following processes (A) to (C).

[0046] (A) The content voice / sound processing unit 4 increases the frequency of the content voice (language voice within the content) so that the content voice can be easily detected.

[0047] (B) When the content voice is detected, the content voice / sound processing unit 4 changes the spectrum of the content voice to emphasize the content voice so that the viewer can more easily hear the content voice.

[0048] (C) The content voice / sound processor 4 reduces the frequency level of content sounds (sounds that interfere with the recognizability of the dialogue) other than the content voice contained in the content.

[0049] As a result, the graph showing the relationship between the frequency of the content voice and the level of the content voice and the graph showing the relationship between the frequency of the content sound and the level of the content sound change from the state shown in Fig. 2A to the state shown in Fig. 2B. Comparing Fig. 2A with Fig. 3A, it can be seen that the content voice and the content sound are distinguishable in a predetermined frequency band, and the frequency level of the content voice increases relative to the frequency level of the content sound. By performing this control, the content voice / sound processing unit 4 reduces the overall volume of the content voice and the content sound, and prevents the content voice from becoming more audible to humans.

[0050] An example of the control of the audibility of the content sound can be achieved by replacing the content voice with the content sound in the above description of the control of the audibility of the content voice. The above will be explained in more detail with reference to FIG. 3.

[0051] Fig. 3 is a third graph for explaining the control of the audibility of the content voice or content sound by the content voice / sound playback device of the content voice according to Embodiment 1. Note that in each of graphs (a), (b), and (c) of Fig. 3, a dashed line is drawn in the 2 kHz frequency band, which is an example of the frequency band of human voice.

[0052] As shown in the graph of FIG. 3(a), the content signal capable of identifying the content includes both a content voice signal capable of identifying the content voice and a content sound signal capable of identifying the content sound. It can be seen from the graph of FIG. 3(a) that the frequency of the content voice and the frequency of the content sound can be distinguished. In the state shown in the graph of FIG. 3(a), neither the content voice nor the content sound increases or decreases, i.e., the increase or decrease of each of the content voice and the content sound is 0 dB. Therefore, there is no overall increase or decrease of the content voice and the content sound, i.e., the overall increase or decrease of the content voice and the content sound is 0 dB.

[0053] When the content genre is news, drama, or the like, the content voice / sound processing unit 4 changes the content voice signal and the content sound signal as shown in the graph of FIG. 3(b). The genres of news and drama are examples of genres in which the importance of the content voice is higher than the importance of the content sound. In this case, as can be seen from the graph of FIG. 3(b), the content voice / sound processing unit 4 increases the frequency level value of the content voice by 2 dB in the 2 kHz frequency band, while decreasing the frequency level value of the content sound by 3 dB. In this case, the content voice / sound processing unit 4 decreases the overall frequency level of the content voice and the content sound by 1 dB in the 2 kHz frequency band.

[0054] On the other hand, when the content genre is sports, music, landscape video, or the like, the content voice / sound processing unit 4 changes the content voice signal and the content sound signal as shown in the graph of FIG. 3(c). Genres such as sports, music, and landscape video are examples of genres in which the importance of the content sound is higher than the importance of the content voice. In this case, as can be seen from the graph of FIG. 3(c), the content voice / sound processing unit 4 increases the frequency level value of the content sound by 2 dB in the 2 kHz frequency band, while decreasing the frequency level value of the content voice by 3 dB. In this case, the content voice / sound processing unit 4 also decreases the sum of the frequency level of the content voice and the frequency level of the content sound by 1 dB in the 2 kHz frequency band.

[0055] 3, the content voice / sound processing unit 4 changes both the content voice and the content sound, thereby lowering the overall volume of the content voice and the content sound and increasing the frequency level of the content voice relative to the frequency level of the content sound.

[0056] FIG. 4 is a flowchart for explaining the process executed by the content voice / sound processing device 10 of this embodiment.

[0057] In step S1, the content voice / sound processing device 10 waits for the detection of a HOTWARD, which instructs the human voice recognition unit 3 to start human voice recognition. Next, in step S2, the content voice / sound processing device 10 determines whether or not a HOTWARD has been detected. If it is not determined in step S2 that a HOTWARD has been detected, the content voice / sound processing device 10 repeats the processes of steps S1 and S2.

[0058] On the other hand, if it is determined in step S2 that HOTWARD has been detected, the human voice recognition unit 3 starts recognizing the human voice input from the human voice input unit 7. In this case, in step S3, the genre determination unit 2 of the content voice / sound processing device 10 determines the genre of the content currently being played based on various information included in the received content signal.

[0059] In step S3, the genre determination unit 2 may determine that the genre of the content is a genre in which the importance of the content voice is higher than the importance of the content sound, such as news or drama. In this case, in step S4, the content voice / sound processing unit 4 enables the audibility control of the content voice.

[0060] In controlling the audibility of content voice, the content voice / sound processing unit 4 increases the gain characteristics of the sound in the mid-frequency band (content voice) where the voice components of a typical person are abundant. In this case, in controlling the audibility of content voice, the content voice / sound processing unit 4 decreases the gain characteristics of the low and high frequency bands (content sound) where the voice components of a typical person are few. However, the content voice / sound processing unit 4 decreases the overall volume of the content voice and content sound.

[0061] In other words, the content voice / sound processing unit 4 reduces the overall volume of the content voice and content sound, and prevents the content voice from becoming less audible to humans. In this case, the content voice / sound processing unit 4 increases the frequency level of the content voice relative to the frequency level of the content sound.

[0062] In this way, when the content genre is a news program or a drama program, the content voice / sound processing unit 4 enables content voice audibility control so that the most important content voice of the anchor or actor is not missed. This makes it possible to make the content voice easier to hear, for example, by lowering the frequency level of the content sound. At this time, the content voice / sound output unit 5 reduces the overall volume of the content voice and content sound. This reduces the error rate of human voice recognition by the human voice recognition unit 3.

[0063] On the other hand, in step S3, the genre determination unit 2 may determine that the genre of the content being played is a genre in which the content sound is more important than the content voice, such as music, movies, or sports. In this case, in step S5, the content voice / sound processing unit 4 disables the audibility control of the content voice.

[0064] However, in step S5, the content voice / sound processing unit 4 may execute control to suppress a decrease in the ease with which the content sound can be heard by humans. To this end, the content voice / sound processing unit 4 may increase the frequency level of the content voice relative to the frequency level of the content sound. In this case, the content voice / sound processing unit 4 also reduces the overall volume of the content voice and content sound emitted by the content voice / sound output unit 5. This reduces the error rate in recognizing human voices by the human voice recognition unit 3. It also reduces the discomfort felt by users who are enjoying content in which content sound is important, such as music, movies, or sports, due to loud content voices. This suppresses a decrease in the user's satisfaction with viewing the content.

[0065] In step S6, the content voice / sound processing device 10 enters a state of waiting for a question or instruction. In step S7, if a human voice signal of a user's question or instruction is received from the microphone serving as the human voice input unit 7, the content voice / sound processing device 10 answers or responds to the question or instruction in step S8. If a human voice signal of a question or instruction is not received in step S7, the content voice / sound processing device 10 executes step S9.

[0066] In step S9, the content voice / sound processing device 10 determines whether the question or instruction has ended. If it is determined in step S9 that the question or instruction has ended, the content voice / sound processing device 10 returns the state of the "sound audibility control" to its original state in step S10. If it is determined in step S9 that the question or instruction has not ended, the content voice / sound processing device 10 repeats the processes of steps S1 to S9.

[0067] (Embodiment 2) The content voice / sound reproducing device 100 of the second embodiment will be described with reference to Figures 5 to 7. Note that the following description will not repeat the same points as those of the content voice / sound reproducing device 100 of the first embodiment. The content voice / sound reproducing device 100 of the present embodiment differs from the content voice / sound reproducing device 100 of the first embodiment in the following points.

[0068] FIG. 5 is a functional block diagram of the content voice / sound reproducing device 100 according to this embodiment.

[0069] 5, the content voice / sound processing unit 4 includes an object-based audio processing unit 43 in addition to a decoder 41 and a channel-based audio processing unit 42. The object-based audio processing unit 43 includes a volume control unit 43A and a rendering unit 43B. The decoder 41 decodes a content signal capable of identifying content and acquires object metadata from the decoded content signal. The decoder 41 also transmits a content voice signal and a content sound signal included in the decoded content signal to the object-based audio processing unit 43.

[0070] In this embodiment, there may be cases where the genre determined by the genre determination unit 2 during the period when the human voice recognition unit 3 is recognizing human voice is a genre in which the importance of the content voice is higher than the importance of the content sound. In this case, in this embodiment, the volume control unit 43A increases the volume of the content voice group relative to the volume of the content sound group.

[0071] Also in this embodiment, there may be cases where the genre determined by the genre determination unit 2 during the period when the human voice recognition unit 3 is recognizing human voice is a genre in which the importance of the content sound is higher than the importance of the content voice. In this case, in this embodiment, the volume control unit 43A increases the volume of the content sound group relative to the volume of the content voice group.

[0072] The signal capable of identifying the content voice and the signal capable of identifying the content sound after the volume change are referred to as a new content voice signal and a new content sound signal, respectively. The volume control unit 43A generates a new content voice signal and a new content sound signal based on the content voice signal and the content sound signal, and transmits them to the rendering unit 43B.

[0073] The rendering unit 43B renders a new content voice signal and a new content sound signal, and transmits the rendered content voice signal and the content sound signal to the content voice / sound output unit 5. The content voice / sound output unit 5 outputs the content voice and the content sound based on the rendered content voice signal and the rendered content sound signal received from the rendering unit 43B.

[0074] While the human voice recognition unit 3 is recognizing human voice, the genre determined by the genre determination unit 2 may be determined to be a genre in which the importance of the content voice is higher than the importance of the content sound. In this case, the content voice / sound processing unit 4 increases the volume of the content voice group relative to the volume of the content sound group.

[0075] In this embodiment, the following three methods (1) to (3) are considered as methods for increasing the volume of a content voice group relative to the volume of a content sound group: (1) Maintain the volume of the content voice group while decreasing the volume of the content sound group; (2) Decrease the overall volume of the content voice and content sound, but increase the volume of the content voice group and decrease the volume of the content sound group; (3) Decrease both the volume of the content voice group and the volume of the content sound, but decrease the volume of the content sound to a greater extent than the decrease in the volume of the content voice.

[0076] On the other hand, while the human voice recognition unit 3 is recognizing human voice, the genre determined by the genre determination unit 2 may be determined to be a genre in which the importance of the content sound is higher than the importance of the content voice. In this case, the content voice / sound processing unit 4 increases the volume of the content sound group relative to the volume of the content voice group.

[0077] In this embodiment, the following three methods (4) to (6) are considered as methods for increasing the volume of the content sound group relative to the volume of the content voice group: (4) Maintain the volume of the content sound group while decreasing the volume of the content voice group; (5) Decrease the overall volume of the content voice and content sound, but increase the volume of the content sound group and decrease the volume of the content voice group; (6) Decrease both the volume of the content voice group and the volume of the content sound, but decrease the volume of the content voice group to a greater extent than the decrease in the volume of the content sound group.

[0078] When the period during which the human voice recognition unit 3 recognizes human voice ends, the content voice / sound processing unit 4 returns the volume of at least one of the content voice and the content sound to the state before the volume of at least one of the content voice and the content sound was changed.

[0079] FIG. 6A is a first diagram for explaining the control of the audibility of the content voice or content sound by the content voice / sound reproducing device 100 according to the present embodiment.

[0080] The audio metadata included in the audio scene information shown in Fig. 6A is described in the specifications of the international standard. As shown in Fig. 6A, the audio metadata (Metadata audio elements) included in the content describes groups or switch groups. A switch group (Switch groups of elements) is a special group that can turn on a selected number of grouped elements.

[0081] The content creator creates a group preset that has a combination of multiple groups or multiple switch groups and their respective setting values, and includes it in the content.

[0082] The user can select one of a plurality of group presets created by the content creator. The user may be allowed to select one of the groups or switch groups.

[0083] FIG. 6B is a second diagram for explaining the control of the audibility of the content voice or content sound by the content voice / sound reproducing device 100 of this embodiment.

[0084] 6B, information such as dialogue groups or background sound groups is described in one data item (Mae_contentKind) in certain audio information (Mae_AudioSceneInfo) in the metadata. The content voice / sound processing unit 4 controls at least one of the content voice and the content sound based on information such as dialogue groups or background sound groups included in the content.

[0085] Therefore, in this embodiment, a group of content voices and a group of content sounds can be distinguished from each other. Specifically, a content voice signal that can identify a content voice and a content sound signal that can identify a content sound belong to different groups in the metadata.

[0086] FIG. 7 is a flowchart for explaining the process executed by the content voice / sound processing device 10 of this embodiment.

[0087] The genre determination unit 2 extracts metadata of groups, switch groups, and presets from the audio scene information included in the object sound content. Based on the extracted metadata, the genre determination unit 2 determines whether the genre of the content being played is news or drama in step S3. That is, in step S3, the genre determination unit 2 determines whether the genre is one in which the importance of the content voice is higher than the importance of the content sound.

[0088] In step S3, the genre determination unit 2 may determine that the genre of the content being played is, for example, a genre in which human voices (content voices) such as news or dramas are more important than sounds other than human voices (content sounds). In this case, in step S4A, the content voice / sound processing unit 4 executes object sound-based audio processing.

[0089] In this case, in the object sound audio processing, the content voice / sound processing unit 4 mutes the sound of the background sound group (content sound) output from the content voice / sound output unit 5. Note that the content voice / sound processing unit 4 may set the volume of the background sound group (content sound) below a predetermined value. In other words, the content voice / sound processing unit 4 increases the volume of the content voice group relative to the volume of the content sound group so as to suppress a decrease in the ease with which people can hear the content voice. However, even in this case, the content voice / sound processing unit 4 lowers the overall volume of the content voice and content sound. This reduces the error rate in recognizing human voice by the human voice recognition unit 3.

[0090] In step S10A, the content voice / sound processing unit 4 returns the volume of the content sound group and the volume of the content sound group to the volume before they were changed.

[0091] On the other hand, in step S3, the genre determination unit 2 may determine that the genre of the content is a genre in which background sounds are more important than dialogue sounds, for example, a genre in which sounds other than human voices are important, such as music, movies, or sports. In this case, in step S5A, the content voice / sound processing unit 4 mutes the dialogue sound group output from the content voice / sound output unit 5, or reduces the volume of the dialogue sound group to a predetermined value or less.

[0092] As a result, the content voice / sound processing unit 4 prevents a decrease in the ease with which people can hear the content sound. In this case, too, the volume of the content sound is increased relative to the volume of the content voice while lowering the overall volume of the content voice and the content sound. This reduces the discomfort felt by users due to the content voice and prevents the human voice recognition unit 3 from misrecognizing human voice.

[0093] In processes other than those described above, the content voice / sound processing device 10 of this embodiment executes processes similar to those executed by the content voice / sound processing device 10 of the first embodiment. [Explanation of symbols]

[0094] 1 Content input section 2. Genre Determination Section 3 Human Voice Recognition Unit 4 Content Voice / Sound Processing Unit 5 Content voice / sound output section 6 Echo cancellation section 7 Human voice input section 10 Content Voice / Sound Processing Device 100 Content Voice / Sound Playback Device

Claims

1. a human voice recognition unit that recognizes a human voice based on a human voice signal input by a human voice input unit; a genre determination unit that determines the genre of content being played back based on a content signal input from a content input unit; a content voice / sound processing unit that controls at least one of a content voice output by a content voice / sound output unit in response to playback of the content and a content sound other than the content voice; the content voice / sound processing unit changes the volume of at least one of the content voice and the content sound so as to suppress a decrease in ease of human hearing of either the content voice or the content sound; Content voice / sound processing device.

2. the content voice / sound processing unit, during a period in which the human voice recognition unit recognizes the human voice, reduces the overall volume of the content voice and the content sound based on the genre determined by the genre determination unit, and changes the volume of at least one of the content voice and the content sound so as to suppress a decrease in ease of human audibility of either the content voice or the content sound; The content voice / sound processing device of claim 1 .

3. The frequencies of the content voice and the frequencies of the content sound can be distinguished; the content voice / sound processing unit, during a period in which the human voice recognition unit is recognizing the human voice, if the genre determined by the genre determination unit is a genre in which the importance of the content voice is higher than the importance of the content sound, reduces the overall volume of the content voice and the content sound, and increases the frequency level of the content voice relatively to the frequency level of the content sound; The content voice / sound processing device of claim 2 .

4. The frequencies of the content voice and the frequencies of the content sound can be distinguished; the content voice / sound processing unit, during a period in which the human voice recognition unit is recognizing the human voice, if the genre determined by the genre determination unit is a genre in which the importance of the content sound is higher than the importance of the content voice, reduces the overall volume of the content voice and the content sound, and increases the frequency level of the content sound relatively to the frequency level of the content voice; The content voice / sound processing device of claim 2 .

5. The group of content voices and the group of content sounds can be distinguished; the content voice / sound processing unit, during a period in which the human voice recognition unit is recognizing the human voice, if the genre determined by the genre determination unit is a genre in which the importance of the content voice is higher than the importance of the content sound, reduces the overall volume of the content voice and the content sound, and increases the volume of the group of content voices relative to the volume of the group of content sounds; The content voice / sound processing device of claim 2 .

6. The group of content voices and the group of content sounds can be distinguished; the content voice / sound processing unit, during a period in which the human voice recognition unit is recognizing the human voice, if the genre determined by the genre determination unit is a genre in which the importance of the content sound is higher than the importance of the content voice, reduces the overall volume of the content voice and the content sound, and increases the volume of the group of content sounds relative to the volume of the group of content voices; The content voice / sound processing device of claim 2 .

7. when a period during which the human voice recognition unit recognizes the human voice ends, the content voice / sound processing unit returns the volume of at least one of the content voice and the content sound to the volume before the volume of at least one of the content voice and the content sound was changed; The content voice / sound processing device of claim 2 .

8. an echo cancellation unit that receives a content voice signal that can identify the content voice and a content sound signal that can identify the content sound from the content voice / sound processing unit, the echo cancellation unit reduces at least one of a voice noise signal corresponding to the content voice mixed in the human voice signal and a sound noise signal corresponding to the content sound mixed in the human voice signal, using at least one of the content voice signal and the content sound signal; the human voice recognition unit recognizes the human voice by processing the human voice signal in which at least one of the voice noise signal and the sound noise signal has been reduced by the echo cancellation unit; The content voice / sound processing device of claim 2 .

9. A content voice / sound processing device according to any one of claims 1 to 8; a speaker as the content voice / sound output unit; a microphone as the human voice input unit. Content voice / sound playback device.

10. a step of recognizing a human voice based on a human voice signal input by a human voice input unit; determining the genre of the content being played back based on the content signal input from the content input unit; and controlling at least one of a content voice and a content sound other than the content voice output by a content voice / sound output unit by playing back the content; In the step of controlling at least one of the content voice and the content sound, the volume of at least one of the content voice and the content sound is changed so as to suppress a decrease in ease of human audibility of the content voice or the content sound. Content voice / sound processing methods.

11. Computer, a human voice recognition unit that recognizes a human voice based on a human voice signal input by a human voice input unit; a genre determination unit that determines the genre of the content being played back based on the content signal input from the content input unit; and a content voice / sound processing program for operating as a content voice / sound processing unit that controls at least one of a content voice output by a content voice / sound output unit in response to playback of the content and a content sound other than the content voice, the content voice / sound processing unit changes the volume of at least one of the content voice and the content sound so as to suppress a decrease in ease of human audibility of the content voice or the content sound; Content voice / sound processing program.

12. A computer-readable recording medium storing the content voice / sound processing program according to claim 10.

Citation Information

Patent Citations

  • Television device and remote controller

    JP2012181374A