Voice control system
The voice control system addresses the challenge of synthesizing multiple audio signals with different levels by adjusting input signal levels and preventing distortion, ensuring optimal audio balance and quality.
Patent Information
- Application Number
- JP2024080644
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-28
AI Technical Summary
Existing audio control systems fail to effectively synthesize and output multiple audio signals with different input levels without causing auditory discomfort or distortion, and require means for storing volume settings.
A voice control system that synthesizes multiple audio signals by adjusting input signal levels, using peak hold processing, analog-to-digital conversion, voice recognition, and output determination to ensure desired audio balance and prevent distortion.
The system adjusts input signal levels to achieve desired audio balance and prevents distortion during output, allowing for seamless synthesis and optimal audio quality across varying input levels.
Smart Images

Figure 2025174350000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a voice control system. [Background technology]
[0002] For devices that handle a plurality of audio signals with different recording levels, there are known devices that have a function of automatically adjusting the output level as a countermeasure against auditory discomfort caused by the different recording levels.
[0003] Conventional methods for automatically adjusting audio output involve automatically adjusting the output by setting a threshold value for the input audio signal level corresponding to each setting position of an audio adjustment control, and compressing or amplifying the signal to eliminate volume differences that arise when switching between multiple audio sources (see Patent Document 1).
[0004] Also, there has been a technology that considers storing the volume of one type of input signal and automatically setting the volume to the same as the stored volume when switching to another input signal (see Patent Document 2). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 3279267 [Patent Document 2] Patent No. 4347292 Summary of the Invention [Problem to be solved by the invention]
[0006] However, in the above Patent Document 1, although it is possible to adjust the audio output of multiple audio signals with different recording levels without perceiving a difference in sound pressure even at a certain volume position, it does not take into consideration the synthesis and output of multiple audio signals. Furthermore, Patent Document 2 has a problem in that it requires a means for storing the volume.
[0007] The present disclosure discloses technology for solving the above-mentioned problems, and aims to provide an audio control system that synthesizes multiple audio signals with different input levels and outputs them from a single speaker, adjusting the input signal level of the audio signal to the desired audio balance before outputting it. [Means for solving the problem]
[0008] The voice control system of the present disclosure synthesizes and outputs a plurality of voice signals, a voice input unit having a peak hold processing unit that executes processing to hold the maximum value of input voice, and an analog-to-digital conversion processing unit that executes analog-to-digital conversion processing on the maximum value; The system is equipped with a voice recognition processing unit that identifies whether the maximum value is greater than a threshold value, and a voice determination unit that has an output determination processing unit that determines voice data that cannot be recognized as voice among the multiple voice signals as silence. [Effects of the Invention]
[0009] According to the audio control system of the present disclosure, the input signal level of the audio signal can be adjusted to a desired audio balance before being output. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram showing an overall image of a voice control system according to a first embodiment. [Figure 2] 3 is a flowchart showing the overall operation of the voice control system according to the first embodiment. [Figure 3] 1 is a schematic block diagram of a voice control system according to a first embodiment. [Figure 4] 2 is a schematic block diagram showing details of a voice input unit and a voice determination unit in the voice control system according to the first embodiment. FIG. [Figure 5]3 is a schematic block diagram showing an example of processing for normalizing audio data and adjusting the volume in the audio control system according to the first embodiment. FIG. [Figure 6] 3 is a schematic block diagram showing an example of processing for synthesizing each piece of audio data and adjusting the output in the audio control system according to the first embodiment. FIG. [Figure 7] 1 is a table illustrating the principle of a voice control system. DETAILED DESCRIPTION OF THE INVENTION
[0011] Embodiment 1 In this embodiment, in an audio control system that synthesizes a plurality of input audio signals and outputs them to a speaker, the input audio is automatically adjusted to optimize the output audio level. The voice control system according to this embodiment will be described below with reference to the drawings, assuming that it is applied to synthesizing and outputting a plurality of voices.
[0012] FIG. 1 is a diagram showing an overall image of a voice control system according to the first embodiment, and FIG. 2 is a flowchart showing the overall operation of the system. In a device (see Figure 1) that combines input signals from multiple sound sources, such as built-in audio signals, external audio signals input from a microphone or AUX (auxiliary) terminal, and audio signals received via a network, and outputs the combined signals from a single speaker, it is not always necessary for the volumes of each input to be the same. For example, if audio signal A is low in volume and audio signal B is high in volume, combining and outputting these audio signals can result in audio signal A being drowned out by audio signal B and becoming inaudible. Furthermore, if the volumes of audio signals A and B are each high, the combined audio can become distorted. Here, "distorted audio" means that the sound waveform changes from its original state, resulting in poor sound quality or a distinctive sound. FIG. 1 shows a case where audio data A, audio data B, microphone A, microphone B, received data A, and received data B are collected in one automatic audio compressor 1000 and sound is generated from a speaker 1001.
[0013] To solve the above-mentioned problems, for example, if there is input from microphone A in FIG. 2 (ST50), noise is removed (ST51), and voice determination is performed (ST52). If the input is accepted as voice (ST53), normalization processing is first performed (ST54). Next, voice settings are performed (ST55), volume adjustment is performed (ST56), voice is synthesized (ST57), and output is performed (ST58). For microphone B, the same operations as above are performed in ST60 to ST66, ST57, and ST58. For received data A, the same operations as above are performed in ST70 to ST76, ST57, and ST58. For received data B, the same operations as above are performed in ST80 to ST86, ST57, and ST58. For voice data A, voice settings are performed (ST90), volume adjustment is performed (ST91), voice is synthesized (ST57), and output is performed (ST58). For the audio data B, audio settings are made (ST100), the volume is adjusted (ST101), the audio is synthesized (ST57), and output (ST58). The operation is described in detail below.
[0014] FIG. 3 is a schematic block diagram of a voice control system according to embodiment 1. The voice control system is composed of a voice input unit 1 to which multiple voices are input, a voice determination unit 2 that determines the input voices, a normalization processing unit 3 that normalizes the determined voices, a volume adjustment unit 4 that adjusts the volume of the normalized voices, a voice synthesis unit 5 that synthesizes the volume-adjusted voices into a single voice, an output adjustment unit 6 that adjusts the output of the synthesized voices and feeds it back to the volume adjustment unit 4, and a voice output unit 7 that outputs the voices.
[0015] FIG. 4 shows a detailed schematic block diagram of the voice input unit 1 and the voice determination unit 2, and is a diagram illustrating an example of using the voice input unit 1 and the voice determination unit 2 to determine whether or not an input voice signal is to be identified as voice. FIG. 5 is a block diagram showing a case where a signal identified as speech is normalized by the normalization processing unit 3, and the balance of the speech to be synthesized is adjusted when the volume adjustment unit 4 attenuates the signal to a volume preset by the user, or when the speech synthesized by the output adjustment unit 6 shown in FIG. 3 is distorted. FIG. 6 is a block diagram showing an example in which volume-adjusted voices are synthesized, and an adjustment factor for adjusting the balance of the synthesized voices is output when the synthesized voices are distorted.
[0016] Next, the operation of this embodiment will be described with reference to the drawings. An example of the overall voice control system is shown in the schematic block diagram of Fig. 3, in which N voices are input, and each voice is subjected to voice determination, normalization processing, and volume adjustment before being synthesized into one voice. An output adjustment unit 6 determines whether the synthesized voice is distorted, and performs final volume adjustment.
[0017] A case where a decision is made as to whether or not to identify a voice signal as voice for a single voice input 8 out of N voice inputs will be described with reference to the block diagram shown in Fig. 4. Fig. 4 is configured with a voice input unit 1 and a voice decision unit 2 in the schematic block diagram shown in Fig. 3. The audio input 8 is branched into two, one of which is directly subjected to analog-to-digital conversion processing by an analog-to-digital conversion processing unit 9, where it is converted into signed binary audio data. Here, "sign" refers to + or -, and if the most significant bit of the binary number is 0, it represents a positive number, and if it is 1, it represents a negative number (see Figure 7).
[0018] The other undergoes processing to hold the maximum value by a peak hold processing unit 10, and analog-to-digital conversion processing is performed on the maximum value by an analog-to-digital conversion processing unit 11. A voice identification processing unit 12 in the voice determination unit 2 determines whether or not the maximum value is greater than a threshold value x, and an output determination processing unit 13 determines whether or not to output. In other words, data that cannot be recognized as voice among multiple voice signals is determined to be silent. The voice determination unit 2 peak-holds the maximum value of the voice signal and determines whether or not it is voice based on that value.
[0019] The threshold value x is the maximum value x that can be expressed in X bits, where X is the maximum number of bits that can be input to the system according to this embodiment, with Z being the number of effective signal bits, and it will silence audio up to the maximum value x. First, it is determined whether it is greater than the threshold value x. As a determination means, if the values of bits X to Z of the binary audio data are all the same, it can be determined that they are smaller than the threshold value x, and if any one of the values of bits X to Z is different, it can be determined that they are greater than the threshold value x.
[0020] The above will be explained in detail with reference to FIG. 7, which is a table showing the principle of the voice control system. For example, when data to be processed by analog-to-digital conversion is expressed as a signed 10-bit (Z-bit) binary number, it can be expressed as numbers ranging from -512 (1000000000) to 511 (0111111111). When determining that the lower two bits and below indicate silence, the thresholds x and x' can be expressed as -4 (1111111100) and 3 (0000000011). In this case, if the upper eight bits other than the lower two bits are the same value and are data of 0 or 1, replacing the lower two bits with 00 will convert the data to the same data as silence, 0 (0000000000). In other words, the output determination processor 13 determines that there is silence.
[0021] Next, the operation of normalizing and adjusting the volume of binary audio data 14 identified as audio by audio determination unit 2 will be described with reference to the block diagram shown in Fig. 5. Fig. 5 is composed of normalization processing unit 3 and volume adjustment unit 4 shown in the schematic block diagram of Fig. 3.
[0022] In the normalization processing unit 3, the amplification factor determination processing unit 15 calculates the difference between the number of effective audio bits Y and the maximum value y expressed in Y bits, and determines the amplification factor. Data reflecting the amplification factor is output by the amplification result output processing unit 16. For example, if the number of effective audio bits is 8 bits (Y bits), in FIG. 7, the minimum value can be expressed as -256 (1100000000) and the maximum value as 255 (0011111111). If the maximum value of the input audio data is 100, a difference of 255 - 100 = 155 occurs with respect to the maximum effective bit value of 255. The audio data is amplified by multiplying all input data by the ratio of this difference (255 / 100) (255 / 100 times). In this way, the normalization processing unit 3 normalizes each piece of audio data to a fixed level. That is, the normalization processing unit 3 has a function of correcting the magnitude of the audio signal to a target value based on the maximum value of the audio signal obtained by the audio determination unit 2.
[0023] In the volume adjustment unit 4, the attenuation rate is set by the attenuation rate setting processing unit 17 in accordance with the volume setting 18 for each audio that has been set in advance by the user. The attenuation result is also output by the attenuation result output processing unit 20, reflecting the adjustment rate for attenuation set by the output adjustment unit 6, which will be described later. Here, the volume setting 18 and the adjustment rate represent values between 0 and 100%, and processing is performed to attenuate the values of all audio data up to the set percentage. The operation of synthesizing multiple pieces of audio data whose volumes have been adjusted by the volume adjustment unit 4 and adjusting the output will be described with reference to the block diagram shown in Fig. 6. Fig. 6 shows a case where the system is configured with the audio synthesis unit 5, output adjustment unit 6, and audio output unit 7 in the schematic block diagram shown in Fig. 3.
[0024] The voice synthesis unit 5 adds all of the voice data from the N attenuation result output processes 20a to 20n output by the volume adjustment unit 4 and synthesizes them into a single voice. In the output adjustment unit 6, the synthesized voice determination processing unit 21 determines whether the added voice is distorted. The output adjustment unit 6 adjusts the level of the synthesized voice data. The synthesized voice determination processing unit 21 determines whether the number of effective voice bits Y is greater than the maximum value y. If it is determined that the number of effective voice bits Y is greater, it is determined that the added voice is distorted. Then, the adjustment rate setting processing unit 23 sets an adjustment rate based on the difference. The adjustment rate setting unit 19 then feeds back the adjustment rate to the volume adjustment unit 4 in the preceding stage, attenuating the entire synthesized voice. If the synthesized voice determination processing unit 21 determines that the number of effective voice bits Y is less than the maximum value y, it is determined that the added voice is not distorted, and the voice is set by the synthesis result output processing unit 22, and the voice is output as voice after being digital-to-analog converted by the voice output unit 7.
[0025] The above will be explained in detail with reference to Figure 7, a table showing the principles of the voice control system. For example, if the number of effective voice bits is 8 bits (Y bits), the minimum value can be expressed as -256 (1100000000) and the maximum value as 255 (0011111111). If the maximum value of the synthesized voice data is, for example, 400, there will be a difference of 255-400=-145 from the maximum effective bit value of 255. By multiplying all input data by the ratio of this difference (255 / 400) (255 / 400 times), the voice data can be attenuated and converted into voice data whose upper limit is the maximum value of the effective bits. That is, the output adjustment unit 6 compares the maximum value of the synthesized sound with the output upper limit value, and feeds back this to the volume adjustment unit 4, thereby automatically adjusting the balance of the output sound.
[0026] As described above, according to this embodiment, when multiple voices are synthesized and output, voices that cannot be determined as voices are muted, only the necessary voices are synthesized, and the synthesized voice can be output without distortion by adjusting the output. Furthermore, if the user sets the balance of each voice in advance, it is possible to synthesize and emphasize a voice that the user wants to emphasize. That is, the user can set the volume adjustment unit 4 to any output level, such as 100% for voice signal A and 80% for voice signal B, in advance using the volume setting (user setting) 18. By setting the balance of each voice, it is also possible to increase the output level of the voice signal that the user wants to emphasize.
[0027] According to the audio control system of this embodiment, in audio control that synthesizes multiple input audio signals and outputs them from a speaker, it is possible to automatically adjust the optimal audio output level regardless of the input level of the audio signals. Furthermore, it is possible to adjust the volume of the output level of each audio signal as needed, and adjust the balance of the audio to be synthesized. Furthermore, when the input level of a voice signal is low and noise that is not voice is input among a plurality of input voice signals, it is possible to process it as silence.
[0028] Although the present disclosure describes various exemplary embodiments and examples, the various features, aspects, and functions described in one or more embodiments are not limited to application to a particular embodiment, but may be applied to the embodiments alone or in various combinations. Therefore, countless variations not exemplified are conceivable within the scope of the technology disclosed in the specification of the present disclosure, including, for example, cases where at least one component is modified, added, or omitted, and cases where at least one component is extracted and combined with components of another embodiment. [Explanation of symbols]
[0029] 1 audio input unit, 2 audio determination unit, 3 normalization processing unit, 4 volume adjustment unit, 5 voice synthesis unit, 6 output adjustment unit, 7 voice output unit, 8 voice input, 9,11 analog-to-digital conversion processing section, 10 peak hold processing section, 12 voice recognition processing unit, 13 output determination processing unit, 14 voice data, 15 amplification factor determination processing unit, 16 amplification result output processing unit, 17 attenuation factor setting processing unit, 18 volume setting, 19 adjustment rate setting unit, 20a to 20n attenuation result output processing, 21 synthetic speech determination processing unit, 22 synthesis result output processing unit, 23 adjustment rate setting processing unit.
Claims
1. An audio control system that synthesizes and outputs a plurality of audio signals, a voice input unit having a peak hold processing unit that executes processing to hold the maximum value of input voice, and an analog-to-digital conversion processing unit that executes analog-to-digital conversion processing on the maximum value; A voice control system comprising a voice recognition processing unit that identifies whether the maximum value is greater than a threshold value, and a voice determination unit having an output determination processing unit that determines voice data among the multiple voice signals that cannot be recognized as voice as silence.
2. a normalization processing unit that corrects the magnitude of the audio signal to a target value based on the maximum value of the audio signal obtained by the audio determination unit; a volume adjustment unit that includes an attenuation rate setting processing unit that sets an attenuation rate in accordance with a volume setting of each audio previously set by a user and that reflects an adjustment rate setting for attenuation, and adjusts the volume of each of the audio data; a voice synthesis unit that adds up all of the voice data of the plurality of attenuation result output processes outputted by the volume adjustment unit and synthesizes them into one voice; 2. The voice control system according to claim 1, further comprising: a synthetic voice determination processing unit that determines whether the added voice is distorted; and an output adjustment unit having an adjustment rate setting unit that sets the adjustment rate.
Citation Information
Patent Citations
AUDIO OUTPUT ADJUSTMENT METHOD AND DEVICE THEREOF
JP3279267B2
audio equipment
JP4347292B2