Volume control device, audio system, and program

The volume control device improves upon conventional systems by using timing differences and signal correlation to accurately determine conversation states within vehicles, enabling effective and keyword-independent audio volume adjustments.

JP2025082921APending Publication Date: 2025-05-30DENSO TEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023196485
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Conventional volume control devices for vehicles struggle to accurately determine the conversation state of passengers and adjust the audio volume accordingly, relying on pre-registered keywords and natural language processing.

Method used

A volume control device that acquires voice signals from multiple sources within the vehicle and determines the conversation state based on the timing differences and correlation between these signals, allowing for automatic adjustment of the audio volume without relying on specific keywords.

Benefits of technology

This approach enables more accurate determination of the conversation state and appropriate volume control, ensuring that conversations between passengers are not disrupted by audio from devices, while also improving the overall user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025082921000001_ABST
    Figure 2025082921000001_ABST
Patent Text Reader

Abstract

To determine the conversation state of occupants with higher accuracy and to realize appropriate volume control in accordance with the conversation state.SOLUTION: A volume control device acquires a voice signal from each sound source in the vehicle interior, determines whether there is a conversation in the vehicle interior on the basis of the relationship state between the acquired voice signals, and controls the playback volume of the audio signal on the basis of the determination result of whether there is a conversation.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments relate to a volume control device, an audio system, and a program.

Background Art

[0002] Conventionally, a volume control device that automatically controls the volume of an audio device or the like is known so as not to interfere with conversations between passengers in a vehicle interior (see, for example, Patent Document 1).

[0003] For example, the technology disclosed in Patent Document 1 performs speech recognition based on natural language processing on the collected vehicle interior speech. As a result, when a pre-registered specific keyword is included in the conversation content, control is performed to automatically lower the volume of the audio device. The specific keyword is, for example, a group of words indicating that a conversation such as "What did you say?" or "I couldn't hear" has not been established.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the above-described conventional technology, there is still room for further improvement in more accurately determining the conversation state of the passengers and realizing appropriate volume control according to the conversation state.

[0006] For example, when using the above-described conventional technology, it is impossible to determine the presence or absence of a conversation unless the passenger utters a registered keyword. Therefore, when using the above-described conventional technology, it is impossible to realize appropriate volume control according to the conversation state.

[0007] One aspect of the embodiment is made in view of the above, and an object is to provide a volume control device, an audio system, and a program that can determine the conversation state of an occupant with higher accuracy and realize appropriate volume control according to the conversation state.

Means for Solving the Problems

[0008] The volume control device according to one aspect of the embodiment acquires voice signals from each sound source in the vehicle interior, determines the presence or absence of conversation in the vehicle interior based on the relationship state between the acquired voice signals, and controls the reproduction volume of the audio signal based on the determination result of the presence or absence of conversation.

Effects of the Invention

[0009] According to one aspect of the embodiment, since the conversation state is determined based on the relationship state between voice signals having a high correlation with the conversation state, for example, the timing difference between voice signals, the conversation state of the occupant can be determined with higher accuracy, and appropriate volume control according to the conversation state can be realized.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Embodiments for Carrying Out the Invention

[0011] Hereinafter, embodiments of the volume control device, the audio system, and the program disclosed in the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the embodiments shown below.

[0012] First, a configuration example of the audio system 1 according to the embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing a configuration example of the audio system 1 according to the embodiment.

[0013] The audio system 1 is an in-vehicle audio system. As shown in FIG. 1, the audio system 1 includes a sensor unit 3, an operation unit 5, a speaker 7, a display 9, and an in-vehicle device 10. The in-vehicle device 10 includes a volume control device 11 and an audio device 12. The volume control device 11 and the audio device 12 are connected to be communicable with each other.

[0014] In addition, the audio system 1 can optionally connect an external audio device such as a smartphone 200 and execute the audio playback application (abbreviation of "application software") of the smartphone 200 via the audio device 12. That is, the audio system 1 can output the playback sound by the audio playback application of the smartphone 200 from the speaker 7.

[0015] The sensor unit 3 includes a plurality of microphones 31-1, 31-2, …, 31-m (m is a natural number of 2 or more). In the following, these microphones 31-1, 31-2, …, 31-m may be collectively referred to as "microphone 31".

[0016] The microphone 31 is provided corresponding to each seat position in the vehicle interior, for example, so that each occupant can acquire the voice uttered. In this case, the aforementioned m corresponds to the number of seats. Note that an installation method of associating one microphone 31 with a plurality of seats is also applicable. The sensor unit 3 including the microphone 31 is connected to the volume control device 11. Note that the sensor unit 3 may include various sensors other than the microphone 31. The sensor unit 3 may include, for example, a camera 32 (see FIG. 15). An example using the camera 32 will be described later with reference to FIG. 15.

[0017] The operation unit 5, the speaker 7, and the display 9 are connected to the audio device 12. The operation unit 5 is an operation component group of the audio device 12. The operation unit 5 is composed of, for example, a switch, a volume, a touch panel, etc., and outputs an operation signal for various operations.

[0018] The speaker 7 is installed at an appropriate position in the vehicle interior and outputs an audio signal reproduced by the audio device 12 as voice. The display 9 is composed of a display device such as a liquid crystal display panel and performs information display related to the audio playback of the audio device 12. The display 9 displays, for example, the name of the music being played, the name of the artist, etc. Note that the operation unit 5 and the display 9 may be integrally provided by a touch panel display or the like.

[0019] The volume control device 11 controls the amplification factor for the audio signal in the audio device 12 and adjusts the playback volume of the audio signal output as sound from the speaker 7. In this specification, unless otherwise specified, the volume does not refer to the instantaneous waveform amplitude in the music (audio signal), but is used in the sense of the amplification factor (attenuation factor) of the amplification circuit that amplifies (attenuates) the audio signal. For example, it corresponds to the set value of the volume adjustment volume in the audio device 12. The volume control device 11 acquires the input signal from each microphone 31 of the sensor unit 3 and the audio signal from the audio device 12. Further, the volume control device 11 executes a conversation determination process for determining whether a conversation is being conducted based on the relationship state between the voices indicated by the acquired signals.

[0020] The "relationship state between the voices" refers to, for example, the conversation situation between the passengers or the similar situation of the input signal from each microphone 31 with respect to the audio signal. In other words, the "relationship state between the voices" is the state regarding the time difference between the voice signals acquired by the plurality of microphones 31. Further, the "relationship state between the voices" is the similar state between the voice signal acquired by the microphone 31 and the audio signal. Thereby, the controller 113 described later can perform a conversation determination based on the feature amount indicating the state regarding the time difference between the voice signals and the feature amount indicating the similar state between the voice signal and the audio signal.

[0021] In addition, the volume control device 11 performs automatic volume control of the audio device 12 according to the determination result in the conversation determination process. For example, when it is determined that a conversation is being conducted, the volume control device 11 performs volume reduction control to reduce the volume of the audio device 12. At this time, the volume control device 11 outputs a volume control command including a volume control value for reducing the volume to the audio device 12, thereby reducing the volume of the audio device 12. Thereby, it is possible to prevent the conversation being conducted between the passengers from being disturbed by the voice output from the audio device 12.

[0022] Note that, in the present embodiment, the volume control device 11 determines whether a conversation is taking place (i.e., the presence or absence of a conversation) using a conversation determination model 112a (see FIGS. 2 and later) that is independent of natural language processing. The conversation determination model 112a is a machine learning model learned using, as explanatory variables, feature amounts indicating the above-described "conversation interaction situation among passengers" or feature amounts indicating the above-described "similarity situation with respect to the audio signals of the input signals from each microphone 31". Further, the conversation determination model 112a is learned such that at least a classification value indicating whether a conversation is taking place is used as the objective variable. The learning method of the conversation determination model 112a will be described later with reference to FIGS. 6 and 13. Further, details of the conversation determination process using the conversation determination model 112a will be described later with reference to FIGS. 2 and later.

[0023] Further, for example, when it is determined that the conversation has ended, the volume control device 11 performs volume restoration control to restore the volume of the audio device 12 to the volume before the conversation (return it to the original state). At this time, the volume control device 11 outputs a volume control command including a volume control value for restoring the decreased volume to the audio device 12, thereby restoring the volume of the audio device 12 to the volume before the conversation.

[0024] The audio device 12 outputs an audio signal to the volume control device 11 and the speaker 7. Further, the audio device 12 changes the level of the audio signal output to the speaker 7 based on the volume control command input from the volume control device 11, and changes the volume of the reproduced audio output from the speaker 7.

[0025] Further, when the smartphone 200 is connected, the audio device 12 outputs the audio signal from the smartphone 200 to the volume control device 11 and the speaker 7. Further, when the smartphone 200 is connected, the audio device 12 controls the volume of the reproduced audio with respect to the audio signal from the smartphone 200 output from the speaker 7 based on the volume control command input from the volume control device 11.

[0026] Next, a configuration example of the volume control device 11 will be described with reference to FIG. 2. FIG. 2 is a block diagram showing a configuration example of the volume control device 11 according to the embodiment. Note that in FIG. 2 and FIG. 8 shown later, only the components necessary for explaining the features of this embodiment are shown, and descriptions of general components are omitted.

[0027] Also, in the description using FIGS. 2 and 8, descriptions of components that have already been explained may be simplified or omitted.

[0028] As shown in FIG. 2, the volume control device 11 includes a communication unit 111, a storage unit 112, and a controller 113. The aforementioned sensor unit 3 is connected to the controller 113.

[0029] The communication unit 111 is realized by a network adapter, a communication bus, or the like. The communication unit 111 is connected to the audio device 12 by wire or wirelessly, and transmits and receives various information to and from the audio device 12.

[0030] The storage unit 112 is realized by a storage device such as a ROM (Read Only Memory), a RAM (Random Access Memory), a flash memory, or a disk device. In the example of FIG. 2, the storage unit 112 stores a conversation determination model 112a, a word determination model 112b, and volume control value information 112c.

[0031] The conversation determination model 112a is a machine learning model used for the aforementioned conversation determination process. The conversation determination model 112a is, for example, a DNN (Deep Neural Network) model. The conversation determination model 112a is loaded into a controller 113 described later and operates as a program to function as a conversation determination AI (Artificial Intelligence). The conversation determination model 112a is pre-trained and generated so that when the input signals from each microphone 31 and the audio signal from the audio device 12 are input, the conversation determination AI can output a classification value indicating whether at least a conversation is taking place. Also, in the present embodiment, the conversation determination model 112a is pre-trained so that it can output each classification value corresponding to the breakdown when no conversation is taking place.

[0032] The word determination model 112b is a natural language processing model used for the word determination process described later. The word determination model 112b is loaded into a controller 113 described later and operates as a program to function as a word determination AI. The word determination model 112b is pre-trained and generated so that when the input signal from each microphone 31 is input, the word determination AI can perform speech recognition and output a classification value indicating whether a specific keyword indicating danger is included.

[0033] The volume control value information 112c is various control values used for volume control. Specifically, it includes a volume decrease value and a volume decrease speed value in volume decrease control during a conversation or the like. In the present embodiment, the volume control device 11 outputs a correction instruction for correcting a volume adjustment value based on a user operation or the like by adding the volume decrease value of the volume control value information 112c as a volume decrease control command in volume decrease control. This volume decrease control command includes a volume decrease speed value.

[0034] In addition, the volume reduction value and the volume reduction speed value include the values in the "normal time" and the values in the "emergency time". The "emergency time" corresponds to the case where an event requiring emergency occurs during the operation of the vehicle. The "normal time" corresponds to the cases other than the "emergency time". In the "emergency time", the volume reduction value and the volume reduction speed value are set such that the volume reduction value is larger and the volume reduction speed is faster than in the "normal time".

[0035] Note that as for the volume control value information 112c, a method using the volume control value by the volume reduction control and the volume control value immediately before the volume reduction control is also applicable. In this case, the volume control device 11 stores the volume control value immediately before the volume reduction control, calculates and stores the volume control value for the volume reduction control, and outputs a volume control command including the volume control value. Then, when performing the volume return control, it outputs a volume control command including the volume control value immediately before the volume reduction control that was stored.

[0036] Note that these data are stored in the storage unit 112 in an updatable manner in the form of a lookup table or the like. The volume control value becomes the attenuation value of the audio signal (attenuation degree of the attenuation circuit), the time constant of the attenuation characteristic, and the like. Note that it is also possible to use the amplification factor of the power amplification circuit (preamplifier for speaker output) of the audio signal as the volume control value.

[0037] The controller 113 corresponds to a so-called processor. The controller 113 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphical Processing Unit), or the like. The controller 113 executes the program according to the illustrated embodiment stored in the storage unit 112, using the RAM as a work area. In addition, the controller 113 can be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0038] The controller 113 executes various information processes described below with reference to FIGS. 3 to 14 (excluding FIGS. 6 and 8).

[0039] FIG. 3 is an explanatory diagram of the conversation determination process. As shown in FIG. 3, the controller 113 executes a conversation determination process using the conversation determination model 112a. The conversation determination model 112a inputs the input signals from each microphone 31 (equivalent to microphones 1 , m ~ microphone m in the figure) and the audio signal from the audio device 12. Then, as an output value for the input, the conversation determination model 112a outputs a classification value corresponding to "conversation", indicating that conversation is taking place among the passengers, or a classification value corresponding to "other than conversation", indicating that no conversation is taking place. Thereby, the controller 113 can realize a conversation determination process for determining whether conversation is taking place among at least the passengers using the conversation determination model 112a.

[0040] In this example, the classification of "other than conversation" is subdivided into five classifications, and the conversation determination model 112a outputs six types of output values for various inputs. Specifically, "Sing-along" is a classification value indicating that the passengers are singing along with the music being played by the audio. "Artificial sound 1 " is a classification value indicating artificial sounds such as the sirens of ambulances, fire trucks, and police cars, and crossing sounds, which require attention during vehicle operation.

[0041] "Artificial sound 2 " is a classification value indicating artificial sounds such as games played by passengers other than the driver, which do not require attention during vehicle operation. "High-pitched voice" is a classification value indicating the voice emitted by any of the passengers in a high tone. "Other than the above" is a classification value indicating monologues and the like, which are voices other than "Sing-along", "Artificial sound 1 ", "Artificial sound 2 ", and "High-pitched voice".

[0042] Note that when the classification value is "Conversation", the controller 113 performs volume control to lower the volume. Here, "Normal time" shown in the figure means other than "Emergency time" described later. Also, when the classification value is "Singalong", since the passengers are only singing along with the music (not having a conversation), the controller 113 does not perform volume reduction control.

[0043] Also, when the classification value is "Artificial sound 1 ", since the passengers (especially the driver) need to pay attention, the controller 113 performs volume control to lower the volume. Also, when the classification value is "Artificial sound 2 ", the controller 113 does not perform volume reduction control. Also, when the classification value is "High-pitched voice", the controller 113 executes "Word determination processing". The "Word determination processing" is a process of determining whether a specific keyword corresponding to an emergency is included in the voice. Details of the "Word determination processing" will be described later with reference to FIG. 7 and the like.

[0044] Note that as shown in FIG. 3, when the controller 113 inputs the input signal from each microphone 31 (hereinafter, appropriately referred to as "microphone input signal") and the audio signal from the audio device 12 to the conversation determination model 112a, it performs "AI input processing".

[0045] FIG. 4 is an explanatory diagram of the AI input processing. As shown in FIG. 4, the controller 113 performs "preprocessing" and "feature amount extraction processing" in the AI input processing. In the "preprocessing", the controller 113 performs, for example, "noise reduction", "power spectrum conversion", "vocal separation", etc. on each microphone input signal and the audio signal from the audio device 12.

[0046] The controller 113 performs "noise reduction" by using, for example, a cut filter of a noise band (depending on the sound source, a cut filter of a high-frequency voice band (low-pass filter), a removal filter of impulse noise (high frequency) (band cut filter of the impulse noise band or low-pass filter)).

[0047] Further, the controller 113 performs "power spectrum conversion" that spectrally analyzes the signal levels of the components in each frequency band in the audio signal by means of processing using, for example, a fast Fourier transform (FFT).

[0048] Further, the controller 113 performs "vocal separation" by using, for example, a pre-trained model of "Open-Unmix" that is publicly available as open source for sound source separation. "Vocal separation" separates only the vocal part from the audio signal to facilitate comparison with each microphone input signal.

[0049] In addition to using this open-source model in "vocal separation", in the case of 3D sound sources, etc., the content data includes additional data indicating the vocal (sound source type), which can be used. Also, in the case of normal sound sources, although complete separation is difficult, there are methods such as extracting the main vocal band of a person (in the case of conversation, components around 1 kHz are often (strong)) using a band-pass filter. Also, in the case of stereo sound sources, there are methods such as an addition process of the LR channels that emphasizes the common components of the left and right channels. The addition process of the LR channels is a process that utilizes the fact that the vocal is often a signal of the same level (similar) in the left and right channels, and other sound sources (instruments, etc.) tend to be unevenly distributed in one channel.

[0050] Then, the controller 113 extracts feature quantities that become explanatory variables in the conversation determination model 112a based on each signal after the "pre-processing".

[0051] FIG. 5 is a diagram showing a specific example of the feature quantity that becomes an explanatory variable in the conversation determination model 112a. As shown in FIG. 5, the controller 113 extracts a feature quantity indicating the "conversation situation between passengers" from each signal. This feature quantity is, for example, "the utterance timing indicated by each microphone input signal", "the time difference between the microphone input signals", "the reaction speed between the microphone input signals", "the degree of overlap of the microphone input signals", and the like.

[0052] The "utterance timing indicated by each microphone input signal" refers to the start point of the utterance of each conversation participant. The "time difference between microphone input signals" refers to the time difference at this start point. If the time difference is large, there may be no conversation. The "response speed between microphone input signals" refers to the response speed of the utterance by other conversation participants other than one person to the utterance of one of the conversation participants. The "degree of overlap of microphone input signals" indicates the degree to which the utterances of each conversation participant overlap in time. When determining the presence or absence of a conversation, ideally, it is desirable that the utterances of each conversation participant are exchanged without overlapping, but in actual conversations, it is often the case that conversation participants speak while covering the other person's utterance (while overlapping in time). The above-mentioned degree of overlap is a feature quantity for learning the degree of overlap allowed as a conversation. Based on the above, the shorter the "time difference between microphone input signals", the faster the "response speed between microphone input signals", and the smaller the "degree of overlap of microphone input signals", the stronger the tendency to indicate that a conversation is taking place. As a result, the controller 113 can perform a conversation determination based on a feature quantity indicating the "conversation exchange situation among the passengers", that is, a time-based feature quantity, without relying on natural language processing such as voice recognition of specific keywords included in the utterance.

[0053] Note that the "utterance timing indicated by each microphone input signal" is extracted as the point in time when the waveform indicating the human utterance in each microphone input signal after preprocessing starts to show a signal level equal to or higher than the threshold value.

[0054] In addition, the "speech timing indicated by each microphone input signal" is extracted as the point in time when the signal level of the human frequency band component in each microphone input signal after preprocessing becomes equal to or higher than a threshold value in temporal change. The "time difference between microphone input signals" is calculated as the time difference at this point in time (speech timing) between the microphone input signals. The "response speed between microphone input signals" is calculated, for example, from the time difference between the end speech timing (the point in time when the signal level of the human frequency band component in the microphone input signal becomes less than the threshold value in temporal change) and the start timing of the next speech. The "degree of overlap of microphone input signals" is calculated, for example, from the overlap of waveforms (the time difference between these timings when the end speech timing is earlier than the next speech timing) when the waveforms indicating the human speech in each microphone input signal are overlapped.

[0055] In addition, the controller 113 extracts a feature amount indicating the "similar situation of each microphone input signal with respect to the audio signal" from each signal. This feature amount is, for example, the "pitch difference with respect to the audio signal" or the like. Note that the smaller the pitch difference, that is, the higher the similarity, the stronger the tendency that the occupant is singing along with the music and not having a conversation. The pitch difference can be calculated, for example, by frequency comparison between the signal waveform of the microphone input signal and the signal waveform of the audio signal. Then, the similarity based on this pitch difference is obtained, for example, by comparing the average value during the determination period of the similarity determination with a threshold value. This calculation process of the similarity may be a moving average process in which the oldest value is deleted and the latest value is added while shifting the determination period with the passage of time. Thereby, the controller 113 can perform conversation determination based on the feature amount indicating the "similar situation of each microphone input signal with respect to the audio signal" without relying on natural language processing.

[0056] Based on such feature amounts, the conversation determination model 112a is pre-learned so that each classification value shown in FIG. 3 can be output. FIG. 6 is an explanatory diagram during pre-learning.

[0057] As shown in FIG. 6, the conversation determination model 112a is learned by the learning device 300. The learning device 300 is provided in a learning environment independent of the acoustic system 1, for example. Note that this learning environment is preferably the same as the vehicle cabin environment in which the learned conversation determination model 112a is used. The environment similar to the vehicle cabin environment corresponds to an example of a "simulated environment of the vehicle cabin".

[0058] The learning device 300 is a computer having a processor (not shown). In addition to the preprocessing and feature extraction processing shown in FIG. 4, the learning control unit 301 in FIG. 6 is provided so as to be able to execute learning processing by deep learning by this processor.

[0059] The conversation determination model 112a is learned by supervised learning by this learning device 300. The input data is the learning voice signals and learning audio signals acquired by the microphones 1 ~ the microphones m provided respectively according to the seat positions of the subjects. The learning device 300 performs the AI input processing shown in FIG. 4 on these learning voice signals and learning audio signals, and uses the extracted respective feature amounts as the input data (explanatory variables) of the conversation determination model 112a. Also, the input values by the subjects (the subjects selectively input the respective classification values shown in FIG. 3) and the like are used as the correct answer data (objective variables). 1 ~ the microphones m The learning control unit 301 appropriately updates the weight values and the like using the error backpropagation method or the like based on the determination data (predicted values) output by the conversation determination model 112a and the correct answer data in the learning process.

[0060] The conversation determination model 112a generated by being pre-learned in this way is installed in the application device after learning. In the present embodiment, the learned conversation determination model 112a is stored in the storage unit 112 at the time of manufacturing the volume control device 11 or the like via communication or a storage medium.

[0061]

[0062] ​Next, the word determination process will be described with reference to FIG. 7. FIG. 7 is an explanatory diagram of the word determination process. The word determination process is a process performed when the classification value is "high-pitched voice" as described above.

[0063] When an event requiring emergency (hereinafter, appropriately referred to as "emergency event") occurs during the operation of the vehicle including the driver, it is assumed that the passengers of the vehicle will utter a specific keyword in a high tone. The word determination process determines whether this specific keyword has been uttered.

[0064] As shown in FIG. 7, in the word determination process, the controller 113 inputs each microphone input signal to the word determination model 112b and performs speech recognition using the word determination model 112b.

[0065] Then, when the output value of the word determination model 112b indicates that the voice includes a specific keyword indicating danger such as "Danger!", "Right!", "Left!", the controller 113 performs volume control to rapidly decrease the volume. This rapid volume decrease during an emergency decreases the volume in a shorter time and to a lower level compared to normal as described above.

[0066] In addition, when the output value of the word determination model 112b indicates that the voice does not include a specific keyword, that is, in the case of "otherwise", the controller 113 regards it as mumbling or the like and does not require a volume decrease.

[0067] Note that the configuration and learning method of the word determination model 112b differ depending on whether the word determination model 112b performs a danger determination. When performing a danger determination as in the example of FIG. 7, the word determination model 112b does not perform word recognition by natural language processing but determines the presence or absence of voice accompanying the occurrence of an emergency event. The learning method is learning using learning data in the same environment as in FIG. 6, where the input data is the voice data acquired by each microphone and the correct answer data is the presence or absence data of voice accompanying the occurrence of an emergency event based on the input value of the subject or the like.

[0068] Also, when the risk determination is not performed, the word determination model 112b is trained to perform word recognition by, for example, natural language processing. The training method is training using training data where the input data is the voice data acquired by each microphone and the correct data is the text data indicating the aforementioned specific keyword. When the risk determination is not performed up to this point, the word determination model 112b outputs the recognized keyword. Then, the controller 113 searches a table in which, for example, word data and data indicating whether the voice is associated with the occurrence of an emergency are associated with each other based on the recognized keyword, and determines the presence or absence of the voice associated with the occurrence of an emergency.

[0069] Next, a configuration example of the audio device 12 will be described with reference to FIG. 8. FIG. 8 is a block diagram showing a configuration example of the audio device 12 according to the embodiment.

[0070] As shown in FIG. 8, the audio device 12 includes a communication unit 121, a storage unit 122, a controller 123, and a short-range wireless communication unit 124. The controller 123 is connected to the aforementioned operation unit 5, speaker 7, and display 9.

[0071] The communication unit 121 is realized by a network adapter, a communication bus, or the like. The communication unit 121 is connected to the volume control device 11 by wire or wirelessly, and transmits and receives various information to and from the volume control device 11.

[0072] The storage unit 122 is realized by a storage device such as a ROM, RAM, flash memory, or disk device. In the example of FIG. 8, the storage unit 122 stores sound source data 122a. The sound source data 122a is a group of audio data that is the target of audio playback.

[0073] The controller 123 corresponds to a so-called processor. The controller 123 is implemented by a CPU, MPU, GPU, etc. The controller 123 executes the program according to the illustrated embodiment stored in the storage unit 122, using the RAM as a working area. Also, the controller 123 can be implemented by an integrated circuit such as an ASIC or FPGA.

[0074] The short-range wireless communication unit 124 performs short-range wireless communication with the smartphone 200. The short-range wireless communication unit 124 connects to the smartphone 200 by, for example, Bluetooth (registered trademark).

[0075] The controller 123 performs various operations based on the operation input from the operation unit 5. For example, the controller 123 reads an audio signal or the like of the audio data to be played from the sound source data 122a based on an input such as an audio playback instruction from the operation unit 5, and outputs voice to the speaker 7 and displays various information such as the music title on the display 9. Also, the controller 123 changes the volume, sound quality, etc. based on a volume control command or the like from the volume control device 11. That is, the controller 123 has an arithmetic processing function (decoding, adjustment processing of volume, sound quality, etc.) of an audio signal (digital signal) and a digital-to-analog conversion function (generation (conversion) of an analog audio signal for driving the speaker 7). Note that the audio device 12 has a power amplification circuit or the like for driving the speaker 7, but since it is a configuration of a general audio device, the description thereof is omitted.

[0076] Also, when the smartphone 200 is connected (during the playback operation of the music stored in the smartphone 200), the controller 123 outputs the audio signal from the smartphone 200 to the volume control device 11 and the speaker 7. Note that when the smartphone 200 is connected (during the playback operation of the music stored in the smartphone 200), the audio device 12 controls the playback volume of the speaker 7 for the audio signal from the smartphone 200 based on the volume control command input from the volume control device 11, in the same manner as the playback of the sound source data 122a in the storage unit 122.

[0077] Next, the processing procedure executed by the volume control device 11 will be described with reference to FIGS. 9 to 12. FIG. 9 is a flowchart (part 1) showing the processing procedure executed by the volume control device 11 according to the embodiment. FIG. 10 is a flowchart (part 2) showing the processing procedure executed by the volume control device 11 according to the embodiment. FIG. 11 is a flowchart (part 3) showing the processing procedure executed by the volume control device 11 according to the embodiment. FIG. 12 is a flowchart (part 4) showing the processing procedure executed by the volume control device 11 according to the embodiment. Note that these processes are executed when the system power supply (audio system 1) is turned on.

[0078] As shown in FIG. 9, the controller 113 of the volume control device 11 executes various initial setting processes of the audio system 1 (step S101), and determines whether or not audio reproduction by the audio device 12 has started (step S102). If audio reproduction has not started (step S102, No), the controller 113 repeats step S102.

[0079] When audio reproduction has started (step S102, Yes), the controller 113 sets the volume reduction control status to 0 (step S103). The volume reduction control status is a status value indicating whether or not volume reduction control is in progress, or if volume reduction control is in progress, whether it is normal or emergency. In this embodiment, when the volume reduction control status is 0, it indicates that volume reduction control is not in progress. When the volume reduction control status is 1 or 2, it indicates that volume reduction control is in progress. When the volume reduction control status is 1, it indicates that normal volume reduction control is in progress, and when it is 2, it indicates that emergency volume reduction control is in progress.

[0080] Next, the controller 113 determines whether there is a microphone input of a certain level or higher in any of the microphones 31 (step S104). If there is a microphone input of a certain level or higher (step S104, Yes), the controller 113 acquires each microphone input signal and the audio signal (step S105). Then, the controller 113 inputs the acquired signals to the conversation determination model 112a (step S106).

[0081] If there is no microphone input of a certain level or higher in step S104 (step S104, No), the controller 113 determines whether the volume reduction control status is 1 or 2 (step S107). For example, after reducing the volume of the audio device 12 during the conversation of the passenger, if the conversation is interrupted, that is, if it can be determined that the conversation has ended, etc., there will be a state where there is no microphone input of a certain level or higher.

[0082] If the volume reduction control status is 1 or 2 (step S107, Yes), the controller 113 performs volume restoration control processing (step S108). The volume restoration control processing is a process of returning the reduced volume to its original state. Thereby, for example, at the end of the conversation, the automatically reduced volume can be returned to its original state (the volume before the volume reduction due to the start of the conversation). The processing procedure of the volume restoration control processing will be described later with reference to FIG. 12.

[0083] After performing the volume restoration control processing, the controller 113 repeats the processing from step S104. If the volume reduction control status is 0 (step S107, No), the controller 113 repeats the processing from step S104.

[0084] Next, as shown in FIG. 10, the controller 113 determines the classification value output from the conversation determination model 112a, which is the result of the input in step S106 (step S109). When the classification value indicates "conversation" or "synthetic voice" 1 "(step S109, "conversation, synthetic voice" 1”), the controller 113 determines whether the volume reduction control status is 0 (step S110).

[0085] When the volume reduction control status is 0 (step S110, Yes), the controller 113 corrects the volume adjustment value by user operation or the like based on the normal volume reduction value and the volume reduction speed value of the volume control value information 112c, and outputs a volume control command corresponding to this normal volume reduction control to the audio device 12 (step S111). Then, the controller 113 sets the volume reduction control status to 1 (step S112). Thereafter, the process proceeds to step S116. Also, when the volume reduction control status is not 0 (step S110, No), the controller 113 does not execute steps S111 and S112. Thereafter, the process proceeds to step S116.

[0086] Also, when the classification value indicates “high - pitched voice” (step S109, high - pitched voice), the controller 113 executes a word determination process (step S113). Note that after the word determination process (step S113), the process proceeds to step S116. As shown in FIG. 11, in the word determination process, the controller 113 inputs the microphone input signal to the word determination model 112b (step S201).

[0087] Then, the controller 113 determines whether the voice contains a specific keyword based on the output value of the word determination model 112b (step S202). When the specific keyword is included (step S202, Yes), the controller 113 corrects the volume adjustment value by user operation or the like based on the emergency volume reduction value and the volume reduction speed value of the volume control value information 112c, and outputs a volume control command corresponding to this emergency volume reduction control to the audio device 12 (step S203). Then, the controller 113 sets the volume reduction control status to 2 (step S204). Also, when the specific keyword is not included (step S202, No), the controller 113 does not execute steps S203 and S204.

[0088] Returning to the explanation of FIG. 10, the controller 113 also determines whether the classification values ​​are “conversation” or “artificial sound.” 1 " or "High-pitched vocalization" (step S109, "≠ conversation, artificial sound" 1 , "high-pitched vocalization"), and it is determined whether the volume reduction control status is 1 or 2 (step S114).

[0089] If the volume reduction control status is 1 or 2 (step S114, Yes), the controller 113 executes the volume restoration control process (step S115). 1 When the volume reduction control status is 1 or 2, it means that the status is other than "conversation," "artificial noise," "high-pitched speech," or "high-pitched speech." 1 This refers to the end of a situation in which the voice was classified as "high-pitched vocalization" or "high-pitched vocalization" and determined to require volume reduction control. This allows controller 113 to use conversation determination model 112a to determine the end of a situation in which the voice was classified as "high-pitched vocalization" and to return the reduced volume to its original level accordingly. After the volume return control process (step S115), the process proceeds to step S116.

[0090] 12, in the volume restoration control process, the controller 113 judges the value of the volume reduction control status judged to be 1 or 2 in step S114 (step S301). If the volume reduction status is 1 (step S301, 1), normal volume reduction control is in progress, so the controller 113 cancels the correction of the volume control command by the normal volume reduction control (step S302).

[0091] On the other hand, if the volume reduction control status is 2 (steps S301, 2), emergency volume reduction control is in progress, so the controller 113 cancels the correction of the volume control command by emergency volume reduction control (step S303).

[0092] Then, the controller 113 outputs a volume control command to the audio device 12 to cancel the correction and return the volume to the volume adjustment value by user operation or the like (step S304). Then, the controller 113 sets the volume decrease control status to 0 (step S305).

[0093] Return to the description of FIG. 10. Also, when the volume decrease control status is 0 (step S114, No), the controller 113 does not execute step S115. After that, the process proceeds to step S116. Then, in step S116, the controller 113 determines whether the audio playback by the audio device 12 has ended.

[0094] If the audio playback has not ended (step S116, No), the controller 113 repeats the process from step S104. Also, if the audio playback has ended (step S116, Yes), the controller 113 determines whether the system power has been turned off (step S117). If the system power has not been turned off (step S117, No), the controller 113 repeats the process from step S102. Also, if the system power has been turned off (step S117, Yes), the controller 113 ends the process.

[0095] Note that the controller 113 can perform self - learning of the conversation determination model 112a. This will be described with reference to FIG. 13. FIG. 13 is an explanatory diagram during self - learning. Note that the configuration shown in FIG. 13 is the same as the configuration shown in FIG. 6. And in the case of the configuration shown in FIG. 13, the volume control device 11 stores a program corresponding to the learning control unit 301 of the learning device 300, and the difference from the configuration shown in FIG. 6 is that the controller 113 executes learning processing based on the input data in the actual vehicle cabin environment.

[0096] In the actual vehicle cabin environment, the controller 113 performs self - learning of the conversation determination model 112a based on the occurrence of an event indicating that a conversation is to be held and the microphone input signal during the actual conversation at that time. For example, when the controller 113 receives, as an event, an operation by a passenger indicating that a conversation is about to start, it shifts to the self - learning mode. In the self - learning mode, the controller 113 uses, as correct data (target variable), input values, etc. (in this case, six classification values of "conversation") by the passenger associated with an operation by the passenger indicating that a conversation is to start.

[0097] Then, the learning control unit 301 of the volume control device 11 1 ~ the passenger m executes learning processing using, as input data (explanatory variables), the learning audio signals and learning audio signals acquired by the microphones 1 ~ the microphones m respectively provided according to the seat positions of the passengers. Then, the learning control unit 301 of the volume control device 11 appropriately updates weight values, etc. using the error backpropagation method or the like based on the determination data (predicted value) output by the conversation determination model 112a and the correct data in this learning process.

[0098] As a result, re - (additional) learning of the conversation determination model 112a is performed based on the learning data corresponding to the conversation state in the actual vehicle cabin environment. Therefore, it is possible to improve the accuracy of the conversation determination model 112a suitable for the actual vehicle cabin environment. Here, an operation by a passenger, etc. is cited as an example of the occurrence of an event indicating that a conversation is to be held, but it may also be an operation associated with a preset conversation. In this case, for example, if the sensor unit 3 includes a camera 32, a preset operation (such as a gesture) may be detected based on the camera image of the camera 32.

[0099] In addition, the controller 113 can also perform a control value learning process for learning the volume control value information 112c. This will be described with reference to FIG. 14. FIG. 14 is a flowchart showing the processing procedure of the control value learning process. Note that this process is repeatedly executed during control value learning (performed based on a learning instruction operation by the user).

[0100] In the control value learning process, the controller 113 corrects each volume control value in the volume control value information 112c according to the actual user reaction. Specifically, as shown in FIG. 14, the controller 113 acquires the user reaction to the performed volume control (step S401).

[0101] The user reaction is, for example, the amount of operation of the operation performed by the user with respect to the volume control. The user operation is, for example, a manual volume adjustment operation or the like. Alternatively, the user operation is, for example, a playback operation after a manual rewind operation within a threshold time (re-listening to the music: can be estimated as re-listening due to insufficient volume) or the like.

[0102] The controller 113 acquires the user operation via the operation unit 5 from the audio device 12. Then, the controller 113 automatically corrects the volume control value in the volume control value information 112c according to the user reaction indicated by the acquired user operation (step S402).

[0103] For example, when the user manually lowers the volume further, the user reaction indicates that the amount of volume decrease in the volume decrease control was insufficient. In this case, the controller 113 corrects the volume decrease value in the volume control value information 112c to be smaller by, for example, the amount of volume the user lowered. Considering that this correction is repeated, a history of the volume decrease values before and after the correction may be stored in the volume control value information 112c, and the volume decrease control may be performed based on, for example, the average value of the corrected volume decrease values in the history. Thereby, the controller 113 can customize the volume control value information 112c according to the preferences of each user.

[0104] Note that here, the amount of manual operation or the like is given as an example of the user's reaction. However, for example, keywords indicating requests such as "I want it to be lowered more" may be identified by voice recognition, and the volume control value may be corrected according to the keywords.

[0105] Also, as described above, the sensor unit 3 may include various sensors other than the microphone 31, for example, the camera 32. Next, an in-vehicle system example including this camera 32 (specific examples such as the arrangement of the speakers and microphones) and the state transition (transition of the volume state) in a typical audio-visual situation will be described with reference to FIG. 15. FIG. 15 is a diagram showing an in-vehicle system example including the camera 32.

[0106] As shown in FIG. 15, in the configuration of this in-vehicle system example, the camera 32 is provided in the front of the passenger compartment C. The camera 32 can image the entire panorama inside the passenger compartment C. Also, the passenger compartment C has four seats S-1 to S-4, and four speakers 7-1 to 7-4 are respectively provided in the vicinity of each of the seats S-1 to S-4 to mainly provide voice to the users U-1 to U-4 sitting on each seat. Also, the microphones 31-1 to 31-4 are provided corresponding to the seats S-1 to S-4. The usage situation of the in-vehicle system is that the users U-1 to U-4 are seated on the seats S-1 to S-4. And in the passenger compartment C, a conversation is being held, and it is assumed that the users U-2 on the seat S-2 and the user U-3 on the seat S-3 are the parties having the conversation. Also, it is assumed that the volume (volume adjustment value) based on the volume adjustment operation of the user U-1 is VOLa.

[0107] In this case, the controller 113 identifies the conversation participants based on the input signals from the microphones 31-1, 31-2, 31-3, 31-4 and the camera images of the camera 32. In the case of this example, the inputs from the microphones 31-2 and 31-3 are in a state that satisfies the condition for determining the aforementioned conversation, and the inputs from the microphones 31-1 and 31-4 are not in a state that satisfies this condition. Also, in the camera image of the camera 32, the operation state is such that it can be determined that the users U-2 and U-3 are having a conversation. Then, the controller 113 controls the volume of each speaker 7 according to the seat positions of the identified participants.

[0108] Therefore, in the example of FIG. 15, the controller 113 targets the speakers 7-2 and 7-3 corresponding to the seat positions of the users U-2 and U-3 who are the conversation participants for volume control (see the dashed closed curves in the figure).

[0109] And, since the users U-2 and U-3 are in conversation, the controller 113 performs volume reduction control to reduce the volume of the audio output from the speakers 7-2 and 7-3. Specifically, until the conversation of the users U-2 and U-3 is detected, the audio device 12 outputs an audio signal from the speakers 7-1 to 7-4 at the volume VOLa adjusted by the user U.

[0110] In that situation, when the controller 113 determines that there is a conversation with the users U-2 and U-3 as the conversation participants based on the audio signals acquired by the microphones 31-2 and 31-3, the controller 113 targets the speakers 7-2 and 7-3 for volume reduction control.

[0111] Then, the controller 113 starts correcting the volume VOLa of the audio signal output from these speakers 7-2 and 7-3 by the normal volume reduction value b stored in the volume control value information 112c. The controller 113 sets the volume of the audio signal output from the speakers 7-2 and 7-3 to the volume (VOLa - b).

[0112] While the conversation between users U-2 and U-3 continues, the controller 113 maintains the state of this volume (VOLa-b). Then, when the controller 113 determines the end of the conversation between users U-2 and U-3 based on the voice signals acquired by microphones 31-2 and 31-3, it returns the volume of the audio signal output from speakers 7-2 and 7-3 to its original state. That is, the controller 113 sets the volume of the audio signal output from speakers 7-2 and 7-3 to volume VOLa. Note that the controller 113 determines that the conversation has ended, for example, when the determination that there is no conversation continues for a threshold time or more.

[0113] Thereby, volume control can be performed according to the seat positions of the users engaged in conversation. That is, at least for the conversation participants, volume control can be performed such that the content of each other's utterances can be heard clearly. Also, when it is determined that the conversation has ended, it can be judged that volume reduction control is no longer necessary, and the volume of the audio signal can be automatically restored to the state before volume reduction control. Note that the volumes of speakers 7-1 and 7-4 for users U-1 and U-4 other than the conversation participants are not affected by the conversation between users U-2 and U-3, so users U-1 and U-4 can continue to listen to audio such as music at an appropriate volume.

[0114] As described above, the volume control device 11 according to the embodiment includes a controller 113. The controller 113 acquires voice signals from each sound source in the vehicle compartment C, determines the presence or absence of conversation in the vehicle compartment C based on the relationship state between the acquired voice signals, and controls the playback volume of the audio signal based on the determination result of the presence or absence of conversation. Therefore, according to the volume control device 11, since the conversation state is determined based on the relationship state between each voice signal having a high correlation with the conversation state, for example, the timing difference between voice signals, the conversation state of the occupant can be determined with higher accuracy, and appropriate volume control according to the conversation state can be realized.

[0115] Note that in the above-described embodiment, an example in which the volume control device 11 and the audio device 12 are separate entities is given, but the volume control device 11 and the audio device 12 may be provided integrally.

[0116] In the above-described embodiment, an example in which the audio device 12 reproduces an audio signal based on the sound source data 122a or the smartphone 200 has been given. However, the audio signal may be an audio signal of a radio broadcast. Further, the in-vehicle device 10 may further include a car navigation device, and the audio signal may be an audio signal output by the car navigation device.

[0117] Further effects and modifications can be easily derived by those skilled in the art. Therefore, a broader aspect of the present invention is not limited to the specific details and representative embodiments described and represented as above. Accordingly, various changes can be made without departing from the spirit or scope of the general inventive concept defined by the appended claims and their equivalents.

Explanation of Reference Numerals

[0118] 1 Acoustic system 3 Sensor unit 5 Operation unit 7 Speaker 9 Display 10 In-vehicle device 11 Volume control device 12 Audio device 31 Microphone 32 Camera 111 Communication unit 112 Storage unit 112a Conversation determination model 112b Word determination model 112c Volume control value information 113 Controller 121 Communication unit 122 Storage unit 122a Sound source data 123 Controller 124 Short-range wireless communication unit 200 Smartphone

Claims

1. Acquire voice signals from each sound source in the vehicle interior, Based on the relationship state between the acquired voice signals, determine the presence or absence of conversation in the vehicle interior, Control the playback volume of the audio signal based on the determination result of the presence or absence of conversation, A volume control device.

2. The voice signals from each sound source are Voice signals acquired by a plurality of microphones provided according to the seat positions in the vehicle interior, The volume control device according to Claim 1.

3. The relationship state between the voice signals is A state related to the time difference between the voice signals acquired by the plurality of microphones, The volume control device according to Claim 2.

4. The relationship state between the voice signals is A similar state between the voice signal acquired by the microphone and the audio signal, The volume control device according to Claim 2.

5. Have an artificial intelligence learned with learning data using the temporal feature amounts in the voice signals in a plurality of learning sound sources as explanatory variables and whether or not the plurality of voice signals are conversations as the objective function, Calculate the temporal feature amounts in the voice signals acquired by the plurality of microphones, Input the temporal feature amounts into the artificial intelligence, and use the output of the artificial intelligence as the determination result of the presence or absence of conversation, The volume control device according to Claim 2 or Claim 3.

6. Have an artificial intelligence learned with learning data using the learning audio signal and the learning voice signal acquired by the learning microphone as explanatory variables and the similar state between the learning audio signal and the learning voice signal as the objective function, Input the voice signal acquired by the microphone provided according to the seat position in the vehicle interior and the audio signal into the artificial intelligence, and use the output of the artificial intelligence as the determination result of the presence or absence of conversation, The volume control device according to Claim 4.

7. The plurality of learning sound sources are Voice signals of subjects seated at the seat positions acquired by a plurality of microphones provided according to the seat positions in the vehicle interior of a vehicle for collecting learning data, The volume control device according to Claim 5.

8. Detect the occurrence of an event accompanied by conversation in the vehicle interior, Use the voice signals acquired by the plurality of microphones at the time of event occurrence as explanatory variables, and self-learn the artificial intelligence with learning data having the determination that there is conversation as the objective variable, The volume control device according to Claim 5.

9. When it is determined that there is conversation in the vehicle interior, lower the volume of the audio signal, When the determination that there is no conversation in the vehicle interior continues for a threshold time or longer, the volume of the audio signal is restored to the volume before the decrease. The volume control device according to claim 1.

10. Having a speaker associated with a seat position installed in the vehicle interior, Controlling the reproduction volume of the audio signal in the speaker associated with the seat on which the parties to the conversation are seated based on the determination result of the presence or absence of the conversation. The volume control device according to claim 2.

11. An in-vehicle device mounted on a vehicle and outputting an audio signal from an in-vehicle speaker as sound, and an external audio device connected to the in-vehicle device and outputting an audio signal to the in-vehicle device. The in-vehicle device is Acquiring voice signals from each sound source in the vehicle interior, Determining the presence or absence of conversation in the vehicle interior based on the relationship state between the acquired voice signals, Controlling the reproduction volume of the voice output from the in-vehicle speaker based on the determination result of the presence or absence of the conversation. An audio system.

12. Acquiring voice signals from each sound source in the vehicle interior, Determining the presence or absence of conversation in the vehicle interior based on the relationship state between the acquired voice signals, Controlling the reproduction volume of the audio signal based on the determination result of the presence or absence of the conversation. A program executed by a computer including the process.

Citation Information

Patent Citations

  • Device and method for automatic sound volume control

    JP2007043356A