Voice processing device, voice processing method, and program

The voice processing device addresses the challenge of dynamically setting call areas in vehicles by using a voice signal processing and recognition system to enhance communication quality through selective audio output control, reducing noise and echo.

JP2025151506APending Publication Date: 2025-10-09PANASONIC AUTOMOTIVE SYST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024052973
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing in-vehicle devices lack the ability to dynamically set a call area for speakers based on conversation modes, leading to inefficient and potentially noisy communication experiences.

Method used

A voice processing device that includes a voice signal input processing unit, a voice recognition unit, and a selection switching unit to selectively switch call areas based on speaker information, allowing for precise control of audio output units to optimize communication quality.

Benefits of technology

Enables setting a call area corresponding to a conversation mode, improving communication quality by reducing echo, noise, and crosstalk, and ensuring high-quality audio transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025151506000001_ABST
    Figure 2025151506000001_ABST
Patent Text Reader

Abstract

To set a speech area of an utterer corresponding to a conversation mode.SOLUTION: A voice processing device includes: a voice signal input processing section for processing voice signals to be input from a plurality of voice signal input sections; a voice recognition section for recognizing the voice of utterers in respective areas so as to associate utterer information with the respective areas based on the voice signals processed in the voice signal input processing section; and a selection changeover section for performing selective changeover into speech areas corresponding to the setting of a conversation mode based on the utterer information. The selection changeover section selects the speech area by the selection of the voice signal to be transmitted from the voice signal input processing section and the selection of a voice signal output section to perform a voice signal output among the plurality of voice signal output sections.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an audio processing device, an audio processing method, and a program. [Background technology]

[0002] 2. Description of the Related Art Conventionally, there is an in-vehicle device that switches on a speaker at a seat of a predetermined passenger for hands-free calling in a vehicle.

[0003] For example, Patent Document 1 discloses an in-vehicle device that performs image recognition of vehicle occupants and turns on a speaker in the seat space of the occupant associated with the other party to be called to output audio. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent Publication No. 2021-034781 Summary of the Invention [Problem to be solved by the invention]

[0005] The present disclosure provides a voice processing device, a voice processing method, and a program that are capable of setting a call area for a speaker corresponding to a conversation mode. [Means for solving the problem]

[0006] The voice processing device of the present disclosure comprises a voice signal input processing unit that processes voice signals input from a plurality of voice signal input units, a voice recognition unit that recognizes the voice of a speaker in each area based on the voice signal processed by the voice signal input processing unit and associates speaker information with each of the areas, and a selection switching unit that selectively switches to a call area corresponding to a conversation mode setting based on the speaker information, wherein the selection switching unit selects the call area by selecting a voice signal transmitted from the voice signal input processing unit and selecting a voice signal output unit from a plurality of voice signal output units that will output a voice signal. [Effects of the Invention]

[0007] According to the present disclosure, it is possible to set a call area for a speaker corresponding to a conversation mode. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of a vehicle communication model to which a voice processing device according to an embodiment is applied. [Figure 2] FIG. 2 is a diagram showing an example of a processing block for audio processing in an in-vehicle device. [Figure 3] FIG. 3 is a diagram illustrating an example of a functional block configuration in which a control unit switches calls. [Figure 4] FIG. 4 is a diagram showing an example of functional blocks of the SR that performs speaker recognition. [Figure 5] FIG. 5 is a flow diagram showing an example of speaker registration processing in an in-vehicle device. [Figure 6] FIG. 6 is a flowchart showing an example of speaker recognition processing performed by the SR of the in-vehicle device in the speaker recognition mode. [Figure 7] FIG. 7 is a flow diagram illustrating an example of a control process of the audio system. [Figure 8] FIG. 8 is a diagram illustrating an example of a setting unit for the conversation mode. [Figure 9] FIG. 9 is a diagram showing an example of seating patterns selected according to the conversation mode. [Figure 10] FIG. 10 is a diagram illustrating an example of a processing block of a voice processing device according to the first modification of the embodiment. [Figure 11] FIG. 11 is a diagram for explaining the operation of the ICC. [Figure 12] FIG. 12 is a flow chart showing an example of a control process of the audio system. [Figure 13] FIG. 13 is a diagram illustrating an example of a configuration of hardware blocks of a voice processing device according to the third modification of the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of a voice processing device, a voice processing method, and a program according to the present disclosure will be described in detail with reference to the accompanying drawings.

[0010] (Embodiment) 1 is a diagram illustrating an example of a vehicle communication model to which a voice processing device according to an embodiment is applied. The plan view of the vehicle 1 illustrated in FIG. 1 shows the arrangement of seats and the arrangement of a voice system. The vehicle 1 is illustrated schematically, including the steering wheel and four wheels.

[0011] Vehicle 1 in Fig. 1 is, as an example, a right-hand drive vehicle, and is a vehicle having a front row of seats (also referred to as front seats) for the driver and passenger, and a row of rear seats (also referred to as rear seats). For the sake of explanation, Fig. 1 shows four passengers in total, one in driver's seat 41 (also referred to as seat 41) and passenger seat 42 (also referred to as seat 42), and one each in seats 43 and 44 of the three rear seats.

[0012] The number of passengers is not limited to this. The number of passengers may be one driver, or may be multiple passengers up to the maximum number of passengers that can be accommodated in the vehicle 1.

[0013] 1 (driver's seat 41, passenger seat 42, rear seats 43, 44) each correspond to the "area" of the speaker. The entire rear seats may be set as one area, or the area may be further divided and the area between seats 43 and 44 may also be set as a seat. In this description, the speaker's area is fixed to driver's seat 41, passenger seat 42, rear seats 43, and rear seats 44.

[0014] The vehicle 1 is equipped with an in-vehicle device 10, microphones (a first microphone 21, a second microphone 22, a third microphone 23, and a fourth microphone 24) and speakers (a first speaker 31, a second speaker 32, a third speaker 33, and a fourth speaker 34) that are communicatively connected to the in-vehicle device 10.

[0015] Here, the audio processing device according to the embodiment is applied to an in-vehicle device 10. The speakers correspond to "plurality of audio signal output units." The microphones correspond to "plurality of audio signal input units."

[0016] The first microphone 21 and the second microphone 22 are oriented to input the voice of the speaker in the front seat. The first microphone 21 is oriented to directly input the voice of the speaker in the driver's seat 41, and the second microphone 22 is oriented to directly input the voice of the speaker in the passenger seat 42. As an example, the first microphone 21 and the second microphone 22 are mounted between the driver's seat 41 and the passenger seat 42. The first microphone 21 may be mounted on the headrest of the driver's seat 41. The second microphone 22 may be mounted on the headrest of the passenger seat 42.

[0017] The third microphone 23 and the fourth microphone 24 are oriented to input the voice of the speaker in the rear seat. The third microphone 23 is oriented to directly input the voice of the speaker in seat 43, and the fourth microphone 24 is oriented to directly input the voice of the speaker in seat 44. As an example, the third microphone 23 and the fourth microphone 24 are mounted on a headrest or the like between the seats 43 and 44.

[0018] The first speaker 31 is a speaker near the driver's seat 41. The second speaker 32 is a speaker near the passenger's seat 42. The third speaker 33 is a speaker near the seat 43. The fourth speaker 34 is a speaker near the seat 44. Here, "near the seat" refers to a speaker corresponding to the seat.

[0019] The number and positions of microphones and speakers are merely examples and are not limited to those shown in FIG. 1. The number and placement of microphones may be any number as long as it is possible to associate each speaker with a seat by the microphone into which the voice is input. Speakers may be provided according to the number of seats in the rear seats. For example, speakers provided in the headrests of the rear seats may be provided to accommodate three people.

[0020] Fig. 2 is a diagram showing an example of a processing block for audio processing in the in-vehicle device 10. The audio processing block shown in Fig. 2 has an audio signal input processing unit 11, an ES / NS / CTS 12, a control unit 13, a transmission unit 14, and an SR 15. The audio signal input processing unit 11 has a first BF 111, a second BF 112, an MC / EC 113, and a CTC 114.

[0021] Here, BF stands for beam former. MC / EC stands for music canceller and echo canceller. CTC stands for cross-talk canceller. SR stands for the voice recognition unit that performs speaker recognition. ES stands for echo suppressor. NS stands for noise suppressor. CTS stands for cross-talk suppressor.

[0022] The voice signal input processing unit 11 performs processing to separate the voice signals of the speakers at each seat from the input signals input from each microphone (first microphone 21, second microphone 22, third microphone 23, and fourth microphone 24).

[0023] First, the first BF 111 and the second BF 112 emphasize the voice of the speaker in each seat. The first BF 111 emphasizes the voice signal of the speaker in the driver's seat 41 from the input signal of the first microphone 21, and emphasizes the voice signal of the speaker in the passenger seat 42 from the input signal of the second microphone 22, and outputs the voice signal of the speaker in the driver's seat 41 and the voice signal of the speaker in the passenger seat 42 in parallel.

[0024] In addition, the second BF 112 emphasizes the voice signal of the speaker in seat 43 from the input signal of the third microphone 23, and emphasizes the voice signal of the speaker in seat 44 from the input signal of the fourth microphone 24, and outputs the voice signals of the speakers in each seat in parallel.

[0025] The MC / EC 113 performs music cancellation and echo cancellation on each of the audio signals (audio signal for seat 41, audio signal for seat 42, audio signal for seat 43, audio signal for seat 44) output in parallel from the first BF 111 and the second BF 112.

[0026] The MC / EC 113 functions as a music canceller and cancels input components corresponding to the music being played back from the speaker set 30 from the audio signals (audio signals from seats 41, 42, 43, and 44).

[0027] The MC / EC 113 acts as an echo canceller and cancels echo components, which are signals obtained by re-inputting the speaker's voice output from the speaker set 30, from the speaker's voice signal directly input to each microphone (first microphone 21, second microphone 22, third microphone 23, and fourth microphone 24). As an echo canceller, the MC / EC 113 removes the echo components by, for example, sampling the voice signal immediately before being output from the speaker set 30 and comparing the samples with a phase shift.

[0028] CTC 114 cancels crosstalk between signal transmission paths by canceling out audio signals transmitted by other signal transmission paths in the signal transmission paths through which each audio signal (audio signal for seat 41, audio signal for seat 42, audio signal for seat 43, audio signal for seat 44) is transmitted.

[0029] After crosstalk cancellation by CTC114, each separated audio signal (audio signal for seat 41, audio signal for seat 42, audio signal for seat 43, audio signal for seat 44) is output to control unit 13 via ES / NS / CTS12.

[0030] The SR15 performs speaker recognition on each voice signal (voice signal from seat 41, voice signal from seat 42, voice signal from seat 43, and voice signal from seat 44) and provides the speaker recognition results of each voice signal (voice signal from seat 41, voice signal from seat 42, voice signal from seat 43, and voice signal from seat 44) to the control unit 13. As an example, the SR15 acquires each separated voice signal (voice signal from seat 41, voice signal from seat 42, voice signal from seat 43, and voice signal from seat 44) after crosstalk cancellation by the CTC 114 and performs speaker recognition. Note that the SR15 may acquire a signal output from the ES / NS / CTS 12 and perform speaker recognition. Alternatively, the SR15 may perform speaker recognition using a microphone at each seat in a separate system. Alternatively, the SR15 may be substituted by a microphone provided in another in-vehicle device.

[0031] The speaker recognition result provided to the control unit 13 is information that associates identification information (corresponding to the seat of the speaker of the voice signal) corresponding to the communication channel (CH) from which the voice signal was acquired with speaker information corresponding to the recognized speaker. The speaker information includes, for example, attribute information that indicates the relationship between registered persons, such as registered information such as family, adult, or child.

[0032] The ES / NS / CTS 12 is a suppressor that handles echo, noise, or crosstalk. The ES / NS / CTS 12 processes unnecessary components that cannot be processed by the audio signal input processing unit 11. For example, as an echo suppressor, the ES / NS / CTS 12 compares the signal strength (volume) of the audio signal after crosstalk cancellation with the audio signal immediately before output from the speaker set 30, and attenuates the one with the lower signal strength (volume). As a noise suppressor, the ES / NS / CTS 12 attenuates noise components such as road noise or wind noise. As a crosstalk suppressor, the ES / NS / CTS 12 attenuates crosstalk components from other seats.

[0033] 3 is a diagram showing an example of the configuration of functional blocks for switching calls by the control unit 13. As shown in FIG. 3, the control unit 13 includes a selection switching unit 131, a selected CH audio mixing unit 132, and a playback speaker selection unit 133.

[0034] The selection switch unit 131 selects the CH of the voice signal of the seat corresponding to the setting of the conversation mode based on the CH of each seat included in the speaker recognition result in the SR 15 and the speaker information of the occupant of each seat. Each CH corresponds to each voice signal output from the ES / NS / CTS 12 (voice signal of seat 41, voice signal of seat 42, voice signal of seat 43, voice signal of seat 44).

[0035] Furthermore, when the combination of selected channels is changed, the selection switching unit 131 resets the MC / EC 113 to prevent abnormal noise when switching the call area. When the selection is changed, the call area changes, and the combination of channels from which audio signals are acquired and the combination of speakers to be turned on also change, so the echo path in the vehicle interior space changes. For this reason, when switching the call area, the MC / EC 113 is reset to stabilize the audio quality during switching. The call area is an area in the vehicle interior space where a call can be made at the seat of one or more people selected as the call recipient, and the call area changes depending on the arrangement of the selected seats and the combination of the selected seats.

[0036] The selected channel audio mix unit 132 mixes (in other words, superimposes) the audio signals of the selected audio channels (audio signal for seat 41, audio signal for seat 42, audio signal for seat 43, and audio signal for seat 44) and outputs them. If only one audio channel is selected, only the selected audio signal is output.

[0037] The playback speaker selection unit 133 selects the speaker of the seat corresponding to the selected audio CH as the playback speaker, turns on the playback speaker, and switches off the others. The signal output from the selected CH audio mix unit 132 is output to both the playback speaker after switching and the transmission unit 14.

[0038] The transmitting unit 14 transmits the signal output from the selected CH audio mixing unit 132 (the audio signal from seat 41, the audio signal from seat 42, the audio signal from seat 43, the audio signal from seat 44, or a mixed signal made by mixing multiple audio signals) to the other party.

[0039] The voice of the other party is reproduced from the speaker selected by the control unit 13.

[0040] 4 is a diagram showing an example of functional blocks of the SR15 that performs speaker recognition. For example, the SR15 selects a channel to acquire each of the voice signals (voice signal from seat 41, voice signal from seat 42, voice signal from seat 43, voice signal from seat 44) separated by the CTC 114 (see FIG. 2), and recognizes the speaker for each of the voice signals (voice signal from seat 41, voice signal from seat 42, voice signal from seat 43, voice signal from seat 44). Note that the speaker recognition processing procedure for each voice signal (voice signal from seat 41, voice signal from seat 42, voice signal from seat 43, voice signal from seat 44) is the same, so the speaker recognition processing procedure will be described in detail using one voice signal as an example.

[0041] The functional blocks of speaker recognition shown in FIG. 4 include a speech acquisition unit 301 , a preprocessing unit 302 , a feature calculation unit 303 , a similarity calculation unit 304 , and a determination processing unit 305 .

[0042] 4, a voice acquisition unit 301 acquires a voice signal. Then, a preprocessing unit 302 performs preprocessing such as calculating a voice section and limiting the passband of the voice section signal.

[0043] Next, the feature calculation unit 303 calculates the feature of the speech signal in the speech section after preprocessing. As an example, the feature calculation unit 303 calculates the speaker feature of the speech signal by applying a speaker feature DNN (Deep Neural Network). The DNN is a trained model generated based on speech data of a huge number of training speakers in the speaker training DB 306. The speech signal in the speech section is input to the speaker feature DNN, and the feature is obtained from the output layer side of the speaker feature DNN. The feature is called a speaker feature because it indicates the characteristics of the speaker's speech.

[0044] Next, the similarity calculation unit 304 calculates the similarity between the speaker features obtained by the feature calculation unit 303 and the speaker features of the registered users. A speaker who has registered speaker features in advance is called a registered user. The speaker features of the registered users are included in the data 307.

[0045] Next, the determination processing unit 305 outputs a determination result that a registered user whose similarity satisfies a predetermined condition is the speaker based on the similarity with the speaker feature of each registered user. If none of the similarities satisfies the predetermined condition, the determination result that the speaker is an unregistered speaker is output.

[0046] 5 is a flow diagram showing an example of a speaker registration process in the in-vehicle device 10. The registration process is started when the SR15 of the in-vehicle device 10 is switched to the speaker registration mode. The switch to the speaker registration mode may be manual or automatic. For example, the switch is made when the user operates the speaker registration button on the in-vehicle device 10.

[0047] 5, the in-vehicle device 10 first prompts the user making the registration to speak (step S1). For example, the in-vehicle device 10 outputs a message via a built-in speaker or display, either by voice or by displaying on a UI screen, encouraging the user making the registration to speak, and waits for voice input from a microphone for a certain period of time. The microphone for inputting the voice may be manually selected, or may be automatically selected by detecting the microphone corresponding to the seat from which the voice was spoken.

[0048] Next, the SR 15 acquires the audio signal from the selected microphone for a predetermined period of time, performs preprocessing, and applies the speaker feature DNN to the preprocessed audio signal to calculate speaker features (step S2).

[0049] Next, the SR 15 associates the calculated speaker features with speaker information and registers them in the data 307 (step S3).

[0050] 6 is a flow diagram showing an example of speaker recognition processing performed by the SR15 of the in-vehicle device 10 in the speaker recognition mode. This processing is processing for identifying passengers in each seat. After the power of the in-vehicle device 10 is turned on, the speaker recognition mode is automatically or manually entered, and the following speaker recognition processing is started.

[0051] First, the SR 15 acquires a voice signal corresponding to each seat (step S11). The SR 15 acquires a voice signal corresponding to each seat by acquiring a voice signal uttered at the seat from an input signal of a microphone corresponding to the seat.

[0052] Next, SR15 performs preprocessing of the audio signals corresponding to each seat, applies the speaker feature DNN to the preprocessed audio signals to calculate speaker features, and calculates the similarity with the speaker features of registered users (step S12).

[0053] Next, for each of the voice signals corresponding to each seat, the SR15 determines whether the speaker at each seat is one of the registered users or an unregistered speaker based on the calculated similarity and identifies the speaker (step S13).

[0054] 7 is a flow chart showing an example of a control process of the audio system, which is started after the power supply of the in-vehicle device 10 is turned on.

[0055] First, the audio system performs audio signal input processing in steps S20 to S22. Specifically, the audio system performs BF processing using the first BF 111 and the second BF 112 (step S20). Next, the audio system performs MC / EC processing using the MC / EC 113 (step S21). Next, the audio system performs CTC processing using the CTC 114 (step S22).

[0056] Next, the voice system performs a selection switching determination process of steps S23 to S27. First, the voice system determines whether speaker recognition has been performed (step S23). Specifically, the voice system determines whether the occupants of each seat have been identified. Immediately after the power supply of the in-vehicle device 10 is turned on, speaker recognition has not yet been performed (step S23: NO), so the voice system performs the speaker recognition process (step S24). The speaker recognition process may be performed at any time after startup, but unless the occupants of each seat speak, the identification of the occupants of each seat cannot be completed before a call. Therefore, the in-vehicle device 10 may complete the voice recognition for identifying the occupants of the seats by outputting a message prompting the occupants to speak by voice or by displaying it on a UI screen.

[0057] After speaker recognition, i.e., identification of the occupant of each seat, is completed (step S23: YES), the voice system performs a channel selection process (step S25). Specifically, the selection switch unit 131 selects a channel of the voice signal of the seat corresponding to the setting of the conversation mode from the channels of each voice signal (voice signal of seat 41, voice signal of seat 42, voice signal of seat 43, voice signal of seat 44) output from the ES / NS / CTS 12, based on the speaker information of the occupant of the seat whose speaker has been recognized.

[0058] Next, if there is a change in the channel due to, for example, a user selection (step S26: YES), the selection switching unit 131 performs a switching process (step S27). For example, the selection switching unit 131 resets the MC / EC 113, switches the channel for mixing audio signals, and instructs switching of the playback speaker.

[0059] If there is no change in the channel due to user selection or the like (step S26: No), or after the switching process (step S27) has been performed, the selected channel audio mix unit 132 mixes the audio signals of the selected channel, and the playback speaker selector 133 selects a playback speaker (step S28). If the selected channel is changed due to the selection of the playback speaker, the speaker corresponding to the selected channel is turned on and the others are turned off.

[0060] Then, the transmitting unit 14 transmits the audio-mixed signal of the selected CH to the other party (step S29).

[0061] Fig. 8 is a diagram showing an example of a setting unit for a conversation mode. Fig. 8 shows, as an example of the setting unit, a UI screen displayed on the in-vehicle device 10. The in-vehicle device 10 is provided with an audio on / off button 151, a hands-free setting button 152, etc., so that the audio on / off and the hands-free on / off can be switched on / off on the in-vehicle device 10.

[0062] The UI screen shown in Fig. 8 has different conversation mode selection buttons 150. The user selects one pattern from the plurality of selection buttons 150 to set the conversation mode.

[0063] The selection buttons 150 show, as examples, "normal mode," "adult mode," "family mode," "business mode," "everyone mode," and "personal mode."

[0064] Of these, "family mode" is a mode in which all seats belonging to family members are selected according to the speaker information of the recognition result. "adult mode" is a mode in which all seats belonging to adults are selected according to the speaker information of the recognition result of the seats belonging to family members. "business mode" is a mode in which one parent is selected. "adult mode" may also be a mode in which all seats belonging to adults are selected according to the speaker information of the recognition result of the seats.

[0065] Furthermore, the "normal mode" is a mode that selects the driver's seat 41. The "all mode" is a mode that selects all seats whose speech is recognized. The "personal mode" is a mode that selects only the person who has started a hands-free call using, for example, a voice command when a call comes in or when making a call.

[0066] In this way, by selecting each conversation mode, the conversation area can be automatically switched to the seat of the target person.

[0067] These are merely examples, and other conversation modes may be provided as appropriate depending on the speaker configuration, seating arrangement, etc.

[0068] Furthermore, the conversation modes may be selected by a contact operation such as touching the selection screen, or the conversation mode may be selected by recognizing a voice command uttered by the user.

[0069] In addition, the appropriate conversation mode may be automatically selected and presented to the user based on the caller's information (such as a phone number) when hands-free communication begins. For example, if a call is from a registered family member, the "family mode" is automatically selected.

[0070] 9 is a diagram showing an example of a seat pattern selected depending on the conversation mode, and shows an example of two different passenger patterns as an example.

[0071] The first pattern 1 is a pattern in which two parents sit in the front seats and two children sit in the back seats. Figure 9 shows an example of a layout in vehicle 1, with parent A1 in driver's seat 41, parent A2 in passenger seat 42, child a1 in seat 43, and child a2 in seat 44. It is assumed that speaker recognition for each seat has been completed.

[0072] The table shows three settings as examples: "Adult Mode," "Family Mode," and "Business Mode."

[0073] In the example passenger pattern shown in Pattern 1, when "adult mode" is set, parent A1 and parent A2 are selected based on the speaker information obtained by speaker recognition for each seat. Since the seats of parent A1 and parent A2 are known, the CHs of the voice signals of parent A1 and parent A2 (CHs corresponding to the first microphone 21 and second microphone 22) are selected, and the respective speakers of parent A1 and parent A2 (first speaker 31 and second speaker 32) are also selected to be on.

[0074] When "family mode" is set, parent A1, parent A2, child a1, and child a2 are similarly selected based on speaker information obtained by speaker recognition at each seat. Since each seat is known, the channels of the voice signals of parent A1, parent A2, child a1, and child a2 (channels corresponding to first microphone 21, second microphone 22, third microphone 23, and fourth microphone 24) are selected, and the speakers of parent A1, parent A2, child a1, and child a2 (first speaker 31, second speaker 32, third speaker 33, and fourth speaker 34) are also selected to be on.

[0075] Similarly, when the "business mode" is set, the parent A1 (driver) is selected based on the speaker information obtained by speaker recognition for each seat. Since each seat is known, the CH of the parent A1's voice signal (CH corresponding to the first microphone 21) is selected, and the parent A1's speaker (first speaker 31) is also turned on.

[0076] The second pattern 2 is a pattern in which one parent sits in the front seat and one parent and one child sit in the back seat. The passenger seat 42 is assumed to be a different person X.

[0077] In this case, when "adult mode" is set, parent A1 and parent A2 are similarly selected based on the speaker information obtained by speaker recognition for each seat. However, because parent A2's seat is different from pattern 1, the CH corresponding to the fourth microphone 24 is selected as the channel for the parent A2's voice signal, and the fourth speaker 34 is selected as the speaker for parent A2.

[0078] When "family mode" is set, three people, parent A1, parent A2, and child a1, are selected. Since the passenger seat 42 is occupied by another person X, the CH corresponding to the second microphone 22 is not selected, and the second speaker 32 is also turned off.

[0079] In the present embodiment, an example is shown in which the voice processing device is applied to an in-vehicle device 10. The in-vehicle device 10 shown as an example may be a communication device with a calling function, or may be a separate voice processing device that is connected to the communication device for communication. Also, a passenger's smartphone may be paired with the voice processing device to enable hands-free calling.

[0080] Furthermore, speaker recognition by the SR15 may be performed for a limited period of time after engine start, during which time speaker recognition for each seat is completed, or it may be performed continuously. Continuous operation allows for the detection and addition of passengers from seats who have not spoken within a certain period of time since engine start. Continuous operation of speaker recognition also allows for adaptability to changes in seating arrangements, even after passengers get in and out of the vehicle without turning off the engine while the vehicle is stopped, or even if a seat is moved while the vehicle is moving. The conversation mode may also be changed after a hands-free call is initiated.

[0081] Furthermore, when switching the call area, the user may be prompted to decide whether to change it. For example, in family mode, if a child who was asleep before the hands-free call started wakes up and starts talking during the hands-free call, the new speaker is determined to be a child of the family, and an additional seat is automatically added in family mode, changing the call area. In such a case, the user may be prompted to decide whether to add another seat to expand the call area.

[0082] In this embodiment, multiple microphones and multiple speakers are used, and the microphone and speaker corresponding to each seat are automatically selected according to the set conversation mode to switch the call area. When the seat call area is switched according to the set conversation mode, the echo path in the vehicle interior changes, causing abnormal noise. However, the abnormal noise can be suppressed by resetting the MC / EC 113, for example. In addition, high-quality voice communication is possible by canceling music and noise.

[0083] Furthermore, recognition is possible even when the face is not visible to the camera, as it is based solely on the voice. Furthermore, even if the camera is dark and the image is poorly captured, or if the person's appearance has changed since the facial image was registered due to wearing glasses, sunglasses, or a mask, recognition is possible based solely on the voice.

[0084] (Variation 1) In the embodiment, the configuration of the audio system has been described taking as an example a case where passengers in each seat are close to each other. In this case, even if multiple people are selected as speakers in a family mode or the like, the voices of each speaker reach each other's ears directly, so that what each speaker says into the microphone can be shared. On the other hand, in a vehicle with three rows of seats, passengers in the first and third rows may be selected. In such a case, because the seats of the speakers are far apart and the voices do not reach each other's ears directly due to music, road noise, etc., it is difficult to hear directly, and it becomes difficult for the speakers to share what they say into the microphone with each other.

[0085] Therefore, in cases where it is difficult for the selected speakers to share what they say into the microphone, for example, because the selected speakers are seated far apart, a modified configuration is shown that includes an in-car conversation support function.

[0086] Fig. 10 is a diagram showing an example of a processing block of an audio processing device according to a first modification of the embodiment. The audio processing block shown in Fig. 10 is provided with an ICC 16 (an example of an audio processing unit) in a configuration having microphones and speakers on the third and subsequent columns. The same reference numerals are used to designate parts corresponding to the audio processing block shown in Fig. 2. Since the description of parts corresponding to the audio processing block shown in Fig. 2 would be redundant, the description will be omitted, and the operation of the ICC 16 will be described in detail.

[0087] The ICC 16 is a processing block that performs in-car communication support. The ICC 16 is a processing block that performs processing that includes the voice signal of the target speaker, and is controlled by the selection switch unit 131 of the control unit 13.

[0088] As an example, the selection switching unit 131 switches ICC16 on when the selected speakers are positioned so that their voices cannot reach each other directly, based on the seating arrangement of the speakers. When ICC16 is turned on, for example, the voice of a speaker in the front seat speaking into a microphone in front is output from a speaker in the rear seat, making it easier for the speaker in the rear seat to hear the voice of the speaker in the front seat. The operation of ICC16 will be described in detail later. Furthermore, the selection switching unit 131 may turn ICC16 off when there is only one selected speaker or when the selected speakers are positioned so that their voices can reach each other directly.

[0089] Whether or not the arrangement is such that it is difficult for speakers to directly reach each other's voices can be determined by setting conditions. For example, it is determined to be such if the seating distance between at least one pair of selected speakers is equal to or greater than a certain value. Distance data indicating the respective distances between the seats may be stored.

[0090] For example, if seats in the first and third rows are selected, it is considered relevant. On the other hand, if seats in the first and second rows are selected, it is considered not relevant. Furthermore, if seats in the first, second, and third rows are included in the selection, the second row is considered not relevant, and ICC16 for the second row is turned off.

[0091] Other conditions may also be used to switch on the ICC 16, such as driving speed, whether the windows are open or closed, the volume of music in the car, the noise level in the car, road conditions, weather information, the volume of the selected speaker's voice, etc.

[0092] For example, ICC16 is switched on when the driving speed is over 100 km / h. ICC16 is also switched on when the windows are open. ICC16 is also switched on when the volume of the music inside the car is over 30. ICC16 is also switched on when the noise level estimated from the microphone inside the car is over 70 dBA. ICC16 is also switched on when the road surface conditions at the vehicle's driving location are poor based on vehicle position information. ICC16 is also switched on when weather information is acquired, such as when there is a lot of rain noise. ICC16 is also switched on to amplify the voice of the selected speaker when the volume of the speaker's voice is below a threshold.

[0093] Although several examples have been shown so far in which the ICC 16 is switched based on the determination of the selection switching unit 131, this is not limitative. The user may also actively turn on the ICC 16 regardless of the determination of the selection switching unit 131.

[0094] If the selected channel combination is changed, the feedback path will also change, so it is recommended that you reset the ICC16 as well as the MC / EC113.

[0095] FIG. 11 is a diagram for explaining the operation of the ICC 16. The processing blocks shown in FIG. 11 illustrate, as an example, processing between two pairs of microphones and speakers located at separate positions in the front and rear. For example, processing is illustrated between two pairs: a pair of microphone 21 and speaker 31 in the front seat and a pair of microphone 23 and speaker 33 in the rear seat. Also illustrated is processing between two pairs: a pair of microphone 22 and speaker 32 in the front seat and a pair of microphone 24 and speaker 34 in the rear seat. Note that this relationship is an example for explanation, and two pairs with each seat and each of the other seats may have a similar processing relationship. The same applies when there are microphones and speakers in the third row and beyond.

[0096] As an example, in the following description, it is assumed that a speaker in the front seat speaks into the front microphone 21, and a speaker in the rear seat speaks into the rear microphone 23.

[0097] 11, an input signal SD1 input from a front microphone 21 is input to a PreEQ 401, and the audio signal from the PreEQ 401 is processed in order by an MC 402, an EC 403, and an HC 404. EQ is an equalizer. The audio signal is processed by dividing it into 16 subbands, for example.

[0098] The MC 402 cancels the reproduced music included in the audio signal. In this example, the MC 402 acquires the music reproduction signal being reproduced by the music player 600 and cancels the music reproduction signal from the audio signal to be transmitted.

[0099] The EC 403 cancels echoes that occur when the voice of a speaker in the rear seat, output from the front speaker 31, is input again to the front microphone 21. In this example, the EC 403 acquires an audio signal immediately before it is output from the front speaker 31, and cancels a signal corresponding to the acquired audio signal from the input signal input from the front microphone 21.

[0100] The HC 404 cancels feedback that occurs when an audio signal input to the front microphone 21 is output from the rear speaker 33 and then input again to the front microphone 21. In this example, the HC 404 acquires the audio signal just before it is output from the rear speaker 33, and cancels the signal corresponding to the acquired audio signal from the input signal input from the front microphone 21. This makes it possible to cancel the audio signal input again from the front microphone 21. Because there is a time lag until the audio signal is input again from the front microphone 21, the acquired audio signal is held in a delay circuit (D) until that timing.

[0101] The audio signal is processed in the order of HC 404, NS 405, AGC 406, LIM 407, and PostEQ 408. Here, AGC stands for Auto Gain Controller, and LIM stands for Limiter.

[0102] The NS 405 extracts frequency components from the audio signal by fast Fourier transform (FFT) and attenuates noise components of the noise input from the microphone 21, such as road noise or wind noise.

[0103] The audio signal output from the NS 405 is then given a predetermined gain by the AGC 406 and output via the LIM 407. Finally, the audio signal output via the PostEQ 408 is mixed with the music playback signal being played by the music player 600, and the resulting signal is output to the speaker 33. In summary, turning on the ICC 16 is realized by outputting an audio signal input to the front microphone 21 from the rear speaker 33 and by operating the HC 404. When the ICC 16 is turned on, the AGC 406 may also operate.

[0104] The same applies to the processing between the rear seat microphone 23 and the front seat speaker 31 shown in Fig. 11. An input signal SD2 input from the rear microphone 23 is input to PreEQ 501, and the audio signal from PreEQ 501 is processed in order by MC 502, EC 503, and HC 504. The audio signal is similarly divided into 16 subbands and processed.

[0105] The MC 502 similarly cancels the reproduced music included in the audio signal. That is, the MC 502 acquires the music reproduction signal being reproduced by the music player 600 and cancels the music reproduction signal from the audio signal to be transmitted.

[0106] The EC 503 cancels an echo that occurs when the audio signal of the front speaker output from the rear speaker 33 is input again to the rear microphone 23. In this example, the EC 503 acquires the audio signal immediately before it is output from the rear speaker 33, and cancels the signal corresponding to the acquired audio signal from the input signal input from the rear microphone 23.

[0107] The HC 504 cancels feedback that occurs when an audio signal input to the rear microphone 23 is output from the front speaker 31 and then input again to the rear microphone 23. In this example, the HC 504 acquires the audio signal immediately before it is output from the front speaker 31, and cancels the signal corresponding to the acquired audio signal from the input signal input from the rear microphone 23. This makes it possible to cancel the audio signal input again from the rear microphone 23. Because there is a time lag until the audio signal is input again from the rear microphone 23, the acquired audio signal is held in a delay circuit (D) until that timing.

[0108] The audio signal is processed by the HC 504, followed by the NS 505, the AGC 506, the LIM 507, and the PostEQ 508 in this order.

[0109] The NS 505 extracts frequency components from the audio signal by fast Fourier transform, and attenuates noise components of the noise input from the microphone 23, such as road noise or wind noise.

[0110] The audio signal output from the NS 505 is then given a predetermined gain by the AGC 506 and output via the LIM 507. Finally, the audio signal output via the PostEQ 508 is mixed with the music playback signal being played by the music player 600, and the resulting signal is output to the speaker 31. In summary, turning on the ICC 16 is realized by outputting an audio signal input to the rear microphone 23 from the front speaker 31 and by operating the HC 504. When the ICC 16 is turned on, the AGC 506 may also operate.

[0111] In this way, when ICC16 is implemented, the voice of a speaker in the front seat speaking into the front microphone 21 (or microphone 22) is output to the speaker 33 (or speaker 34) in the rear seat. Therefore, even if it is difficult for the speakers to hear each other directly, the content of the speech can be heard through the speaker. Furthermore, while this may cause feedback, the effects of this can be suppressed by providing a feedback canceller.

[0112] Furthermore, if the speaker's voice is quiet, the AGC 406 can be used to amplify and output the voice.

[0113] Fig. 12 is a flow diagram showing an example of a control process of an audio system according to the first modification of the embodiment. This process includes an ICC control process after step S26 in the process shown in Fig. 7. The ICC control process will be described in detail below, and illustrations and descriptions of other processes similar to those in Fig. 7 will be omitted as appropriate.

[0114] The ICC control process corresponds to the processes of steps S31 and S32 after step S26. First, in step S31, the selection switching unit 131 determines whether ICC processing is necessary based on the seats corresponding to the conversation mode setting or the ICC control information. From the seats corresponding to the conversation mode setting, it determines whether ICC processing is necessary based on the distance between the selected seats. From the ICC control information, it determines whether ICC processing is necessary based on the audio level of the input audio (including the influence of ambient noise, etc.).

[0115] When the selection switching unit 131 determines that ICC processing is necessary (step S31: YES), it selects ICC processing and selects the HC and playback speaker settings to be subjected to ICC processing (step S32). For example, if the distance between the selected seats is equal to or greater than a certain distance, it selects the HC between the seats and the setting for audio playback by the playback speaker.

[0116] If the selection switching unit 131 determines that ICC processing is not necessary, the result in step S31 is NO.

[0117] Next, the selection switching unit 131 determines whether there is a channel change (step S33). For example, if there is a channel change due to a user selection (step S33: YES), a switching process is performed (step S34). For example, the selection switching unit 131 resets MC / EC / ICC, switches the mix channel, and issues instructions to switch the playback speaker.

[0118] The subsequent processing in steps S35 and S36 corresponds to steps S8 and S29 shown in FIG. 7, respectively, and therefore description thereof will be omitted here.

[0119] (Variation 3) FIG. 13 is a diagram illustrating an example of a configuration of hardware blocks of a voice processing device according to the third modification of the embodiment.

[0120] 13 includes a CPU 201, a memory 202, a touch panel 203, a display 204, a storage device 205, a communication IF (interface) 206, and a connection IF 207. Each unit is interconnected via a bus.

[0121] The CPU 201 is a central processing unit (CPU) that executes a predetermined program stored in the memory 202 to control each unit and perform processing.

[0122] The memory 202 is a read-only memory (ROM) or a random access memory (RAM). The memory 202 stores predetermined programs and data. The memory 202 also has a work area that the CPU 201 uses for processing.

[0123] The touch panel 203 is a sensor that detects a touch position on the screen of the display 204 .

[0124] The display 204 is a display such as a liquid crystal display.

[0125] The storage device 205 is a storage such as a hard disk drive (HDD) or a solid state drive (SSD).

[0126] The communication IF 206 is a communication interface that communicates with external devices, and for example, connects to a predetermined network (such as the Internet) via wireless communication.

[0127] The connection IF 207 is an interface for connecting to an external device via a wired or wireless connection. The connection IF 207 is, for example, an interface such as Bluetooth. The connection IF 207 is communicably connected to external devices such as a microphone and a speaker. In addition to the microphone and the speaker, a camera or the like may also be connected.

[0128] The camera has an imaging device such as a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS), and captures an image of an object such as the interior of a vehicle.

[0129] The speaker is a speaker set for outputting predetermined sounds (operation sounds, notification sounds, music, etc.) and voices (voices of the other party on the call or the speaker in the car, etc.) played by the CPU 201, and corresponds to a speaker set 30 having multiple speakers such as a first speaker 31, a second speaker 32, a third speaker 33, and a fourth speaker 34.

[0130] The microphones are microphones that convert the voices of each seat into audio signals and input them, and there are a plurality of microphones such as a first microphone 21, a second microphone 22, a third microphone 23, and a fourth microphone 24.

[0131] The CPU 201 may execute a predetermined program stored in the memory 202 to realize some or all of the functions of the processing blocks shown in the embodiment and the first modification.

[0132] In addition to speaker recognition based on voice signals, face image recognition using a camera may also be performed.

[0133] The present disclosure can be realized in software, hardware, or software in conjunction with hardware.

[0134] The present disclosure may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium, or as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium. A program product is a computer-readable medium on which a computer program is recorded.

[0135] In addition, a program recording some or all of the procedures may be provided by recording it on a recording medium, or stored in ROM and provided as a computer-configured information processing device, or the program may be downloaded via a network and executed by a computer. The CPU of the computer performs processing by reading and executing the program.

[0136] Although the embodiments have been described above with reference to the drawings, the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications or alterations within the scope of the claims. Such modifications or alterations are also considered to fall within the technical scope of the present disclosure. Furthermore, the components in the embodiments may be combined in any manner without departing from the spirit of the present disclosure.

[0137] (Addendum) Aspects of the present disclosure are, for example, as follows. (Item 1) an audio signal input processing unit that processes audio signals input from a plurality of audio signal input units; a voice recognition unit that recognizes the voice of a speaker in each area based on the voice signal processed by the voice signal input processing unit and associates speaker information with each area; a selection / switching unit that selectively switches to a call area corresponding to a setting of a conversation mode based on the speaker information; and the selection switching unit selects the call area by selecting the audio signal transmitted from the audio signal input processing unit and selecting an audio signal output unit to output the audio signal from among a plurality of audio signal output units. Audio processing device. (Item 2) the audio signal input processing unit has an echo canceller that cancels an echo component of the audio signal output from the audio signal output unit that is re-input from the audio signal input unit, the selection switching unit resets the echo canceller when changing the selection of the audio signal output unit. Item 1. A voice processing device according to item 1. (Item 3) when a voice of a new speaker is recognized based on the voice signal processed by the voice signal input processing unit, the voice recognition unit associates speaker information corresponding to the new speaker with an area corresponding to the new speaker; when the speaker information of the new speaker is added during a call in the set conversation mode and the speaker information of the new speaker corresponds to the set conversation mode, the selection switching unit switches the call area by including in the selection the voice signal transmitted from the voice signal input processing unit and the voice signal output unit corresponding to the new speaker. Item 3. A voice processing device according to item 1 or 2. (Item 4) a voice processing unit that outputs a voice signal input to a first voice signal input unit that is the voice signal input unit corresponding to a first speaker to a second voice signal output unit that is the voice signal output unit corresponding to a second speaker, and outputs a voice signal input to the second voice signal input unit that is the voice signal input unit corresponding to the second speaker to a first voice signal output unit that is the voice signal output unit corresponding to the first speaker, 4. A voice processing device according to any one of items 1 to 3. (Item 5) the voice processing unit is selected when a distance between an area of ​​the first speaker and an area of ​​the second speaker is equal to or greater than a certain value; 5. A voice processing device according to any one of items 1 to 4. (Item 6) The voice processing unit is selected when a state in which it is difficult to hear the voices of the first speaker and the second speaker is detected. 6. A voice processing device according to any one of items 1 to 5. (Item 7) the voice acquisition unit is selected when noise estimated from the voice signal input unit is equal to or greater than a certain value; 7. A voice processing device according to any one of items 1 to 6. (Item 8) A setting unit is provided for setting the conversation mode from among a plurality of conversation modes. 8. A voice processing device according to any one of items 1 to 7. (Item 9) the speaker information includes attribute information indicating a relationship between speakers, the voice signal input unit and the voice signal output unit corresponding to an area of ​​a speaker including the attribute information corresponding to the set conversation mode are selected; 9. A voice processing device according to any one of items 1 to 8. (Item 10) a transmitting unit that transmits to a call partner a signal on which each audio signal from the one or more audio signal input units selected by the selection switching unit is superimposed; 10. A voice processing device according to any one of items 1 to 9. (Item 11) the voice recognition unit associates the registered user's speaker information with an area corresponding to the voice signal input unit to which voice information corresponding to the registered user's voice information has been input; 8. A voice processing device according to any one of items 1 to 7. (Item 12) 1. A method of processing audio in an audio system having a plurality of audio signal input units and a plurality of audio signal output units, comprising: processing audio signals input from the plurality of audio signal input units; a step of recognizing the voice of a speaker in each area based on the processed voice signal and associating speaker information with each area; selecting, based on the speaker information, a voice signal to be transmitted from the plurality of voice signal input units and a voice signal output unit to be caused to output a voice signal from the plurality of voice signal output units, thereby switching to a call area corresponding to a setting of a conversation mode; An audio processing method comprising: (Item 13) A computer in which a plurality of audio signal input units and a plurality of audio signal output units are communicatively connected, an audio signal input processing unit that processes audio signals input from the plurality of audio signal input units; a voice recognition unit that recognizes the voice of a speaker in each area based on the voice signal processed by the voice signal input processing unit and associates speaker information with each area; a selection / switching unit that selectively switches to a call area corresponding to a setting of a conversation mode based on the speaker information, Make it work, selecting the call area by selecting the audio signal transmitted from the audio signal input processing unit and selecting an audio signal output unit to output the audio signal from among the plurality of audio signal output units; program. [Explanation of symbols]

[0138] 1 vehicle Seats 41, 42, 43, and 44 21 First Mic 22 Second Microphone 23 The Third Mic 24 The Fourth Mic 31 First Speaker 32 Second Speaker 33 Third Speaker 34 Fourth Speaker 10 In-vehicle equipment 11 Audio signal input processing section 12 ES / NS / CTS 13 Control Unit 14 Transmitter 15SR 111 First BF 112 Second BF 113 MC / EC 114 CTC

Claims

1. an audio signal input processing unit that processes audio signals input from a plurality of audio signal input units; a voice recognition unit that recognizes the voice of a speaker in each area based on the voice signal processed by the voice signal input processing unit and associates speaker information with each area; a selection / switching unit that selectively switches to a call area corresponding to a setting of a conversation mode based on the speaker information; and the selection switching unit selects the call area by selecting the audio signal transmitted from the audio signal input processing unit and selecting an audio signal output unit to output the audio signal from among a plurality of audio signal output units. Audio processing device.

2. the audio signal input processing unit has an echo canceller that cancels an echo component of the audio signal output from the audio signal output unit that is re-input from the audio signal input unit, the selection switching unit resets the echo canceller when changing the selection of the audio signal output unit. The audio processing device according to claim 1 .

3. when a voice of a new speaker is recognized based on the voice signal processed by the voice signal input processing unit, the voice recognition unit associates speaker information corresponding to the new speaker with an area corresponding to the new speaker; when the speaker information of the new speaker is added during a call in the set conversation mode and the speaker information of the new speaker corresponds to the set conversation mode, the selection switching unit switches the call area by including in the selection the voice signal transmitted from the voice signal input processing unit and the voice signal output unit corresponding to the new speaker.

3. The audio processing device according to claim 1.

4. The voice processing unit further includes a voice processing unit that outputs a voice signal input to a first voice signal input unit that is the voice signal input unit corresponding to a first speaker to a second voice signal output unit that is the voice signal output unit corresponding to a second speaker, and outputs a voice signal input to the second voice signal input unit that is the voice signal input unit corresponding to the second speaker to a first voice signal output unit that is the voice signal output unit corresponding to the first speaker.

3. The audio processing device according to claim 1.

5. The voice processing unit is selected when any one of the following conditions is satisfied: when the distance between the area of ​​the first speaker and the area of ​​the second speaker is equal to or greater than a certain value; when a state in which it is difficult to hear the voices of the first speaker and the second speaker is detected; or when the volume of noise included in the voice signal input by the voice signal input unit is equal to or greater than a certain value. The audio processing device according to claim 4 .

6. A setting unit is provided for setting the conversation mode from among a plurality of conversation modes.

3. The audio processing device according to claim 1.

7. a transmitting unit that transmits each of the voice signals from the one or more voice signal input units selected by the selection switching unit to a call partner; 3. The audio processing device according to claim 1.

8. the voice recognition unit associates the registered user's speaker information with an area corresponding to the voice signal input unit to which voice information corresponding to the registered user's voice information has been input; 3. The audio processing device according to claim 1.

9. 1. A method of processing audio in an audio system having a plurality of audio signal input units and a plurality of audio signal output units, comprising: processing audio signals input from the plurality of audio signal input units; a step of recognizing the voice of a speaker in each area based on the processed voice signal and associating speaker information with each area; selecting a voice signal to be transmitted from the plurality of voice signal input units and a voice signal output unit to be caused to output a voice signal from the plurality of voice signal output units based on the speaker information, and switching to a call area corresponding to a setting of a conversation mode; An audio processing method comprising:

10. A computer in which a plurality of audio signal input units and a plurality of audio signal output units are communicatively connected, an audio signal input processing unit that processes audio signals input from the plurality of audio signal input units; a voice recognition unit that recognizes the voice of a speaker in each area based on the voice signal processed by the voice signal input processing unit and associates speaker information with each area; a selection / switching unit that selectively switches to a call area corresponding to a setting of a conversation mode based on the speaker information, Make it work, selecting the call area by selecting the audio signal transmitted from the audio signal input processing unit and selecting an audio signal output unit to output the audio signal from among the plurality of audio signal output units; program.

Citation Information

Patent Citations

  • Sound system

    JP2021034781A