Information processing device

The information processing apparatus addresses the challenge of canceling multiple speaker sounds by calculating sound ratios and generating a reference signal based on loudness, enhancing speech recognition by effectively eliminating interference.

JP7841514B2Active Publication Date: 2026-04-07TOYOTA JIDOSHA KK
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing systems struggle to effectively cancel multiple sounds output from multiple speakers in a vehicle, which interfere with speech recognition due to insufficient bandwidth for transmitting all output signals and the need to mix some signals, leading to ineffective cancellation.

Method used

An information processing apparatus calculates the ratio of sound magnitudes from multiple speakers, sets a mixing ratio for generating a reference signal to cancel out these sounds, and uses this signal to effectively eliminate the interfering sounds based on their loudness.

Benefits of technology

The solution allows for effective cancellation of multiple speaker sounds by emphasizing louder sounds and reducing the processing load on the system, thereby improving speech recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007841514000001
    Figure 0007841514000001
  • Figure 0007841514000002
    Figure 0007841514000002
  • Figure 0007841514000003
    Figure 0007841514000003
Patent Text Reader

Abstract

To effectively cancel a plurality of sounds output from a plurality of speakers.SOLUTION: A control part of an information processing device acquires a plurality of first sounds output from a plurality of speakers. The control part of the information processing device calculates a rate of a magnitude of the plurality of first sounds acquired. Then, the control part sets a mixing ratio by applying a ratio of a magnitude of the plurality of first sounds calculated as the mixing ratio that is a ratio for mixing a plurality of sound signals, the mixing ratio used for generating a reference signal in order to cancel, from the sound acquired from a microphone, a plurality of output sounds output from the plurality of speakers in accordance with the plurality of sound signals.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to an information processing apparatus.

Background Art

[0002] Patent Document 1 discloses an in-vehicle voice recognition apparatus that performs echo cancellation processing for removing echo components included in input voice. The in-vehicle voice recognition apparatus disclosed in Patent Document 1 removes echo components included in input voice collected via a microphone and signals from a plurality of sound sources. Further, the in-vehicle voice recognition apparatus performs voice recognition on an output from which the echo components have been removed. Further, the in-vehicle voice recognition apparatus monitors the input voice collected via the microphone and the signals from the plurality of sound sources, determines whether the voice recognition rate is improved by removing the echo components, and controls the echo cancellation processing.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of this disclosure is to effectively cancel a plurality of sounds output from a plurality of speakers.

Means for Solving the Problems

[0005] The information processing apparatus according to this disclosure acquires a plurality of first sounds output from a plurality of speakers, calculates a ratio of the magnitudes of the acquired plurality of first sounds, A mixing ratio is a ratio for mixing multiple sound signals, and the mixing ratio is set by applying the calculated ratio of the magnitudes of the multiple first sounds to the mixing ratio used to generate a reference signal for canceling out the multiple output sounds output from the multiple speakers in response to the multiple sound signals from the sound acquired by the microphone. It includes a control unit configured to perform the following actions. [Effects of the Invention]

[0006] This disclosure makes it possible to effectively cancel out multiple sounds output from multiple speakers. [Brief explanation of the drawing]

[0007] [Figure 1] Figure 1 is a diagram showing the schematic configuration of the speech recognition system according to this embodiment. [Figure 2] Figure 2 shows the arrangement of multiple speakers and in-vehicle equipment within a vehicle. [Figure 3] Figure 3 is a block diagram schematically showing an example of the functional configuration of an in-vehicle device. [Figure 4] Figure 4 shows an example of the table structure of percentage information held in a percentage information database. [Figure 5] Figure 5 is a flowchart of the first process performed by the control unit of the in-vehicle device. [Figure 6] Figure 6 is a flowchart of the second process performed by the control unit of the in-vehicle device. [Modes for carrying out the invention]

[0008] Let's consider a scenario where sound is acquired using a microphone. In this case, speakers surrounding the microphone may be emitting sound. As a result, the sound acquired by the microphone will be mixed with the sound emitted by the speakers.

[0009] Furthermore, there are cases where multiple speakers output multiple sounds. In such cases, a reference signal is used to cancel out the multiple sounds output from the multiple speakers, thereby canceling out the multiple sounds from the sound acquired by the microphone. The information processing device according to this disclosure aims to cancel out the multiple sounds output from multiple speakers in such cases.

[0010] The control unit of the information processing device according to this disclosure acquires multiple first sounds output from multiple speakers. The control unit of the information processing device calculates the ratio of the loudness of the acquired multiple first sounds. The control unit then sets the mixing ratio by applying the calculated ratio of the loudness of the multiple first sounds as the mixing ratio. Here, the mixing ratio is the ratio for mixing multiple sound signals, and is the ratio used to generate a reference signal to cancel out the multiple output sounds output from multiple speakers in response to the multiple sound signals from the sound acquired by the microphone.

[0011] As explained above, the mixing ratio is specified by the information processing unit. This allows mixing to be performed according to the volume of the first sound output by each speaker, and a reference signal is generated. Then, using the reference signal, the multiple sounds output by the multiple speakers are canceled out from the sound acquired from the microphone.

[0012] In this process, the reference signal mixes multiple sound signals according to the loudness of the first sound. As a result, in the multiple output sounds, the louder the output sound, the greater the proportion of mixing in the reference signal. Therefore, more emphasis can be placed on the louder output sounds among the multiple output sounds. Consequently, multiple sounds output from multiple speakers can be effectively canceled out.

[0013] The following describes specific embodiments of this disclosure with reference to the drawings. Unless otherwise specified, the hardware configurations, module configurations, functional configurations, etc., described in each embodiment are not intended to limit the technical scope of the disclosure to those configurations alone.

[0014] <Embodiment> The speech recognition system 1 in this embodiment will be described based on FIGS. 1 to 2. FIG. 1 is a diagram showing the schematic configuration of the speech recognition system 1 according to this embodiment. The speech recognition system 1 includes an in-vehicle device 100 and an audio device 200. In the speech recognition system 1, the in-vehicle device 100 and the audio device 200 are interconnected by an in-vehicle network via an interface (I / F).

[0015] (In-vehicle device) The in-vehicle device 100 is a device mounted on the vehicle 10. The in-vehicle device 100 provides various services such as providing information to the user in response to the speech of the user boarding the vehicle 10. Also, the in-vehicle device 100 outputs a signal for outputting sound such as music or white noise in the vehicle 10. Specifically, the in-vehicle device 100 includes a source device 11, a microphone 12, and a canceller 13.

[0016] The source device 11 is a device that outputs an audio signal for causing the audio device 200 to reproduce sound such as music. The source device 11 is, for example, a media reading device such as a CD drive. The source device 11 reads a media such as a CD and outputs an audio signal. Also, the source device 11 is a device that outputs information such as music or white noise held in advance as an audio signal. The microphone 12 is a microphone mounted on the vehicle 10. The microphone 12 acquires the speech of the user boarding the vehicle 10.

[0017] The canceller 13 is a device for canceling other sounds from the sound including the speech of the user acquired by the microphone 12. Details of the method by which the canceller 13 cancels other sounds from the sound including the speech of the user will be described later.

[0018] The in-vehicle device 100 is configured to include a computer having a processor 110, a main memory unit 120, and an auxiliary storage unit 130. The processor 110 is, for example, a CPU (Central Processing Unit) or a DSP (Digital Signal Processor). The main memory unit 120 is, for example, a RAM (Random Access Memory). The auxiliary storage unit 130 is, for example, a ROM (Read Only Memory). Also, the auxiliary storage unit 130 is, for example, an HDD (Hard Disk Drive), or a disk recording medium such as a CD-ROM, a DVD disk, or a Blu-ray disk. Further, the auxiliary storage unit 130 may be a removable medium (portable storage medium). Here, examples of the removable medium include a USB memory or an SD card.

[0019] In the in-vehicle device 100, an operating system (OS), various programs, and various information tables, etc. are stored in the auxiliary storage unit 130. Also, in the in-vehicle device 100, the processor 110 loads the program stored in the auxiliary storage unit 130 into the main memory unit 120 and executes it, whereby various functions as described later can be realized. However, some or all of the functions in the in-vehicle device 100 may be realized by a hardware circuit such as an ASIC or an FPGA. Note that the in-vehicle device 100 does not necessarily have to be realized by a single physical configuration, and may be configured by a plurality of computers that cooperate with each other.

[0020] (Audio device) The audio device 200 is a device mounted on the vehicle 10. The audio device 200 outputs sounds such as music using the audio signal received from the in-vehicle device 100. The audio device 200 is configured to include a microcomputer 21, an amplifier 22, a plurality of speakers 23, and a mixer 24. Note that the audio device 200, like the in-vehicle device 100, is configured to include a computer.

[0021] The microcontroller 21 is a device that converts the audio signal received from the vehicle 10 into an output signal for the audio device 200 (speaker 23) to output sound. The microcontroller 21 transmits the converted output signal to the amplifier 22. The amplifier 22 is a device that amplifies the output signal received from the vehicle 10 and outputs it to the speaker 23. In this embodiment, the output signal is a signal whose volume is specified by the magnitude of its voltage.

[0022] Furthermore, speaker 23 is a speaker installed in vehicle 10. Multiple speakers 23 are installed in vehicle 10. Speaker 23 uses the amplified output signal received from amplifier 22 to output sound within vehicle 10. Here, microcontroller 21 sends multiple channels of output signals to the multiple speakers 23 to output sound. The number of output signals sent by microcontroller 21 may be equal to the number of speakers 23, or it may be less than the number of speakers 23. If the number of output signals is less than the number of speakers 23, at least two speakers 23 will output sound using the same output signal.

[0023] In this case, the user of vehicle 10 may speak into the in-vehicle device 100 (microphone 12). In this case, it is assumed that the microphone 12 will acquire not only the speech from vehicle 10 but also sounds such as music output by speaker 23. If this occurs, the sounds such as music output by speaker 23 may interfere with the speech recognition of the vehicle 10 user by the in-vehicle device 100.

[0024] Therefore, it is assumed that the audio device 200 transmits an output signal to the in-vehicle device 100, and the in-vehicle device 100 refers to the output signal and cancels out the sound output by the audio device 200 (speaker 23) from the sound acquired by the microphone 12. On the other hand, in this embodiment, the in-vehicle device 100 and the audio device 200 are connected by their respective interfaces (I / F). At this time, the audio signal transmitted by the in-vehicle device 100 and the signal used to cancel out the sound output by the audio device 200 (speaker 23) are transmitted and received simultaneously between the in-vehicle device 100 and the audio device 200. In this case, there may be insufficient bandwidth for transmitting and receiving signals. Thus, because it is not possible to transmit all of the multiple output signals to the in-vehicle device 100 due to insufficient bandwidth, it becomes necessary to mix at least some of the output signals.

[0025] Figure 2 shows the arrangement of multiple speakers 23 and the in-vehicle device 100 in the vehicle 10. As shown in Figure 2, eight speakers 23 (speaker 23A, speaker 23B, speaker 23C, speaker 23D, speaker 23E, speaker 23F, speaker 23G, and speaker 23H) are installed inside the vehicle 10. The in-vehicle device 100 is installed in the front right of the vehicle 10. The in-vehicle device 100 is installed, for example, in front of the driver's seat of the vehicle 10.

[0026] In this configuration, speakers 23C and 23D are positioned relatively close to each other compared to the other speakers 23. Therefore, the direction from the in-vehicle device 100 to speakers 23C and 23D is similar to that of the other speakers 23. Furthermore, the difference in distance between speakers 23C and 23D and the in-vehicle device 100 is smaller than the difference in distance between the other two speakers 23. Therefore, the positional relationship between speakers 23C and 23D and the in-vehicle device 100 is similar to that between the other speakers 23 and the in-vehicle device 100.

[0027] Therefore, the way in which the sound output by speakers 23C and 23D is reflected by the influence of walls, etc., before reaching the in-vehicle device 100 is more similar than the way in which the sound output by the other speakers 23 is reflected before reaching the in-vehicle device 100. Also, the amount by which the sound output by speakers 23C and 23D is attenuated before reaching the in-vehicle device 100 is more similar than the amount by which the sound output by the other speakers 23 is attenuated before reaching the in-vehicle device 100. In other words, the sound output by speakers 23C and 23D undergoes more similar changes than the other speakers 23 before reaching the in-vehicle device 100. For this reason, in this embodiment, speakers 23C and 23D are treated as a single group.

[0028] Furthermore, in the diagram shown in Figure 2, speakers 23G and 23H are also located relatively close to the other speakers 23. Therefore, speakers 23G and 23H can be treated as a single group. The mixer 24 then mixes the output signals output to speakers 23C and 23D, and the output signals output to speakers 23G and 23H, as one signal each. This allows, for example, eight output signals to be reduced to six signals. In this embodiment, two speakers 23 are treated as one group, but two or more speakers 23 can be treated as a single group. These may be treated as a single group. In this way, multiple speakers 23 located within a predetermined distance from each other are treated as a single group, and the output signals for these multiple speakers 23 are mixed.

[0029] The mixer 24 then outputs a mixed signal (sometimes referred to as the "reference signal") to the in-vehicle device 100. The mixer 24 outputs the output signals of other speakers 23 that do not belong to any of the other output signals as reference signals to the in-vehicle device 100 without mixing them with the other output signals. Specifically, the mixer 24 outputs the reference signal to the canceller 13 via the I / F. This allows the canceller 13 to use the reference signal to cancel out the sound output by the speaker 23. The mixer 24 outputs the reference signal (output signal) to the in-vehicle device 100 in real time.

[0030] At this time, it is necessary to determine the ratio (hereinafter sometimes referred to as the "mixing ratio") by which the audio device 200 mixes the output signals. Therefore, the in-vehicle device 100 causes each speaker 23 belonging to one group to output white noise. The in-vehicle device 100 also uses the microphone 12 to acquire the loudness (sound pressure or volume, etc.) of the white noise output from each speaker 23. Then, the in-vehicle device 100 specifies the ratio of the loudness of the white noise output from each speaker 23 as the mixing ratio of the output signals for each speaker 23.

[0031] With the mixing ratio determined in this way, the reference signal will be mixed with the sounds output by multiple speakers 23 belonging to a single group, according to the loudness of the white noise. In other words, among the sounds output by multiple speakers 23 belonging to a single group, the sound output by the speaker 23 with a louder white noise will be mixed to a larger extent in the reference signal. Therefore, more weight can be given to the sounds with louder volumes among the sounds output by multiple speakers 23 belonging to a single group. As a result, multiple sounds output by multiple speakers 23 belonging to a single group can be effectively canceled out. In addition, the sound output by speakers 23 that do not belong to a group is also output as a reference signal without being mixed. This allows the sound output by speakers 23 that do not belong to a group to also be canceled out.

[0032] (Functional Configuration) Next, the functional configuration of the in-vehicle device 100 that constitutes the voice recognition system 1 will be explained based on Figures 3 and 4. Figure 3 is a block diagram that schematically shows an example of the functional configuration of the in-vehicle device 100.

[0033] The in-vehicle device 100 comprises a control unit 101, an input / output unit 102, and a ratio information database 103 (ratio information DB 103). The control unit 101 has the function of performing calculation processing for controlling the in-vehicle device 100. The control unit 101 can be implemented by the processor 110 in the in-vehicle device 100. The input / output unit 102 has the function of connecting the in-vehicle device 100 to the audio device 200 and inputting and outputting various signals. The input / output unit 102 can be implemented by the interface in the in-vehicle device 100.

[0034] The control unit 101 outputs an audio signal to the audio device 200 via the input / output unit 102. The control unit 101 also outputs an audio signal to the audio device 200 via the input / output unit 102 for each speaker 23 to output white noise. Specifically, the control unit 101 causes the source device 11 to output an audio signal. Here, the audio signal for outputting white noise includes a signal specifying which speaker 23 should output the white noise. The control unit 101 groups these signals into one unit. The control unit 101 outputs an audio signal to cause white noise to be output to the multiple speakers 23 (in the example shown in Figure 2, speakers 23C and 23D, and speakers 23G and 23H). Here, the control unit 101 outputs the audio signal to each speaker 23 sequentially so that the multiple speakers 23 grouped together do not output white noise simultaneously.

[0035] In this case, the control unit 101, for example, causes multiple speakers 23 grouped together to output white noise of the same magnitude. Also, if the control unit 101 has set values ​​such that the volume of sound output to each speaker 23 differs for the same audio signal, it outputs sound according to those set values.

[0036] When the control unit 101 outputs an audio signal to generate white noise, it uses the microphone 12 to acquire the white noise output from each speaker 23. At this time, the control unit 101 acquires the loudness of the white noise output from each speaker 23 that has been acquired by the microphone 12 (hereinafter sometimes simply referred to as "white noise loudness"). The control unit 101 then calculates the ratio of the loudness of the white noise output from each speaker 23. Here, the control unit 101 calculates the ratio of the loudness of the white noise output from multiple speakers 23 belonging to each group.

[0037] The control unit 101 specifies the ratio of the calculated white noise sound intensity as the mixing ratio. The control unit 101 then stores the ratio information indicating the specified mixing ratio in the ratio information DB 103. The ratio information DB 103 has the function of storing ratio information. The ratio information DB 103 can be implemented by the auxiliary storage unit 130 in the in-vehicle device 100. Figure 4 is a diagram showing an example of the table configuration of the ratio information held in the ratio information DB 103.

[0038] As shown in Figure 4, the percentage information includes a group ID field, a speaker ID field, and a percentage field. The group ID field stores an identifier (group ID) to identify the group to which each speaker 23 belongs. The speaker ID field stores an identifier (speaker ID) to identify speaker 23. Here, if multiple speakers belong to one group, multiple speaker IDs are stored in the group ID field corresponding to one group ID.

[0039] The ratio field stores the mixing ratio (the ratio of the loudness of the white noise) calculated by the control unit 101. Here, the mixing ratio for each speaker 23 belonging to a group is normalized so that the sum of the mixing ratios of multiple speakers 23 is 1. In other words, the mixing ratios stored in the multiple ratio fields corresponding to a group all add up to 1.

[0040] The control unit 101 refers to the ratio information and transmits instruction information to the audio device 200 instructing it to mix the speakers 23 belonging to each group in the specified ratio. In other words, the control unit 101 transmits instruction information to the audio device 200 specifying the mixing ratio for each of the multiple speakers 23 belonging to each group. Upon receiving the instruction information, the audio device 200 performs mixing according to the mixing ratio for each group. Specifically, the control unit 101 transmits the instruction information via the interface between the two devices. The audio device 200 then instructs the mixer 24 to mix the output signals according to the specified mixing ratio for each group. This allows the audio device 200 to mix the output signals and enables the in-vehicle device 100 to receive the reference signal.

[0041] The control unit 101 determines whether or not the user of the vehicle 10 has spoken. If the user of the vehicle 10 has spoken, the control unit 101 performs speech recognition to identify the content of the user's speech. At this time, the control unit 101 acquires a reference signal about the sound output by the speaker 23. The control unit 101 then causes the canceller 13 to cancel out the sound output by the speaker 23 from the sound acquired by the microphone 12. Here, canceling out the sound output by the speaker 23 from the sound acquired by the microphone 12 (hereinafter sometimes referred to as "input sound") includes canceling out the sound output by the speaker 23 as noise from the input sound. Furthermore, canceling out the sound output by the speaker 23 from the input sound also includes canceling out the echo of the sound output by the speaker 23 from the input sound.

[0042] At this time, the reference signal acquired by the control unit 101 is the reference signal at the time (period) when the user of the vehicle 10 speaks. In other words, the control unit 101 acquires reference information about the sound output by the multiple speakers 23 when the user of the vehicle 10 speaks, and causes the canceller 13 to cancel out the sound from the multiple speakers 23 from the input sound including the user's speech. At this time, the control unit 101 may amplify or process the reference signal as appropriate before canceling out the sound from the multiple speakers 23. Note that the method by which the control unit 101 cancels out sound using the reference signal can be a known method. For example, the control unit 101 may cancel out the sound output by the multiple speakers 23 by superimposing a sound with a phase opposite to the phase indicated by the reference signal.

[0043] (First process) The first process performed by the control unit 101 in the in-vehicle device 100 in the voice recognition system 1 will be explained with reference to Figure 5. Figure 5 is a flowchart of the first process performed by the control unit 101. The first process shown in Figure 5 is the process of outputting white noise to the speaker 23, determining the mixing ratio, and giving instructions to the audio device 200.

[0044] In the first process, first, in S101, an audio signal that outputs white noise is output to the audio device 200 using the source device 11. Next, in S102, the white noise is acquired using the microphone 12. Next, in S103, the volume of the white noise from each speaker 23 belonging to a group is specified as the mixing ratio. Next, in S105, ratio information indicating the specified mixing ratio is generated and stored in the ratio information DB 103. Next, in S106, instruction information is transmitted to the audio device 200. Then, the first process is completed.

[0045] (Second process) Next, the second process performed by the control unit 101 in the in-vehicle device 100 in the speech recognition system 1 will be explained with reference to Figure 6. Figure 6 is a flowchart of the second process performed by the control unit 101. The second process shown in Figure 6 is a process for speech recognition of the user's speech in the vehicle 10. The second process shown in Figure 6 is executed repeatedly at predetermined intervals. This allows the second process to perform real-time speech recognition of the user's speech in the vehicle 10. The second process starts after the first process has been executed.

[0046] In the second process, first, in S201, it is determined whether or not the microphone 12 has captured speech from the user of vehicle 10. If the determination in S201 is negative, it means that the user of vehicle 10 has not spoken. Therefore, there is no need to perform speech recognition, and the second process is terminated.

[0047] If a positive result is obtained in S201, the input sound is acquired in S202. In S203, a reference signal is acquired. Then, in S204, a cancellation process is performed to cancel out the sound output by speaker 23 from the input sound. Here, in S203, the control unit 101 acquires the reference signal from the audio device 200 at the time (period) when the user of vehicle 10 spoke. Then, the cancellation process is performed using the acquired reference signal, making it possible to cancel out the sound that was output by speaker 23 when the user of vehicle 10 was speaking.

[0048] Next, in S205, speech recognition processing is performed on the sound that has undergone cancellation processing. This makes it possible to perform speech recognition on the sound obtained by the microphone 12, which includes the speech of the vehicle 10 user and the sounds output by the multiple speakers 23, with the sounds output by the multiple speakers 23 canceled out.

[0049] As explained above, a mixing ratio is specified in the speech recognition system 1. This allows mixing to be performed according to the volume of the sounds output by the multiple speakers 23, and a reference signal is generated. Then, using the reference signal, the sounds output by the multiple speakers 23 are canceled out from the sound acquired by the canceller 13. In this way, multiple sounds output from multiple speakers can be effectively canceled out.

[0050] (Variation 1) In this embodiment, the in-vehicle device 100 and the audio device 200 are installed inside the vehicle 10. However, devices similar to the in-vehicle device 100 and the audio device 200 may be installed in locations other than the vehicle 10. For example, devices similar to the in-vehicle device 100 and the audio device 200 may be installed in any location, such as inside or outside a vehicle. This embodiment can also be applied in such cases.

[0051] (Modification 2) In this embodiment, multiple speakers 23 located within a predetermined distance from each other are treated as a single group, and the output signals are mixed. However, in this modified example, all output signals from all speakers 23 in the vehicle 10 may be mixed and output to the in-vehicle device 100 as a single reference signal. In this way, the in-vehicle device 100 can perform cancellation processing using a single reference signal without acquiring multiple reference signals. As a result, the processing load on the in-vehicle device 100 (canceller 13) can be reduced. Even in this way, multiple sounds output from multiple speakers 23 can be effectively canceled out.

[0052] (Variation 3) The mixer 24 may be provided in the in-vehicle device 100. In this case, the microcontroller 21 in the audio device 200 outputs an output signal to the mixer 24 provided in the in-vehicle device 100. The mixer 24 then mixes the output signals and outputs a reference signal to the canceller 13. This makes it possible to cancel out the sounds output by multiple speakers 23 while reducing the number of reference signals that the canceller 13 has to process. As a result, by reducing the processing load on the in-vehicle device 100 (canceller 13), multiple sounds output from multiple speakers 23 can be effectively canceled out.

[0053] (Modification 4) The first process shown in Figure 5 may be configured to be executed when an input to start execution is received for the in-vehicle device 100. In other words, the first process may be configured to be executed by the person who designs the model of the vehicle 10 (designer) or by the user of the vehicle 10 before the vehicle 10 is manufactured. Let's consider the case where the designer has the first process executed. In this case, the designer executes the first process when designing the model of the vehicle 10 before the vehicle 10 is manufactured. By pre-setting the calculated mixing ratio for the audio device 200, it becomes unnecessary to perform the first process again in the same model as vehicle 10, which has the same interior, etc. Here, pre-setting the calculated mixing ratio for the audio device 200 includes pre-storing ratio information indicating the calculated mixing ratio in the ratio information DB 103 and transmitting instruction information to the audio device 200. Furthermore, pre-setting the calculated mixing ratio for the audio device 200 also includes the designer directly specifying the calculated mixing ratio to the audio device 200 (mixer 24) in advance.

[0054] Furthermore, consider the case where the user of vehicle 10 executes the first process. In this case, the user of vehicle 10 may configure the system to execute the first process when the interior of vehicle 10 is changed. Here, changing the interior of vehicle 10 means, for example, changing the material or shape of the seats or upholstery inside vehicle 10. By changing the material or shape of the seats or upholstery inside vehicle 10, the way the sound is reflected inside vehicle 10 when the speaker 23 outputs sound will be different. As a result, it is expected that the sound acquired by the in-vehicle device 100 will change. Also, changing the interior of vehicle 10 may mean changing the speaker 23. It is expected that the sound acquired by the in-vehicle device 100 will change if the model or manufacturer of the speaker 23 is changed. By having the user of vehicle 10 have the in-vehicle device 100 execute the first process at any time, even if the sound acquired by the in-vehicle device 100 changes, multiple sounds output from multiple speakers 23 can be effectively canceled out.

[0055] (Variation 5) In this embodiment, the microphone 12 is provided in the in-vehicle device 100. That is, the microphone 12 is directly connected to the control unit 101. However, the connection configuration between the microphone 12 and the control unit 101 is not limited to this example. In another example, the control unit 101 may be indirectly connected to the microphone 12 via one or more computers, such as relay devices.

[0056] <Other Embodiments> The embodiments described above are merely examples, and this disclosure may be modified as appropriate without departing from its essence. Furthermore, the processes and means described in this disclosure may be freely combined and implemented as long as no technical inconsistencies arise.

[0057] Furthermore, a process described as being performed by a single device may be divided and executed by multiple devices. Conversely, a process described as being performed by different devices may be executed by a single device. In a computer system, the hardware configuration (server configuration) by which each function is implemented can be flexibly changed.

[0058] The present disclosure can also be realized by supplying a computer program implementing the functions described in the embodiments above to a computer, and having one or more processors in the computer read and execute the program. Such a computer program may be provided to the computer by a non-temporary computer-readable storage medium that can be connected to the computer's system bus, or it may be provided to the computer via a network. The non-temporary computer-readable storage medium includes any type of disk, such as magnetic disks (floppy disks or hard disk drives (HDDs), etc.), optical disks (CD-ROMs, DVDs, or Blu-ray discs, etc.), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards, flash memory, or optical cards, and any other type of medium suitable for storing electronic instructions. [Explanation of Symbols]

[0059] 1. Voice recognition system 10. Vehicles 100...In-vehicle equipment 11. Source equipment 12. Microphone 13. Canceller 101. Control Unit 102...Input / output section 103. Percentage Information Database 200 Audio Equipment 21. Microcontroller 22. Amplifier 23 speakers 24 Mixer

Claims

1. To acquire multiple first tones output from multiple speakers, The ratio of the loudness of the multiple first tones obtained is calculated, A mixing ratio is a ratio for mixing multiple sound signals, and this mixing ratio is set by applying the calculated ratio of the magnitudes of the multiple first sounds to a mixing ratio used to generate a reference signal for canceling out the multiple output sounds output from the multiple speakers in response to the multiple sound signals from the sound acquired by the microphone. A control unit configured to perform the following actions: Information processing device.

2. The control unit is further configured to output the set mixing ratio. The information processing apparatus according to claim 1.

3. Outputting the mixing ratio includes instructing the mixer to mix at the mixing ratio. The information processing apparatus according to claim 2.

4. The aforementioned plurality of speakers are arranged within a predetermined distance from each other. The information processing apparatus according to claim 1.

5. The control unit, The input sound, which includes multiple second tones output from the multiple speakers, is acquired via the microphone. The plurality of signals for the output of the plurality of second tones are mixed at the mixing ratio and the resulting signal is obtained as the reference signal, Using the acquired reference signal, the plurality of second tones are canceled out from the input sound, It is configured to perform further actions. The information processing apparatus according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • On-vehicle speech recognition device

    JP2009181025A

  • Echo cancellation method, echo cancellation device, speech processing unit, and program

    JP2018170564A

  • Echo canceler device

    WO2016024345A1