Sound processing device and sound processing program

By forming direct and reflected sound directivities based on distance relationships, the sound processing device enhances sound collection and recognition accuracy in voice processing devices, addressing the inefficiencies of existing technologies.

JP2025099403APending Publication Date: 2025-07-03DENSO CORP +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023216042
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing voice processing devices face a decrease in voice recognition accuracy due to the exclusion of direct sound with relatively small sound pressure when reflected sound is prioritized, leading to inefficient sound collection and recognition.

Method used

The sound processing device forms direct and reflected sound directivities based on distances from the sound source to the microphone array and reflectors, enabling effective collection and recognition of both direct and reflected sounds, even when direct sound pressure is low.

Benefits of technology

This approach improves the Signal Noise Ratio (SNR) of both direct and reflected sounds, ensuring accurate recognition of all sound components, thereby suppressing a decrease in recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099403000001_ABST
    Figure 2025099403000001_ABST
Patent Text Reader

Abstract

To provide a sound processing device and a sound processing program that suppress a degradation in the accuracy of sound recognition.SOLUTION: The sound processing device forms the directivity for direct sounds to enable a microphone array 20 to collect direct sounds, on the basis of the distance from a sound source to the microphone array 20, the distance from the sound source to a reflector, and the distance from the microphone array 20 to the reflector, and forms directivity for reflected sounds to enable the microphone array 20 to collect reflected sounds, on the basis of the distance from the sound source to the microphone array 20 and the distance from the sound source to the reflector. The sound processing device then extracts the direct and reflected sounds from the data of sounds collected by the microphone array 20 while the directivity for direct sounds and the directivity for reflected sounds are formed, and carries out recognition with respect to the direct and reflected sounds.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a sound processing device and a sound processing program.

Background Art

[0002] Conventionally, as described in Patent Document 1, there is known a voice processing device including a sound source separation unit that separates an audio signal input from a sound collection unit into a direct sound from a sound source and a reflected sound, and a voice recognition unit that recognizes the voice of the separated direct sound.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the voice processing device described in Patent Document 1, the reflected sound separated from the sound collected in a room such as a vehicle interior is excluded from the recognition target of the voice recognition unit. However, depending on the direction of the sound source, for example, when the sound source is not facing the sound collection unit directly, the direct sound becomes less likely to reach the sound collection unit than the reflected sound, so the sound pressure of the reflected sound becomes greater than that of the direct sound. As a result, if the reflected sound is excluded from the recognition target of the voice recognition unit as in the voice processing device described in Patent Document 1, the direct sound with a relatively small sound pressure will be recognized, leading to a decrease in voice recognition accuracy.

[0005] An object of the present disclosure is to provide a sound processing device and a sound processing program that suppress a decrease in the recognition accuracy of sound.

Means for Solving the Problems

[0006] The invention according to claim 1 forms a direct - sound directivity for causing a sound collection unit to collect direct sound, which is sound propagating directly from a sound source toward the sound collection unit, based on a value (d) related to the distance from the sound source (12) to the sound collection unit (20), a value (x1) related to the distance from the sound source to a reflector (14) that reflects the sound from the sound source, and a value (x2) related to the distance from the sound collection unit to the reflector. Also, based on the value (d) related to the distance from the sound source to the sound collection unit and the value (x1) related to the distance from the sound source to the reflector, it forms a reflected - sound directivity for causing the sound collection unit to collect reflected sound, which is sound that is reflected by the reflector from the sound source and propagates to the sound collection unit. The sound processing apparatus includes a directivity - forming unit (S106), an extraction unit (S112) that extracts direct sound and reflected sound from the data of the sound collected by the sound collection unit in a state where the direct - sound directivity and the reflected - sound directivity are formed, and a recognition unit (S118) that performs recognition on the direct sound and the reflected sound. Also, the invention according to claim 10 is a sound - processing program that causes a sound - processing apparatus to function as a directivity - forming unit (S106) that forms a direct - sound directivity for causing a sound collection unit to collect direct sound, which is sound propagating directly from a sound source toward the sound collection unit, based on a value (d) related to the distance from the sound source (12) to the sound collection unit (20), a value (x1) related to the distance from the sound source to a reflector (14) that reflects the sound from the sound source, and a value (x2) related to the distance from the sound collection unit to the reflector. Also, based on the value (d) related to the distance from the sound source to the sound collection unit and the value (x1) related to the distance from the sound source to the reflector, it forms a reflected - sound directivity for causing the sound collection unit to collect reflected sound, which is sound that is reflected by the reflector from the sound source and propagates to the sound collection unit. The program also causes the apparatus to function as an extraction unit (S112) that extracts direct sound and reflected sound from the data of the sound collected by the sound collection unit in a state where the direct - sound directivity and the reflected - sound directivity are formed, and a recognition unit (S118) that performs recognition on the direct sound and the reflected sound.

[0007] As a result, since the direct - sound directivity and the reflected - sound directivity are formed, the SNR of the direct sound and the reflected sound is improved. Also, even when the sound pressure of the direct sound is relatively small, the reflected sound is not excluded and is recognized as a signal for the reflected sound. Therefore, the recognition of relatively small sounds is suppressed. Consequently, the decrease in the sound recognition accuracy is suppressed. Note that SNR is the abbreviation of Signal Noise Ratio.

[0008] Note that the reference numerals in parentheses attached to each component etc. show an example of the correspondence relationship between the component etc. and the specific components etc. described in the embodiments described later.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Best Mode for Carrying Out the Invention

[0010] Hereinafter, embodiments will be described with reference to the drawings. In the following embodiments, parts that are the same or equivalent to each other are denoted by the same reference numerals, and the description thereof will be omitted.

[0011] (First Embodiment) The sound processing device of this embodiment suppresses a decrease in the recognition accuracy of sound. This sound processing device is used, for example, in a vehicle. First, this vehicle will be described.

[0012] As shown in FIG. 1, a passenger 12 is on board a vehicle 10. Here, the mouth of the passenger 12 is regarded as a sound source. Further, the vehicle 10 includes a reflector 14 that reflects the voice of the passenger 12. The reflector 14 is an inner wall of a window or a door of the vehicle 10 or the like. Therefore, the reflector 14 includes glass or resin. Further, as shown in FIG. 2, the vehicle 10 includes a microphone array 20, a sensor 25, and a sound processing device 30.

[0013] The microphone array 20 corresponds to a sound collection unit. For example, as shown in FIG. 1, it is arranged in front of the passenger 12. Further, the microphone array 20 has a plurality of arranged microphones, and thus collects the sound in the passenger compartment of the vehicle 10 including the voice of the passenger 12 or the like. Further, returning to FIG. 2, the microphone array 20 outputs the data of the collected sound to a sound processing device 30 described later.

[0014] The sensor 25 has, for example, an image sensor such as a camera. Further, the sensor 25 detects the position and orientation of the sound source, here, the position and orientation of the mouth of the passenger 12, by using image recognition, teaching data, machine learning, and the like. Further, the sensor 25 outputs a signal corresponding to the detected position and orientation of the mouth of the passenger 12 to a sound processing device 30 described later.

[0015] Here, as shown in FIG. 1, let the shortest distance from the mouth of the occupant 12 to the center of the microphone array 20 be the array distance d. Let the distance from the mouth of the occupant 12 to the reflector 14 in the direction orthogonal to the direction from the mouth of the occupant 12 toward the microphone array 20 be the first distance x1. Let the distance from the center of the microphone array 20 to the reflector 14 in the direction orthogonal to the direction from the mouth of the occupant 12 toward the microphone array 20 be the second distance x2. Here, the first distance x1 and the second distance x2 are the distances in the left-right direction of the vehicle 10.

[0016] Returning to FIG. 2, the sound processing device 30 is mainly composed of a microcomputer or the like, and includes a CPU, a ROM, a flash memory, a RAM, an I / O, an A / D converter, and a bus line connecting these components. Further, the sound processing device 30 executes a program stored in the ROM. Thereby, based on the position of the mouth of the occupant 12 detected by the sensor 25, the preset position of the center of the microphone array 20, and the preset position of the reflector 14, the sound processing device 30 calculates the array distance d, the first distance x1, and the second distance x2. Here, the positions of the center of the microphone array 20 and the reflector 14 are preset, but are not limited to being preset, and may be detected by the sensor 25. The sound processing device 30 may calculate the array distance d, the first distance x1, and the second distance x2 based on the positions of the mouth of the occupant 12, the center of the microphone array 20, and the reflector 14 detected by the sensor 25.

[0017] Furthermore, based on the calculated array distances d, the first distance x1, and the second distance x2, the sound processing device 30 forms a direct sound directivity and a reflected sound directivity. Note that the direct sound is the sound that propagates directly from the sound source toward the microphone array 20. The direct sound directivity is the directivity for collecting the direct sound by the microphone array 20. Here, the directivity refers to the characteristic of in which direction the microphone array 20 can collect sound. The reflected sound is the sound that is reflected by the reflector 14 from the sound from the sound source and propagates to the microphone array 20. The reflected sound directivity is the directivity for collecting the reflected sound by the microphone array 20. Also, the number of reflections of the reflected sound is set to three or less here due to the sound attenuation and the difficulty in specifying the direction of the sound. Therefore, the sound when four or more reflections occur is regarded as noise here.

[0018] Furthermore, in the state where the direct sound directivity and the reflected sound directivity are formed, the sound processing device 30 extracts the direct sound and the reflected sound from the sound data collected by the microphone array 20. Also, the sound processing device 30 performs recognition on the extracted direct sound and reflected sound. Note that the details of the processing of the sound processing device 30 will be described later.

[0019] As described above, the vehicle 10 including the sound processing device 30 of the first embodiment is configured. Next, the processing when the program of the sound processing device 30 is executed will be described with reference to the flowchart of FIG. 3. Note that the program of the sound processing device 30 is executed, for example, when the power of the vehicle 10 is turned on. Furthermore, the period of a series of operations from the start of the processing in step S100 of the sound processing device 30 to the return to the processing in step S100 is defined as the control cycle of the sound processing device 30.

[0020] In step S100, the sound processing device 30 acquires the position and orientation of the sound source, here, the position and orientation of the mouth of the occupant 12, from the sensor 25.

[0021] Subsequently, in step S102, the sound processing device 30 compares the position of the mouth of the occupant 12 acquired in step S100 with the position of the mouth of the occupant 12 in the previous control cycle. Thereby, the sound processing device 30 determines whether the position of the sound source has changed.

[0022] For example, when the change amount between the mouth position coordinates of the occupant 12 in the previous control cycle and the mouth position coordinates of the occupant 12 in the current control cycle is equal to or greater than the change threshold, since the position change of the sound source is large, the sound processing device 30 determines that the position of the sound source has changed. At this time, the process of the sound processing device 30 proceeds to step S104. Also, when the change amount between the mouth position coordinates of the occupant 12 in the current control cycle and the mouth position coordinates of the occupant 12 in the previous control cycle is less than the change threshold, since the position change of the sound source is small, the sound processing device 30 determines that the position of the sound source has not changed. At this time, the process of the sound processing device 30 proceeds to step S106. Note that the above change threshold is set by experiments, simulations, etc. so as to determine whether the position of the sound source has changed.

[0023] In step S104 following step S102, since the position of the sound source has changed, in order to re-form the directivity described later, the sound processing device 30 calculates the array distance d, the first distance x1, and the second distance x2.

[0024] Specifically, the sound processing device 30 calculates the array distance d from the position of the mouth of the occupant 12 acquired in step S100 and the position of the center of the microphone array 20 set in advance. Further, the sound processing device 30 calculates the first distance x1 from the position of the mouth of the occupant 12 acquired in step S100 and the position of the reflector 14 set in advance. Also, the sound processing device 30 calculates the second distance x2 from the position of the center of the microphone array 20 set in advance and the position of the reflector 14 set in advance. As described above, the position of the center of the microphone array 20 and the position of the reflector 14 may be detected by the sensor 25. The sound processing device 30 may calculate the array distance d, the first distance x1, and the second distance x2 based on the positions of the mouth of the occupant 12, the center of the microphone array 20, and the reflector 14 detected by the sensor 25.

[0025] In step S106, when the position of the sound source is changed, the sound processing device 30 forms a direct sound directivity and a reflected sound directivity based on the array distance d, the first distance x1, and the second distance x2 calculated in step S104 of the current control cycle. Also, when the position of the sound source has not been changed, the sound processing device 30 uses the direct sound directivity and the reflected sound directivity formed in step S106 of the previous control cycle.

[0026] Specifically, the sound processing device 30 substitutes the array distance d, the first distance x1, and the second distance x2 into the following relational expression (1-1). As a result, as shown in FIG. 4, the sound processing device 30 calculates a direct sound angle Φd, which is an angle related to the directivity for picking up the direct sound. Further, the sound processing device 30 forms a direct sound directivity using the calculated direct sound angle Φd and beamforming such as a delay-and-sum beamformer. Here, the reflected sound is assumed to propagate from the mirror image of the sound source, that is, the mirror image of the occupant 12, to the microphone array 20. Then, the sound processing device 30 substitutes the array distance d and the first distance x1 into the following relational expression (1-2). As a result, the sound processing device 30 calculates a reflected sound angle Φr, which is an angle related to the directivity for picking up the reflected sound. Further, the sound processing device 30 forms a reflected sound directivity using the calculated reflected sound angle Φr and beamforming such as a delay-and-sum beamformer. In FIG. 4, the direct sound angle Φd is set to 0 radians.

[0027] [Number]

[0028] Returning to the flowchart of FIG. 3, subsequently, in step S108, the sound processing device 30 acquires the sound data picked up by the microphone array 20 in a state where the direct sound directivity and the reflected sound directivity in step S106 are formed.

[0029] Subsequently, in step S110, the sound processing device 30 uses SED such as VAD to detect a section in which sound is being emitted from the sound data acquired in step S108. Note that VAD is an abbreviation for Voice Activity Detection. SED is an abbreviation for Sound Event Detection.

[0030] Here, the time when the direct sound reaches the microphone array 20 and the time when the reflected sound reaches the microphone array 20 are different. Further, this time difference varies depending on the direction of the sound source, here, the direction of the mouth of the occupant 12.

[0031] Therefore, in step S112 following step S110, the sound processing device 30 extracts direct sound and reflected sound, as shown in FIG. 5, from, for example, the direction of the sound source acquired in step S100 and the time of the sound section detected in step S110.

[0032] Returning to the flowchart of FIG. 3, subsequently, in step S114, the sound processing device 30 calculates the equivalent sound pressure level of the direct sound extracted in step S112 using the following relational expression (1-3). Further, the sound processing device 30 calculates the equivalent sound pressure level of the reflected sound extracted in step S112 using the following relational expression (1-4). Furthermore, the sound processing device 30 calculates the average of the sound pressures of the direct sound and the reflected sound extracted in step S112, for example, the arithmetic mean. Also, the sound processing device 30 calculates the equivalent sound pressure level of the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound using the following relational expression (1-5). In the following relational expressions, s is the number of samples or time. po is the reference sound pressure, for example, 20 μPa in the case of air. Lp_d is the equivalent sound pressure level of the direct sound. sd1 is the number of samples or time when the direct sound starts. sd2 is the number of samples or time when the direct sound ends. pd(s) is the sound pressure of the direct sound with respect to s. Lp_r is the equivalent sound pressure level of the reflected sound. sr1 is the number of samples or time when the reflected sound starts. sr2 is the number of samples or time when the reflected sound ends. pr(s) is the sound pressure of the reflected sound with respect to s. Lp_ave is the equivalent sound pressure level of the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound. s1 is the number of samples or time when the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound starts. s2 is the number of samples or time when the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound ends. pave(s) is the sound pressure of the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound with respect to s.

[0033]

Equation

[0034] Also, the sound processing device 30 compares the equivalent sound pressure levels of the direct sound, the reflected sound, and the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound calculated above. Thereby, the sound processing device 30 selects the sound for performing sound recognition described later. Specifically, the sound processing device 30 selects the sound with the largest equivalent sound pressure level among the direct sound, the reflected sound, and the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound. For example, when the equivalent sound pressure level of the direct sound is equal to or higher than the equivalent sound pressure level of the reflected sound and the equivalent sound pressure level of the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound, the sound processing device 30 selects the direct sound. Also, when the equivalent sound pressure level of the reflected sound is equal to or higher than the equivalent sound pressure level of the direct sound and the equivalent sound pressure level of the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound, the sound processing device 30 selects the reflected sound. Further, when the equivalent sound pressure level of the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound is equal to or higher than the equivalent sound pressure level of the direct sound and the equivalent sound pressure level of the reflected sound, the sound processing device 30 selects the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound.

[0035] Subsequently, in step S116, the sound processing device 30 performs source separation such as BSS on the sound selected in step S114 using NMF or the like. Thereby, the sound processing device 30 removes the noise included in the sound selected in step S114. Note that NMF is an abbreviation for Nonnegative Matrix Factorization. BSS is an abbreviation for Blind Source Separation.

[0036] Subsequently, in step S118, the sound processing device 30 performs sound recognition on the sound for which source separation was performed in step S116. For example, the sound processing device 30 converts the sound for which source separation was performed in step S116 into character data by using a speech recognition engine or the like. Also, the sound processing device 30 outputs the converted character data to a display (not shown). Thereby, characters corresponding to the sound such as the voice of the occupant 12 in the vehicle interior are displayed on the display (not shown). Thereafter, the processing of the sound processing device 30 returns to step S100.

[0037] As described above, the sound processing device 30 performs processing. Next, suppression of a decrease in sound recognition accuracy by the sound processing device 30 will be described.

[0038] Here, as shown in FIG. 4, an angle formed by a straight line connecting the mouth of the occupant 12 and the center of the microphone array 20 and a straight line extending in the direction in which the occupant 12 emits sound is defined as a sound source angle θ. The direction in which sound is emitted is, for example, the direction in which the main energy of the sound is radiated, or the direction of the component with the highest sound pressure. Also, for example, it is assumed that the array distance d is twice the first distance x1 and the direct sound angle Φd is 0 radians. At this time, as shown in FIG. 6, as the sound source angle θ changes, the equivalent sound pressure levels of the direct sound and the reflected sound change. Furthermore, since there is a range of the sound source angle θ in which the direct sound becomes less likely to reach the microphone array 20 than the reflected sound, there is a range of the sound source angle θ in which the equivalent sound pressure level of the reflected sound becomes greater than the equivalent sound pressure level of the direct sound. As a result, when the reflected sound is excluded from the recognition target of the speech recognition unit as in the speech processing device described in Patent Document 1, the direct sound with a relatively low sound pressure is recognized, so the speech recognition accuracy decreases.

[0039] On the other hand, the sound processing device 30 of the present embodiment forms a direct sound directivity based on the array distance d, the first distance x1, and the second distance x2 in step S106. Also, the sound processing device 30 serves as a directivity forming unit that forms a reflected sound directivity based on the array distance d and the first distance x1. Furthermore, the sound processing device 30 serves as an extraction unit that extracts the direct sound and the reflected sound from the sound data collected by the microphone array 20 in a state where the direct sound directivity and the reflected sound directivity are formed in step S112. Also, the sound processing device 30 serves as a recognition unit that performs recognition on the direct sound and the reflected sound in step S118.

[0040] As a result, since the direct sound directivity and the reflected sound directivity are formed, the SNR of the direct sound and the reflected sound is improved. Further, even when the sound pressure of the direct sound is relatively small, the reflected sound is not excluded and is recognized as a signal for the reflected sound. For this reason, recognition of relatively small sounds is suppressed. Therefore, a decrease in the sound recognition accuracy is suppressed.

[0041] Also, here, as described in Japanese Patent Application Laid-Open No. 2019-176430, a voice recognition device mounted on a vehicle is known. This voice recognition device selects a microphone to be used for voice input based on the sound pressure input to a plurality of microphones and a signal from a seating sensor, and turns on a lamp corresponding to the selected microphone. Further, this voice recognition device causes the speaker to direct their face toward the microphone by lighting the lamp.

[0042] However, the lighting of the lamp may be an obstruction or cause discomfort to a person who wants to enjoy the scenery outside the vehicle or a person who is relaxing in the seat. Also, even if the lamp is lit, the speaker does not always speak toward the microphone.

[0043] On the other hand, in the sound processing device 30 of the present embodiment, since the direct sound directivity and the reflected sound directivity are formed, the SNR is improved regardless of the direction of the sound source. For this reason, it is not necessary to direct the sound source toward the microphone as in the voice recognition device described in Japanese Patent Application Laid-Open No. 2019-176430.

[0044] Also, in the first embodiment, the following effects are also achieved.

[0045] [1-1] The sound processing device 30 serves as a selection unit that selects, in step S114, the sound having the largest equivalent sound pressure level among the direct sound, the reflected sound, and the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound.

[0046] This suppresses recognition of sounds with relatively low sound pressure. Therefore, a decrease in sound recognition accuracy is suppressed. Also, since the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound is used, when the noise is white noise, the noise component is reduced by being smoothed. For this reason, the sound to be recognized is likely to be emphasized.

[0047] [1-2] The occupant 12 corresponding to the sound source and the microphone array 20 are arranged in an interior space such as a vehicle interior.

[0048] As a result, compared with the case where the sound source and the microphone array 20 are arranged outdoors, sound is more likely to be reflected by the reflector 14. For this reason, the microphone array 20 is more likely to pick up the reflected sound. Therefore, a decrease in recognition accuracy regarding the reflected sound is suppressed.

[0049] [1-3] The reflector 14 that reflects sound includes glass and resin. As a result, compared with the case where the reflector 14 is formed of rubber or the like that easily absorbs sound, the reflector 14 is more likely to reflect sound. For this reason, the microphone array 20 is more likely to pick up the reflected sound. Therefore, a decrease in recognition accuracy regarding the reflected sound is suppressed.

[0050] (Modification 1) In the above-described first embodiment, the sound processing device 30 selects the sound having the highest equivalent sound pressure level among the direct sound, the reflected sound, and the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound. In contrast, the sound processing device 30 may select from the direct sound and the reflected sound without using the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound.

[0051] Specifically, when the equivalent sound pressure level of the direct sound is equal to or higher than the equivalent sound pressure level of the reflected sound in step S114, the sound processing device 30 selects the direct sound. Also, when the equivalent sound pressure level of the direct sound is less than the equivalent sound pressure level of the reflected sound, the sound processing device 30 selects the reflected sound. Even in such a form, the same effects as those of the first embodiment are achieved.

[0052] Alternatively, without performing the selection in step S114, the sound processing device 30 may directly perform sound source separation and recognition on each of the direct sound, the reflected sound, and the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound. That is, the process of step S114 may be omitted. Even in such a form, the same effects as those of the first embodiment are achieved.

[0053] (Modification 2) Here, as shown in FIG. 6, there is a range of the sound source angle θ in which the equivalent sound pressure level of the reflected sound is greater than the equivalent sound pressure level of the direct sound. For this reason, in step S114, when the sound source angle θ is equal to or greater than the first threshold value and equal to or less than the second threshold value, the sound processing device 30 performs recognition on the reflected sound. Even in such a form, the same effects as those of the first embodiment are achieved. The first threshold value and the second threshold value are set by experiments, simulations, etc. so that the range of the sound source angle θ in which the equivalent sound pressure level of the reflected sound is greater than the equivalent sound pressure level of the direct sound is set. For example, when the array distance d is twice the first distance x1 and the direct sound angle Φd is 0 radians, the first threshold value is a value near 1 / 4×π radians. The second threshold value is a value near π radians.

[0054] (Second Embodiment) In the second embodiment, the form of the reflector 14 that reflects sound is different from that of the first embodiment. Also, the posture of the sound source is different from that of the first embodiment. Otherwise, it is the same as the first embodiment.

[0055] Instead of the window or the inner wall of the door of the vehicle 10, the reflector 14 is, as shown in FIG. 7, the ceiling or the sunroof of the vehicle 10. For this reason, here, the first distance x1 and the second distance x2 are the distances in the vertical direction of the vehicle 10. Also, since the backrest of the seat of the vehicle 10 is reclined, the occupant 12 corresponding to the sound source is lying down. In FIG. 7, the direct sound angle Φd is set to 0 radians.

[0056] As described above, the sound processing device 30 of the second embodiment is configured. Also in this second embodiment, the same effects as those of the first embodiment are achieved.

[0057] (Third Embodiment) In the third embodiment, the form of the reflector 14 that reflects sound is different from that of the first embodiment. Also, the processing of the sound processing device 30 is different from that of the first embodiment. Otherwise, it is the same as the first embodiment.

[0058] As shown in FIG. 8, the reflector 14 includes a first reflector 141 and a second reflector 142. The first reflector 141 is located at a position relatively close to the occupant 12. The first distance x1 corresponds to the distance from the mouth of the occupant 12 to the first reflector 141 in a direction orthogonal to the direction from the mouth of the occupant 12 toward the microphone array 20. The second distance x2 corresponds to the distance from the center of the microphone array 20 to the first reflector 141 in a direction orthogonal to the direction from the mouth of the occupant 12 toward the microphone array 20. In FIG. 8, the direct sound angle Φd is set to 0 radians.

[0059] The second reflector 142 is located at a position farther from the occupant 12 than the first reflector 141, for example, on the side opposite to the first reflector 141. Here, the distance from the mouth of the occupant 12 to the second reflector 142 in a direction orthogonal to the direction from the mouth of the occupant 12 toward the microphone array 20 is defined as the third distance x3.

[0060] As described above, the sound processing device 30 of the third embodiment is configured. Next, the processing of the sound processing device 30 in the third embodiment will be described.

[0061] The processing from step S100 to step S104 of the sound processing device 30 is performed in the same manner as in the first embodiment described above.

[0062] In step S106, the sound processing device 30 forms the direct sound directivity and the first reflected sound directivity in the same manner as in the first embodiment. Further, the sound processing device 30 forms the second reflected sound directivity. Note that the first reflected sound directivity corresponds to the reflected sound directivity. The first reflected sound is the sound that is reflected from the sound source by the first reflector 141 and propagates to the microphone array 20. The second reflected sound directivity is the directivity for collecting the second reflected sound by the microphone array 20. The second reflected sound is the sound that is reflected from the sound source by the second reflector 142 and propagates to the microphone array 20.

[0063] Specifically, the sound processing device 30 substitutes the array distance d and the third distance x3 into the following relational expression (2). Thereby, as shown in FIG. 8, the sound processing device 30 calculates the second reflected sound angle Φs, which is the angle related to the directivity for collecting the second reflected sound. Further, the sound processing device 30 forms the second reflected sound directivity by using the calculated second reflected sound angle Φs and beamforming such as a delay-and-sum beamformer.

[0064]

Equation

[0065] The processing from step S108 to step S110 following step S106 is performed in the same manner as in the first embodiment.

[0066] In step S112 following step S110, the sound processing device 30 extracts the direct sound, the first reflected sound, and the second reflected sound from the sound section detected in step S110.

[0067] Subsequently, in step S114, the sound processing device 30 calculates the equivalent sound pressure levels of the direct sound and the first reflected sound in the same manner as in the first embodiment. Further, the sound processing device 30 calculates the equivalent sound pressure level of the second reflected sound. Also, the sound processing device 30 calculates the equivalent sound pressure level of the sound obtained by averaging the sound pressure of the direct sound, the sound pressure of the first reflected sound, and the sound pressure of the second reflected sound.

[0068] Furthermore, the sound processing device 30 selects the sound with the highest equivalent sound pressure level among the direct sound, the first reflected sound, the second reflected sound, and the sound obtained by averaging the sound pressure of the direct sound, the sound pressure of the first reflected sound, and the sound pressure of the second reflected sound.

[0069] The processes from step S116 to step S118 following step S114 are performed in the same manner as in the first embodiment.

[0070] As described above, the sound processing device 30 of the third embodiment performs processing. Even in such a third embodiment, the same effects as those of the first embodiment are achieved.

[0071] (Fourth Embodiment) In the fourth embodiment, as shown in FIG. 9, the vehicle 10 further includes an array moving device 35. Also, the processing of the sound processing device 30 is different from that of the first embodiment. Other than these, it is the same as the first embodiment.

[0072] The array moving device 35 has a linear guide, a motor, etc., and changes the position and orientation of the microphone array 20 based on a signal from the sound processing device 30 described later.

[0073] Next, the processing of the sound processing device 30 of the fourth embodiment will be described with reference to the flowchart of FIG. 10.

[0074] In step S100, the sound processing device 30 acquires the positions of the occupant 12 and the reflector 14 from the sensor 25. Here, the position of the reflector 14 is detected by the sensor 25, but it is not limited to this and may be preset.

[0075] In step S120 following step S100, the sound processing device 30 controls the array moving device 35 based on the positions of the occupant 12 and the reflector 14. Thereby, the sound processing device 30 changes the position of the microphone array 20 based on the positions of the occupant 12 and the reflector 14. For example, the sound processing device 30 changes the position of the microphone array 20 to a position where the microphone array 20 can easily pick up the reflected sound.

[0076] In step S104 following step S120, the sound processing device 30 calculates each distance from the sound source acquired in step S100 and the position of the reflector 14, and the position of the microphone array 20 moved in step S120. Each distance is the array distance d, the first distance x1, and the second distance x2.

[0077] The processing from step S106 to step S118 following step S104 is performed in the same manner as in the first embodiment.

[0078] As described above, the sound processing device 30 of the fourth embodiment performs the processing. Even in such a fourth embodiment, the same effects as those of the first embodiment are achieved. Further, in the fourth embodiment, the following effects are also achieved.

[0079] [2] The sound processing device 30 serves as a control unit that changes the position of the microphone array 20 based on the positions of the sound source and the reflector 14.

[0080] Thereby, even if the positions of the sound source and the reflector 14 change, the position of the microphone array 20 can be changed to a position where the microphone array 20 can easily pick up the reflected sound. For this reason, the microphone array 20 can easily pick up the reflected sound. Therefore, a decrease in recognition accuracy regarding the reflected sound is suppressed.

[0081] (Fifth Embodiment) In the fifth embodiment, the form of the reflector 14 is different from that of the first embodiment. Otherwise, it is the same as the first embodiment.

[0082] As shown in FIGS. 11 and 12, the reflector 14 curves convexly toward the outside of the vehicle compartment of the vehicle 10 on the side opposite to the occupant 12 and the microphone array 20, and has a curved surface 145 that reflects sound from a sound source. In this case, the sound processing device 30 calculates the direct sound angle Φd using the curvature of the curved surface 145 in addition to the array distance d, the first distance x1, and the second distance x2. For this reason, in FIGS. 11 and 12, the direct sound angle Φd is set to 0 radians.

[0083] As described above, the sound processing device 30 of the fifth embodiment is configured. Also in this fifth embodiment, the same effects as those of the first embodiment are achieved. Further, in the fifth embodiment, the following effects are also achieved.

[0084] [3] Due to the curved surface 145 of the reflector 14, the sound from the sound source is more likely to be reflected by the reflector 14 than when the reflector 14 is planar. For this reason, the microphone array 20 is more likely to pick up the reflected sound. Therefore, a decrease in the recognition accuracy of the reflected sound is suppressed.

[0085] (Other Embodiments) The present disclosure is not limited to the above-described embodiments, and can be appropriately modified with respect to the above-described embodiments. Also, in each of the above-described embodiments, it goes without saying that the elements constituting the embodiments are not necessarily essential, except in cases where it is explicitly stated that they are essential or when they are considered to be clearly essential in principle.

[0086] The directivity forming unit, extraction unit, recognition unit, selection unit, control unit, etc. and the method described in the present disclosure may be implemented by a dedicated computer provided by configuring a processor and a memory programmed to execute one or more functions embodied by a computer program. Alternatively, the directivity forming unit, extraction unit, recognition unit, selection unit, control unit, etc. and the method described in the present disclosure may be implemented by a dedicated computer provided by configuring a processor with one or more dedicated hardware logic circuits. Or, the directivity forming unit, extraction unit, recognition unit, selection unit, control unit, etc. and the method described in the present disclosure may be implemented by one or more dedicated computers configured by a combination of a processor and a memory programmed to execute one or more functions and a processor configured by one or more hardware logic circuits. Further, the computer program may be stored in a computer-readable non-transitory tangible recording medium as instructions to be executed by a computer.

[0087] In each of the above embodiments, the sound processing device 30 is used for the sound in the vehicle 10. In contrast, the sound processing device 30 is not limited to being used for the sound in the vehicle 10. For example, the sound processing device 30 may be used for the sound in a building such as a house, and the sound source and the microphone array 20 may be arranged indoors in the building.

[0088] In each of the above embodiments, the equivalent sound pressure level is cited as the value related to the sound pressure used when selecting sound. In contrast, the value related to the sound pressure is not limited to the equivalent sound pressure level, and simply the sound pressure or the like may be used.

[0089] In each of the above embodiments, the sensor 25 detects the position of the mouth of the occupant 12 as the sound source position by means of an image sensor such as a camera. In contrast, the means for detecting the sound source position is not limited to an image sensor such as a camera. For example, the sensor 25 includes a seating sensor, a seat position detection sensor, and a reclining angle sensor. The seating sensor detects on which seat the occupant 12 is seated. The seat position detection sensor detects the position coordinates of the seat of the occupant 12. The reclining angle sensor detects the angle of the backrest portion of the seat of the occupant 12. Then, the sensor 25 detects the position of the head of the occupant 12 from the position coordinates of the seat of the occupant 12 and the angle of the backrest portion of the seat detected thereby. Further, the sensor 25 may output the detected position of the head of the occupant 12 to the sound processing device 30 as the sound source position. Furthermore, in this case, since there is an error in the sound source position, the sound processing device 30 may form a direct sound directivity with, for example, the direct sound angle Φd being zero.

[0090] Each of the above embodiments and each modification may be combined as appropriate.

Explanation of Reference Numerals

[0091] 10 Vehicle 12 Sound Source 14 Reflector 20 Microphone Array 30 Sound Processing Device

Claims

1. Based on a value (d) related to the distance from the sound source (12) to the sound collection unit (20), a value (x1) related to the distance from the sound source to a reflector (14) that reflects the sound from the sound source, and a value (x2) related to the distance from the sound collection unit to the reflector, a direct sound directivity for causing the sound collection unit to collect the direct sound that propagates directly from the sound source toward the sound collection unit is formed, and based on the value (d) related to the distance from the sound source to the sound collection unit and the value (x1) related to the distance from the sound source to the reflector, a reflected sound directivity for causing the sound collection unit to collect the reflected sound that is the sound from the sound source reflected by the reflector and propagates to the sound collection unit is formed, a directivity forming unit (S106); An extraction unit (S112) that extracts the direct sound and the reflected sound from the data of the sound collected by the sound collection unit in a state where the direct sound directivity and the reflected sound directivity are formed; A recognition unit (S118) that performs recognition on the direct sound and the reflected sound; A sound processing apparatus comprising the above.

2. The sound processing apparatus selects the direct sound when a value related to the sound pressure of the direct sound is greater than or equal to a value related to the sound pressure of the reflected sound, further comprises a selection unit (S114) that selects the reflected sound when a value related to the sound pressure of the direct sound is less than a value related to the sound pressure of the reflected sound, The recognition unit performs recognition on the sound selected by the selection unit. The sound processing apparatus according to Claim 1.

3. The selection unit selects the sound with the largest value among a value related to the sound pressure of the direct sound, a value related to the sound pressure of the reflected sound, and a value related to the sound pressure of the sound obtained by averaging the sound pressure of the direct sound and the sound pressure of the reflected sound. The sound processing apparatus according to Claim 2.

4. The recognition unit performs recognition on the reflected sound when an angle (θ) formed by a straight line connecting the sound source and the sound collection unit and a straight line extending in the direction in which the sound source emits sound is greater than or equal to a first threshold value and less than or equal to a second threshold value. The sound processing apparatus according to Claim 1.

5. The sound source and the sound collection unit are arranged indoors. The sound processing apparatus according to any one of Claims 1, 2, and 4.

6. The reflector includes glass. The sound processing apparatus according to any one of Claims 1, 2, and 4.

7. The reflector includes resin. The sound processing apparatus according to any one of Claims 1, 2, and 4.

8. The sound processing device according to claim 1, further comprising a control unit (S120) that changes the position of the sound collection unit based on the positions of the sound source and the reflector.

9. The sound processing device according to claim 1, wherein the reflector is curved convexly toward the side opposite to the sound source and the sound collection unit, and has a curved surface (145) that reflects sound from the sound source.

10. A sound processing device, Based on a value (d) related to the distance from the sound source (12) to the sound collection unit (20), a value (x1) related to the distance from the sound source to a reflector (14) that reflects sound from the sound source, and a value (x2) related to the distance from the sound collection unit to the reflector, a direct sound directivity for causing the sound collection unit to collect the direct sound that propagates directly from the sound source toward the sound collection unit is formed, and based on the value (d) related to the distance from the sound source to the sound collection unit and the value (x1) related to the distance from the sound source to the reflector, a reflected sound directivity for causing the sound collection unit to collect the reflected sound that is the sound from the sound source reflected by the reflector and propagates to the sound collection unit is formed (S106), An extraction unit (S112) that extracts the direct sound and the reflected sound from the data of the sound collected by the sound collection unit in a state where the direct sound directivity and the reflected sound directivity are formed, and A sound processing program that functions as a recognition unit (S118) for recognizing the direct sound and the reflected sound.

Citation Information

Patent Citations

  • Audio processing device, audio processing method, and audio processing program

    JP6703460B2