Sound source separation device, sound source separation method, and program

The sound source separation device effectively isolates audio within a target area by subdividing regions with a microphone array and using MVDR beamformers, addressing the challenge of capturing audio from specific areas in environments with multiple sound sources.

JP7761208B2Active Publication Date: 2025-10-28HONDA MOTOR CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022197924
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-10-28
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

Conventional sound source separation methods struggle to extract only the audio signal within a target two-dimensional area, such as when zooming in on a specific person in a shooting environment, as they treat surface sound sources as point sources, leading to the capture of audio from other individuals.

Method used

A sound source separation device and method that uses a microphone array to subdivide a desired two-dimensional region into smaller areas using sub-beams, applying a beamforming method to extract and separate sound signals within these regions, employing a Minimum Variance Distortionless Response (MVDR) beamformer to enhance separation efficiency.

Benefits of technology

Enables the extraction of audio signals within a target two-dimensional region, ensuring that only the desired audio is captured while minimizing interference from other sources, maintaining separation performance and processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007761208000013
    Figure 0007761208000013
  • Figure 0007761208000014
    Figure 0007761208000014
  • Figure 0007761208000015
    Figure 0007761208000015
Patent Text Reader

Abstract

To provide a sound source separation device capable of extracting only a sound signal in a two-dimensional region of an object, a sound source separation method, and a program.SOLUTION: A sound source separation device comprises: a microphone array having M microphones (M is an integer of two or more) arranged in a first interval for collecting a sound of an acoustic signal; and a sound source separation part that segments a desired two-dimensional region by an interval Δθ of an azimuth angle direction and an interval Δφ of an elevation angle direction as a second interval, extracts the acoustic signal collected by the microphone array to each segmented region by separating them by a beam forming method by using a sub-beam corresponding to the secondary region surrounded by the interval Δθ and the interval Δφ of the segmented region, and separates the acoustic signal of a desired two-dimensional source region by adding the extracted acoustic signal. The sound source separation part fixes the number of the sub-beams to a predetermined number, and performs processing.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound source separation device, a sound source separation method, and a program. [Background technology]

[0002] By performing processing such as beamforming on acoustic signals collected by a microphone array, it is possible to perform sound source separation, extracting only specific sound sources from observed signals containing a mixture of multiple sound sources (see, for example, Patent Document 1).

[0003] These sound source separation processes are based on the theory that the sound source is a point source. Since normal sound sources are surface sources, separation processes have traditionally been performed by treating surface sound sources as point sources. Conventionally, methods that use a wide beam (directivity), such as delay-and-sum beamforming and echo cancellation, have been used to simulate sound source separation of surface sound sources.

[0004] Generally, when a video is shot with a video camera or the like, an audio signal is also recorded. The video camera or the like can narrow down the area to be shot using a zoom lens, etc. However, there is a need to extract and record only the audio within the area to be shot. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-46759 Summary of the Invention [Problem to be solved by the invention]

[0006] However, with conventional technology, it was difficult to extract only the audio signal within the target area. For example, if there were five people in a shooting environment and you zoomed in on one person, you could extract only that person from the image, but the audio of the other four people would also be picked up.

[0007] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a sound source separation device, a sound source separation method, and a program that can extract only audio signals within a target two-dimensional area. [Means for solving the problem]

[0008] (1) In order to achieve the above object, a sound source separation device according to one embodiment of the present invention includes a microphone array having M microphones arranged at a first interval (M is an integer greater than or equal to 2) that collect sound signals, and a sound source separation unit that subdivides a desired two-dimensional region at second intervals, i.e., an azimuth interval Δθ and an elevation interval Δφ, and separates and extracts, for each of the subdivided regions, sound signals collected by the microphone array by a beamforming method using sub-beams that correspond to the two-dimensional region bounded by the interval Δθ and the interval Δφ of the subdivided region, and adds up the extracted sound signals to separate the sound signals of the desired two-dimensional region, wherein the sound source separation unit performs processing by fixing the number of sub-beams to a predetermined number.

[0009] (2) In the sound source separation device described in (1) above, a parameter for extracting the desired two-dimensional region is R, and the parameter R satisfies |θ|≦R, |φ|≦R. The sound source separation unit extracts the desired two-dimensional region from the sound source by dividing the desired two-dimensional region by −R in the azimuth angle θ direction. θ is the lower limit and R θ is the upper limit, and -R φ is the lower limit and R φ may be specified as the upper limit.

[0010] (3) In the sound source separation device described in (1) or (2) above, the spatial filter used for sound source separation is expressed by the following formula:

number

[0011] (4) In the sound source separation device according to any one of (1) to (3) above, the sub-beams may be formed by a Minimum Variance Distortionless Response (MVDR) beamformer.

[0012] (5) In the sound source separation device described in (1) above, the number N of two-dimensional regions enclosed by an azimuth angle θ and an elevation angle φ that subdivide the desired two-dimensional region at a second interval may be 10×10 or more.

[0013] (6) In order to achieve the above object, a sound source separation method according to one embodiment of the present invention is a sound source separation method in which an acoustic signal is collected by a microphone array having M (M is an integer greater than or equal to 2) microphones arranged at a first interval, a sound source separation unit subdivides a desired two-dimensional region at a second interval, and for each of the subdivided regions, separates and extracts the acoustic signals collected by the microphone array by a beamforming method using sub-beams corresponding to two-dimensional regions bounded by an azimuth angle θ and an elevation angle φ of the subdivided region, and separates the acoustic signals of the desired two-dimensional region by adding the extracted acoustic signals, and the sound source separation unit performs processing by fixing the number of sub-beams to a predetermined number.

[0014] (7) In order to achieve the above object, a program according to one embodiment of the present invention is a program that causes a computer of a sound source separation device to collect acoustic signals using a microphone array having M (M is an integer greater than or equal to 2) microphones arranged at a first interval, subdivide a desired two-dimensional region at a second interval, separate and extract the acoustic signals collected by the microphone array for each of the subdivided regions using a beamforming method using sub-beams corresponding to two-dimensional regions bounded by an azimuth angle θ and an elevation angle φ of the subdivided region, separate the acoustic signals of the desired two-dimensional region by adding the extracted acoustic signals, and perform processing with the number of sub-beams fixed to a predetermined number. [Effects of the Invention]

[0015] According to the above (1) to (7), it is possible to extract only the audio signal within the target two-dimensional region. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 illustrates an example of a one-dimensional SSB technique. [Figure 2] FIG. 2 is a diagram for explaining a coordinate system. [Figure 3] FIG. 1 is a diagram for explaining the target direction of SSB in a two-dimensional area. [Figure 4] FIG. 1 is a schematic diagram showing an area to be extracted in three-dimensional coordinates. [Figure 5] 1A to 1C are schematic diagrams for explaining a method according to an embodiment. [Figure 6] FIG. 10 is a schematic diagram illustrating a case where zooming out is performed in synchronization with video shooting. [Figure 7] 1A and 1B are conceptual diagrams of images captured and audio signals collected by the sound source separation processing according to the embodiment; [Figure 8] 1 is a diagram illustrating an example of the configuration of a sound source separation system according to an embodiment. [Figure 9] 1 is a flowchart of a processing procedure performed by a sound source separation device according to an embodiment. [Figure 10]FIG. 10 is a diagram showing the arrangement of a microphone array used in evaluation. [Figure 11] FIG. 10 is a diagram showing the results of the average SDR, the number of data below the threshold, and the average processing time for 120 patterns when the fixed value of N is changed. [Figure 12] FIG. 10 is a diagram showing an example of evaluation results of the audio processing speed when the fixed value of N is changed. [Figure 13] FIG. 10 is a diagram showing an example of change in average SDR with respect to region size R. DETAILED DESCRIPTION OF THE INVENTION

[0017] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component is appropriately changed so that each component can be recognized. In all the drawings for explaining the embodiments, the same reference numerals are used for components having the same functions, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX that has been calculated or processed. "XX" is any element (for example, any information).

[0018] [About the Scan-and-Sum Beamformer Method] Beamforming is a signal processing method using an array, and can be applied to sound source separation. However, general beamformers are only capable of separating point sound sources. In contrast to this, the Scan-and-Sum Beamformer (hereinafter referred to as the SSB method) has been proposed as a method for separating area-based surface sound sources (see Reference 1, Japanese Patent Publication No. 2021-197566). In the SSB method, a surface sound source is decomposed into a large number of point sound sources concentratedly distributed within a certain area, and the results of conventional beamforming designed for each point sound source are integrated. In the SSB method, each constituent beamformer is called a sub-beamformer (Figure 1). Figure 1 is a diagram showing an example of a one-dimensional SSB method. In the case of this SSB method, Although it is possible to separate planar sound sources in one-dimensional regions, it does not disclose how to separate sound sources in two-dimensional regions.

[0019] Reference 1; Zhi Zhong, Muhammad Shakeel, Katsutoshi Itoyama, Kenji Nishida, and Kazuhiro Nakadai, “Assessment of a beamforming implementation developed for surface sound source separation”, In 2021 IEEE / SICE International Symposium on System Integration (SII), pp. 369-374, 2021.

[0020] For this reason, in this embodiment, the above-described Scan-and-Sum Beamformer method is extended so that two-dimensional variable regions can be extracted.

[0021] FIG. 2 is a diagram for explaining a coordinate system. As shown in FIG. 2, the horizontal plane is the xy plane, and the vertical direction perpendicular to the xy plane is the z-axis direction. The angle in the xy plane is θ (azimuth angle), and the angle between the xy plane and the z axis is φ (elevation angle). A black circle represents, for example, a sound source. In two-dimensional sound source separation, it is necessary to express the TDoA (Time Difference of Arrival) in all directions in space, so an expression including the elevation angle direction φ is used.

[0022] [2D domain-extended SSB filter design] In this embodiment, R (degrees) is set as a parameter of the variable region to be extracted. Next, in this embodiment, a two-dimensional variable region to be extracted is designated as a parameter R that satisfies the relationship of the following equation (1).

[0023]

number

[0024] The two-dimensional variable region is defined as shown in FIG. 3. FIG. 3 is an image diagram of the extraction region. Symbol o is the origin, which is, for example, the position of the sound pickup unit. In this embodiment, as shown in FIG. 3, the region is scanned in both directions of elevation angle θ and elevation angle φ with a width of Δ, and expanded to two dimensions by taking the sum of all the scans. FIG. 4 is a schematic diagram showing the region to be extracted in three-dimensional coordinates. In FIG. 5, symbol g11 indicates the two-dimensional region to be extracted. The time delay of the microphone sensor position (x, y, z) relative to the origin for a sound source coming from the (θ, φ) direction is expressed by the following equation (2), assuming that the sound is a plane wave. In equation (2), c is the speed of sound.

[0025]

number

[0026] Assuming an environment without reverberation and free space with no amplitude attenuation, and using a microphone array with M microphones, the arrival time difference at the mth microphone relative to the reference point (origin) is τ m Then, the signal from the (θ, φ) direction at the microphone is expressed by the following equation (3).

[0027]

number

[0028] In equation (3), z m (t) represents the observed signal at the mth microphone, and s(t) represents the source signal. In addition, in the SSB design, the target direction of each sub-beamformer is -R θ is the lower limit, and R θ is the upper limit, and in the φ direction, -R φ is the lower limit, and R φ is set as the upper limit and specified at intervals of Δ. i =-R θ +Δ(i-1), φ i =-R φ +Δ(i-1), (θ1,φ1),(θ2,φ1),…,(θ i ,φ j ),…,(θ Nθ ,φ Nφ ) and the target direction of the sub-beamformer. θ , N φ are the numbers of sub-beamformers in the θ and φ directions, respectively. The total number of sub-beamformers is N = N θ N φ When a two-dimensional SSB is constructed using these, the spatial filter becomes the following equation (4).

[0029]

number

[0030] In equation (4), w θi,φi is (θ i,φ j ) is the filter of the sub-beamformer whose target sound source direction is θi,φj is (θ i ,φ j ) is the weight of the sub-beamformer whose target sound source direction is the SSB. Figure 3 is a diagram for explaining the target direction of the SSB in a two-dimensional region. Note that the SSB spatial filter w in one-dimensional region extraction is Σ n=1 N (b n ·w n ) is expressed as w n represents the spatial filter of the nth sub-beamformer, and b n represents the weight for the nth sub-beamformer.

[0031] Next, by substituting equation (4) into equation (3), the output of the beamformer in the frequency domain is expressed as equation (5) below. Note that the superscript H is the Hermitian conjugate. Note that Y(w) represents the Fourier transform of the output y(t) of the beamformer in the time domain.

[0032]

number

[0033] In this embodiment, an SSB is configured using an MVDR (Minimum Variance Distortionless Response) beamformer as a sub-beamformer, for example, as an example of W in equation (5). In this case, the filter of the MVDR beamformer is expressed by the following equation (6). Note that the MVDR beamformer minimizes the output for the entire observed signal while leaving the signal in the target direction undistorted. Furthermore, the MVDR beamformer uses observed values ​​in which the target signal and noise are mixed, and minimizes the output power of the beamformer while ensuring all-pass characteristics in the target sound source direction under constraint conditions, thereby minimizing the power of noise without removing the target signal. Note that the beamformer used for W in equation (5) is not limited to the MVDR beamformer, and other beamformers may also be applied.

[0034]

number

[0035] In equation (6), R is the covariance matrix of the observations, and R=E[zz H ]. E[K] represents the expected value of K. a θ,φ is called the array manifold vector of the used array in the (θ, φ) direction, and is expressed by the following equation (7).

[0036]

number

[0037] When one-dimensional region extraction is performed using the SBM method, the beam pattern for that region is realized by synthesizing multiple sub-beamformers with directionality in the azimuth angle θ direction, as in Reference 1 and Patent Document (JP 2021-197566 A). On the other hand, when the two-dimensional extraction region of this embodiment is performed by the SBM method, Δθ i ×Δφ j By combining the sub-beamformers for each area enclosed by the square, a beam pattern for a desired two-dimensional area can be realized.

[0038] Variable R in the variable domain θ ,R φ By introducing the above, the scanning interval Δ of the sub-beamformer and the number N of the sub-beamformers, which were previously fixed values, become variable. From the definition of the target direction of each sub-beamformer described above, the total number N of sub-beamformers is expressed by the following equation (8).

[0039]

number

[0040] From equation (8), the size of the variable region becomes larger. For this reason, in this embodiment, R θ ,R φ To make Equation (8) valid when becomes large, the number N of sub-beamformers is fixed and the scanning interval Δ of the sub-beamformers is changed. Therefore, in equation (8), N is a constant and Δ is R θ ,R φ is defined as a variable determined by. As shown in Fig. 5, when the extraction region is expanded, Δ increases while N remains constant. Fig. 5 is a schematic diagram for explaining the method according to this embodiment. In Fig. 5, symbol g20 represents the state before expansion, and symbol g30 represents the state after expansion with N fixed. Symbol g21 represents the sound source.

[0041] In addition, when N is solidified, the total number of sub-beamformers does not change even if the extraction region is expanded, so a constant processing time can be expected regardless of the size of the extraction region.

[0042] FIG. 6 is a schematic diagram of zooming out synchronized with video capture. Reference symbol g40 represents the state before zooming out, and reference symbol g50 represents the state after zooming out. Note that FIG. 6 is a simplified, one-dimensional illustration. In the example of FIG. 6, there are a total of 22 sub-beamformers, and the total number of sub-beamformers does not change, as shown by reference symbols g40 and g40. In contrast, the angle of the sub-beamformer scanning interval Δθ is larger after zooming out than before zooming out. Conversely, when zooming in, the angle of the sub-beamformer scanning interval Δθ decreases, as the reference symbol g50 changes to g40. Note that FIG. 6 is a one-dimensional illustration, so only the azimuth angle scanning interval Δθ is shown. However, to extract a two-dimensional region, as in FIGS. 3 and 4, the sub-beamformers scan each area enclosed by a rectangle with the azimuth angle scanning interval Δθ and the elevation angle scanning interval Δφ.

[0043] When equation (3) is Fourier transformed, Z m (w)=e- jωτm Since it can be written as S(w), z(w) = aS(w). Therefore, if we assume free space with no attenuation, the array manifold vector can be said to be the transfer function from the sound source signal to the observed sound. This allows us to formulate the filter design for SSB with two-dimensional domain expansion from the microphone array arrangement and observed sound.

[0044] FIG. 7 is a conceptual diagram of an image captured and an audio signal collected by the sound source separation processing according to this embodiment. When photographing and recording audio of two people (hu1, hu2) having a conversation, as in symbol g60, an image including the first speaker hu1 and the second speaker hu2 is captured, and the first audio signal of the first speaker hu1 and the second audio signal of the second speaker hu2 are recorded. As shown in symbol g70, when two people (hu1, hu2) are having a conversation and the first speaker hu1 is photographed and recorded, an image of the first speaker hu1 is photographed and the voice of the first speaker hu1 is recorded. As described above, according to this embodiment, in conjunction with zooming in or out of an image, it is possible to extract audio signals corresponding to people in the captured image from the collected audio signals.

[0045] [Example of sound source separation system configuration] 8 is a diagram showing an example of the configuration of a sound source separation system according to this embodiment. The sound source separation system 1 includes a sound collection unit 2, a sound source separation device 3, and an imaging unit 4. The sound collection unit 2 includes M (M is an integer of 2 or more) microphones 21-1, ..., microphone 21-N. In the following description, when one of the microphones 21-1, ..., microphone 21-N is not specified, it will be referred to as microphone 21. The sound source separation device 3 includes a sound acquisition unit 31, a transfer function storage unit 32, a beam pattern storage unit 33, a sound source separation unit , an output unit 35, an operation unit , a region control unit 37, and an image acquisition unit . The sound source separation unit 34 includes a separation unit 341, an evaluation unit 342, and a selection unit 343 (evaluation unit).

[0046] The sound source separation system 1 is a device that can simultaneously record video and audio, such as a video camera or a smartphone.

[0047] The image capturing unit 4 is a device that captures moving images or a series of still images. The image capturing unit 4 includes, for example, a zoom lens 41 that is configured with multiple lenses. The image capturing unit 4 varies the image capturing area in accordance with an area instruction output by the area control unit 37. Note that the zoom function is not limited to multiple optical lenses, and may be, for example, a digital zoom that captures images using a portion of the image sensor.

[0048] The sound collection unit 2 is a microphone array including M microphones 21 arranged at a first interval. The sound collection unit 2 collects acoustic signals emitted by a sound source and outputs the collected acoustic signals of m (m is an integer between 2 and M) channels to the voice acquisition unit 31. The position of each microphone 21 is known.

[0049] The operation unit 36 ​​is, for example, a mechanical switch or a software switch that selects or changes the zoom magnification or the shooting area. The operation unit 36 ​​detects the result of the user's operation and outputs the detected result of the operation to the area control unit 37.

[0050] The area control unit 37 generates an area instruction for varying the shooting area and the sound collection area in accordance with the operation result output by the operation unit 36. The area control unit 37 outputs the generated area instruction to the shooting unit 4 and the sound source separation unit 34.

[0051] The sound acquisition unit 31 acquires analog m-channel acoustic signals output by the sound collection unit 2 and converts the acquired analog acoustic signals into digital acoustic signals. The acoustic signals output by each of the m microphones 21 of the sound collection unit 2 are sampled using signals with the same sampling frequency. The sound acquisition unit 31 outputs the converted digital acoustic signals to the sound source separation unit 34.

[0052] The transfer function storage unit 32 stores, for each microphone 21 included in the sound collection unit 2, a transfer function modeled by expressing it as a function that uses the direction of arrival as an argument.

[0053] The beam pattern storage unit 33 may store sub-beam patterns.

[0054] The sound source separation unit 34 separates the acoustic signals of the desired region in accordance with the region instruction output by the region control unit 37, and outputs the separated acoustic signals of the desired region to the output unit 35. The sound source separation unit 34 may evaluate the beam pattern used for the separation. Based on the evaluation result, the sound source separation unit 34 may select the number of microphones 21 and the interval at which the desired region is divided. The desired region is a region that includes a two-dimensional region in which a surface sound source to be separated exists.

[0055] The separation unit 341 subdivides the desired region into a predetermined number (N) of regions at equal intervals. The sound source separation unit 34 uses a sub-beamformer for each subdivided region to extract acoustic signals for the subdivided region from the collected acoustic signal by a beamforming method. The sound source separation unit 34 separates the desired planar sound source by adding the acoustic signals extracted for each subdivided region. The separation unit 341 sets the number of microphones 21 and the intervals for dividing the desired region to initial values ​​stored in the separation unit 341. The separation unit 341 may update the number of microphones 21 and the intervals for dividing the desired region based on the selection result output by the selection unit 343.

[0056] The evaluation unit 342 may evaluate the quality of the selected beam pattern using, for example, a signal-to-distortion ratio (SDR) and a threshold value. The evaluation unit 342 may output the evaluation result to the selection unit 343.

[0057] The selection unit 343 may select the number of microphones 21 and the interval at which the desired region is divided based on the evaluation result obtained by the evaluation unit 342. The selection unit 343 may represent the cost function J, the number of microphones 21, and the interval at which the desired region is divided in a three-dimensional graph, as will be described later, and select the number of microphones 21 and the interval at which the desired region is divided by detecting the minimum value in this graph. The selection unit 343 may output the selected evaluation result to the separation unit 341.

[0058] The output unit 35 outputs the audio signal acquired by the sound source separation unit 34 or the acoustic signal of the desired region separated by the sound source separation unit 34, and the image output by the output unit 35 to an external device. The external device is, for example, a speaker and an image display device.

[0059] The image acquisition unit 38 acquires the image captured by the image capture unit 4 and outputs it to the output unit 35.

[0060] [Example of processing procedure] An example of the processing procedure performed by the sound source separation device 3 will be described. Fig. 9 is a flowchart of the processing procedure performed by the sound source separation device according to this embodiment. Note that the following processing is performed, for example, when video and audio recording is being performed.

[0061] (Step S1) The sound source separation unit 34 acquires an audio signal.

[0062] (Step S2) The sound source separation unit 34 acquires an area instruction.

[0063] (Step S3) The sound source separation unit 34 determines whether the acquired area instruction is a zoom-in instruction. If the acquired area instruction is not a zoom-in instruction (Step S3; NO), the sound source separation unit 34 proceeds to the process of Step S4. If the acquired area instruction is a zoom-in instruction (Step S3; YES), the sound source separation unit 34 proceeds to the process of Step S6.

[0064] (Step S4) The sound source separation unit 34 determines whether the acquired area instruction is a zoom-out instruction. If the acquired area instruction is not a zoom-out instruction (Step S4; NO), the sound source separation unit 34 proceeds to the processing of Step S5. If the acquired area instruction is a zoom-out instruction (Step S4; YES), the sound source separation unit 34 proceeds to the processing of Step S7.

[0065] (Step S5) The image acquisition unit 38 acquires the captured image. After processing, the image acquisition unit 38 proceeds to the processing of step S10.

[0066] (Step S6) The image acquisition unit 38 acquires the image captured by zooming in. After the process, the image acquisition unit 38 proceeds to the process of step S8.

[0067] (Step S7) The image acquisition unit 38 acquires the image captured after zooming out. After the process, the image acquisition unit 38 proceeds to the process of step S10.

[0068] (Step S8) The sound source separation unit 34 subdivides the desired region at second intervals (azimuth angle interval Δθ, elevation angle interval Δφ). That is, the sound source separation unit 34 sets the azimuth angle interval Δθ and the elevation angle interval Δφ of the sub-beamformers that are suited to the desired two-dimensional region.

[0069] (Step S9) If zoom-in is selected, the sound source separation unit 34 separates and extracts the sound signals collected by the microphone array for each subdivided region by beamforming using sub-beams corresponding to the subdivided region. Furthermore, the sound source separation unit 34 separates the sound signals of the desired region by adding the extracted sound signals. If zoom-in is selected, the sound source separation unit 34 performs general sound source localization processing and sound source separation processing to extract the sound sources contained in the collected audio signals. After processing, the sound source separation unit 34 proceeds to processing at step S10.

[0070] (Step S10) The output unit 35 outputs the acquired audio signal and image, or the separated audio signal of the desired two-dimensional area and the zoomed-in image, to an external device.

[0071] The azimuth angle interval Δθ and elevation angle interval Δφ of the sub-beamformers aligned to the desired two-dimensional area may be set to predetermined values ​​or may be selectable. For example, in the following evaluation, a search was conducted under the assumption that Δθ = Δφ, and a value of 4 degrees was used, which resulted in a processing time that was approximately six times faster without degrading performance. Note that Δθ and Δφ may be the same value or different values.

[0072] 9 is an example, and is not limiting. For example, some processes may be performed simultaneously or in parallel, and the order of some process steps may be changed.

[0073] [evaluation] In the evaluation, an MVDR beamformer was used as a sub-beamformer, and the target direction of each sub-beamformer was specified by dividing the extraction area. Observed signals were obtained, and filters were designed using these. In addition, the weights of SSB b θ,φ are all set to 1 and constructed as a simple sum.

[0074] As evaluation indices, performance was evaluated using SDR and threshold value, and the performance of computational processing speed was evaluated using the audio processing time. Here, SDR represents the separation performance of sound source separation. SDR is expressed on a log scale of the ratio of the target sound source component of the extracted sound source to other components, and is given by the following equation (9). Note that the higher the SDR value, the better the separation performance.

[0075]

number

[0076] The threshold is used as an index to detect signals with SDR outliers. In the evaluation, the threshold was set to the mean SDR minus the standard deviation of the SDR, so that results where the performance after separation is significantly lower than the mean value can be detected. The reason for using the threshold as an index of separation performance, in addition to SDR, is that the calculation is performed as the proportion of SSBs that have poor separation performance.

[0077] FIG. 10 is a diagram showing the arrangement of the microphone array used in the evaluation. The microphone array used was a 6-channel spherical microphone array with six microphones arranged at a first interval. The radius of the sphere was 1.5 cm. Note that, although six microphones were used in the microphone array in the evaluation, the number of microphones is not limited to this.

[0078] The source signal used for the evaluation was speech audio data from Libri Speech. Libri Speech is a dataset of approximately 1,000 hours of English speech audio, but it includes many subsets. For the evaluation, we used the development clean dataset from among the subsets. Furthermore, for the evaluation, we used source signals that were divided into approximately 6-second segments. This audio signal includes signals with quiet speech and long pauses. To exclude these, we only used signals whose total absolute amplitude was in the top 30 percent of all audio signals. For the evaluation, we created 10 patterns using this audio signal, randomly positioning two sound sources at positions 0≦|θ|, |φ|≦5 and 30≦|θ|, |φ|≦35.

[0079] When these two sound sources s1(t) and s2(t) are placed at (θ1, φ1) and (θ2, φ2), the Fourier transform Z(w) of the audio signal is expressed as in the following equation (10).

[0080]

number

[0081] By performing an inverse Fourier transform on equation (10), the sound source signal z(t) that constitutes the two sound sources can be generated.

[0082] For simplicity, the parameter representing the size of the region is R θ =Rφ=R as the square area An evaluation was conducted. Using 10 patterns of audio with two sound sources arranged, SSB audio separation was performed with R set to 12 sizes in 5-step increments from 5° to 60°, and the performance was evaluated for each. The fixed N was then varied to determine the average value of 120 patterns (10 patterns x 12 sizes), the number of data below the threshold for the same sound source and same area size, and the average processing time as performance for a fixed value of N.

[0083] (Evaluation results) Figure 11 shows the results of the average SDR, the number of data points below the threshold, and the average processing time for 120 patterns when the fixed value of N is changed. In Figure 11, the horizontal axis represents the fixed value of N, and the vertical axis represents the average SDR and the number of data points below the threshold. Symbol g101 represents the number of data points below the threshold. Line g102 represents the average SDR when the value of N is changed. As shown in Figure 11, only N=5×5 has a low average SDR, and above that, the average SDR increases gradually with N. Data below the threshold also rarely appears above N=10×10. Therefore, the evaluation results show that performance is ensured above N=10×10.

[0084] Figure 12 shows an example of the evaluation results for audio processing speed when the fixed value of N is changed. The horizontal axis of Figure 12 shows the fixed value of N, and the vertical axis shows the average processing time per second. The average processing time increases as N increases. From the above analysis, in order to achieve high processing speed while maintaining performance, it can be said that the evaluation results show that selecting N = 10 × 10 is the best option, as it allows you to keep the value of N as small as possible while maintaining performance.

[0085] 13 is a diagram showing an example of change in average SDR with respect to region size R. The horizontal axis represents region size R (degrees), and the vertical axis represents average SDR. Note that evaluation is performed with N=5×5. As shown in Figure 13, the SDR value attenuates as R increases. As the region expands, the scanning density of the sub-beamformer decreases, resulting in a decrease in separation performance. In the evaluation, the sound sources were positioned at 30 ≤ |θ| and |φ| ≤ 35, so the SDR decreased near 30 degrees. For this reason, it is preferable to set the boundary lines of the region so that the sound sources do not exist on the boundary lines. For example, the sound source separation system 1 may perform sound source localization processing in advance to estimate the sound source direction, and then set the boundary lines so that they do not overlap with the sound source direction.

[0086] In the above-described embodiments, examples of simultaneously recording video and audio have been described, but the present invention is not limited to this. For example, in a device having a camera function and a microphone function, such as a smartphone, it is possible to zoom in on an object whose image you want to capture audio from and record only the separated audio signal.

[0087] In the above example, an example of a signal to be collected and separated is described using an audio signal, but the signal is not limited to this. The signal to be collected and separated may be any acoustic signal, such as the chirping of birds, the chirping of insects, or the sounds of animals.

[0088] Note that a program for realizing all or part of the functions of the sound source separation device 3 of the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the sound source separation device 3. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage providing environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also includes devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that serves as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line.

[0089] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0090] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0091] 1...sound source separation system, 2...sound collection unit, 3...sound source separation device, 21, 21-1,... 21-M...microphone, 4...photographing unit, 31...sound acquisition unit, 32...transfer function memory unit, 33...beam pattern memory unit, 34...sound source separation unit, 35...output unit, 36...operation unit, 37...area control unit, 38...image acquisition unit, 341...separation unit, 342...evaluation unit, 343...selection unit, 41...zoom lens

Claims

1. a microphone array having M microphones (M is an integer equal to or greater than 2) arranged at a first interval to pick up an acoustic signal; an imaging unit that captures moving images or continuous still images and whose imaging area is variable in response to an operation instruction from a user; an operation unit that detects an operation result of the user's operation of selecting or varying a zoom magnification or a shooting area, and outputs the detected operation result; an area control unit that generates an area instruction for varying the shooting area and the sound collection area in accordance with the operation result output by the operation unit, and outputs the generated area instruction; a sound source separation unit that subdivides a desired two-dimensional region at second intervals, Δθ in the azimuth angle direction and Δφ in the elevation angle direction, in accordance with the region instruction output by the region control unit, separates and extracts, for each of the subdivided regions, acoustic signals collected by the microphone array by a beamforming method using sub-beams corresponding to two-dimensional regions surrounded by the intervals Δθ and Δφ of the subdivided region, and adds up the extracted acoustic signals to separate the acoustic signals of the desired two-dimensional region, The sound source separation unit fixes the number of the sub-beams to a predetermined number and separates and extracts the acoustic signals collected by the microphone array using a beamforming method.

2. A parameter for extracting the desired two-dimensional region is R, The parameter R satisfies |θ|≦R, |φ|≦R, The sound source separation unit The desired two-dimensional area is defined by −R θ is the lower limit and R θ is the upper limit, and -R φ is the lower limit and R φ is specified as the upper limit, The sound source separation device according to claim 1 .

3. The spatial filter used for sound source separation is: [Equation 1] w θi,φi is (θ i , φ j ) is a filter of the sub-beamformer with the target sound source direction as the θi,φj is (θ i , φ j ) is the weight of the sub-beamformer with the target sound source direction, The sound source separation device according to claim 1 or 2.

4. The sub-beams are MVDR (Minimum Variance Distortionless Response) beamformers. The sound source separation device according to claim 1 or 2.

5. the number N of two-dimensional regions enclosed by the azimuth angle θ and the elevation angle φ that subdivide the desired two-dimensional region at second intervals is 10 × 10 or more; The sound source separation device according to claim 1 or 2.

6. Acquiring an acoustic signal with a microphone array having M microphones (M is an integer of 2 or more) arranged at a first interval; an operation unit detects an operation result of an operation by the user to select or change a zoom magnification or a shooting area on an imaging unit that captures moving images or successive still images and changes a shooting area in response to an operation instruction from the user, and outputs the detected operation result; an area control unit generates an area instruction for varying the shooting area and the sound collection area in accordance with the operation result output by the operation unit, and outputs the generated area instruction; a sound source separation unit subdivides a desired two-dimensional region at a second interval in accordance with the region instruction output by the region control unit, separates and extracts, for each of the subdivided regions, an acoustic signal collected by the microphone array by a beamforming method using sub-beams corresponding to a two-dimensional region surrounded by an azimuth angle θ and an elevation angle φ of the subdivided region, and separates the acoustic signal of the desired two-dimensional region by adding the extracted acoustic signals; the sound source separation unit fixes the number of the sub-beams to a predetermined number and separates and extracts the acoustic signals collected by the microphone array by a beamforming method. Sound source separation method.

7. The computer of the sound source separation device collecting an acoustic signal with a microphone array having M microphones (M is an integer of 2 or more) arranged at a first interval; Detecting an operation result of a user's operation to select or vary a zoom magnification or a shooting area with respect to a shooting unit that shoots moving images or successive still images and whose shooting area is variable in response to a user's operation instruction, and outputting the detected operation result; generating an area instruction for varying the shooting area and the sound collection area in accordance with the output operation result, and outputting the generated area instruction; subdivide the desired two-dimensional area at a second interval in accordance with the output area instruction, separate and extract, for each of the subdivided areas, the acoustic signals collected by the microphone array by a beamforming method using sub-beams corresponding to two-dimensional areas surrounded by an azimuth angle θ and an elevation angle φ of the subdivided area, and add up the extracted acoustic signals to separate the acoustic signals of the desired two-dimensional area; the number of the sub-beams is fixed to a predetermined number, and the acoustic signals collected by the microphone array are separated and extracted by a beamforming method; program.

Citation Information

Patent Citations

  • Sound collection device

    JP2007013400A

  • Beamforming processor and beamforming method

    JP2015046759A

  • Sound source separation device, sound source separation method, and program

    JP2021197566A