Audio processing device, imaging device, audio processing system, control method and program for audio processing device
Patent Information
- Application Number
- JP2025023348
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2026-08-27
AI Technical Summary
【0012】 本発明によれば、例えば複数の被写体のうちの主被写体が発する声等の音声をできる限り鮮明に記録することができる。
Smart Images

Figure 2026137316000001_ABST
Abstract
Description
Technical Field
[0004] , , , ,
[0003]
[0001] The present invention relates to an audio processing device, an imaging device, an audio processing system, a control method for an audio processing device, and a program.
Background Art
[0002] In recent years, Bluetooth-Low-Energy (hereinafter referred to as "BLE"), which can reduce power consumption compared to conventional Bluetooth (registered trademark), is known. And electronic devices having a wireless transceiver corresponding to this BLE standard have become widespread. In addition, Auracast-Broadcast-Audio (hereinafter referred to as "Auracast (registered trademark)") has been announced by the standardization organization "Bluetooth-Special-Interest-Group". Auracast is a system in which an Auracast receiver can participate in the broadcast transmission from an Auracast transmitter without limitation when the Auracast receiver and the Auracast transmitter are arranged within a predetermined range.
[0003] Non-Patent Document 1 describes an overview of Auracast (described as "Broadcast" in this document). The Auracast described in this Non-Patent Document 1 is composed of an Auracast Transmitter, an Auracast Assistant, and an Auracast Receiver. Hereinafter, the Auracast Transmitter is referred to as a "transmitter", the Auracast Assistant is referred to as an "assistant", and the Auracast Receiver is referred to as a "receiver". The transmitter is a transmitter for transmitting audio data or audio files, as well as advertisements, to unspecified receivers. The audio data or audio files transmitted by the transmitter are provided from a device serving as a sound source connected to the transmitter, a device serving as a sound source incorporated in the transmitter, or a device storing the audio file.
[0004] The assistant is installed in a receiver that scans for advertisements from transmitters. The assistant provides an interface that allows the user of the receiver to receive audio data transmitted from the transmitter once they have selected a scanned source transmitter. Examples of receivers with the assistant include personal computers, tablet devices, smartphones, and other mobile devices. The assistant provides the necessary information to receive audio data or audio files transmitted from the transmitter to another receiver that is pre-paired with the assistant, or to a receiver installed in the same receiver as the assistant. Examples of receivers include audio output devices such as headphones, earphones, and speakers, or recording devices that record audio data or audio files. In this way, audio data or audio files from the transmitter can be received by the receiver.
[0005] Potential applications for Auracast include receiving audio from televisions installed in public areas, and directly receiving announcements at airports and train stations. Another potential application is with imaging devices such as cameras. However, in this case, when filming subjects far from the photographer or through glass, the camera's built-in microphone cannot pick up sounds near the subject. Therefore, by installing a sound pickup device and transmitter near the subject and using the camera to acquire audio data transmitted from the transmitter, it becomes possible to capture immersive video. This makes it suitable for filming events such as sporting events in stadiums, theme parks like zoos and aquariums, and children's recitals and sports days.
[0006] Here, we will explain a conventional shooting method using Auracast in a stadium with reference to Figure 7. Figure 7 is a diagram illustrating a conventional shooting method using Auracast in a stadium. As shown in Figure 7, multiple athletes are playing soccer on the stadium grounds, surrounded by a large number of spectators 790. The photographer 770 wants to receive the voice of one of the athletes, the main subject 780, and the sounds of the game using the imaging device 7100 via Auracast, while simultaneously recording video of the main subject 780 with the imaging device 7100, amidst the cheering of the spectators 790. In addition, multiple microphones 7301 are placed in the stadium. From these microphones 7301, the microphone 7301 closest to the main subject 780 can be selected to record the voice of the main subject 780 and the sounds of the game in video. The assistant and receiver are mounted on the imaging device 7100. The display unit of the imaging device 7100 displays the results of scanning advertisements from the transmitter (see, for example, Figure 3). The photographer 770 can select a microphone closest to the main subject 780 from the scan results and acquire audio data with the imaging device 7100. Although audio data can be acquired in this way, the voice of the main subject 780 and the sounds of the competition included in the audio data may be drowned out by the cheers of the spectators 790. Patent Document 1 describes a method of using multiple directional microphones to target sound from a specific direction and remove sound from other directions. The method described in Patent Document 1 can be applied to the state shown in Figure 7. Specifically, when shooting video with the imaging device 7100, multiple directional microphones can be arranged in an array, and the cheers of the spectators 790, which become noise, can be removed based on the delay time of the voice of the main subject 780 reaching each directional microphone and the difference in the angle of incidence of the sound. This makes it possible to suppress the voice of the main subject 780 and the sounds of the competition from being drowned out by the cheers of the spectators 790. [Prior art documents] [Patent Documents]
[0007] [Patent Document 1] Japanese Patent Publication No. 2016-82432 [Non-patent literature]
[0008] [Non-Patent Document 1] Nick Hunn, "Introducing BluetoothTM LE Audio", p.277-294 [Overview of the project] [Problems that the invention aims to solve]
[0009] However, in the direction the directional microphone is pointed, there are also many spectators 790 beyond the main subject 780, and noise from cheering is generated from multiple spectators 790 simultaneously. Therefore, it is difficult to remove the noise based on the difference in audio delay time and incident angle. As a result, the video that the cameraman 770 expects to be is supposed to be the voice of the main subject 780 and the sounds of the competition, but the actual video is buried under the cheering of the spectators 790.
[0010] The present invention has been made in view of the above-mentioned problems. The present invention aims to provide an audio processing device, an imaging device, an audio processing system, a control method for the audio processing device, and a program that can record sounds such as the voice of a main subject among multiple subjects as clearly as possible. [Means for solving the problem]
[0011] To achieve the above objective, the present invention provides an audio processing device capable of processing sound collected by a plurality of microphones arranged around a target for sound acquisition, comprising: a first information acquisition means for acquiring location information relating to the location of the target for sound acquisition; a second information acquisition means for acquiring location information relating to the location of the audio processing device; a third information acquisition means for acquiring identification information for identifying each of the microphones; and, based on the location information of the target for sound acquisition, the location information of the audio processing device, and the identification information of each of the microphones, selecting from the plurality of microphones the target for sound acquisition The system is characterized by comprising: a determination means for determining a first microphone facing the sound acquisition target and having directionality toward the sound acquisition target, and a second microphone facing the opposite side of the sound acquisition target and having directionality in the same direction as the direction toward the first microphone; a first sound acquisition means for wirelessly acquiring the sound collected by the first microphone as first sound data; a second sound acquisition means for wirelessly acquiring the sound collected by the second microphone as second sound data; and a noise reduction means for performing noise reduction processing to reduce noise contained in the first sound data with the second sound data. [Effects of the Invention]
[0012] According to the present invention, for example, it is possible to record sounds such as the voice of the main subject among multiple subjects as clearly as possible. [Brief explanation of the drawing]
[0013] [Figure 1] This is a block diagram showing an example of the hardware configuration of a voice processing system. [Figure 2] This is a block diagram showing a variation of the hardware configuration of an audio processing system. [Figure 3] This figure shows an example of the usage status of the voice processing system. [Figure 4] Figure 3 illustrates an example of the content of microphone advertisements in the usage state of the audio processing system shown. [Figure 5]This figure shows a variation of the usage state of the voice processing system. [Figure 6] This flowchart shows the processes performed by the imaging device. [Figure 7] This is a diagram illustrating the conventional method of using Auracast for filming in a stadium. [Modes for carrying out the invention]
[0014] Embodiments of the present invention will be described in detail below with reference to the drawings. However, the configurations described in the following embodiments are merely illustrative, and the scope of the present invention is not limited to the configurations described in the embodiments. For example, each part constituting the present invention can be replaced with any configuration that can perform a similar function. In addition, any components may be added.
[0015] Figure 1 is a block diagram showing an example of the hardware configuration of a voice processing system. As shown in Figure 1, the voice processing system 1000A includes an imaging device 100, microphones (hereinafter simply referred to as "microphones") 1301a, 1301b, 1301c, 1301d, and a wireless module 1302. The imaging device 100 includes a control unit 101, an imaging unit (imaging means) 102, a display unit 103, an operating member 104, a recording medium (storage means) 105, a microphone 106, and a wireless module 107. In the voice processing system 1000A, the imaging device 100 can also be, for example, a notebook personal computer, a tablet terminal, or a smartphone. The wireless module 10 includes an assistant 108, a radio wave intensity detection unit 109, and a receiver 110. The assistant 108 and the receiver 110 may be configured separately or as an integrated unit. When the assistant 108 and the receiver 110 are configured as an integrated unit, the configuration will have the function of both the receiver 110 and the assistant 108. The imaging device 100 also includes a storage unit 112, a feature extraction unit 113, a noise reduction unit 114, a GPS 115, a speaker 116, and a motion detection unit 117.
[0016] The control unit 101 is a CPU (computer) that controls each hardware component of the imaging device 100. The imaging unit 102 is configured to be able to capture images using, for example, an optical lens, an image sensor, an imaging signal processing circuit, and the like. Note that the optical lens may be configured to be detachable from the imaging device 100. Also, the image sensor is not particularly limited and is, for example, composed of a photoelectric conversion element such as a CMOS sensor or a CCD. In the imaging unit 102, the light information within the angle of view of the imaging device 100 imaged on the image sensor by the optical lens is photoelectrically converted by the image sensor and then amplified, A / D converted, and compression processed by the signal processing circuit. The image data thus processed is output to the control unit 101. The control unit 101 causes this image data to be displayed on the display unit 103. Thereby, an image within the angle of view of the imaging device 100 can be confirmed. Note that the image may be a moving image or a still image. Also, in addition to the image data being displayed on the display unit 103, for example, various setting items of the imaging device 100 and any transmitter information such as advertisements are also displayed. Thus, the display unit 103 is also used as a Graphical User Interface (GUI).
[0017] Image data is acquired as a moving image or a still image by operating the operating member 104, and written to the recording medium 105 in a file format such as MPEG or JPG for storage. Various programs are also stored on the recording medium 105. These programs include, for example, a program that causes the control unit 101 to execute each process (control method for the audio processing device) described later. The operating member 104 is composed of, for example, a push button, a dial, a slide switch, or a touch panel. The recording medium 105 is composed of non-volatile memory such as flash memory. The recording medium 105 may be fixed to the imaging device 100 or it may be detachable. When recording image data as a moving image, audio data synchronized with the moving image is written to the recording medium 105 for storage. Audio data can be acquired by a microphone 106 or Auracast (registered trademark) described later. Audio data can also be acquired by an external microphone (not shown) connected to a connection terminal (not shown) of the imaging device 100. The microphone 106 and the external microphone may be either a microphone capable of digital output of acquired audio data or a microphone capable of analog output.
[0018] The imaging device 100 is wirelessly connected to microphones 1301a to 1301d via a wireless module 1302. This allows the imaging device 100 to process the audio collected by microphones 1301a to 1301d. Thus, the imaging device 100 functions as an audio processing device. Audio processing will be described later.
[0019] FIG. 2 is a block diagram showing a modified example of the hardware configuration of the voice processing system. The voice processing system 1000B shown in FIG. 2 includes an imaging device 100, a wireless microphone device 1300a, a wireless microphone device 1300b, a wireless microphone device 1300c, and a wireless microphone device 1300d. The wireless microphone device 1300a includes a microphone 1301a and a wireless module 1302a, and these are connected to each other in a wired manner so as to be communicable. Similarly, the wireless microphone device 1300b includes a microphone 1301b and a wireless module 1302b, and these are connected to each other in a wired manner so as to be communicable. Also, the wireless microphone device 1300c includes a microphone 1301c and a wireless module 1302c, and these are connected to each other in a wired manner so as to be communicable. The wireless microphone device 1300d includes a microphone 1301d and a wireless module 1302d, and these are connected to each other in a wired manner so as to be communicable. The imaging device 100 is wirelessly communicably connected to the microphone 1301a via the wireless module 1302a. Similarly, the imaging device 1 is wirelessly communicably connected to the microphone 1301b via the wireless module 1302b. Also, the imaging device 100 is wirelessly communicably connected to the microphone 1301c via the wireless module 1302c. The imaging device 100 is wirelessly communicably connected to the microphone 1301d via the wireless module 1302d.
[0020] The audio processing systems 1000A and 1000B can be used under similar operating conditions. Below, the operating conditions of the audio processing system 1000A will be described as representative. Figure 3 shows an example of the operating conditions of the audio processing system. As shown in Figure 3, the audio processing system 1000A is used in the soccer stadium 3000. The audio processing system 1000A can also be used in facilities other than the soccer stadium 3000 (for example, concert venues, zoos, etc.). In the soccer stadium 3000, a soccer match is played on the soccer field 3001. In addition, a large number of spectators 390 watch the soccer match while cheering in the spectator stands 3002 surrounding the soccer field 3001. The photographer 370 sits in the spectator stands 3002 with the spectators 390 and takes photographs with the imaging device 100. The photographer 370 photographs one player from among the players playing soccer on the soccer field 3001 as the main subject 380. The main subject 380 is also the target sound acquisition object for sound acquisition by the imaging device 100. Microphones 1301a to 1301d are placed around the soccer field 3001 (main subject 380). Microphone 1301a is positioned facing the north side of the spectator stands 3002 and has a directivity that allows it to preferentially acquire sound from that north side. Microphone 1301b is positioned facing the south side of the soccer field 3001 and has a directivity that allows it to preferentially acquire sound from that south side. Microphone 1301c is positioned facing the north side of the soccer field 3001 and has a directivity that allows it to preferentially acquire sound from that north side. Microphone 1301d is positioned facing the south side of the spectator stands 3002 and has a directivity that allows it to preferentially acquire sound from that south side. The directivity of microphones 1301a to 1301d is indicated by the direction pointed by the arrow placed next to the microphone. In the configuration shown in Figure 3, the audio processing system 1000A has microphones 1301a to 1301d, as well as microphones 1301e to 1301h and microphones 1301i to 1301l, which are arranged in the same way as microphones 1301a to 1301d. Note that the number of microphones is not limited to the number shown in the configuration in Figure 3.
[0021] Here, the method for acquiring audio data using Auracast will be explained using the usage state of the audio processing system 1000A shown in Figure 3 as an example. Microphones 1301b and 1301c are main audio microphones (first microphones) for collecting audio from the main subject 380, respectively. Microphones 1301a and 1301d are noise reference microphones (second microphones) for reducing noise contained in the audio data (audio signal) acquired by microphones 1301b and 1301c, respectively. The audio data from microphones 1301a to 1301d is assigned a code to identify each microphone. The wireless module 1302 is a transmitter. The wireless module 1302 advertises information from microphones 1301a to 1301d and transmits audio data to the imaging device 100. The advertisement includes, for example, the existence and name of Auracast. Other information includes the IDs of microphones 1301a to 1301d, microphone location information, microphone type, microphone directivity, and microphone pair information.
[0022] In the imaging device 100, the control unit 101 controls the assistant 108 of the wireless module 107 to perform a scan for advertisements from the wireless module 1302. The results of this scan are displayed on the display unit 103. In this embodiment, it is assumed that the assistant 108 is capable of receiving advertisements from microphones 1301a to 1301d. The contents of the advertisements are temporarily stored in the storage unit 112, which is composed of volatile memory and non-volatile memory.
[0023] In this embodiment, if the photographer 370 intends to acquire audio data via Auracast and the imaging device 100 is configured to enable such audio data acquisition, the audio data from the wireless module 1302 is acquired by the receiver 110. The receiver 110 is paired with the assistant 108 via a wired connection or wirelessly (BLE), and they can communicate with each other. The receiver 110 and the assistant 108 are assumed to be connected via a wired connection. In order for the receiver 110 to acquire audio data from the wireless module 1302, information necessary to acquire one audio data from the audio data received by the assistant 108 from the wireless module 1302 is provided. Based on this information, the receiver 110 can acquire audio data from the wireless module 1302. Note that the receiver 110 can only acquire audio data from microphones 1301a to 1301d received by the assistant 108. Therefore, it is necessary to automatically or manually select one pair of microphones (one set) from among microphones 1301a to 1301d that were able to receive advertisements by Assistant 108.
[0024] When selecting a pair of microphones, the system detects the main subject 380 within the field of view of the imaging device 100 and acquires the position information of the main subject 380. The method for acquiring this position information is not particularly limited and may include methods based on the position information of the imaging device 100, the lens direction of the imaging device 100, the distance from the imaging device 100 to the main subject 380, etc. Another method may involve processing image data of the entire soccer stadium 3000 and using the characteristics of the soccer stadium 3000 extracted from the processing results. The acquired position information is then stored, for example, by downloading via wired cable connection or wireless communication connection, or on portable media such as memory.
[0025] Figure 4 is a diagram illustrating an example of the content of a microphone advertisement in the usage state of the audio processing system shown in Figure 3. In Figure 4, "ID001" is the identification code for microphone 1301a. In addition to the identification code, the identification information for identifying microphone 1301a includes location information regarding the position of microphone 1301a, type information regarding the type of microphone 1301a, directional information regarding the directivity of microphone 1301a, and pair information. The location information for microphone 1301a is "North Side," indicating that it is located on the north side of soccer field 3001. The type information for microphone 1301a is "Noise Reference Microphone." The directional information for microphone 1301a is "North," indicating that it has a directivity that can preferentially acquire sound from the north. The pair information for microphone 1301a is "Microphone Pair A." Microphone 1301a forms a pair with another microphone that has the pair information "Microphone Pair A." In this embodiment, the microphone that forms a pair with microphone 1301a is microphone 1301c.
[0026] "ID002" is the identification code for microphone 1301b. The location information for microphone 1301b is "North Side," indicating that it is located on the north side of the soccer field 3001. The type information for microphone 1301b is "Main Audio Microphone." The directional information for microphone 1301b is "South," indicating that it has a directional pattern that can preferentially acquire sound from the south. The pair information for microphone 1301b is "Microphone Pair B." Microphone 1301b forms a pair with another microphone that has the pair information "Microphone Pair B." In this embodiment, the microphone that forms a pair with microphone 1301b is microphone 1301d. "ID003" is the identification code for microphone 1301c. The location information for microphone 1301c is "South Side," indicating that it is located on the south side of the soccer field 3001. The type information for microphone 1301c is "Main Audio Microphone." The directional information for microphone 1301c is "North," indicating that it has a directional pattern that allows it to preferentially acquire sound from the north. The pair information for microphone 1301c is "Microphone Pair A," the same as the pair information for microphone 1301a. "ID004" is the identification code for microphone 1301d, which is a noise reference microphone placed on the south side of soccer field 3001. The location information for microphone 1301d is "South Side," indicating that it is placed on the south side of soccer field 3001. The type information for microphone 1301d is "Noise Reference Microphone." The directional information for microphone 1301d is "South," indicating that it has a directional pattern that allows it to preferentially acquire sound from the south. The pair information for microphone 1301d is "Microphone Pair B," the same as the pair information for microphone 1301b. In this embodiment, the identification information for microphones 1301a to 1301d includes, but is not limited to, identification code, position information, species information, directional information, and pair information; at least one of these pieces of information is sufficient. Furthermore, the contents shown in Figure 4 may be displayed on the display unit 103 of the imaging device 100.
[0027] Next, the operation (first operation) of selecting a suitable pair of microphones from the pair of main audio microphones and noise reference microphones will be described. First, the control unit 101 identifies the main subject 380 intended by the photographer 370 based on the photographer's line of sight direction and the focus position of the imaging device 100. The control unit 101 acquires the position information of the main subject 380 based on the position of the imaging device 100, the distance from the imaging device 100 to the main subject 380, and the direction of the main subject 380 relative to the imaging device 100. The position information of the main subject 380 can also be acquired by analyzing the image surrounding the main subject 380 by performing image recognition processing using machine learning in the feature extraction unit 113, and based on the results of the analysis. The assistant 108 also receives advertisement information associated with microphones 1301a to 1301d and acquires the position information of each microphone mentioned above based on the advertisement information. Then, based on the position information of the main subject 380 and the position information of each microphone, the microphone facing the main subject 380 and having a directivity toward the main subject 380 is selected from among microphones 1301a to 1301d as the main audio microphone. In particular, among the microphones that could be the main audio microphone, the microphone closest to the main subject 380 (microphone 1301c in the state shown in Figure 3) is selected as the main audio microphone. On the other hand, similar to the selection of the main audio microphone, the noise reference microphone is selected based on the position information of each microphone. Specifically, the microphone facing the opposite side of the main subject 380 and having a directivity in the same direction as the main audio microphone (microphone 1301a in the state shown in Figure 3) is selected as the noise reference microphone.
[0028] Figure 5 shows a modified example of the usage state of the audio processing system. Here, we will mainly explain the differences from the usage state shown in Figure 3, and similar matters will be omitted from the explanation. As shown in Figure 5, microphones 1301a to 1301d are placed around the soccer field 3001. Microphone 1301a is positioned facing northeast of the spectator seats 3002 and has a directivity that can preferentially acquire sound from that northeast direction. Microphone 1301b is positioned facing southwest of the soccer field 3001 and has a directivity that can preferentially acquire sound from that southwest direction. Microphone 1301c is positioned facing northeast of the soccer field 3001 and has a directivity that can preferentially acquire sound from that northeast direction. Microphone 1301d is positioned facing southwest of the spectator seats 3002 and has a directivity that can preferentially acquire sound from that southwest direction.
[0029] Next, the operation (second operation) of selecting a suitable pair of main audio microphones and noise reference microphones in the modified example shown in Figure 5 will be described. First, the control unit 101 identifies the main subject 380 intended by the photographer 370 based on the photographer's line of sight direction and the focus position of the imaging device 100. The control unit 101 acquires the position information of the main subject 380 based on the position of the imaging device 100, the distance from the imaging device 100 to the main subject 380, and the direction of the main subject 380 relative to the imaging device 100. The GPS 115 also acquires the position information of the imaging device 100. The assistant 108 receives advertisement information associated with microphones 1301a to 1301d and acquires the position information of each microphone based on the advertisement information. Then, based on the position information of the main subject 380 and the position information of each microphone, one microphone from microphones 1301a to 1301d is selected as the main audio microphone. Furthermore, the radio wave strength of microphones 1301a to 1301d, which receive advertisements from assistant 108, is detected by the radio wave strength detection unit 109. Each radio wave strength is used to estimate the distance from the imaging device 100 to microphones 1301a to 1301d, i.e., the positions of microphones 1301a to 1301d. This provides positional information for microphones 1301a to 1301d. For example, based on each radio wave strength and the advertisement information, it is possible to determine the microphone closest to the imaging device 100 (microphone 1301c in the state shown in Figure 3) as the main audio microphone. On the other hand, a microphone facing the opposite side of the main subject 380 and having the same directivity as the direction the main audio microphone is facing is determined as the noise reference microphone. Specifically, among microphones 1301a and 1301d, which can serve as noise reference microphones, the microphone located on a straight line passing between microphone 1301c, which is the main audio microphone, and the main subject 380 (microphone 1301a in the state shown in Figure 3) is determined to be the noise reference microphone.
[0030] Figure 6 is a flowchart showing the processing performed by the imaging device. As shown in Figure 6, in step S601, the control unit 101 of the imaging device 100 determines whether the power supply of the imaging device 100 has been turned ON by an operation on the operating member 104. If the control unit 101 determines that the power supply has been turned ON as a result of the determination in step S601, the process proceeds to step S602. On the other hand, if the control unit 101 determines that the power supply is not ON, i.e., OFF as a result of the determination in step S601, the process remains in a waiting state at step S601.
[0031] In step S602, the control unit 101 activates the imaging unit 102. This makes imaging possible with the imaging unit 102.
[0032] In step S603, the control unit 101 displays the live image captured by the imaging unit 102 on the display unit 103. This live image is maintained unless the display on the display unit 103 is switched by operating the operating member 104.
[0033] In step S604, the control unit 101 controls the assistant 108 to perform a scan of external devices being advertised. Then, the control unit 101 temporarily stores at least one external device in the storage unit 112 according to the result of this process. Note that the external devices that can be received by the assistant 108 include not only Auracast-compatible devices but also devices that do not support Auracast. In this embodiment, we will focus on Auracast-compatible devices and omit devices that do not support Auracast. Therefore, in this embodiment, the external devices that can be received by the assistant 108 are microphones 1301a to 1301d, which are Auracast-compatible devices.
[0034] In step S605, the control unit 101 controls the radio wave strength detection unit 109 to detect the strength of the radio waves transmitted from microphones 1301a to 1301d that were scanned in step S604 (third information acquisition step). As mentioned above, each radio wave strength is also used as location information (identification information) for microphones 1301a to 1301d. Furthermore, each radio wave strength is associated with microphones 1301a to 1301d and temporarily stored in the storage unit 112. Thus, in this embodiment, the radio wave strength detection unit 109 functions as a means (third information acquisition means) for wirelessly acquiring location information for microphones 1301a to 1301d.
[0035] In step S606, the control unit 101 determines whether or not the acquisition of audio data by Auracast is selected as enabled on the screen (not shown) displayed on the display unit 103. If the control unit 101 determines that audio data acquisition is selected as enabled as a result of the determination in step S606, the process proceeds to step S607. On the other hand, if the control unit 101 determines that audio data acquisition is not selected as enabled, that is, that audio data acquisition is disabled as a result of the determination in step S606, the process proceeds to step S619. Note that if audio data acquisition is selected as enabled, it reflects the photographer 370's intention to acquire audio data by Auracast. In this case, sound collection by the microphone 106 of the imaging device 100 is restricted, and sound collection by microphones 1301a to 1301d becomes possible. Conversely, if audio data acquisition is disabled as a result of the determination, sound collection by microphones 1301a to 1301d is restricted, and sound collection by microphone 106 of the imaging device 100 becomes possible.
[0036] In step S607, the control unit 101 determines the number of microphones scanned in step S604. In this embodiment, the determination in step S607 is whether the number of microphones is 2 or more. If the control unit 101 determines that the number of microphones is 2 or more as a result of the determination in step S607, the process proceeds to step S608. On the other hand, if the control unit 101 determines that the number of microphones is not 2 or more, i.e., that the number of microphones is less than 2, the process proceeds to step S619. Note that if the number of microphones is less than 2, the noise reduction process described later becomes difficult. Also, if the number of microphones is less than 2, this fact may be displayed on the display unit 103. This allows the photographer 370 to understand that noise reduction processing in the imaging device 100 is difficult and to consider countermeasures.
[0037] In step S608, the control unit 101 controls the GPS 115 to acquire positional information regarding the position of the imaging device 100 (second information acquisition step). Thus, in this embodiment, the GPS 115 functions as a means for acquiring positional information of the imaging device 100 (second information acquisition means).
[0038] In step S609, the control unit 101 detects the main subject 380 based on, for example, the direction of the photographer's line of sight or the focus position of the imaging device 100.
[0039] In step S610, the control unit 101 acquires positional information regarding the position of the main subject 380 detected in step S609 (first information acquisition step). Specifically, the control unit 101 acquires positional information of the main subject 380 based on, for example, the positional information of the imaging device 100, the lens direction of the imaging device 100, and the distance from the imaging device 100 to the main subject 380. Thus, in this embodiment, the control unit 101 also functions as a means for acquiring positional information of the main subject 380 (first information acquisition means). Note that the imaging device 100 may be provided with means for acquiring positional information of the main subject 380 separately from the control unit 101. Also, in the flowchart shown in Figure 6, the acquisition of positional information of microphones 1301a to 1301d, the acquisition of positional information of the imaging device 100, and the acquisition of positional information of the main subject 380 are performed in this order, but are not limited to this order. For example, the acquisition of the position information of the main subject 380, the acquisition of the position information of the imaging device 100, and the acquisition of the position information of microphones 1301a to 1301d may be performed in this order.
[0040] In step S611, the control unit 101 acquires identification information (microphone information) for all microphones (microphones 1301a to 1301d) that were scanned in step S604.
[0041] In step S612, the control unit 101 determines whether to use an automatic selection method or a manual selection method to select the main audio microphone and the noise reference microphone from microphones 1301a to 1301d. The "automatic selection method" is a method in which the control unit 101 makes the selection automatically. The "manual selection method" is a method in which the selection is made by operating on the screen displayed on the display unit 103. The determination in step S612 is made based on the selection result on a selection screen displayed on the display unit 103, which includes the options "automatic selection method" and "manual selection method". If the control unit 101 determines in step S612 that the manual selection method is used, the process proceeds to step S613. On the other hand, if the control unit 101 determines in step S612 that the automatic selection method is used, the process proceeds to step S615.
[0042] In step S613, the control unit 101 displays the identification information acquired in step S611 on the display unit 103. The identification information displayed on the display unit 103 is not particularly limited and can be, for example, the content shown in Figure 4.
[0043] In step S614, if the control unit 101 performs an operation to select a main voice microphone and a noise reference microphone from the identification information displayed in step S613, i.e., microphones 1301a to 1301d, it temporarily stores each microphone in the storage unit 112. This determines the main voice microphone and the noise reference microphone (determination step). Thus, in this embodiment, the display unit 103 also functions as a selection operation means for performing an operation to select a main voice microphone and a noise reference microphone. Furthermore, the control unit 101 also functions as a determination means for determining the main voice microphone and the noise reference microphone based on the selection result of the selection operation means. After step S614 is executed, the process proceeds to step S616.
[0044] In step S615, the control unit 101 selects a main audio microphone and a noise reference microphone from among microphones 1301a to 1301d based on the position information of microphones 1301a to 1301d, the position information of the imaging device 100, and the position information of the main subject 380. The control unit 101 then temporarily stores each selected microphone in the storage unit 112. This determines the main audio microphone and the noise reference microphone (determination step). Thus, in this embodiment, the control unit 101 also functions as a selection means for selecting the main audio microphone and the noise reference microphone. Furthermore, the control unit 101 also functions as a determination means for determining the main audio microphone and the noise reference microphone based on the selection result of the selection means. After step S615 is executed, the process proceeds to step S616. In step S616 and beyond, it is assumed that microphone 1301c is determined as the main audio microphone and microphone 1301a is determined as the noise reference microphone.
[0045] In step S616, the control unit 101 determines whether the microphone selection in steps S614 and S615 was performed correctly. If the control unit 101 determines that the microphone selection was performed correctly as a result of the determination in step S616, the process proceeds to step S618. On the other hand, if the control unit 101 determines that the microphone selection was not performed correctly, that is, that the microphone selection was abnormal, the process proceeds to step S617.
[0046] In step S617, the control unit 101 displays on the display unit 103 that there was an abnormality in the microphone selection. In this case, sound collection by the microphone 106 becomes possible. After step S617 is executed, the process proceeds to step S619.
[0047] In step S618, the control unit 101 controls the receiver 110 of the wireless module 107 to acquire the sound collected by the microphone 1301c as first audio data from the wireless module 1302 (first audio acquisition step). Following this acquisition, the control unit 101 acquires the sound collected by the microphone 1301a as second audio data from the wireless module 1302 (second audio acquisition step). Thus, in this embodiment, the receiver 110 functions as both a first audio acquisition means for acquiring the first audio data and a second audio acquisition means for acquiring the second audio data. The control unit 101 then performs noise reduction processing to reduce (remove) the noise contained in the first audio data using the second audio data (noise reduction step). Thus, in this embodiment, the control unit 101 also functions as a noise reduction means for reducing noise. The noise reduction processing is performed using a known method. For example, noise reduction processing is performed using the difference between the sound pressure level L2 at a distance D2 from the noise source (typically referred to as "source 302" in Figure 3) to microphone 1301c and the sound pressure level L1 at a distance D1 from source 302 to microphone 1301a. Specifically, the logarithmic attenuation rate δ satisfies L1-L2=-20log(D2 / D1). In noise reduction processing, noise can be corrected and reduced using this relationship.
[0048] In step S619, the control unit 101 determines whether or not the recording operation is set using the operating member 104. If, as a result of the determination in step S619, the control unit 101 determines that the recording operation is set, that is, the recording operation setting is ON, the process proceeds to step S620. As a result, the imaging unit 102 can capture a video including the main subject 380, that is, a video of the soccer match. On the other hand, if, as a result of the determination in step S619, the control unit 101 determines that the recording operation is not set, that is, the recording operation setting is OFF, the process proceeds to step S624.
[0049] In step S620, the control unit 101 acquires the first audio data, which has undergone noise reduction processing in step S618, and the video of the soccer match captured by the imaging unit 102.
[0050] In step S621, the control unit 101 records the first audio data and soccer match images acquired in step S620 onto the recording medium 105.
[0051] In step S622, the control unit 101 determines whether or not a recording stop instruction has been issued by the operating member 104. If, as a result of the determination in step S622, the control unit 101 determines that a recording stop instruction has been issued, that is, the recording stop instruction is in the ON state, the process proceeds to step S623. On the other hand, if, as a result of the determination in step S622, the control unit 101 determines that there is no recording stop instruction, that is, the recording stop instruction is in the OFF state, the process returns to step S606 and the subsequent steps are executed in order.
[0052] In step S623, the control unit 101 stops recording the first audio data and the soccer match image.
[0053] In step S624, the control unit 101 determines the state of the recording stop switch operated by the operating member 104. If the control unit 101 determines that the recording stop switch is in the ON state as a result of the determination in step S624, the process proceeds to step S625. On the other hand, if the control unit 101 determines that the recording stop switch is in the OFF state as a result of the determination in step S624, the process returns to step S603 and executes the subsequent steps in order.
[0054] In step S625, the control unit 101 determines whether the power to the imaging device 100 has been turned OFF by the operation of the operating member 104. If the control unit 101 determines that the power has been turned OFF as a result of the determination in step S625, the process ends. On the other hand, if the control unit 101 determines that the power has not been turned OFF, i.e., that it is ON as a result of the determination in step S625, the process returns to step S603 and the subsequent steps are executed in order.
[0055] As described above, the imaging device 100 can process the audio, such as the voice of the main subject 380, to be as clear as possible through the noise reduction processing in step S618, that is, to reduce noise in the audio as much as possible. As a result, clear audio can be recorded, and thus the clarity of the audio is maintained during playback. In this embodiment, the third information acquisition means, the first audio acquisition means, and the second audio acquisition means described above are collectively composed of a unit having an assistant 108 and a receiver 110 in Auracast®-Broadcast-Audio. In step S606, it is possible to switch the enabled or disabled acquisition by this unit (switching operation means). This allows for quick switching between sound collection by microphones 1301a to 1301d and sound collection by microphone 106 of the imaging device 100.
[0056] Furthermore, if the main audio microphone and the noise reference microphone are predetermined, the processing from steps S610 to S617 may be omitted. It is also preferable to extract the characteristics of the main subject 380 in consideration of the case where the main subject 380 moves out of the field of view. This allows the main subject 380 to be tracked by adjusting the orientation of the imaging device 100, and thus the determined state of the main audio microphone and the noise reference microphone can be maintained. When the main subject 380 moves out of the field of view, the motion detection unit 117 (Figure 1) of the imaging device 100 determines whether the cause is the movement of the main subject 380 itself or a change in the orientation of the imaging device 100. The motion detection unit 117 is a motion sensor such as an angular velocity sensor. If the motion detection unit 117 does not detect any movement of the imaging device 100, or if the movement is not such that the main subject 380 moves out of the field of view, it is determined that the main subject 380 has moved out of the field of view. Furthermore, the motion detection unit 117 is not limited to being composed of a motion sensor, and may be configured to detect motion based on the amount of change between frames over time in the image acquired by the imaging unit 102. Also, if the main subject 380 moves out of the field of view due to a change in the orientation of the imaging device 100, and the main audio microphone for the main subject 380 has already been determined, the determination of the main audio microphone may be canceled if the main subject 380 does not return to the field of view within a predetermined time. This is effective, for example, when the photographer 370 is tracking the moving main subject 380 with the imaging device 100, and is unable to keep track of the main subject 380, causing it to temporarily move out of the field of view.
[0057] While preferred embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above, and various modifications and changes are possible within the scope of its gist. The present invention provides a program that implements one or more of the functions of the above embodiments to a system or device via a network or storage medium. It can also be implemented by one or more general-purpose processors (ASICs) of the computer of the system or device reading and executing the program. Furthermore, the present invention can also be implemented by a dedicated processor (e.g., an ASIC or FPGA) that implements one or more functions. Moreover, the present invention can also be implemented by a combination of a general-purpose processor and a dedicated processor. Here, "processor" refers to a processor in a broad sense and includes both general-purpose processors and dedicated processors. Furthermore, the process of implementing the present invention may be executed by only one processor, or by the cooperation of multiple processors located in physically separate locations.
[0058] This embodiment includes the following configurations, methods, and programs. (Configuration 1) A sound processing device capable of processing sound collected by multiple microphones placed around a target sound acquisition object, A first information acquisition means for acquiring location information relating to the location of the audio acquisition target, A second information acquisition means for acquiring location information relating to the location of the aforementioned audio processing device, A third information acquisition means for acquiring identification information to identify each of the aforementioned microphones, A determination means that determines, based on the location information of the voice acquisition target, the location information of the voice processing device, and the identification information of each microphone, from among the plurality of microphones, a first microphone facing the voice acquisition target and having a directivity toward the voice acquisition target, and a second microphone facing the opposite side of the voice acquisition target and having a directivity in the same direction as the direction the first microphone is facing. A first voice acquisition means wirelessly acquires the sound collected by the first microphone as first voice data, A second audio acquisition means wirelessly acquires the audio collected by the second microphone as second audio data, An audio processing device comprising: noise reduction means for performing noise reduction processing to reduce noise contained in the first audio data with the second audio data. (Configuration 2) The audio processing device according to Configuration 1, characterized in that the determination means determines the microphone closest to the audio acquisition target among the microphones that can be the first microphone as the first microphone. (Configuration 3) The audio processing device according to Configuration 2, characterized in that the determination means determines, as the second microphone, a microphone that lies on a straight line passing through the first microphone and the audio acquisition target, from among the microphones that could be the second microphone. (Configuration 4) The control unit of the audio processing device selects the first microphone and the second microphone from among the plurality of microphones, The audio processing device according to any one of configurations 1 to 3, characterized in that the determination means determines the first microphone and the second microphone based on the selection result of the selection means. (Configuration 5) The configuration includes a selection operation means for selecting the first microphone and the second microphone from among the multiple microphones, The audio processing device according to any one of configurations 1 to 4, characterized in that the determination means determines the first microphone and the second microphone based on the selection result of the selection operation means. (Configuration 6) The audio processing device according to any one of Configurations 1 to 5, characterized in that the third information acquisition means is configured to acquire identification information of each microphone wirelessly. (Configuration 7) The audio processing device according to any one of Configurations 1 to 6, characterized in that the third information acquisition means acquires at least one of the following as identification information for each microphone: the ID of each microphone, location information relating to the position of each microphone, type information relating to the type of each microphone, and directional information relating to the directivity of each microphone. (Configuration 8) The audio processing device according to any one of Configurations 1 to 7, characterized in that the noise reduction process is performed using the difference between the sound pressure level at the distance from the noise source to the first microphone and the sound pressure level at the distance from the noise source to the second microphone. (Configuration 9) The audio processing apparatus according to any one of Configurations 1 to 8, comprising a storage means for storing the first audio data on which the noise reduction processing has been performed. (Configuration 10) The audio processing apparatus according to Configuration 9, characterized in that the storage means is capable of storing the image of the audio acquisition target together with the first audio data. (Configuration 11) The audio processing device according to any one of Configurations 1 to 10, characterized in that the third information acquisition means, the first audio acquisition means, and the second audio acquisition means are collectively composed of a unit having an assistant and a receiver in Auracast®-Broadcast-Audio. (Configuration 12) The audio processing device according to Configuration 11, further comprising a switching operation means for performing an operation to switch between enabling and disabling acquisition in Auracast-broadcast-audio. (Configuration 13) An imaging means for performing imaging, An imaging device comprising an audio processing device described in any one of configurations 1 to 12. (Configuration 14) Multiple microphones placed around the target object from which to acquire sound, An imaging device comprising an audio processing device described in any one of configurations 1 to 12. (Method 1) A method for controlling an audio processing device capable of processing audio collected by multiple microphones placed around a target audio acquisition object, A first information acquisition step involves acquiring location information relating to the location of the audio acquisition target, A second information acquisition step involves acquiring location information relating to the location of the aforementioned audio processing device, A third information acquisition step involves acquiring identification information to identify each of the aforementioned microphones, A determination step in which, based on the location information of the voice acquisition target, the location information of the voice processing device, and the identification information of each microphone, a first microphone facing the voice acquisition target and having a directivity toward the voice acquisition target, and a second microphone facing the opposite side of the voice acquisition target and having a directivity in the same direction as the first microphone are facing, A first audio acquisition step in which the audio collected by the first microphone is acquired wirelessly as first audio data, A second audio acquisition step in which the audio collected by the second microphone is acquired wirelessly as second audio data, A control method for an audio processing device, characterized by comprising: a noise reduction step of performing a noise reduction process to reduce noise contained in the first audio data using the second audio data. (Program 1) A program characterized by causing a computer to execute the control method described in Method 1. [Explanation of Symbols]
[0059] 100 Imaging device 101 Control Unit 102 Imaging Unit 108 Assistants 110 Receiver 301a~301d Microphone 380 Main subject 1000A Voice Processing System
Claims
1. A sound processing device capable of processing sound collected by multiple microphones placed around a target object from which sound is to be acquired, A first information acquisition means for acquiring location information relating to the location of the audio acquisition target, A second information acquisition means for acquiring location information relating to the location of the aforementioned audio processing device, A third information acquisition means for acquiring identification information to identify each of the aforementioned microphones, A determination means that determines, from among the plurality of microphones, a first microphone facing the voice acquisition target and having a directivity toward the voice acquisition target, and a second microphone facing the opposite side of the voice acquisition target and having a directivity in the same direction as the first microphone, based on the location information of the voice acquisition target, the location information of the voice processing device, and the identification information of each microphone, A first audio acquisition means wirelessly acquires the audio collected by the first microphone as first audio data, A second audio acquisition means wirelessly acquires the sound collected by the second microphone as second audio data, An audio processing device comprising: noise reduction means for performing noise reduction processing to reduce noise contained in the first audio data with the second audio data.
2. The voice processing device according to claim 1, characterized in that the determination means determines the microphone closest to the voice acquisition target among the microphones that can be the first microphone as the first microphone.
3. The voice processing apparatus according to claim 2, characterized in that the determination means determines, among the microphones that could be the second microphone, a microphone that lies on a straight line passing through the first microphone and the voice acquisition target as the second microphone.
4. The control unit of the audio processing device selects the first microphone and the second microphone from among the plurality of microphones, The audio processing apparatus according to claim 1, characterized in that the determination means determines the first microphone and the second microphone based on the selection result of the selection means.
5. The system includes a selection operation means for selecting the first microphone and the second microphone from among the multiple microphones, The audio processing apparatus according to claim 1, characterized in that the determination means determines the first microphone and the second microphone based on the selection result of the selection operation means.
6. The audio processing device according to claim 1, characterized in that the third information acquisition means is configured to wirelessly acquire identification information of each microphone.
7. The audio processing device according to claim 1, characterized in that the third information acquisition means acquires at least one of the following as identification information for each microphone: the ID of each microphone, location information relating to the position of each microphone, type information relating to the type of each microphone, and directional information relating to the directivity of each microphone.
8. The audio processing device according to claim 1, characterized in that the noise reduction process is performed using the difference between the sound pressure level at the distance from the noise source to the first microphone and the sound pressure level at the distance from the noise source to the second microphone.
9. The audio processing device according to claim 1, further comprising a storage means for storing the first audio data on which the noise reduction processing has been performed.
10. The audio processing apparatus according to claim 9, characterized in that the storage means is capable of storing the image of the audio acquisition target together with the first audio data.
11. The audio processing device according to claim 1, characterized in that the third information acquisition means, the first audio acquisition means, and the second audio acquisition means are collectively composed of a unit having an assistant and a receiver in Auracast®-Broadcast-Audio.
12. The audio processing device according to claim 11, further comprising a switching operation means for performing an operation to switch between enabling and disabling acquisition in Auracast-Broadcast-Audio.
13. An imaging means for performing imaging, An imaging device comprising the sound processing device described in claim 1.
14. Multiple microphones are placed around the target object from which to acquire audio, A voice processing system characterized by comprising the voice processing device described in claim 1.
15. A method for controlling an audio processing device capable of processing audio collected by multiple microphones placed around a target object from which to acquire audio, A first information acquisition step involves acquiring location information relating to the location of the audio acquisition target, A second information acquisition step involves acquiring location information relating to the location of the aforementioned audio processing device, A third information acquisition step involves acquiring identification information to identify each of the aforementioned microphones, A determination step in which, based on the location information of the voice acquisition target, the location information of the voice processing device, and the identification information of each microphone, a first microphone facing the voice acquisition target and having a directivity toward the voice acquisition target, and a second microphone facing the opposite side of the voice acquisition target and having a directivity in the same direction as the first microphone are facing, A first audio acquisition step in which the audio collected by the first microphone is acquired wirelessly as first audio data, A second audio acquisition step in which the audio collected by the second microphone is acquired wirelessly as second audio data, A control method for an audio processing device, characterized by comprising: a noise reduction step of performing a noise reduction process to reduce noise contained in the first audio data using the second audio data.
16. A program characterized by causing a computer to execute the control method described in claim 15.
Citation Information
Patent Citations
Microphone system, noise removal method, and program
JP2016082432A