Audio processing device, imaging device, audio processing system, control method and program for audio processing device

JP2026142784APending Publication Date: 2026-09-08CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025029979
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2026-09-08

AI Technical Summary

Benefits of technology

【0014】 本発明によれば、音声取得対象からの音声を迅速に集音することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026142784000001_ABST
    Figure 2026142784000001_ABST
Patent Text Reader

Abstract

The objective is to provide an audio processing device, an imaging device, an audio processing system, a control method for the audio processing device, and a program that can rapidly collect audio from an audio acquisition target. [Solution] In the voice processing system 1000, the portable terminal 10, which is a voice processing device, includes a storage means (storage unit 103) in which first identification information for identifying a voice acquisition target is stored in advance, an information acquisition means (transmitting / receiving unit 104) for acquiring second identification information for identifying candidate targets that may be voice acquisition targets, a matching means (matching unit 105) for matching the first identification information and the second identification information, and a determination means (control unit 109) for determining a specific microphone from among the microphone units 1030 to 1035 that is capable of preferentially collecting sound from the voice acquisition target based on the matching result by the matching means.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The present invention relates to an audio processing device, an imaging device, an audio processing system, a method for controlling an audio processing device, and a program. [[Background Art]]

[0002] In recent years, Bluetooth Low Energy (hereinafter referred to as "BLE"), which enables lower power consumption that reduces power consumption compared to conventional Bluetooth (registered trademark), has been known. Electronic devices having a wireless transmission / reception unit compatible with the BLE standard have become widespread. Further, Auracast Broadcast Audio (hereinafter referred to as "Auracast (registered trademark)") has been announced by the standardization organization Bluetooth Special Interest Group. Auracast is a system in which, when an Auracast receiver and an Auracast transmitter are disposed within a predetermined range, the Auracast receiver can unlimitedly join the broadcast transmission from the Auracast transmitter.

[0003] Non-Patent Document 1 describes an outline of Auracast (referred to as "Broadcast" in this document). Auracast described in Non-Patent Document 1 includes an Auracast Transmitter, an Auracast Assistant, and an Auracast Receiver. Hereinafter, the Auracast Transmitter is referred to as a "transmitter", the Auracast Assistant is referred to as an "assistant", and the Auracast Receiver is referred to as a "receiver". The transmitter is a transmitter for transmitting audio data or an audio file and an advertisement to unspecified receivers. The audio data or audio file transmitted by the transmitter is provided from a device serving as an audio source connected to the transmitter, a device serving as an audio source incorporating the transmitter, or a device storing the audio file.

[0004] The assistant is installed in a receiver that scans for advertisements from transmitters. The assistant provides an interface that allows the user of the receiver to receive audio data transmitted from the transmitter once they have selected a scanned source transmitter. Examples of receivers with the assistant include personal computers, tablet devices, smartphones, and other mobile devices. The assistant provides the necessary information to receive audio data or audio files transmitted from the transmitter to another receiver that is pre-paired with the assistant, or to a receiver installed in the same receiver as the assistant. Examples of receivers include audio output devices such as headphones, earphones, and speakers, or recording devices that record audio data or audio files. In this way, audio data or audio files from the transmitter can be received by the receiver.

[0005] Potential applications for Auracast include receiving audio from televisions installed in public areas, and directly receiving announcements at airports and train stations. Another potential application is with imaging devices such as cameras. However, in this case, when filming subjects far away from the photographer or through glass, the camera's built-in microphone cannot pick up sounds near the subject. Therefore, by installing a sound pickup device and transmitter near the subject and using the camera to acquire audio data transmitted from the transmitter, it becomes possible to capture immersive video. This makes it suitable for filming events such as sporting events in stadiums, theme parks like zoos and aquariums, and children's recitals and sports days.

[0006] Here, we will explain the method of shooting using Auracast in a stadium with reference to Figure 6. Figure 6 is a diagram illustrating the method of shooting using Auracast in a stadium. As shown in Figure 6, spectators 601 and 602 are sitting in the stands of the stadium, respectively, watching a tennis match between tennis player 603 and tennis player 604. Spectators 601 and 602 are each wearing earphones (headphones) 120 on their heads. Spectator 601 is also watching the tennis match while holding a camera (imaging device) 640 and shooting it as video. When shooting video, the sound acquired by the microphone built into the camera 640 tends to be sound from around spectator 601. As a result, it becomes difficult to acquire sufficient sound from the vicinity of at least one of the main subjects (the objects of recording), tennis player 603 and tennis player 604, which may result in a video with a low sense of realism.

[0007] In recent years, multiple microphones (microphones 630 to 635) have been placed around the tennis court located in the center of the stadium. Microphones 630 to 635 each have a transmitter function. This allows microphones 630 to 635 to each transmit audio to unspecified photographers, for example, via Auracast. In this case, camera 640 can receive the audio from microphones 630 to 635 and superimpose it onto the video. This makes it possible to obtain a video with a sense of realism. Spectator 602 is watching the game while carrying a mobile device 610. The mobile device 610 is an assistant and receives advertisements transmitted via Auracast from microphones 630 to 635. In addition, earphone 120 receives the audio transmitted from microphones 630 to 635. This allows spectator 602 to experience a greater sense of realism.

[0008] For example, let's consider a case where spectator 601 focuses on tennis player 603 and films a video with tennis player 603 as the main subject. In this case, spectator 601 wants to receive audio data collected by a microphone near tennis player 603 (microphone 633 in the state shown in Figure 6) with camera 640 in order to gain a greater sense of realism. Spectator 602 also has the same desire. However, if microphones 630 to 635 have transmitter functions, camera 640 cannot receive audio data from only one microphone. Therefore, one microphone near tennis player 603 must be selected. To this end, tennis players 603 and 604 are provided with wireless tags (not shown). Meanwhile, camera 640 is registered with device information such as the identification code of the wireless tag carried by tennis player 603, which is the player whose audio data is to be acquired. Camera 640 determines the location of the wireless tag carried by tennis player 603 and the locations of microphones 630 to 635. Based on these positioning results, spectator 601 can select the microphone closest to the wireless tag carried by tennis player 603. Patent Document 1 describes a configuration in which a terminal (corresponding to camera 640 or mobile terminal 610) identifies a candidate device according to its distance. [Prior art documents] [Patent Documents]

[0009] [Patent Document 1] Japanese Patent Publication No. 2023-48041 [Non-patent literature]

[0010] [Non-Patent Document 1] Nick Hunn, "Introducing BluetoothTM LE Audio", p.277-294 [Overview of the project] [Problems that the invention aims to solve]

[0011] However, in the configuration described in Patent Document 1, if there are multiple potential target devices, it is necessary to detect the distance to each device, which means that time is spent in identifying the device.

[0012] The present invention has been made in view of the above-mentioned problems. The object of the present invention is to provide an audio processing device, an imaging device, an audio processing system, a control method for the audio processing device, and a program that can rapidly collect audio from an audio acquisition target. [Means for solving the problem]

[0013] To achieve the above objective, the present invention provides an audio processing device capable of processing sound collected by a plurality of microphones arranged around a target sound acquisition object, the device comprising: a storage means for which first identification information for identifying the sound acquisition object is stored in advance; an information acquisition means for acquiring second identification information for identifying candidate targets that may become the sound acquisition object; a matching means for comparing the first identification information and the second identification information; and a determination means capable of determining a specific microphone from the plurality of microphones that is capable of preferentially collecting sound from the sound acquisition object based on the matching result of the matching means. [Effects of the Invention]

[0014] According to the present invention, it is possible to quickly collect sound from a sound acquisition target. [Brief explanation of the drawing]

[0015] [Figure 1A] This is a block diagram showing an example of the hardware configuration of the audio processing system according to the first embodiment. [Figure 1B] This figure shows an example of the usage status of the voice processing system. [Figure 1C] This figure shows an example of a screen displayed on the display unit of a mobile device. [Figure 2] It is a flowchart showing processing executed by a microphone unit. [Figure 3] It is a flowchart showing processing executed by a mobile terminal. [Figure 4] It is a block diagram showing an example of the hardware configuration of an audio processing system according to a second embodiment. [Figure 5] It is a flowchart showing processing executed by an imaging device. [Figure 6] It is a diagram for explaining an imaging method using Auracast in a stadium.

Mode for Carrying Out the Invention

[0016] Hereinafter, each embodiment of the present invention will be described in detail with reference to the drawings. However, the configurations described in the following embodiments are merely illustrative, and the scope of the present invention is not limited by the configurations described in the embodiments. For example, each part constituting the present invention can be replaced with one having any configuration that can exert the same function. In addition, any component may be added. Further, any two or more configurations (features) among the embodiments can be combined.

[0017] <First Embodiment> The first embodiment will be described below with reference to Figures 1A to 3. Figure 1A is a block diagram showing an example of the hardware configuration of the voice processing system according to the first embodiment. As shown in Figure 1, the voice processing system 1000 includes a mobile terminal 10 and an earphone (receiver) 120. The voice processing system 1000 also includes microphone (hereinafter simply referred to as "microphone") units 1030, 1031, 1032, 1033, 1034, and 1035. The voice processing system 1000 also includes wireless tags 150 and 151. The mobile terminal 10 is not particularly limited and can be, for example, a smartphone, a tablet, or a notebook personal computer. The mobile terminal 10 also has at least the assistant function of the Auracast function. The mobile terminal 10 includes an operation unit 101, a display unit (notification means) 102, a storage unit (storage means) 103, a transmitting / receiving unit (information acquisition means) 104, a matching unit (matching means) 105, a selection unit 106, a microphone 107, an audio processing unit 108, and a control unit 109. These hardware components of the mobile terminal 10 are connected to each other in a way that allows them to communicate with one another.

[0018] The operation unit 101 is constituted by, for example, a button switch, a touch panel provided on the surface of the display unit 102, or the like. The display unit 102 is constituted by, for example, a liquid crystal display. For example, a setting menu, a remaining battery level, statuses such as wireless connection, warnings and the like are displayed on the display unit 102. In addition, when the mobile terminal 10 includes an image capturing unit capable of capturing an image, an image captured by the image capturing unit is also displayed on the display unit 102. The storage unit 103 is constituted by a RAM which is a volatile memory, a ROM which is a non-volatile memory, or the like. Various types of information and various programs are stored in the storage unit 103. This program includes, for example, a program that causes a control unit 109 to execute each step (a control method for the mobile terminal 10) described later. In addition, when the mobile terminal 10 also has an Auracast (registered trademark) receiver function, audio data received by the receiver from a transmitter is stored in the storage unit 103. In the present embodiment, the transmission / reception unit 104 is constituted by a unit including an assistant and a receiver for Auracast (registered trademark) broadcast audio. This enables the transmission / reception unit 104 to receive advertisement packets wirelessly transmitted from the microphone units 1030 to 1035, and perform wireless transmission and reception with the paired earphone 120. Information necessary for the earphone 120 to acquire audio data distributed via Auracast from the microphone selected by the selection unit 106 is transmitted to the earphone 120. Further, the transmission / reception unit 104 acquires, from a predetermined area of the received advertisement packet from the microphone, wireless tag information including an ID of a wireless tag that the microphone has succeeded in receiving, in association with an identification code of the microphone included in the advertisement packet.

[0019] The matching unit 105 extracts and matches the wireless tag ID temporarily stored in the storage unit 103 with the tag information contained in the advertisement packets wirelessly transmitted from microphone units 1030 to 1035. Based on this matching result, the selection unit 106 selects one microphone from among microphone units 1030 to 1035. The control unit 109 controls the transmitting / receiving unit 104 to send information to the earphone 120 necessary for the earphone 120 to acquire the audio data distributed via Auracast from the selected microphone. Microphone 107 is a built-in microphone in the mobile terminal 10. Microphone 107 collects sound from the vicinity of the mobile terminal 10 and outputs it to the audio processing unit 108. The audio processing unit 108 performs various processes on the sound output from microphone 107, such as gain control, noise reduction, and encoding, and outputs it to the control unit 109. In this way, the mobile terminal 10 functions as an audio processing device capable of processing the sound collected by microphone units 1030 to 1035. Audio processing will be described later. Microphone 107 is activated when no microphone is selected by the selection unit 106. In this case, the control unit 109 omits transmitting the information necessary to acquire the audio data distributed by Auracast, and transmits the audio data collected by microphone 107 and processed by the audio processing unit 108 to the earphone 120 via the transmission / reception unit 104. The control unit 109 is a computer that controls each piece of hardware in the mobile terminal 10.

[0020] Microphone units 1030 to 1035 each have a microphone 131, an audio processing unit 132, a transmitting unit (transmitting / receiving means) 133, a receiving unit (transmitting / receiving means) 134, an information acquisition unit 135, and a control unit 136, respectively. Microphone units 1030 to 1035 also each function as transmitters. The receiving unit 134 can receive advertisement packets transmitted by wireless tags 150 and 151. The advertisement packets of wireless tags received by the receiving unit 134 include wireless tag information (wireless tag ID). The information acquisition unit 135 extracts (acquires) the wireless tag ID. The control unit 136 generates an advertisement packet for Auracast transmission that includes the identification code of the microphone unit having the control unit 136 and the wireless tag information extracted by the information acquisition unit 135. The advertisement packet for Auracast transmission is wirelessly transmitted from the transmitting unit 133. Furthermore, if the reception of an advertisement packet from the wireless tag fails, a blank (zero) is written to a predetermined area of ​​the wireless tag information. This makes it possible to determine that the wireless tag has not been received. The voice processing unit 132 performs various processes such as gain control, noise reduction, and coding on the voice collected by the microphone 131 to generate voice data. This voice data is distributed via Auracast through the transmission unit 133 under the control of the control unit 136.

[0021] The earphone 120 functions as a receiver and is pre-paired with the mobile terminal 10. The earphone 120 includes a transceiver 121, an audio processing unit 122, a speaker 123, and a control unit 124. The transceiver 121 receives various information from the transceiver 104 of the mobile terminal 10. This information includes, for example, information necessary to receive audio data distributed via Auracast from a microphone unit selected by the mobile terminal 10. Based on this information, the transceiver 121 can receive audio data from the microphone unit selected by the mobile terminal 10. The audio data received by the transceiver 121 is processed by the audio processing unit 122, controlled by the control unit 124, for example, by decoding, D / A conversion, and output as audio by the speaker 123. In this embodiment, the earphone 120 is configured separately from the mobile terminal 10, but it is not limited to this configuration and may, for example, be configured integrally with the mobile terminal 10.

[0022] Wireless tags 150 and 151 can each generate an advertisement packet containing at least an identification code such as a wireless tag ID, and can continue wireless transmission. The communication between wireless tags 150 and 151 is not particularly limited as long as wireless transmission and reception is possible, such as BLE or WiFi. In this embodiment, transmission and reception are performed using BLE.

[0023] Figure 1B shows an example of the voice processing system in use. Figure 1C shows an example of the screen displayed on the display unit of a mobile terminal. As shown in Figure 1B, the voice processing system 1000 can be used at the tennis court 10000. At the tennis court 10000, a tennis match is played on the tennis court 10001 by tennis players 1603 and 1604. Tennis player 1603 is wearing a wireless tag 150, and tennis player 1604 is wearing a wireless tag 151. In addition, in the spectator stands 10002 surrounding the tennis court 10001, many spectators watch the tennis match while cheering. Spectator 1601 watches the match wearing an imaging device 40 and earphones 120 that are communicatively connected to the imaging device 40. Spectator 1601 will be described in the second embodiment described later. Spectator 1602 watches the match while wearing a mobile device 10 and earphones 120 that are connected to the mobile device 10 for communication. The earphones 120 are audio output devices that receive audio wirelessly from a specific microphone unit (described later) and output that audio. Spectator 1602 can choose one of the tennis players 1603 and 1604 playing tennis on the tennis court 10001 (for example, tennis player 1603) as the target for audio acquisition.

[0024] To select the audio acquisition target, the selection screen 1 shown in Figure 1C is used. On selection screen 1, a list of players entered in a tennis match at the tennis court 10000 is displayed. In this embodiment, selection screen 1 includes multiple (multiple types) of options: "Player 1", "Player 2", "Player 3", "Player 4", etc. These are player identification information (first identification information) that identifies the tennis players and are stored in the storage unit 103 in advance (storage step). The storage of player identification information in the storage unit 103 is performed, for example, by a spectator 1602 through manual operation. Advertising media such as the tennis match homepage and brochures also publish the names and photos of the tennis players, as well as the wireless tag information of the wireless tags carried by each player. The spectator 1602 can store the player identification information in the storage unit 103 by acquiring each wireless tag information using a predetermined method (e.g., a barcode) on their mobile terminal 10. For example, if "Player 1" on selection screen 1 is tennis player 1603 and "Player 2" is tennis player 1604, the user checks the checkbox for "Player 1" and then presses the OK button 11. This determines that the voice acquisition target is tennis player 1603. This decision is stored in the memory unit 103. In this way, selection screen 1 functions as an operating means that allows the user to select one player (the target for matching in the matching unit 105, which will be described later) from among "Player 1", "Player 2", "Player 3", "Player 4", etc.

[0025] As mentioned above, tennis player 1603 wears wireless tag 150, and tennis player 1604 wears wireless tag 151. Tennis players 1603 and 1604 are also potential targets for voice acquisition. Wireless tag 150 functions as a transmission means to transmit wireless tag information (second identification information) identifying tennis player 1603 to microphone units 1030-1035. Wireless tag 151 functions as a transmission means to transmit wireless tag information identifying tennis player 1604 to microphone units 1030-1035. Microphone units 1030-1035 are arranged at intervals around the tennis court 10001 (tennis players 1603 and 1604). Note that the number of microphones is not limited to the configuration shown in Figure 1B.

[0026] Figure 2 is a flowchart showing the processing performed by the microphone unit. Here, we will typically describe the processing performed by microphone unit 1030 among microphone units 1030 to 1035. The program based on the flowchart shown in Figure 2 is pre-stored in the memory unit (not shown) of microphone unit 1030 and executed by the control unit 136. Furthermore, the execution of this program starts when the power to wireless tags 150 and 151 is turned ON by operating the operating members (not shown) of wireless tags 150 and 151. As shown in Figure 2, in step 201, the control unit 136 of microphone unit 1030 clears the write area of ​​the wireless tag information contained in the advertisement packet wirelessly transmitted by microphone unit 1030 to return it to its initial state. The method of returning it to the initial state is not particularly limited, and examples include writing blanks or zeros to the write area.

[0027] In step 202, the control unit 136 determines whether the receiving unit 134 has acquired (received wirelessly) the wireless tag ID, which is the device code transmitted from the wireless tag 150 or wireless tag 151, i.e., the wireless tag information. If, as a result of the determination in step 202, the control unit 136 determines that the wireless tag information has been acquired, the process proceeds to step S203. On the other hand, if, as a result of the determination in step 202, the control unit 136 determines that the wireless tag information has not been acquired, the process proceeds to step S205.

[0028] In step S203, the control unit 136 acquires positional information regarding the positional relationship between the wireless tag 150 or wireless tag 151, which was determined to have acquired wireless tag information in step S202, and the microphone unit 1030. That is, the control unit 136 detects the position of the wireless tag 150 or wireless tag 151 relative to the microphone unit 1030. For example, the position (positional information) of the wireless tag 150 is detected based on the intensity of the signal received by the receiving unit 134 when a wireless transmission from the wireless tag 150 is received. Note that the position detection method is not limited to this, and for example, the receiving unit 134 may be provided with at least two antennas, and the method may be based on the phase difference of the reception angle when the wireless transmission from the wireless tag 150 is received by each antenna of the receiving unit 134. This method is the AoA (Angle of Arrival) method supported from Bluetooth 5.1.

[0029] In step S204, the control unit 136 updates the contents of the write area of ​​the advertisement packet wirelessly transmitted by the microphone unit 1030 by writing the wireless tag information determined to have been acquired in step S202. The wireless tag information includes the location information acquired in step S203.

[0030] In step S205, the control unit 136 wirelessly transmits the advertisement packet from the transmission unit 133. If the advertisement packet was updated in step S204, that advertisement packet is transmitted in step S205.

[0031] In step S206, the control unit 136 determines whether the power supply to wireless tags 150 and 151 has been turned OFF. If the control unit 136 determines that the power supply has been turned OFF as a result of the determination in step S206, the process ends. On the other hand, if the control unit 136 determines that the power supply is not OFF, i.e., remains ON, the process returns to step S201 and the subsequent steps are executed in order.

[0032] Figure 3 is a flowchart of the process executed on the mobile terminal. As an example, this describes the case where "Player 1," which represents tennis player 1603, is selected as the target for voice acquisition on selection screen 1. As shown in Figure 3, in step S301, the control unit 109 of the mobile terminal 10 determines whether the wireless tag information (wireless tag ID) of the wireless tag carried by the tennis player being watched by the spectator 1602 has already been registered (stored) in the storage unit 103. Here, the tennis players are tennis player 1603 and tennis player 1604, and the wireless tag information is player identification information. If, as a result of the determination in step S301, the control unit 109 determines that the wireless tag information has been registered, the process proceeds to step S302. On the other hand, if, as a result of the determination in step S301, the control unit 109 determines that the wireless tag information has not been registered, the process proceeds to step S310.

[0033] In step S302, the control unit 109 determines whether or not it has received an advertisement packet (information acquisition step) wirelessly transmitted from the wireless tag 150 carried by the tennis player 1603 via the microphone units 1030 to 1035 at the transmitting / receiving unit 104. The advertisement packet contains wireless tag information from the wireless tag 150, etc. If, as a result of the determination in step S302, the control unit 109 determines that it has received the advertisement packet, the process proceeds to step S303. On the other hand, if, as a result of the determination in step S302, the control unit 109 determines that it has not received the advertisement packet, the process proceeds to step S309.

[0034] In step S303, the control unit 109 controls the matching unit 105 to compare the player identification information (wireless tag information) already registered in the storage unit 103 with the wireless tag information contained in the advertised packet determined in step S302 (matching step).

[0035] In step S304, the control unit 109 determines the number of matches, i.e., one of the following three states, based on the matching results in step S303. The first state is when the player identification information matches one wireless tag (when "=1" is obtained in step S304). The second state is when the player identification information matches multiple wireless tag information (when ">1" is obtained in step S304). The third state is when the player identification information and wireless tag information do not match at all, i.e., there is no match (when "=0" is obtained in step S304). If the control unit 109 determines in step S304 that the state is the first state, the process proceeds to step S305. If the control unit 109 determines in step S304 that the state is the second state, the process proceeds to step S306. Furthermore, if the control unit 109 determines, as a result of the determination in step S304, that the state is in the third state, the process proceeds to step S309.

[0036] In step S305, the control unit 109 selects a specific microphone unit from among microphone units 1030 to 1035 that is capable of preferentially collecting the voice from the tennis player 1603 (voice acquisition target) (determination step). Specifically, the control unit 109 controls the selection unit 106 to select a microphone unit from among microphone units 1030 to 1035 that transmits wireless tag information that matches the player identification information. Then, the control unit 109 determines the microphone unit selected by the selection unit 106 as the specific microphone unit. After step S305 is executed, the process proceeds to step S307.

[0037] In step S306, similar to step S305, the control unit 109 determines a specific microphone unit (determination step). Specifically, the control unit 109 controls the selection unit 106 to select the microphone unit closest to the tennis player 1603 (wireless tag 150) based on the positional information of microphone units 1030 to 1035. Then, the control unit 109 determines the microphone unit selected by the selection unit 106 as the specific microphone unit. Thus, in this embodiment, the control unit 109 also functions as a determination means to determine a specific microphone unit from among microphone units 1030 to 1035 based on the matching result in step S303. Note that in the mobile terminal 10, the part that functions as a determination means may be provided separately from the control unit 109. After step S306 is executed, the process proceeds to step S307.

[0038] In step S307, the control unit 109 transmits specific microphone unit information, determined in step S305 or step S306, from the transmitting / receiving unit (information transmission means) 104 to the earphone 120. This establishes a communication-enabled connection between the specific microphone unit and the earphone 120.

[0039] In step S308, the control unit 109 permits the acquisition of audio data at the earphone 120 in accordance with user operation on the operation unit 101. As a result, in step S308 after step S307, audio data from a specific microphone unit is acquired at the earphone 120. This audio data is mainly the voice data emitted from the tennis player 1603 (the voice acquisition target). This allows the voice from the tennis player 1603 to be quickly collected, so that the spectator 1602 can listen to the voice while watching the match with a sense of presence. Also, in step S308 after step S310, audio data from the microphone 107 is acquired at the earphone 120. This audio data is mainly the voice data from the vicinity of the mobile terminal 10.

[0040] In step S309, the control unit 109 displays a warning on the display unit 102. In step S309 after step S302 has been executed, the display unit 102 displays a warning indicating that the transmitting / receiving unit 104 is unable to acquire advertised packets. Also, in step S309 after step S304 has been executed, the display unit 102 displays a warning indicating that the system is in the third state.

[0041] In step S310, the control unit 109 controls the selection unit 106 to select the microphone 107 as the microphone for voice acquisition. After step S310 is completed, the process proceeds to step S308.

[0042] In step S311, the control unit 109 determines whether or not to stop acquiring audio data from the earphone 120. This determination is based on whether or not an operation to stop audio data acquisition has been performed by the operation unit 101. If an operation to stop audio data acquisition has been performed, it is determined that audio data acquisition has been stopped; if no such operation has been performed, it is determined that audio data acquisition will continue. If, as a result of the determination in step S311, the control unit 109 determines that audio data acquisition has been stopped, the process ends. On the other hand, if, as a result of the determination in step S311, the control unit 109 determines that audio data acquisition will continue, the process returns to step S301 and the subsequent steps are executed in order.

[0043] As described above, the mobile terminal 10 can quickly determine a specific microphone unit based on the matching result from the matching unit 105. As a result, the voice from the tennis player 1603 is prioritized for pickup by the specific microphone unit.

[0044] <Second Embodiment> The second embodiment will be described below with reference to Figures 4 and 5, focusing on the differences from the previously described embodiment, and omitting explanations of similar matters. Figure 4 is a block diagram showing an example of the hardware configuration of the audio processing system according to the second embodiment. As shown in Figure 4, the audio processing system 1000 of this embodiment includes an imaging device 40, earphones 120, microphone units 1030 to 1035, wireless tags 150 and 151. The imaging device 40, like the mobile terminal 10, has operation units 101 to 109 and functions as an audio processing device capable of processing audio collected by microphone units 1030 to 1035. The imaging device 40 also includes an imaging unit (imaging means) 401, an image identification unit 402, a wireless tag detection unit (determination means) 403, a main subject detection unit 404, and a receiving unit 405. These hardware components of the imaging device 40 are connected to each other in a communicative manner.

[0045] The imaging unit 401 is configured to capture video and still images, and includes, for example, an optical block, an image sensor, and an image processing block. The optical block consists of multiple lenses. The image sensor is composed of, for example, a CMOS, CCD, etc. The image processing block performs gain adjustment, noise and color correction, and compression processing on the image output from the image sensor. For example, when recording live images or videos, the video captured by the imaging unit 401 is displayed on the display unit 102. Also, when recording videos, audio is superimposed on the video and recorded in the storage unit 103. The audio is output from the speaker 123 of the earphone 120 regardless of whether the imaging device 40 is operating in video recording or live display, including the shooting standby state. The image identification unit 402 identifies the position and characteristics of subjects (objects) present in the image captured by the imaging unit 401. Characteristic information regarding the characteristics of subjects is pre-stored in the storage unit 103.

[0046] The wireless tag detection unit 403 determines the presence or absence of a wireless tag in the image captured by the imaging unit 401 based on additional information (feature information) that matches the characteristics of the subject identified by the image recognition unit 402, that is, it determines whether the subject is capable of transmitting wireless tag information. "Additional information" refers to feature information stored in the storage unit 103 and associated with the wireless tag information (wireless tag ID). The main subject detection unit 404 detects the focus position in the image captured by the imaging unit 401. The subject located at this focus position becomes the main subject. Furthermore, the imaging device 40 can identify the wireless tag carried by the main subject based on the identification result from the image recognition unit 402 and the detection result from the main subject detection unit 404. It is preferable that the microphone unit selected as the specific microphone unit is one, regardless of the number of wireless tags detected by the wireless tag detection unit 403. The purpose of detection by the main subject detection unit 404 is to narrow down the wireless tag information registered in the storage unit 103 to one before identifying the microphone unit. The narrowed-down wireless tag information is then notified to the matching unit 105 by the control unit 109. The receiving unit 405 is a receiver for Auracast. In the first embodiment, it is sufficient for the earphones 120 to have a receiver function in order to enjoy listening to the audio data acquired by Auracast. On the other hand, in the second embodiment, since the storage unit 103 of the imaging device 40 has the function of acquiring and recording audio, the imaging device 40 requires a receiving unit 405. If the imaging device 40 has a receiving unit 405, the earphones 120 can be used via a wired connection in addition to being used via a wireless connection. Audio can also be listened to via unicast communication between the imaging device 40 and the earphones 120.

[0047] Here, an example of the usage state of the audio processing system 1000 of this embodiment will be described with reference to Figure 1B. As shown in Figure 1B, spectator 1601 watches the match while wearing an imaging device 40 and earphones 120 that are communicatively connected to the imaging device 40. Spectator 1601 captures images of the tennis match with the imaging device 40. In addition, spectator 1601, like spectator 1602, can designate one of the tennis players 1603 and 1604 as the audio acquisition target for acquiring their voice. This embodiment is effective when the tracked tennis player (audio acquisition target) changes during video recording or while waiting for recording. As a result, even if the audio acquisition target switches from tennis player 1603 to tennis player 1604, or from tennis player 1604 to tennis player 1603, audio can still be acquired from that audio acquisition target. Furthermore, the imaging device 40 is capable of displaying video and can register multiple wireless tag ID information. As a result, each wireless tag transmitting the registered wireless tag information will be present within the video display area. Furthermore, in order to determine whether multiple wireless tags are present inside or outside the video display area, or their location within the video display area, it is necessary to know which tennis player is carrying each wireless tag. In addition, the aforementioned additional information is acquired by the predetermined method (e.g., barcode) and stored in the storage unit 103.

[0048] The imaging device 40, configured as described above, performs image recognition from the image data output as a video display area. The imaging device 40 then compares this with additional information temporarily registered in the storage unit 103. This allows the system to determine whether the desired wireless tag exists either inside or outside the video display area, or its location within the video display area.

[0049] Figure 5 is a flowchart showing the processes performed by the imaging device. Unlike the flowchart shown in Figure 3, the flowchart in Figure 5 executes steps S501 to S504 instead of step S304. Also, step S505 is executed between steps S304 and S306. Furthermore, after the execution of step S505, step S506 is executed depending on the result of the determination in step S505. As shown in Figure 5, step S501 is executed after the execution of step S302. In step S501, the control unit 109 of the imaging device 40 determines whether the main subject is carrying a wireless tag that has one of the wireless tag information (wireless tag ID) stored in the storage unit 103. If, as a result of the determination in step S501, the control unit 109 determines that the main subject is carrying a wireless tag, the process proceeds to step S503. On the other hand, if, as a result of the determination in step S501, the control unit 109 determines that the main subject is not carrying a wireless tag, the process proceeds to step S502.

[0050] In step S502, the control unit 109 determines whether any person (object) other than the main subject within the field of view of the imaging device 40 is carrying a wireless tag that has one of the wireless tag information (wireless tag ID) stored in the memory unit 103. If, as a result of the determination in step S502, the control unit 109 determines that one of the people other than the main subject is carrying a wireless tag (resulting in "=1" in step S502), the process proceeds to step S503. If, as a result of the determination in step S502, the control unit 109 determines that multiple people other than the main subject are carrying wireless tags (resulting in ">1" in step S502), the process proceeds to step S504. If, as a result of the determination in step S502, the control unit 109 determines that none of the people other than the main subject are carrying wireless tags (resulting in "=0" in step S502), the process proceeds to step S309.

[0051] In step S503, the control unit 109 controls the matching unit 105 to compare the wireless tag information determined in step S501 or step S502 with the wireless tag information contained in the advertisement packets from microphone units 1030 to 1035. After step S503 is executed, the process proceeds to step S304.

[0052] In step S504, the control unit 109 controls the matching unit 105 to compare the multiple wireless tag information determined in step S502 with the wireless tag information contained in the advertisement packets from microphone units 1030 to 1035. In step S504, the control unit 109 does not consider the matching by the matching unit 105 to be successful unless all of the multiple wireless tag information is included in the wireless tag information contained in the advertisement packet. After step S504 is executed, the process proceeds to step S304.

[0053] Then, if the control unit 109 determines, based on the judgment in step S304, that it is in the first state (i.e., "=1" in step S304), the process proceeds to step S305. Also, if the control unit 109 determines, based on the judgment in step S304, that it is in the second state (i.e., ">1" in step S304), the process proceeds to step S505. Furthermore, if the control unit 109 determines, based on the judgment in step S304, that it is in the third state (i.e., "=0" in step S304), the process proceeds to step S309.

[0054] In step S505, the control unit 109 controls the selection unit 106 to determine the matching in step S503 and the matching in step S504. This determination is used to determine how to select the microphone unit after step S505 is executed. If, as a result of the determination in step S505, the control unit 109 determines that it has matched one wireless tag information with the wireless tag information from the microphone unit 1030, etc. (matching in step S503) (resulting in "=1" in step S505), the process proceeds to step S306. If, as a result of the determination in step S505, the control unit 109 determines that it has matched multiple wireless tag information with the wireless tag information from the microphone unit 1030, etc. (matching in step S504) (resulting in ">1" in step S505), the process proceeds to step S506.

[0055] In step S506, the control unit 109 controls the selection unit 106 to select the microphone unit that is closest to the target wireless tags and whose distances are approximately the same. For example, in the state shown in Figure 1B, tennis players 1603 and 1604 are within the field of view of the imaging device 40 (in the captured image). Tennis players 1603 and 1604 within this field of view can each be targets for voice acquisition. Tennis players 1603 and 1604 each carry wireless tags with wireless tag information registered in the storage unit 103. Spectator 1601 has set the focus point to the center of the court (umpire) in order to capture the entire tennis court 10001 within the field of view. In this case, neither tennis player 1603 nor tennis player 1604 are the main subjects. In this state, first, the selection unit 106 extracts microphone units 1034 and 1035, which are approximately the same distance from tennis players 1603 and 1604, as candidates for the specific microphone unit. Next, the selection unit 106 selects the microphone unit closest to microphone units 1034 and 1035 as the specific microphone unit (microphone unit 1034). After step S506 is executed, the process proceeds to step S307.

[0056] As described above, in this embodiment, while imaging tennis players 1603 and 1604, it is possible to quickly collect sound from either tennis player 1603 or tennis player 1604, which is the target of sound acquisition. As a result, spectators 1601 can watch the match while viewing a video that includes the sound.

[0057] While preferred embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above, and various modifications and changes are possible within the scope of its gist. The present invention provides a program that implements one or more functions of the embodiments described above to a system or device via a network or storage medium. It can also be implemented by one or more general-purpose processors (ASICs) of the computer of the system or device reading and executing the program. Furthermore, the present invention can also be implemented by a dedicated processor (e.g., an ASIC or FPGA) that implements one or more functions. Moreover, the present invention can also be implemented by a combination of a general-purpose processor and a dedicated processor. Here, "processor" refers to a processor in a broad sense and includes both general-purpose processors and dedicated processors. Furthermore, the process of implementing the present invention may be executed by only one processor, or by the cooperation of multiple processors located in physically separate locations.

[0058] Each embodiment of the disclosure includes the following configurations, methods, and programs. (Configuration 1) A sound processing device capable of processing sound collected by multiple microphones placed around a target sound acquisition object, A storage means in which first identification information for identifying the voice acquisition target is stored in advance, Information acquisition means for acquiring second identification information that identifies a candidate target that may be subject to voice acquisition, A matching means for comparing the first identification information and the second identification information, An audio processing device comprising a determination means capable of determining a specific microphone from among the plurality of microphones that is capable of preferentially collecting sound from the sound acquisition target, based on the matching result of the matching means. (Configuration 2) The target candidate is provided with a transmission means for transmitting the second identification information, Each of the aforementioned microphones is capable of wirelessly transmitting positional information relating to the positional relationship between the microphone and the target candidate, and is equipped with a transmitting and receiving means capable of wirelessly transmitting and receiving the second identification information transmitted from the transmitting means. The audio processing device according to configuration 1, characterized in that the information acquisition means is capable of wirelessly acquiring the location information and the second identification information transmitted from the transmitting and receiving means. (Configuration 3) The audio processing device according to Configuration 2, characterized in that, if the matching result of the matching means matches the first identification information with one of the second identification information, the microphone that transmits the second identification information from the transmitting and receiving means is determined to be the specified microphone. (Configuration 4) The voice processing device according to Configuration 3, characterized in that, if the matching result of the matching means matches the first identification information with the plurality of second identification information, the determination means determines, based on the position information, the microphone closest to the target candidate that transmits the second identification information from the transmitting means among the plurality of microphones that transmit each of the second identification information from the transmitting means as the specified microphone. (Configuration 5) The voice processing device according to any one of Configurations 1 to 4, characterized in that it is provided with a notification means for notifying that the second identification information cannot be obtained by the information acquisition means. (Configuration 6) The voice processing device according to any one of Configurations 1 to 5, characterized in that it is provided with a notification means for notifying the user if the matching result of the matching means finds that the first identification information and the second identification information do not match. (Configuration 7) The storage means has a plurality of types of the first identification information stored in it beforehand, The audio processing device according to any one of configurations 1 to 6, characterized in that it includes an operating means for performing an operation to select one of the multiple types of first identification information to be matched by the matching means. (Configuration 8) The voice processing device according to any one of Configurations 1 to 7, characterized in that the information acquisition means acquires the ID of the target candidate as the second identification information. (Configuration 9) The audio processing device according to any one of Configurations 1 to 8, characterized in that the information acquisition means is composed of a unit having an assistant and a receiver in Auracast®-Broadcast-Audio. (Configuration 10) The audio processing device is a device used in which the audio processing device is connected in a communicative manner to an audio output device that wirelessly receives the audio from the specific microphone and outputs the audio, The audio processing device according to any one of configurations 1 to 9, further comprising an information transmission means for transmitting specific microphone information relating to a specific microphone determined by the determination means to the audio output device. (Configuration 11) An imaging means for performing imaging, An imaging device comprising an audio processing device described in any one of configurations 1 to 10. (Configuration 12) The subject in the captured image taken by the imaging means is used as the audio acquisition target, The imaging apparatus according to configuration 11, characterized in that characteristic information relating to the characteristics of the subject is stored in the storage means in advance. (Configuration 13) The imaging device according to Configuration 12, further comprising a determination means for determining whether the subject in the captured image is a subject capable of transmitting the second identification information, based on the characteristic information. (Configuration 14) Multiple microphones placed around the target object from which to acquire sound, A voice processing system characterized by comprising a voice processing device described in any one of configurations 1 to 10. (Method 1) A method for controlling an audio processing device capable of processing audio collected by multiple microphones placed around a target audio acquisition object, A storage step in which first identification information for identifying the voice acquisition target is stored in advance, An information acquisition step involves acquiring second identification information to identify potential target candidates that may be subject to voice acquisition, A matching step of comparing the first identification information with the second identification information, A control method for an audio processing device, characterized by comprising a determination step that, based on the matching results in the matching step, determines a specific microphone from among the plurality of microphones that is capable of preferentially collecting sound from the sound acquisition target. (Program 1) A program characterized by causing a computer to execute the control method described in Configuration 15. [Explanation of Symbols]

[0059] 10 Mobile devices 40 Imaging device 102 Display section 103 Storage section 104 Transmitter / Receiver 105 Verification section 1030~1035 Microphone Unit 1000 Voice Processing Systems

Claims

1. A sound processing device capable of processing sound collected by multiple microphones placed around a target object from which sound is to be acquired, A storage means in which first identification information for identifying the voice acquisition target is stored in advance, Information acquisition means for acquiring second identification information that identifies target candidates that may be subject to voice acquisition, A matching means for comparing the first identification information and the second identification information, An audio processing device comprising a determination means capable of determining a specific microphone from among the plurality of microphones that is capable of preferentially collecting sound from the sound acquisition target, based on the matching result of the matching means.

2. The target candidate is provided with a transmission means for transmitting the second identification information, Each of the aforementioned microphones is capable of wirelessly transmitting positional information relating to the positional relationship between the microphone and the target candidate, and is equipped with a transmitting and receiving means capable of wirelessly transmitting and receiving the second identification information transmitted from the transmitting means. The voice processing device according to claim 1, characterized in that the information acquisition means is capable of wirelessly acquiring the location information and the second identification information transmitted from the transmitting and receiving means.

3. The voice processing device according to claim 2, characterized in that, if the matching result of the matching means matches the first identification information with one of the second identification information, the microphone that transmits the second identification information from the transmitting / receiving means is determined to be the specified microphone.

4. The voice processing device according to claim 3, characterized in that, if the matching result of the matching means matches the first identification information with the plurality of second identification information, the determination means determines, based on the position information, the microphone closest to the target candidate that transmits the second identification information from the transmitting means among the plurality of microphones that transmit each of the second identification information from the transmitting means as the specified microphone.

5. The voice processing device according to claim 1, further comprising a notification means for notifying that the second identification information cannot be obtained by the information acquisition means.

6. The voice processing device according to claim 1, further comprising a notification means for notifying the user if the matching result of the matching means finds that the first identification information and the second identification information do not match.

7. Multiple types of the first identification information are pre-stored in the storage means. The voice processing device according to claim 1, further comprising an operating means for performing an operation to select one of the multiple types of first identification information to be used as the matching target by the matching means.

8. The voice processing device according to claim 1, characterized in that the information acquisition means acquires the ID of the target candidate as the second identification information.

9. The audio processing device according to claim 1, characterized in that the information acquisition means is comprised of a unit having an assistant and a receiver in Auracast (registered trademark) - Broadcast - Audio.

10. The aforementioned audio processing device is a device used in which the audio processing device is wirelessly connected to an audio output device that receives the audio from the specified microphone and outputs the audio, and is used in a manner that allows communication between the two devices. The audio processing device according to claim 1, further comprising an information transmission means for transmitting specific microphone information relating to a specific microphone determined by the determination means to the audio output device.

11. An imaging means for performing imaging, An imaging device comprising the sound processing device described in claim 1.

12. The subject in the captured image taken by the imaging means is used as the target for audio acquisition. The imaging apparatus according to claim 11, characterized in that characteristic information relating to the characteristics of the subject is stored in the storage means in advance.

13. The imaging device according to claim 12, further comprising a determination means for determining whether the subject in the captured image is a subject capable of transmitting the second identification information, based on the characteristic information.

14. Multiple microphones are placed around the target object from which to acquire audio, A voice processing system characterized by comprising the voice processing device described in claim 1.

15. A method for controlling an audio processing device capable of processing audio collected by multiple microphones placed around a target audio acquisition object, wherein the audio is to be acquired, A storage step in which first identification information for identifying the voice acquisition target is stored in advance, An information acquisition step involves acquiring second identification information to identify potential target candidates that may be subject to voice acquisition, A matching step of comparing the first identification information with the second identification information, A control method for an audio processing device, characterized by comprising a determination step that, based on the matching results in the matching step, determines a specific microphone from among the plurality of microphones that is capable of preferentially collecting sound from the sound acquisition target.

16. A program characterized by causing a computer to execute the control method described in claim 15.

Citation Information

Patent Citations

  • Information processing device and program

    JP2023048041A