Sound collection control method and sound collection control device

The sound collection control method addresses the issue of directing sound beams based on user intent by setting a beam focus on the intended speaking position, improving user experience in audio conferences.

JP2025146261APending Publication Date: 2025-10-03YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024046937
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing sound collection systems do not direct sound collection beams in accordance with a user's intention to speak before speech is made.

Method used

A sound collection control method that includes accepting a selection operation on a user terminal, obtaining position information, and setting a sound collection beam based on this information to focus on the user's intended speaking position.

Benefits of technology

Enables directing a sound collection beam in accordance with the user's intention to speak before speaking, enhancing user experience in audio conferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025146261000001_ABST
    Figure 2025146261000001_ABST
Patent Text Reader

Abstract

To provide a sound collection control method capable of directing a sound collection beam according to the user's intention to speak before speaking.SOLUTION: A sound collection control method for an audio conference system includes a microphone and at least one user terminal corresponding to the microphone. The method receives a selection operation for the user terminal, acquires position information of the user terminal that received the selection operation, and sets a first sound collection beam of the microphone on the basis of the acquired position information.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to a sound collection control method and a sound collection control device. [Background technology]

[0002] Patent Document 1 discloses an invention that adjusts the level of a transmitted voice signal depending on whether or not a near-end talker is speaking. In Patent Document 1, the sound collection beam is directed toward the talker. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2023 / 149254 Summary of the Invention [Problem to be solved by the invention]

[0004] The device of Patent Document 1 does not direct a sound collection beam before a speech is made, and therefore does not direct a sound collection beam in accordance with the user's intention to speak before the speech is made.

[0005] One object of this embodiment is to provide a sound collection control method that can direct a sound collection beam in accordance with the user's intention to speak before speaking. [Means for solving the problem]

[0006] The sound collection control method is a sound collection control method for an audio conference system that includes a microphone and at least one user terminal corresponding to the microphone, and includes accepting a selection operation on the user terminal, obtaining position information of the user terminal that has received the selection operation, and setting a first sound collection beam of the microphone based on the obtained position information. [Effects of the Invention]

[0007] According to one embodiment of the present invention, it is possible to direct a sound collection beam in accordance with the user's intention to speak before speaking. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram showing the configuration of an audio conference system 1. FIG. [Figure 2] This is a schematic elevation view of the interior. [Figure 3] FIG. 2 is a block diagram showing the configuration of a signal processing device 20. [Figure 4] FIG. 2 is a block diagram showing the configuration of a user terminal 40. [Figure 5] 4 is a flowchart showing the operation of the processor 14. [Figure 6] FIG. 10 is a diagram showing an example of a GUI for inputting position information of a user terminal 40 in advance. [Figure 7] FIG. 10 is a diagram showing an example of a GUI for inputting position information of a user terminal 40 in advance. [Figure 8] FIG. 10 is a diagram showing an example of a GUI for inputting position information of a user terminal 40 in advance. [Figure 9] 10 is a schematic diagram of an interior elevation view according to Modification 1. [Figure 10] 10 is a flowchart showing the operation of a processor 14 according to a first modification. [Figure 11] FIG. 10 is a block diagram showing the configuration of an audio conference system 1 according to a third modification. [Figure 12] 10 is a flowchart showing the operation of the processor 14 according to the third modification. DETAILED DESCRIPTION OF THE INVENTION

[0009] 1 is a block diagram showing the configuration of an audio conference system 1 according to this embodiment. The audio conference system 1 includes a display device 10, a signal processing device 20, a personal computer (PC) 30, and a user terminal 40. The signal processing device 20 is connected to the display device 10, the PC 30, and the user terminal 40. The signal processing device 20 is connected to a PC in a remote location via a network. The user terminal 40 and the PC 30 may be connected to the signal processing device 20 via the network.

[0010] FIG. 2 is a schematic elevation view of a room. As an example, the display device 10 is installed on a wall in the room. As an example, the signal processing device 20 is attached to the ceiling. There are multiple users (users u1 and u2) around a desk. A user terminal 40 is installed on the desk in front of user u2.

[0011] 3 is a block diagram showing the configuration of the signal processing device 20. The signal processing device 20 includes a camera 11, a speaker 12, a plurality of microphones 13, a processor 14, a memory 15, and an interface (I / F) 16.

[0012] The memory 15 is a storage medium that stores an operation program for the processor 14. The processor 14 reads the operation program from the memory 15 and performs various operations. The program does not have to be stored in the memory 15. For example, the program may be stored in a storage medium of an external device such as a server. In this case, the processor 14 simply reads the program from the server and executes it each time.

[0013] The processor 14 receives sound signals acquired by the multiple microphones 13. The processor 14 performs predetermined sound processing on the sound signals acquired by the multiple microphones 13. For example, the processor 14 performs beamforming on the sound signals acquired by the multiple microphones 13. Beamforming is a process of forming a sound collection beam with directionality in a predetermined direction by adding delays to sound signals acquired by the multiple microphones 13 and combining them. The sound collection beam can also be formed with directivity that focuses on a predetermined position. When a speaker's voice is detected, the processor 14 forms, for example, a sound collection beam that focuses on the speaker's position. The processor 14 detects the speaker's position by, for example, calculating a correlation value of the sound signals from the multiple microphones 13 to determine the difference in voice acquisition timing (phase difference). The processor 14 can uniquely determine the speaker's position by calculating the difference in voice acquisition timing between three or more microphones 13. The processor 14 can acquire the speaker's voice with high sensitivity by forming directivity that focuses on the determined speaker's position. Multiple sound collection beams can also be formed simultaneously. When a user terminal 40 is selected, the processor 14 forms, for example, a sound collection beam that is focused on the position of the selected user terminal 40. The processor 14 adds a delay to each microphone 13 in advance to align the phase so that, for example, a sound collection beam with increased sensitivity is formed in the direction of the selected user terminal 40.

[0014] At least two microphones 13 can form a sound collection beam.

[0015] The processor 14 outputs a sound signal related to the sound collection beam to the I / F 16. The I / F 16 is a communication I / F such as LAN, USB, HDMI (registered trademark), or Bluetooth (registered trademark). The I / F 16 is connected to the PC 30, for example, via USB. The I / F 16 is also connected to the display device 10, for example, via HDMI (registered trademark). The I / F 16 is also connected to the user terminal 40 via Bluetooth (registered trademark). The I / F 16 is also connected to a network via a LAN.

[0016] 4 is a block diagram of a user terminal 40. The user terminal 40 includes a communication I / F 21, a control unit 22, a flash memory 23, a RAM 24, a user I / F 25, and an LED 26, which is an example of a display. The control unit 22 controls the overall operation of the user terminal 40 by reading an operating program from the flash memory 23 to the RAM 24. The program does not need to be stored in the flash memory 23 of the user terminal 40 itself. The control unit 22 may download the program from a server or the like and read it into the RAM 24 whenever necessary. The control unit 22 accepts user operations via the user I / F 25.

[0017] The user I / F 25 has, for example, an operator for accepting a selection operation, an operator for accepting a volume change operation for the speaker 12, an operator for accepting a framing change operation for the camera 11, or an operator for accepting a microphone mute operation. These operators may be physically separate operators, or they may be the same operator. For example, the operator for accepting a selection operation and the operator for accepting a mute operation may be the same operator. For example, the user terminal 40 accepts a selection operation when it accepts a single press operation on the operator within a short period of time, and accepts a mute operation when it accepts a continuous long press operation on the operator for a predetermined period of time or more. Note that the operation of the user terminal 40 may be executed on the PC 30 as an application program of the PC 30. Furthermore, the selection can be canceled by operating the selected user terminal.

[0018] The signal processing device 20 transmits a sound signal relating to the sound collection beam to a PC at a remote location via a network.

[0019] The signal processing device 20 receives an audio signal from a remote PC via a network. The processor 14 outputs the audio signal received via the network to the speaker 12. The speaker 12 emits the audio signal received from the processor 14.

[0020] This allows the user of the signal processing device 20 to hold an audio conference with a user in a remote location.

[0021] Processor 14 also receives a video signal related to the video captured by camera 11. Processor 14 performs predetermined signal processing on the video signal captured by camera 11. The signal processing is, for example, framing processing using pan, tilt, or zoom. Processor 14 transmits the video signal after sound signal processing to a PC at a remote location via a network.

[0022] The signal processing device 20 receives a video signal from a remote PC via a network. The signal processing device 20 outputs the received video signal to the display device 10. The display device 10 receives an input of the video signal from the signal processing device 20. The display device 10 displays an image based on the received video signal.

[0023] This allows the user of the signal processing device 20 to hold a video conference with a user in a remote location.

[0024] It should be noted that the camera 11 and the speaker 12 are not essential to the present invention. It is also not essential to install the camera 11, the speaker 12, and the microphone 13 on the ceiling. For example, in the present invention, the microphone 13 may be installed on the ceiling, and the camera 11 and the speaker 12 may be set on a desk.

[0025] 5 is a flowchart showing the operation of the processor 14. First, the processor 14 accepts a selection operation on the user terminal 40 (S11). In the example of FIG. 2, the processor 14 is installed near the user u2, and the user u2 performs a selection operation on the user terminal 40.

[0026] The processor 14 acquires the location information of the user terminal 40 that has accepted the selection operation (S12). The location information of the user terminal 40 corresponds to the relative position of the user terminal 40 with respect to the microphone 13 in the signal processing device 20. The location information is stored in advance in the memory 15, for example. The user inputs the location information of the user terminal 40 in advance.

[0027] 6, 7, and 8 are diagrams illustrating an example of a GUI for inputting position information of the user terminal 40 in advance. The GUI is executed on the PC 30, for example, as an application program of the PC 30. The GUI first displays a screen for receiving height information of the microphone 13, as shown in FIG. 6, and receives the height information of the microphone 13. The height information of the microphone 13 corresponds to the distance from the desk on which the user terminal 40 is placed to the ceiling on which the signal processing device 20 is installed. Next, the GUI displays a screen for receiving planar position information of the microphone 13, as shown in FIG. 7, and receives the planar position information of the microphone 13. For example, the GUI displays a square image corresponding to the microphone 13. An administrator of the audio conference system drags the image to input the planar position information of the microphone 13. The GUI also displays a screen for receiving planar position information of the user terminal 40, as shown in FIG. 8, and receives the planar position information of the user terminal 40. For example, the GUI displays a round image corresponding to the user terminal 40. The user drags the image to input the planar position information of the user terminal 40. This allows the user to input relative position information of the user terminal 40 with respect to the microphone 13. The memory 15 stores the input position information. The processor 14 reads the position information of the user terminal 40 from the memory 15. The user can also change the position information if the user changes the position of the user terminal.

[0028] The position of the user terminal 40 may also be determined based on an image acquired by the camera 11, for example. The processor 14 recognizes the user terminal 40 from the image acquired by the camera 11 using a trained algorithm, such as a neural network, that is trained to recognize the user terminal 40 from the image acquired by the camera 11. The processor 14 determines the position information of the user terminal 40 using control information of the camera 11 that recognized the user terminal 40. Alternatively, the processor 14 may determine the position information of the user terminal 40 by referencing a table or function that defines the relationship between the image of the user terminal 40 in the image acquired by the camera 11 and the position information of the user terminal 40. If the microphone 13 and the camera 11 are positioned differently, the processor 14 also obtains relative position information of the microphone 13 and the camera 11, and determines the position information of the user terminal 40 based on the positional relationship between the microphone 13 and the user terminal 40.

[0029] The processor 14 sets the first sound collection beam b1 of the microphone 13 based on the acquired position information of the user terminal 40 (S13). The processor 14 adds a delay to each microphone 13 in advance to align the phase so as to form a sound collection beam with high sensitivity in the direction of the selected user terminal 40, for example. As a result, the first sound collection beam b1 is directed toward the user terminal 40. Therefore, the first sound collection beam b1 is directed toward user u2 who is near the user terminal 40 and has performed a selection operation on the user terminal 40. User u2 has not spoken, but has the intention to speak. User u2 performs a selection operation on the user terminal 40 when he or she has the intention to speak. When the selection operation on the user terminal 40 is performed, the first sound collection beam b1 is directed toward user u2. As a result, the sound collection beam is directed toward user u2 in advance before the user u2 starts speaking. User u2 can enjoy a new customer experience in which the sound collection beam can be directed toward himself or herself according to his or her intention to speak before speaking.

[0030] (Variation 1) Fig. 9 is a schematic elevation view of the interior of a room according to Modification 1. Fig. 10 is a flowchart showing the operation of the processor 14 according to Modification 1.

[0031] The processor 14 determines whether the user terminal 40 is currently selected (S20). If the user terminal 40 is currently selected (YES in S20), the processor 14 executes the operation of the flowchart shown in FIG. 5. If the user terminal 40 is not currently selected (NO in S20), the processor 14 detects the speaker's voice (S21). The processor 14 analyzes the sound signal acquired from the microphone 13 to estimate the direction of arrival of the sound (S22). The sound signal analysis method may be any method, such as a cross-correlation method, a delay-and-sum method, or a MUSIC (Multiple Signal Classification) method. In the cross-correlation method, the processor 14 calculates the cross-correlation of sound signals from multiple microphones, for example. The processor 14 determines the cross-correlation peak of sound signals from, for example, two microphones. Furthermore, the processor 14 determines the cross-correlation peak of sound signals from another two microphones. The processor 14 estimates the direction of arrival of the sound based on the multiple cross-correlation peaks calculated in this manner.

[0032] The processor 14 directs the second sound collection beam b2 in the direction from which the voice of the detected speaker comes (S23). Also, the processor 14 according to the first modification acquires the position information of the user terminal 40 that has accepted the selection operation, and sets the first sound collection beam b1 of the microphone 13 based on the acquired position information of the user terminal 40.

[0033] According to the configuration of Modification 1, the first sound collection beam b1 is directed in the direction of the selected terminal while the second sound collection beam b2 tracks the speaker. User u2 can direct the sound collection beam toward himself / herself according to his / her intention to speak before speaking, and user u1 who is not operating user terminal 40 can also direct the sound collection beam toward himself / herself by speaking if no user terminal is selected.

[0034] (Variation 2) Fig. 11 is a block diagram showing the configuration of an audio conference system 1 according to Modification 2. Components common to those in Fig. 1 are given the same reference numerals, and descriptions thereof will be omitted. In the audio conference system 1 according to Modification 2, a plurality of user terminals (user terminal 40 and user terminal 41) are connected to a signal processing device 20.

[0035] Fig. 12 is a flowchart showing the operation of the processor 14 according to Modification 2. The same processes as those in Fig. 5 are denoted by the same reference numerals, and the description thereof will be omitted.

[0036] The processor 14 of the second modification acquires identification information for identifying the user terminal that has accepted the selection operation (S101). For example, the identification information is information unique to each user terminal, such as the serial number or MAC address of the user terminal. The user terminal that has accepted the selection operation (user terminal 40 or user terminal 41) transmits the identification information to the signal processing device 20.

[0037] Processor 14 acquires the location information of the corresponding user terminal based on the acquired identification information (S102). For example, processor 14 refers to a database that stores identification information and location information in association with each other, and acquires the location information of the user terminal that has accepted the selection operation. The database is stored in memory 15 in advance, for example. The user inputs the identification information and location information of user terminal 40 and user terminal 41 in advance. Processor 14 reads the location information of user terminal 40 or user terminal 41 that corresponds to the identification information from memory 15. The user can also change the location information if he or she changes the location of the user terminal.

[0038] Alternatively, processor 14 may use camera 11 to acquire specific information about the user terminal that accepted the selection operation, and may acquire location information about the user terminal that accepted the selection operation. When user terminal 40 and user terminal 41 accept a selection operation, they light or blink LED 26 in a specific display mode. When user terminal 40 and user terminal 41 accept a selection operation, they light or blink LED 26 in orange, for example. Such a specific display mode of LED 26 corresponds to the specific information.

[0039] The processor 14 recognizes a user terminal that is lit or flashing in the specific display mode from the image of the camera 11. The processor 14 obtains location information corresponding to the recognized user terminal. For example, the processor 14 obtains location information of the user terminal using control information of the camera 11 that recognized the user terminal. Alternatively, the processor 14 may obtain the location information of the user terminal by referencing a table or function that defines the relationship between the image of the user terminal and the location information of the user terminal. If the positions of the microphone 13 and the camera 11 are different, the processor 14 also obtains relative location information of the microphone 13 and the camera 11, and obtains location information of the user terminal 40 based on the positional relationship between the microphone 13 and the user terminal 40.

[0040] The processor 14 sets the first sound collection beam b1 of the microphone 13 based on the acquired position information of the user terminal 40 (S13). The processor 14 adds a delay in advance to align the phase so as to form a sound collection beam with high sensitivity in the direction of the selected user terminal 40, for example. As a result, the first sound collection beam b1 is directed in the direction of the user terminal that has accepted the selection operation among multiple user terminals. Therefore, in the audio conference system of the second modification, even when multiple users use multiple user terminals, it is possible to reflect the intention of the users to speak in the sound collection beam.

[0041] (Variation 3) The user terminal 40 according to the third modification changes the display mode of the LED 26 according to the state of the first sound collection beam b1. The user terminal 41 shown in the third modification may also perform the same operation as the user terminal 40 according to the third modification.

[0042] The first sound collection beam b1 includes a first state in which it faces the direction of the user terminal 40, a second state in which it does not face the direction of the user terminal 40, and a third state in which it transitions from the second state to the first state.

[0043] The first state is a state in which the voice of the user who performed the selection operation on the user terminal 40 is acquired. For example, in the first state, the user terminal 40 lights up the LED 26 in green. The second state is a state in which the voice of the user who performed the selection operation on the user terminal 40 is not acquired. For example, in the second state, the user terminal 40 lights up the LED 26 in red. For example, in the third state, the user terminal 40 flashes the LED 26 in orange.

[0044] This allows the user who performed the selection operation to easily determine from the display state of the LED 26 whether their own voice is available for acquisition, whether it is not available for acquisition, or whether the state is currently changing.

[0045] (Variation 4) If a predetermined time or more has passed without any speech after accepting the selection operation, the processor 14 may cancel the first sound collection beam b1 and redirect it from the direction of the user terminal 40 to the direction of another speaker. This makes it possible to prevent the sound collection beam from being continuously directed at a user who is not speaking.

[0046] When the processor 14 receives a special operation (e.g., double touch) as a selection operation, the processor 14 may set the user terminal 40 to a fixed mode in which the first sound collection beam b1 continues to be directed toward the user terminal 40 (and is not directed toward other speakers) even after a predetermined time has elapsed. When the user terminal 40 is in the fixed mode, the user terminal 40 may notify the user that the fixed mode is in effect by, for example, lighting up the display mode of the LED 26 in white. This allows priority to be given to the speech of a user who has operated the terminal with the intention of speaking.

[0047] (Variation 5) The processor 14 of variant example 5 has a first mode in which the first sound collection beam b1 is directed in the direction of the user terminal 40 that has received the selection operation, and a second mode in which the first sound collection beam b1 is directed in a direction other than the user terminal 40 that has received the selection operation.

[0048] Furthermore, processor 14 may set the first mode when a first operation (for example, a short-time press) is received as a selection operation, and may set the second mode when a second operation (for example, a long-time press) is received.

[0049] In the above embodiment and each modified example, the processor 14 directed the first sound collection beam b1 in the direction of the user terminal 40 that accepted the selection operation. However, the processor 14 may direct the first sound collection beam b1 in a direction other than the user terminal 40 that accepted the selection operation, so as not to acquire the voice of a user who is near the user terminal 40.

[0050] In variant example 5, a user who has no intention of speaking and does not want their voice to be recorded can prevent the sound collection beam from being directed at them by selecting the user terminal 40, for example by pressing the user terminal 40 for a long time.

[0051] (Variation 6) In the audio conference system 1 according to the sixth modification, a plurality of user terminals (for example, the user terminal 40 and the user terminal 41 shown in FIG. 11) are connected to the signal processing device 20. In FIG.

[0052] When the number of user terminals that have received a selection operation is multiple, the processor 14 directs the first sound collection beam b1 toward any one of the user terminals based on a predetermined priority. For example, the processor 14 directs the first sound collection beam b1 toward the user terminal that received the selection operation first, and disables the selection operation of the user terminal that received the selection operation next for a certain period of time.

[0053] Alternatively, the administrator of the audio conference system may input the priority of the user terminals in advance. For example, the administrator of the audio conference system may set a higher priority for a user terminal located near the conference chairperson.

[0054] Furthermore, the administrator of the audio conference system may set in advance for each user terminal whether to enable or disable the selection operation. This allows the administrator of the audio conference system to disable the selection operation of a user terminal installed near an observer who only listens to the conference, for example, so that the first sound collection beam b1 is not directed thereto.

[0055] (Variation 7) In the audio conference system 1 according to the seventh modification, a plurality of user terminals are connected to the signal processing device 20. The plurality of user terminals include a first user terminal that belongs to a certain group and a second user terminal that does not belong to the group. The processor 14 directs the first sound collection beam b1 only in the direction of the first user terminal, and does not direct the first sound collection beam b1 in the direction of the second user terminal.

[0056] An administrator of the audio conference system sets multiple user terminals as first user terminals or second user terminals in advance. For example, the administrator of the audio conference system sets a user terminal installed near an observer who will only listen to the conference as the second user terminal. This allows the administrator of the audio conference system to disable the selection operation of the user terminal installed near the observer so that the first sound collection beam b1 is not directed at the user terminal. Alternatively, the processor 14 may identify, among the multiple connected user terminals, a user terminal that can be recognized in the image of the camera 11 as the first user terminal, and identify a user terminal that cannot be recognized in the image of the camera 11 as the second user terminal. This allows the processor 14 to direct the first sound collection beam b1 only at conference participants who are captured by the camera 11.

[0057] The description of the present embodiment should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention includes the scope equivalent to the claims. [Explanation of symbols]

[0058] 1: Audio conference system, 10: Display device, 11: Camera, 12: Speaker, 13: Microphone, 14: Processor, 15: Memory, 16: I / F, 20: Signal processing device, 40: User terminal, 41: User terminal

Claims

1. A sound pickup control method for an audio conference system including a microphone and at least one user terminal corresponding to the microphone, comprising: Accepting a selection operation on the user terminal, acquiring location information of the user terminal that has accepted the selection operation; setting a first sound collection beam of the microphone based on the acquired position information; Sound pickup control method.

2. The microphone can set sound collection beams in multiple directions, Detecting a speaker and directing a second sound collection beam in the direction of the detected speaker; When the selection operation is accepted, the first sound collection beam is directed toward the user terminal. The sound collection control method according to claim 1 .

3. acquiring identification information for identifying the user terminal that accepted the selection operation; referencing a database that stores the specific information and the location information in association with each other, and acquiring the location information of the user terminal that has accepted the selection operation; The sound collection control method according to claim 1 or 2.

4. acquiring, using a camera, identification information for identifying the user terminal that has accepted the selection operation, and acquiring location information of the user terminal that has accepted the selection operation; The sound collection control method according to claim 1 or 2.

5. The first sound collection beam is a first state facing the user terminal; a second state not facing the user terminal; a third state in which the second state transitions to the first state; Including, changing a display mode of a display device according to the first state, the second state, and the third state; The sound collection control method according to claim 1 or 2.

6. When a predetermined time or more has elapsed, the first sound collection beam is cancelled. The sound collection control method according to claim 1 or 2.

7. a first mode in which the first sound collection beam is directed toward the user terminal that has received the selection operation; a second mode in which the first sound collection beam is directed in a direction other than the user terminal that has received the selection operation; 3. The sound collection control method according to claim 1, further comprising:

8. the number of the user terminals is plural, When the number of the user terminals that have received the selection operation is plural, the first sound collection beam is directed toward any one of the user terminals based on a predetermined priority order. The sound collection control method according to claim 1 or 2.

9. the user terminals include a first user terminal belonging to a certain group and a second user terminal not belonging to the group; directing the first sound collection beam in the direction of the first user terminal; The sound collection control method according to claim 8.

10. A sound pickup control device for an audio conference system including a microphone and at least one user terminal corresponding to the microphone, Accepting a selection operation on the user terminal, acquiring location information of the user terminal that has accepted the selection operation; setting a first sound collection beam of the microphone based on the acquired position information; A sound pickup control device equipped with a control unit.

11. The microphone can set sound collection beams in multiple directions, The control unit Detecting a speaker and directing a second sound collection beam in the direction of the detected speaker; When the selection operation is accepted, the first sound collection beam is directed toward the user terminal. The sound collection control device according to claim 10.

12. The control unit acquiring identification information for identifying the user terminal that accepted the selection operation; referencing a database that stores the specific information and the location information in association with each other, and acquiring the location information of the user terminal that has accepted the selection operation; The sound collection control device according to claim 10 or 11.

13. The control unit acquires, using a camera, identification information for identifying the user terminal that has accepted the selection operation, and acquires location information of the user terminal that has accepted the selection operation. The sound collection control device according to claim 10 or 11.

14. The first sound collection beam is a first state facing the user terminal; a second state not facing the user terminal; a third state in which the second state transitions to the first state; Including, the control unit changes a display mode of the display device according to the first state, the second state, and the third state. The sound collection control device according to claim 10 or 11.

15. The control unit cancels the first sound collection beam when a predetermined time or more has elapsed. The sound collection control device according to claim 10 or 11.

16. a first mode in which the first sound collection beam is directed toward the user terminal that has received the selection operation; a second mode in which the first sound collection beam is directed in a direction other than the user terminal that has received the selection operation; The sound collection control device according to claim 10 or 11, comprising:

17. the number of the user terminals is plural, When the number of the user terminals that have received the selection operation is plural, the control unit directs the first sound collection beam in the direction of any one of the user terminals based on a predetermined priority order. The sound collection control device according to claim 10 or 11.

18. the user terminals include a first user terminal belonging to a certain group and a second user terminal not belonging to the group; The control unit directs the first sound collection beam in a direction toward the first user terminal. The sound collection control device according to claim 17.

Citation Information

Patent Citations

  • Voice signal processing device, voice signal processing method, and voice signal processing program

    WO2023149254A1