Audio signal processing method, terminal, audio signal processing system, management device
The audio signal processing system allows terminals to independently perform localization processing by acquiring control information, addressing the lack of platform-dependent localization mechanisms in existing systems and achieving accurate audio-visual positioning in online meetings.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- YAMAHA CORP
- Filing Date
- 2026-02-09
- Publication Date
- 2026-04-28
AI Technical Summary
Existing audio-visual localization processes in online meetings rely on distribution platforms that lack a localization control mechanism, making it difficult to achieve appropriate sound localization without platform dependency.
An audio signal processing system where terminals independently acquire localization control information to perform audio-visual localization processing, allowing terminals to determine their own positions and output localized audio signals without relying on the distribution platform.
Enables appropriate sound localization processing without needing a distribution platform, ensuring accurate audio-visual positioning in online meetings.
Smart Images

Figure 2026071366000001_ABST
Abstract
Description
Technical Field
[0001] One embodiment of this invention relates to an audio signal processing system, an audio signal processing method in the audio signal processing system, a terminal that executes the audio signal processing method, and a management device.
Background Art
[0002] Conventionally, a configuration in which a distribution platform such as a server that manages an online meeting performs audio-visual localization is known. For example, Patent Document 1 describes a configuration in which a management device (communication server) that manages an online meeting controls the audio-visual localization of each terminal.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, when there is no localization control mechanism on the existing distribution platform side, the localization process as in Patent Document 1 cannot be realized.
[0005] In consideration of the above circumstances, one aspect of the present disclosure aims to provide an audio signal processing method that can realize appropriate audio-visual localization processing without depending on a distribution platform.
Means for Solving the Problems
[0006] The audio signal processing method is used in an audio signal processing system composed of a plurality of terminals that output audio signals. Each of the plurality of terminals acquires localization control information that determines the audio-visual localization position of its own terminal in the audio signal processing system, performs localization processing on the audio signal of its own terminal based on the acquired localization control information, and outputs the audio signal after the localization processing. [Effects of the Invention]
[0007] One embodiment of this invention can achieve appropriate sound localization processing without depending on the distribution platform. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram showing the configuration of the sound signal processing system 1. [Figure 2] This is a block diagram showing the configuration of terminal 11A. [Figure 3] This is a flowchart showing the operation of terminal 11A. [Figure 4] This is a flowchart showing the operation of the control device 12. [Figure 5] This figure shows an example of localization control information. [Figure 6] This is a flowchart showing the operation of terminal 11A in the modified example 1. [Figure 7] This is a flowchart showing the operation of terminal 11A in the modified example 3. [Figure 8] This flowchart shows the operation of the control device 12 according to the modified example 3. [Figure 9] This is a block diagram illustrating the concept of the video signals transmitted by each device in the sound signal processing system 1. [Figure 10] This is a block diagram illustrating the concept of the sound localization position of each terminal in the sound signal processing system 1A according to modified example 5. [Modes for carrying out the invention]
[0009] Figure 1 is a block diagram showing the configuration of the sound signal processing system 1. The sound signal processing system 1 comprises a plurality of terminals (terminal 11A, terminal 11B, and terminal 11C) and a management device 12.
[0010] Terminals 11A, 11B, 11C, and the management device 12 are connected via a network 13. The network 13 includes a LAN (Local Area Network) or the Internet.
[0011] Terminals 11A, 11B, and 11C are information processing devices such as personal computers.
[0012] Figure 2 is a block diagram showing the configuration of terminal 11A. Although Figure 2 shows the configuration of terminal 11A as a representative example, terminals 11B and 11C have the same configuration and functions.
[0013] Terminal 11A includes a display 201, a user interface 202, a CPU 203, RAM 204, a network interface 205, flash memory 206, a microphone 207, a speaker 208, and a camera 209. The microphone 207, speaker 208, and camera 209 may be built into terminal 11A or connected as external devices.
[0014] The CPU 203 is a control unit that reads a program stored in the flash memory 206, which is a storage medium, into the RAM 204 to perform a predetermined function. Note that the program read by the CPU 203 does not necessarily have to be stored in the flash memory 206 within the device itself. For example, the program may be stored in the storage medium of an external device such as a server. In this case, the CPU 203 can simply read the program from the server into the RAM 204 and execute it each time.
[0015] Flash memory 206 stores the application program for online meetings. CPU 203 loads the application program for online meetings into RAM 204.
[0016] The CPU 203 outputs the sound signal acquired by the microphone 207 to the management device 12 via the network I / F 205 according to the function of the application program. The CPU 203 outputs a two-channel (stereo channel) sound signal. Also, the CPU 203 outputs the video signal acquired by the camera 209 to the management device 12 via the network I / F 205.
[0017] The management device 12 receives sound signals and video signals from the terminals 11A, 11B, and 11C. The management device 12 mixes the sound signals received from the terminals 11A, 11B, and 11C. Also, the management device 12 synthesizes the video signals received from the terminals 11A, 11B, and 11C into one video signal. The management device 12 distributes the mixed sound signal and the synthesized video signal to the terminals 11A, 11B, and 11C.
[0018] Each CPU 203 of the terminals 11A, 11B, and 11C outputs the sound signal distributed from the management device 12 to the speaker 208. Also, the CPU 203 outputs the video signal distributed from the management device 12 to the display 201. Thereby, the users of each terminal can hold an online meeting.
[0019] Figure 3 is a flowchart showing the operation at the start of an online meeting of the terminal 11A. Figure 4 is a flowchart showing the operation at the start of an online meeting of the management device 12. The terminals 11B and 11C perform the same operations as the terminal 11A.
[0020] First, the terminal 11A transmits its Mac address, as an example of the unique identification information of its own terminal, to the management device 12 (S11). Similarly, the terminals 11B and 11C transmit their Mac addresses, as an example of the unique identification information of their own terminals, to the management device 12. The management device 12 receives the Mac addresses from the terminals 11A, 11B, and 11C respectively (S21). Then, the management device 12 generates localization control information. The localization control information is information for determining the audio-visual localization positions of each terminal in the sound signal processing system 1.
[0021] Figure 5 shows an example of localization control information. Localization control information associates terminal identification information with localization position information for each terminal. In this example, the terminal identification information is the MAC address. The identification information may also be the username or email address of each terminal, or a unique ID assigned by the management device 12 in online meetings.
[0022] In this example, the information indicating the localization position is the information indicating the panning parameters (volume balance of the L channel and R channel). For example, the localization control information for terminal 11A indicates a volume balance of 80% for the L channel and 20% for the R channel. In this case, the sound signal from terminal 11A is localized to the left. The localization control information for terminal 11B indicates a volume balance of 50% for the L channel and 50% for the R channel. In this case, the sound signal from terminal 11B is localized to the center. The localization control information for terminal 11C indicates a volume balance of 20% for the L channel and 80% for the R channel. In this case, the sound signal from terminal 11C is localized to the right.
[0023] The management device 12 determines the localized position based on the order in which MAC addresses were received, as an example. In other words, the management device 12 determines the localized position based on the order in which users connected to the online meeting.
[0024] In this example, the management device 12 positions the terminals in the order they joined the online meeting, from left to right. For example, if three terminals join the online meeting, the management device 12 positions the first terminal to join on the left, the next terminal in the center, and the last terminal to join on the right. Terminal 11A connects to the management device 12 first and sends its MAC address, then terminal 11B connects to the management device 12 and sends its MAC address, and finally terminal 11C connects to the management device 12 and sends its MAC address. Therefore, the management device 12 positions terminal 11A on the left, terminal 11B in the center, and terminal 11C on the right.
[0025] Of course, generating such localization control information is just one example. For instance, the management device 12 may localize the terminal that first joined the online meeting to the right, the next terminal to the center, and the last terminal to the left. The number of terminals participating in the online meeting is also not limited to this example. For example, if two terminals are participating in the online meeting, the management device 12 may localize the terminal that first joined to the meeting to the right, and the next terminal to the left. In any case, the management device 12 localizes each of the multiple terminals participating in the online meeting to different positions.
[0026] Furthermore, the positioning control information may be generated based on the unique identification information of each terminal. For example, if the identification information is a MAC address, the management device 12 may determine the positioning in ascending order of MAC addresses. For example, in the case of Figure 5, the management device 12 positions terminal 11A, which has the smallest MAC address, on the left, then terminal 11B, which has the smallest MAC address, in the center, and terminal 11C on the right.
[0027] Furthermore, the positioning control information may be generated based on the attributes of the users of each terminal. For example, each user of a terminal has an account level in the online meeting as an attribute. The positioning control information is determined in ascending order of account level. The management device 12, for example, positions users with higher account levels in the center, and users with lower account levels at the far left or far right.
[0028] The management device 12 distributes the localization control information generated as described above to terminals 11A, 11B, and 11C (S23). Terminals 11A, 11B, and 11C each acquire the localization control information (S12). Then, terminals 11A, 11B, and 11C each apply localization processing to the sound signals acquired by microphone 207 (S13). For example, terminal 11A applies panning processing to the volume balance of the stereo channel sound signals acquired by microphone 207 so that the L channel is 80% and the R channel is 20%. Terminal 11B applies panning processing to the volume balance of the stereo channel sound signals acquired by microphone 207 so that the L channel is 50% and the R channel is 50%. Terminal 11C performs panning processing on the volume balance of the stereo channel audio signals acquired by microphone 207 so that the left channel is 20% and the right channel is 80%.
[0029] Terminals 11A, 11B, and 11C each output an audio signal after localization processing (S14). The management device 12 receives the audio signals from terminals 11A, 11B, and 11C and mixes them (S24), then distributes the mixed audio signal to terminals 11A, 11B, and 11C (S25).
[0030] In this manner, the sound signal processing system 1 of this embodiment outputs sound signals that have undergone localization processing at each terminal participating in the online meeting. Therefore, the management device 12, which is the distribution platform for the online meeting, does not need to perform localization processing. Thus, the sound signal processing system 1 of this embodiment can achieve appropriate sound image localization processing independently of the distribution platform, even if the existing distribution platform does not have a localization control mechanism.
[0031] (Variation 1) In the above embodiment, an example was shown in which the management device 12 generates localization control information. However, localization control information may also be generated at each terminal. Figure 6 is a flowchart showing the operation of terminal 11A according to Modification 1. Operations common to Figure 3 are denoted by the same reference numerals and their explanations are omitted. Terminals 11B and 11C perform the same operations as terminal 11A.
[0032] Terminal 11A obtains a participant list from the management device 12 (S101). The participant list includes the time each terminal joined the online meeting, and identification information for each terminal (e.g., MAC address, username, email address, or a unique ID assigned by the management device 12 in the online meeting).
[0033] Terminal 11A generates localization control information based on the acquired participant list (S102). The generation rules for localization control information based on the participant list are the same for all terminals of the sound signal processing system 1. For example, the generation rule establishes a one-to-one correspondence between the order in which terminals joined the online meeting and their localization positions. For example, if three terminals are participating in the online meeting, the generation rule localizes the terminal that joined first to the left, the terminal that joined next to the center, and the terminal that joined last to the right.
[0034] In the modified example 1, the sound signal processing system 1 generates and acquires localization control information at each terminal, eliminating the need for the management device 12 to generate localization control information. The management device 12 only needs to have a participant list and distribute 2-channel (stereo channel) sound signals, and does not need to perform any processing related to localization. Therefore, the configuration and operation of the sound signal processing system 1 of this embodiment can be realized on any platform that has a participant list and distributes 2-channel (stereo channel) sound signals.
[0035] (Modification 2) In the above embodiment, the information indicating the localization position was information indicating the panning parameters (volume balance of the L channel and R channel). However, the localization control information may be, for example, an HRTF (Head Related Transfer Function). An HRTF represents a transfer function from a certain virtual sound source position to the user's right and left ears. For example, the localization control information of terminal 11A indicates an HRTF that localizes to the left of the user. In this case, terminal 11A performs binaural processing by convolving an HRTF that localizes to the left of the user into each of the L channel and R channel sound signals. Also, for example, the localization control information of terminal 11B indicates an HRTF that localizes behind the user. In this case, terminal 11B performs binaural processing by convolving an HRTF that localizes behind the user into each of the L channel and R channel sound signals. Also, for example, the localization control information of terminal 11C indicates an HRTF that localizes to the right of the user. In this case, terminal 11C performs binaural processing by convolving an HRTF that localizes the sound to the user's right side into the respective audio signals of the L channel and R channel.
[0036] The panning parameter is the left-right volume balance, and the localization control information is one-dimensional (left-right position) information. Therefore, with the panning parameter, when there are many participants in an online meeting, the localization positions of each user's voice become close together, making it difficult to localize each user's voice to a different position. However, the localization control information of HRTF is three-dimensional information. Therefore, the sound signal processing system 1 of Modification 2 can localize each user's voice to a different position even when there are more participants in an online meeting.
[0037] (Variation 3) The sound signal processing system 1 according to Modification 3 is an example in which the management device 12 or each terminal generates localization control information based on the video signal. Figure 7 is a flowchart of the operation of terminal 11A according to Modification 3. Operations common to Figure 3 are denoted by the same reference numerals and their explanations are omitted. Terminals 11B and 11C perform the same operations as terminal 11A. Figure 8 is a flowchart of the operation of the management device 12 according to Modification 3. Operations common to Figure 4 are denoted by the same reference numerals and their explanations are omitted. Figure 9 is a block diagram illustrating the concept of the video signal transmitted by each device in the sound signal processing system 1.
[0038] Terminals 11A, 11B, and 11C output the video signal acquired by the camera 209 to the management device 12. At this time, terminals 11A, 11B, and 11C superimpose identification information onto the video signal (S201). For example, terminals 11A, 11B, and 11C encode some of the pixels in the video signal with the identification information.
[0039] Terminals 11A, 11B, and 11C each encode identification information using multiple pixels from the origin (0,0), which is the top-leftmost pixel of the video signal acquired by camera 209. For example, terminals 11A, 11B, and 11C encode the RGB values of pixels with identification information, treating white (R,G,B=255,255,255) as 1 bit data and black (R,G,B=0,0,0) as 0 bit data. If the number of pixels in the video signal is, for example, 1280×720, terminals 11A, 11B, and 11C encode identification information using 1280 pixels in one line (0,0~1279,0) of the video signal where the Y=0 coordinate is.
[0040] The management device 12 receives video signals from terminals 11A, 11B, and 11C (S301) and decodes the identification information (S302). The management device 12 may combine the video signals received from terminals 11A, 11B, and 11C as they are, or it may remove 1280 pixels of one line that corresponds to the Y=0 coordinate before combining them. Alternatively, the management device 12 may replace all 1280 pixels of one line that corresponds to the Y=0 coordinate with white (R,G,B=255,255,255) or black (R,G,B=0,0,0) before combining the video signals.
[0041] When the management device 12 directly combines the video signals received from terminals 11A, 11B, and 11C, the video of each participant displayed during the online meeting will consist of pixels encoded only on the top line, as shown in Figure 9. However, since only the top line of the video is encoded, this does not hinder the viewing of the video during the online meeting.
[0042] The audio signal processing system 1 in Modification 3 is an example in which each terminal can transmit identification information via a video signal. Therefore, the audio signal processing system 1 in Modification 3 can acquire the identification information of each terminal even if the online meeting platform does not have a means to receive identification information such as MAC addresses.
[0043] The identification information may be decrypted at each terminal. In this case, each terminal generates localization control information based on the decrypted identification information. In this case, the rules for generating localization control information based on the identification information are the same for all terminals of the sound signal processing system 1. In this case, the management device 12 does not need to decrypt the identification information. Therefore, the sound signal processing system 1 of the modified example 3 does not require the management device 12 to manage identification information such as MAC addresses, and can be implemented with any distribution platform that distributes 2-channel (stereo channel) sound signals.
[0044] Furthermore, when decoding the identification information at each terminal, it is preferable that each terminal encodes the RGB values of multiple (e.g., 4x4) pixels in the video signal into bit data of 1 (R,G,B=255,255,255) or bit data of 0 (R,G,B=0,0,0). This ensures that even if the management device 12 reduces the size of each terminal's video signal to, for example, 1 / 4 and combines them, the encoded pixels remain. Therefore, each terminal can properly decode the identification information.
[0045] (Modification 4) In the modified example 4 of the sound signal processing system 1, each terminal performs a process to add indirect sound to the sound signal. By adding indirect sound to the sound signal, each terminal in the modified example 4 of the sound signal processing system 1 can reproduce a sound field that simulates a conversation taking place in a predetermined acoustic space such as a conference room or hall.
[0046] Indirect sound is added, for example, by convolving an impulse response, measured in advance in a predetermined acoustic space that is the target of sound field reproduction, into the sound signal. Indirect sound includes early reflections and rear reverberation. Early reflections are clearly defined reflections of sound from the direction of arrival, while rear reverberation is a reflection whose direction of arrival is not determined. Therefore, each terminal may perform binaural processing by convolving an HRTF onto the sound signal acquired by each terminal such that the sound image is localized at the position indicated by the position information of each sound source of the early reflections. Alternatively, early reflections may be generated based on information indicating the position and level of each sound source of the early reflections. Each terminal performs delay processing on the sound signal acquired by each terminal according to the position of each sound source of the early reflections, and controls the level of the sound signal based on the level information of each sound source of the early reflections. As a result, each terminal can clearly reproduce the early reflections in a predetermined acoustic space.
[0047] Furthermore, each terminal may reproduce the sound field of a different acoustic space. The user of each terminal specifies the acoustic space to be reproduced. Each terminal obtains spatial information indicating the specified acoustic space from the management device 12, etc. The spatial information includes impulse response information. Each terminal adds indirect sound to the sound signal using the impulse response of the specified spatial information. Note that the spatial information may also be information indicating the size of a predetermined acoustic space such as a conference room or hall, or the reflectivity of the walls. Each terminal lengthens the reverberation as the size of the acoustic space increases. Also, each terminal increases the level of the initial reflections as the reflectivity of the walls increases.
[0048] (Variation 5) Figure 10 is a block diagram illustrating the concept of the sound localization position of each terminal in the sound signal processing system 1A according to the modified example 5. In the sound signal processing system 1A of modified example 5, users of terminals 11A, 11B, and 11C perform a remote ensemble (remote session). Terminals 11A, 11B, and 11C each acquire sound signals from musical instruments via microphones or signal lines such as audio cables. Terminals 11A, 11B, and 11C each apply localization processing to the acquired sound signals based on localization control information. Terminals 11A, 11B, and 11C output the localized sound signals to the first management device 12A.
[0049] The localization control information is the same as in the various examples described above. However, in the localization control information of Modification Example 5, it is preferable to generate it based on attributes. In this example, the attribute is the type of sound (instrument). For example, the localization position for singing sounds (vocals) is determined to be the front center, the localization position for string instruments such as guitars is the left side, the localization position for percussion instruments such as drums is the rear center, and the localization position for keyboard instruments such as electronic pianos is determined to be the right side.
[0050] For example, terminal 11A acquires vocal and guitar sound signals. The vocal signal is acquired by a microphone, and the guitar signal is acquired via line (audio cable). Terminal 11A performs binaural processing on the vocal signal by convolving an HRTF that localizes the sound to the center in front of the user. Terminal 11A also performs binaural processing on the guitar signal by convolving an HRTF that localizes the sound to the left of the user.
[0051] Terminal 11B acquires the sound signal from the electronic piano. The sound signal from the electronic piano is acquired via a line (audio cable). Terminal 11B performs binaural processing on the sound signal from the electronic piano by convolving an HRTF that localizes the sound to the user's right side.
[0052] Terminal 11C acquires the drum sound signal. The drum sound signal is acquired by a microphone. Terminal 11C performs binaural processing on the drum sound signal by convolving an HRTF that localizes the sound to the center behind the user.
[0053] Of course, in variation 5 as well, the localization processing is not limited to binaural processing; panning processing may also be used. In this case, the localization control information indicates the left and right localization positions (left and right volume balance).
[0054] Terminals 11A, 11B, and 11C output the audio signals, which have undergone localization processing as described above, to the first management device 12A. The first management device 12A has the same configuration and functions as the management device 12 described above. The first management device 12A mixes the audio signals received from terminals 11A, 11B, and 11C. The first management device 12A may also receive video signals from terminals 11A, 11B, and 11C and combine them into a single video signal. The first management device 12A distributes the mixed audio signal and the combined video signal to the listener.
[0055] This allows listeners viewing a remote session to perceive the sound of each instrument as arriving from different locations. In the modified example 5, the first management device 12A only needs to distribute two channels (stereo channels) of audio signals. Therefore, the configuration and operation of the audio signal processing system 1A in the modified example 5 can be realized on any platform that distributes two channels (stereo channels) of audio signals.
[0056] Furthermore, terminals 11A, 11B, and 11C output the audio signals before localization processing to the second management device 12B. The second management device 12B has the same configuration and functions as management device 12 and the first management device 12A. The second management device 12B receives the audio signals that have not undergone localization processing from terminals 11A, 11B, and 11C and mixes them. The second management device 12B distributes the mixed audio signals to terminals 11A, 11B, and 11C.
[0057] As a result, users conducting remote sessions on terminals 11A, 11B, and 11C can hear sounds that have not undergone localization processing, making it easier to monitor the sound of each user. The second management device 12B only needs to distribute two channels (stereo channels) of audio signals. As a result, with a platform that distributes two channels (stereo channels) of audio signals, listeners viewing the remote session can hear the sound of each instrument as if it were coming from a different location, and users conducting remote sessions on terminals 11A, 11B, and 11C can hear sounds that are easy to monitor.
[0058] (Experimental variation 6) Each terminal in Modification 6 performs the same process as in Modification 4, adding indirect sound to the sound signal. However, each terminal generates a first sound signal with indirect sound added and a second sound signal without indirect sound added. The first sound signal is, for example, a sound signal that has undergone localization processing as described above. The second sound signal is, for example, a sound signal that has not undergone localization processing as described above.
[0059] This allows listeners viewing remote sessions to hear immersive sounds, such as those found in a concert hall, and enables users conducting remote sessions on terminals 11A, 11B, and 11C to hear sounds that are easy to monitor.
[0060] Furthermore, it is preferable that the indirect sound simulates the same acoustic space for all terminals. This allows users of terminals 11A, 11B, and 11C (performers in the remote session) located in remote locations to perceive that they are performing a live concert in the same acoustic space. For example, in the example shown in Figure 10, each terminal may output an audio signal with indirect sound from a large concert hall to the first management device 12A, and an audio signal with indirect sound from a small live venue to the second management device 12B. In this case, the first management device 12A distributes the audio signal with indirect sound from a large concert hall, and the second management device 12B distributes the audio signal with indirect sound from a small live venue. A listener may receive the audio signal distributed by the first management device 12A and listen to a remote session that reproduces the acoustics of a large concert hall, or receive the audio signal distributed by the second management device 12B and listen to a remote session that reproduces the acoustics of a small live venue.
[0061] (Example 7) Terminals 11A, 11B, and 11C may further perform processing to add ambient sound to their respective audio signals. Ambient sound includes ambient noise, listener cheers, applause, calls, shouts, singing, or other environmental sounds. This allows listeners viewing the remote session to also hear sounds such as the audience at the live venue, resulting in a more immersive listening experience.
[0062] Preferably, each terminal adds ambient sound to the first tone signal, but does not add ambient sound to the second tone signal. This allows listeners viewing the remote session to hear a realistic sound, and users conducting remote sessions on terminals 11A, 11B, and 11C can hear a sound that is easy to monitor.
[0063] In actual live venues, ambient sounds are generated randomly. Therefore, terminals 11A, 11B, and 11C may each be assigned different ambient sounds. This allows the ambient sounds to be generated randomly, enabling listeners to experience a more immersive sound.
[0064] Furthermore, ambient sounds such as cheers, shouts, and murmurs may differ for each performer participating in a remote session. For example, a terminal outputting a vocal sound signal may be given high-frequency and high-level cheers, shouts, and murmurs. A terminal outputting a drum sound signal may be given low-frequency and low-level cheers, shouts, and murmurs. Generally, in live performances, the frequency and level of cheers, shouts, and murmurs are high for the main performer, the vocalist, while they are low for other instruments (e.g., drums). Therefore, a terminal outputting a sound signal corresponding to the main performer in a live performance can reproduce a higher level of realism by giving high-frequency and high-level cheers, shouts, and murmurs.
[0065] The description of this embodiment is illustrative in all respects and not restrictive. The scope of the invention is indicated by the claims, rather than by the embodiments described above. Furthermore, the scope of the invention is intended to include all modifications within the meaning and scope equivalent to the claims. [Explanation of Symbols]
[0066] 1.1A…Sound signal processing system 11A, 11B, 11C… terminals 12…Management device 12A…1st management device 12B…Second management device 13…Network 201…Display unit 203…CPU 204...RAM 205…Network I / F 206... Flash memory 207... Mike 208...Speaker 209... Camera
Claims
1. A sound signal processing method used in a sound signal processing system consisting of multiple terminals that output sound signals, Each of the aforementioned multiple terminals is: It outputs a video signal containing unique identification information for each terminal. Based on the aforementioned identification information, localization control information is obtained that determines the sound image localization position of the terminal in the sound signal processing system. Based on the acquired localization control information, localization processing is applied to the sound signal of the terminal. The audio signal after the aforementioned localization processing is output. Audio signal processing method.
2. The aforementioned localization control information includes information that determines the left and right localization positions. The aforementioned positioning process includes panning. The sound signal processing method according to claim 1.
3. The aforementioned localization control information includes information that determines the three-dimensional localization position. The aforementioned localization processing includes binaural processing. The sound signal processing method according to claim 1 or claim 2.
4. The aforementioned localization control information is generated based on the user attributes of each terminal. The sound signal processing method according to claim 1 or claim 2.
5. By acquiring spatial information that indicates the acoustic space, Further processing is performed to add indirect sound corresponding to the acoustic space indicated by the spatial information to the sound signal of the terminal. The sound signal processing method according to claim 1 or claim 2.
6. A first sound signal with the aforementioned indirect sound added and a second sound signal without the aforementioned indirect sound are generated, and the first sound signal and the second sound signal are output, respectively. The sound signal processing method according to claim 5.
7. Further processing is performed to add ambient sound to the audio signal of the aforementioned terminal. The sound signal processing method according to claim 1 or claim 2.
8. The aforementioned ambient sound is different for each of the multiple terminals. The sound signal processing method according to claim 7.
9. In an audio signal processing system composed of multiple terminals, including the terminal itself, a video signal containing the terminal's unique identification information is output. Based on the aforementioned identification information, localization control information is generated to determine the sound image localization position of the terminal, Based on the acquired localization control information, localization processing is applied to the sound signal of the terminal. The audio signal after the aforementioned localization processing is output. A terminal equipped with a control unit.
10. The aforementioned localization control information includes information that determines the left and right localization positions. The aforementioned positioning process includes panning. The terminal according to claim 9.
11. The aforementioned localization control information includes information that determines the three-dimensional localization position. The aforementioned localization processing includes binaural processing. The terminal according to claim 9 or claim 10.
12. The aforementioned localization control information is generated based on the user attributes of each terminal. The terminal according to claim 9 or claim 10.
13. The control unit acquires spatial information indicating the acoustic space, Further processing is performed to add indirect sound corresponding to the acoustic space indicated by the spatial information to the sound signal of the terminal. The terminal according to claim 9 or claim 10.
14. The control unit generates a first sound signal with the indirect sound added and a second sound signal without the indirect sound added, and outputs the first sound signal and the second sound signal, respectively. The terminal according to claim 13.
15. Further processing is performed to add ambient sound to the audio signal of the aforementioned terminal. The terminal according to claim 9 or claim 10.
16. The aforementioned ambient sound is different for each of the multiple terminals. The terminal according to claim 15.
17. A sound signal processing system comprising multiple terminals and a management device, The aforementioned control device is The system receives video signals from the aforementioned multiple terminals, each terminal containing its own unique identification information. Localization control information is generated based on the identification information to determine the sound image localization position of each of the multiple terminals. Each of the aforementioned multiple terminals is: The positioning control information is acquired, Based on the acquired localization control information, localization processing is applied to the sound signal of the terminal. The sound signal after the localization processing described above is output, The management device mixes the audio signals output from each of the multiple terminals and distributes them to the multiple terminals. Audio signal processing system.
Citation Information
Patent Citations
Acoustic image localization control system, communication server, multipoint connection unit, and acoustic image localization control method
JP2013017027A