Information processing method, information processing system, and program

The information processing method optimizes sound quality by determining event states and applying suitable audio processing parameters, addressing the lack of adaptive sound quality adjustment in existing systems.

JP2025111222APending Publication Date: 2025-07-30YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024005520
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Existing technologies do not adjust sound quality of microphones to optimal levels according to the current state of an event during its progress.

Method used

An information processing method that receives audio signals from microphones, determines the event state, and applies appropriate audio processing parameters such as beamforming, echo cancellation, and sound localization to optimize sound quality based on the event's state.

Benefits of technology

Provides an optimal sound quality experience during events by automatically adjusting sound parameters to suit presentation or discussion states, enhancing user experience without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111222000001_ABST
    Figure 2025111222000001_ABST
Patent Text Reader

Abstract

To provide an information processing method that can perform adjustment to optimum microphone sound quality depending on a state of an event during progress of the event to provide an optimum sound quality experience depending on the present state.SOLUTION: In an information processing system in which a signal processor comprising a camera, a loudspeaker, a plurality of microphones, a processor, a memory, and an interface (I / F) are connected to a display and a PC, a sound processing method includes: receiving a sound signal from a microphone; receiving output information output from an information processing terminal; on the basis of the received output information, outputting a determination result indicating a first state or a second state; on the basis of the determination result, determining a sound processing parameter; on the basis of the sound processing parameter, applying sound processing to the sound signal; and outputting the sound signal to which the sound processing is applied.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One embodiment of the present invention relates to an information processing method, an information processing system, and a program.

Background Art

[0002] Patent Document 1 discloses a remote conferencing system that performs acoustic processing using environmental data stored in an environmental database from conference reservation data.

[0003] Patent Document 2 discloses a video and audio signal processing apparatus that determines the characteristics of a video scene from an image output from a video decoder and adjusts the sound field of the decoded audio output according to the characteristics of the video scene.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] Neither Patent Document 1 nor 2 adjusts the sound quality of the microphone to the optimal sound quality according to the current state of the event during the progress of the event using a microphone.

[0006] An object of one embodiment of the present invention is to provide an information processing method that provides an optimal sound quality experience according to the current state during the progress of an event.

Means for Solving the Problems

[0007] An information processing method according to an embodiment of the present invention receives an audio signal from a microphone, receives output information output from an information processing terminal, outputs a determination result indicating a first state or a second state based on the received output information, determines an audio processing parameter based on the determination result, performs audio processing on the audio signal based on the audio processing parameter, and outputs the audio signal on which the audio processing has been performed.

Advantages of the Invention

[0008] According to an embodiment of the present invention, it is possible to provide an optimal sound quality experience according to the current state during the progress of an event.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Embodiments for Carrying Out the Invention

[0010] FIG. 1 is a block diagram showing the configuration of the information processing system 1 of the present embodiment. The information processing system 1 includes a display 10, a signal processing device 20, and a personal computer (PC) 30. The signal processing device 20 is connected to the display 10 and the PC 30. The PC 30, which is the first information processing terminal, is connected to a second information processing terminal at a remote location via a network.

[0011] FIG. 2 is an elevation schematic diagram of the interior of a room. The display 10 is installed on the wall surface of the room as an example. The signal processing device 20 is placed on the upper surface of the display 10 as an example. There are a plurality of participants (participants u1, u2) around the desk. The PC 30 is installed on the desk in front of the participant u2.

[0012] FIG. 3 is a block diagram showing the configuration of the signal processing apparatus 20. The signal processing apparatus 20 includes a camera 11, a speaker 12, a plurality of microphones 13, a processor 14, a memory 15, and an interface (I / F) 16.

[0013] The memory 15 is a storage medium that stores the operation program of the processor 14. The processor 14 reads the operation program from the memory 15 and performs various operations. Note that the program does not necessarily have to be stored in the memory 15. For example, the program may be stored in a storage medium of an external device such as a server. In this case, the processor 14 may read the program from the server and execute it each time.

[0014] The processor 14 receives the sound signals acquired by the plurality of microphones 13. The processor 14 performs predetermined sound processing on the sound signals acquired by the plurality of microphones 13. For example, the processor 14 performs beamforming on the sound signals acquired by the plurality of microphones 13. Beamforming is a process of forming a sound collection beam having directivity in a predetermined direction by adding a delay to and synthesizing the sound signals acquired by the plurality of microphones 13. The sound collection beam can also form a directivity that focuses on a predetermined position. The processor 14 forms, for example, a sound collection beam that focuses on the position of the speaker. The processor 14, for example, obtains the difference (phase difference) in the acquisition timing of the sound by obtaining the correlation value of the sound signals of the plurality of microphones 13, and detects the position of the speaker. The processor 14 can uniquely determine the position of the speaker by obtaining the difference in the acquisition timing of the sound in three or more microphones 13. The processor 14 can acquire the voice of the speaker with high sensitivity by forming a directivity that focuses on the obtained position of the speaker.

[0015] A plurality of sound collection beams can be formed simultaneously. Therefore, the processor 14 can also acquire the voices of a plurality of speakers with high sensitivity.

[0016] Note that in this example, the number of microphones is six, but the number of microphones is not limited to six. As long as there are at least two or more microphones 13, the directivity can be changed by beamforming.

[0017] The processor 14 outputs the sound signal related to the sound collection beam to the I / F 16. The I / F 16 is a communication I / F such as USB, HDMI (registered trademark), or Bluetooth (registered trademark). The I / F 16 is connected to the PC 30 by, for example, USB. Also, the I / F 16 is connected to the display 10 by, for example, HDMI (registered trademark). The I / F 16 outputs the sound signal related to the sound collection beam to the PC 30. The PC 30 transmits the sound signal to a second information processing terminal at a remote location.

[0018] The PC 30 receives a sound signal from a second information processing terminal at a remote location. The PC 30 transmits the sound signal to the signal processing device 20 via the I / F 16. The processor 14 outputs the sound signal received from the PC 30 via the I / F 16 to the speaker 12. The speaker 12 emits the sound signal received from the processor 14.

[0019] Thereby, the user of the signal processing device 20 can hold a voice conference with the user at a remote location.

[0020] Also, the processor 14 receives the video signal related to the video captured by the camera 11. The processor 14 performs predetermined signal processing on the video signal captured by the camera 11. The signal processing is, for example, framing processing by panning, tilting, or zooming. The processor 14 outputs the video signal after the sound signal processing to the I / F 16. The I / F 16 outputs the video signal to the PC 30. The PC 30 transmits the video signal to a second information processing terminal at a remote location.

[0021] The PC 30 receives a video signal from a second information processing terminal at a remote location. The PC 30 outputs the received video signal to the signal processing device 20 via the I / F 16. The signal processing device 20 receives the input of the video signal from the PC 30. The signal processing device 20 outputs the received video signal to the display 10. The display 10 receives the input of the video signal from the signal processing device 20. The display 10 displays a video based on the received video signal. Note that the PC 30 may output the received video signal directly to the display 10.

[0022] Thereby, the user of the signal processing device 20 can also hold a video conference with the user at the remote location.

[0023] FIG. 4 is a flowchart showing the operation of the signal processing method of the present embodiment. The signal processing method shown in FIG. 4 is executed by the signal processing device 20.

[0024] The signal processing device 20 receives audio signals from a plurality of microphones 13 (S11). Also, the signal processing device 20 receives output information from the PC 30 (S12). The output information is an output image for display on the display 10. That is, the output information is, for example, a video signal received from the PC 30 via HDMI (registered trademark), and is image information transmitted and received from the second information processing terminal to the first information processing terminal.

[0025] Note that the output information may also be an audio signal received from the PC 30 (audio information transmitted and received from the second information processing terminal to the first information processing terminal), or a video signal received from the camera 11.

[0026] Also, the processing of S11 and the processing of S12 may be performed in either order. Also, the processing of S11 and S12 may be performed simultaneously. Also, the processing of S11 may be performed after the processing of S13 or after the processing of S14, which will be described later.

[0027] The signal processing device 20 outputs a determination result indicating the first state or the second state based on the received output information (S13). The first state is, for example, a presentation state in which a specific participant gives a presentation. The second state is, for example, a discussion state in which a plurality of participants including the near-end side and the far-end side have a conversation. The signal processing device 20 analyzes, for example, the video signal received via HDMI (registered trademark) and outputs a determination result indicating the first state or the second state. More specifically, the signal processing device 20 recognizes, for example, the face images of the participants from the video signal, and determines whether each recognized face image of the participants is a near-end side participant or a far-end side participant. The signal processing device 20 compares, for example, the position of the speaker obtained based on the correlation value of the sound signals of the plurality of microphones 13 with the position of the face images of the participants recognized from the video signal. The signal processing device 20 determines that the face image of the participant that matches the position of the speaker obtained based on the sound signal is the face image of the near-end side participant, and determines that the image of the speaker that does not match is the face image of the far-end side participant.

[0028] The signal processing device 20 outputs a determination result indicating the first state, for example, when an image of meeting materials is included in the output image. On the other hand, the signal processing device 20 outputs a determination result indicating the second state, for example, when a face image of a far-end side participant is included in the output image.

[0029] In addition, the signal processing device 20 may detect the speaker based on the output information and make a determination based on the state of the detected speaker. In a presentation, only one participant speaks, and in a discussion, a plurality of participants speak. That is, the first state corresponds to a single speaker, and the second state corresponds to a plurality of speakers. The signal processing device 20 outputs a determination result indicating the first state, for example, when only one participant has an unmute indication (or no mute indication) in the output image and all other participants have a mute indication. Also, the signal processing device 20 outputs a determination result indicating the second state, for example, when a plurality of participants have an unmute indication (no mute indication for a plurality of participants) in the output image.

[0030] Next, the signal processing device 20 determines acoustic processing parameters based on the determination result (S14). The acoustic processing parameters are, for example, parameters of beamforming processing and correspond to the delay amounts applied to the respective acoustic signals of the plurality of microphones 13. In the first state, the signal processing device 20 determines parameters for forming a narrow-directional pickup beam that focuses on the position of the speaker. Alternatively, in the first state, the signal processing device 20 determines parameters for forming one pickup beam. In the second state, the signal processing device 20 determines parameters for forming a non-focused wide-directional pickup beam. Alternatively, in the second state, the signal processing device 20 determines parameters for forming a plurality of pickup beams.

[0031] Then, the signal processing device 20 performs acoustic processing on the acoustic signal based on the determined acoustic processing parameters (S15). The signal processing device 20 gives the delays indicated by the acoustic processing parameters to the acoustic signals acquired by the plurality of microphones 13 and synthesizes them to form a pickup beam. Finally, the signal processing device 20 outputs the acoustic signal after the acoustic processing (S16).

[0032] As described above, in the information processing system 1 of the present embodiment, during an event using a microphone in a meeting or the like, it is possible to adjust the sound quality of the microphone to be optimal according to the current state of the event. More specifically, in the first state (presentation state) where only one participant is speaking, a narrow-directional sound collection beam that focuses on the position of the one participant is formed. Therefore, voices and noises other than the speaker are not collected, and the participants can listen to the speaker's voice with a sound quality optimal for the presentation. In the second state (discussion state) where multiple participants are speaking, an omnidirectional or multiple sound collection beams for collecting the voices of multiple participants are formed. Therefore, the voices of multiple speakers are collected, and the participants can listen to the speaker's voice with a sound quality optimal for the discussion. In this way, the user of the information processing system 1 of the present embodiment can obtain a new customer experience of being able to listen to the speaker's voice with an optimal sound quality suitable for each of the presentation state or the discussion state without the need to manually adjust the sound processing parameters.

[0033] (Modification Example 1) In the process of S12 in FIG. 4, the signal processing apparatus 20 of Modification Example 1 further outputs a determination result indicating the first state or the second state based on the sound signal received from the PC 30 (the sound information transmitted and received between the second information processing terminal and the first information processing terminal).

[0034] When the output image includes an image of meeting materials and the sound signal received from the PC 30 is in a silent state, the signal processing apparatus 20 outputs a determination result indicating the first state.

[0035] On the other hand, when the output image includes an image of meeting materials and the sound signal received from the PC 30 is in a sounding state, the signal processing apparatus 20 outputs a determination result indicating the second state. That is, the signal processing apparatus 20 determines that multiple participants are having a discussion while referring to the image even if the output image includes an image of meeting materials.

[0036] As a result, the signal processing device 20 can further analyze the sound information to determine the event state with higher accuracy.

[0037] (Modification 2) In the process of S12 in FIG. 4, the signal processing device 20 of Modification 2 further outputs a determination result indicating the first state or the second state based on the video signal received from the camera 11.

[0038] When the output image includes an image of meeting materials and the faces of all the participants detected in the video signal are facing the display 10, the signal processing device 20 outputs a determination result indicating the first state.

[0039] On the other hand, when the output image includes an image of meeting materials but the faces of a plurality of participants detected in the video signal are facing other participants, the signal processing device 20 outputs a determination result indicating the second state. That is, when the output image includes an image of meeting materials, the signal processing device 20 determines that a plurality of participants are having a discussion while referring to the image.

[0040] As a result, the signal processing device 20 can further analyze the video signal received from the camera 11 to determine the event state with higher accuracy.

[0041] (Modification 3) In the process of S13 shown in FIG. 4, the signal processing device 20 of Modification 3 further determines whether the current state is either the first state or the second state based on the sound signal received from the microphone 13.

[0042] The signal processing device 20 detects the speaker position based on the sound signal received from the microphone 13 and forms a sound pickup beam that focuses on the speaker's position. When the output image includes an image of meeting materials and the speaker position continues to be at the same position for a predetermined time or more, the signal processing device 20 outputs a determination result indicating the first state.

[0043] On the other hand, even if the output image includes an image of meeting materials, when the speaker position is at different positions within a predetermined time, that is, when the speaker has changed, the signal processing device 20 outputs a determination result indicating the second state. That is, even if the output image includes an image of meeting materials, the signal processing device 20 determines that a plurality of participants are having a discussion while referring to the image.

[0044] Thereby, the signal processing device 20 can further analyze the sound signal acquired by the microphone 13 to perform a more accurate determination of the event state.

[0045] (Modification Example 4) The sound processing is not limited to beamforming processing. The sound processing includes echo cancellation processing, equalizer processing, reverb processing, or noise cancellation processing, etc. The echo cancellation processing is a process of canceling the sound signal that loops back from the speaker 12 to the microphone 13 after being output. The equalizer processing is a process of adjusting the frequency characteristics of the sound signal acquired by the microphone 13. The reverb processing is a process of adding an indirect sound component to the sound signal acquired by the microphone 13. The noise cancellation processing is a process of canceling the noise component from the sound signal acquired by the microphone 13. The sound processing in Modification Example 4 is echo cancellation processing. The signal processing device 20 determines the parameters of the echo cancellation processing in the process of S14 in FIG. 4.

[0046] The echo cancellation processing is, for example, a filter processing that uses the sound signal output to the speaker 12 as a reference signal to estimate the components of the sound signal that loops back to the microphone 13, and removes the estimated components from the sound signal received from the microphone 13. The filter is an adaptive filter processing that performs an adaptation process so that the residual of the sound signal after the echo cancellation processing becomes small. The adaptation algorithm may be an adaptation algorithm based on the sound signal on the time axis or an adaptation algorithm based on the sound signal on the frequency axis.

[0047] In the first state, the signal processing device 20 sets the intensity of the echo cancellation processing to the maximum. For example, the signal processing device 20 sets the gain of the filter processing to the maximum. Also, when the signal processing device 20 uses an adaptive algorithm based on the sound signal on the frequency axis, it makes the number of frames in the process of converting the sound signal on the time axis to the frequency axis (fast Fourier transform processing) as large as possible. As a result, the listener in the remote location can clearly hear only the voice of the presenter of the presentation.

[0048] In the second state, the signal processing device 20 sets the intensity of the echo cancellation processing to the minimum. For example, the signal processing device 20 sets the gain of the filter processing to the minimum. Also, when the signal processing device 20 uses an adaptive algorithm based on the sound signal on the frequency axis, it makes the number of frames in the fast Fourier transform processing as small as possible. As a result, even if the participants on the near side and the participants in the remote location speak simultaneously, their respective voices will not be interrupted, making it easier to have a discussion.

[0049] (Modification Example 5) The sound processing includes processing for the sound signal output to the speaker 12 (the sound information transmitted and received from the second information processing terminal to the first information processing terminal). The signal processing device 20 performs sound image localization processing on the sound signal output to the speaker 12, for example. In this case, the speaker 12 has a plurality of speaker units. The sound image localization processing is panning processing for controlling the level ratio of the sound signals supplied to the plurality of speaker units, or binaural processing for imparting frequency characteristics based on the head-related transfer function to the sound signals supplied to the plurality of speaker units.

[0050] In the first state, the signal processing device 20 performs sound image localization processing in which the sound image is localized at the center of the display 10. In the second state, the signal processing device 20 performs sound image localization processing in which the sound image is localized at the position of the speaker. The position of the speaker is the position of the speaker of the participant on the far end side detected based on the output image received from the PC 30.

[0051] As a result, the user of Modification 5 does not need to manually adjust the sound processing parameters, can concentrate on the materials on the screen in the presentation state, can perceive that the sound can be heard from the position of the speaker in the discussion state, and can obtain a new customer experience of being able to listen to the speaker's voice with the optimal sound quality according to each state.

[0052] (Modification 6) The sound processing parameters corresponding to the second state are preferably sound processing with lower latency than the sound processing parameters corresponding to the first state. For example, in the second state, the signal processing device 20 may stop the equalizer processing, reverberation processing, or noise cancellation processing, etc., and perform beamforming processing. Also, when the signal processing device 20 performs a process of converting a sound signal on the time axis to the frequency axis (fast Fourier transform process) in various processes, the number of frames may be reduced. Thereby, the participants can quickly hear the voice from a remote location, making it easier to have a conversation.

[0053] Also, the sound processing parameters corresponding to the first state may be sound processing with higher load than the sound processing parameters corresponding to the second state. For example, in the first state, the signal processing device 20 performs echo cancellation processing, equalizer processing, reverberation processing, and noise cancellation processing, etc. Also, when the signal processing device 20 performs a process of converting a sound signal on the time axis to the frequency axis (fast Fourier transform process) in various processes, the number of frames is increased. Thereby, the voice of the presenter in the presentation becomes high-quality and is easier for the listeners to hear.

[0054] (Modification 7) The sound processing includes a process of amplifying the voice based on the sound signal received from the microphone 13 from the speaker 12. The signal processing device 20 performs a process of amplifying the voice based on the sound signal received from the microphone 13 from the speaker 12 in the first state, and does not amplify the voice based on the sound signal received from the microphone 13 from the speaker 12 in the second state.

[0055] As a result, in the presentation, the voice of the presenter is amplified at the near-end venue, so that even in a large venue, the listeners at the near-end can easily hear the voice of the presenter. On the other hand, in the discussion, since the voice of the near-end participants is not output from the speaker 12, the conversation in the discussion is not hindered.

[0056] (Modification Example 8) The sound processing parameters of Modification Example 8 include a first sound processing parameter corresponding to the first state, a second sound processing parameter corresponding to the second state, and a third sound processing parameter corresponding to an intermediate state between the first state and the second state.

[0057] The intermediate state is, for example, a case where the sound signal received from the PC 30 is in an audible state even if the output image includes an image of the meeting materials. Alternatively, the intermediate state may be a case where even if the output image includes an image of the meeting materials, the faces of a plurality of participants detected in the video signal received from the camera 11 are facing other participants. Alternatively, the intermediate state may be a case where even if the output image includes an image of the meeting materials, the speaker position changes to different positions within a predetermined time, that is, the speaker is frequently switched. Alternatively, the intermediate state may be a case where even if the output image includes an image of the meeting materials, for example, when text information (chat post) is received from the PC 30, or when reception of display of reaction information such as emojis and information determined to be a question item is received. Further, the intermediate state is a state where, as in Modification Example 7 above, the voice based on the sound signal received from the microphone 13 is amplified from the speaker 12, and a speaker other than the presenter is detected (for example, when the positions of a plurality of speakers are detected based on the correlation values of the sound signals of a plurality of microphones 13).

[0058] In these intermediate states, the signal processing device 20 forms a sound collection beam with a directivity intermediate between the narrow directivity in the first state and the wide directivity in the second state, for example, in beamforming processing. Alternatively, the signal processing device 20 performs all of the processing such as echo cancellation processing, equalizer processing, reverberation processing, and noise cancellation processing in the first state, stops all of these processes in the second state, and executes only a part of these processes in the intermediate state. Further, the signal processing device 20 may make the strength of these processes weaker (stronger than the second state) than in the first state in the intermediate state. For example, the signal processing device 20 may set the strength of the echo cancellation processing to be weaker than in the first state in the intermediate state.

[0059] (Other examples) In the above embodiment, the sound processing parameters are determined based on the determination result indicating the first state or the second state. However, the signal processing device 20 may determine the control parameters for controlling the camera 11 based on the determination result.

[0060] In the first state, for example, the signal processing device 20 performs framing processing such as panning, tilting, or zooming so that only the speaker appears in the video of the camera 11. In the second state, for example, the signal processing device 20 performs framing processing so that all the participants appear in the video of the camera 11.

[0061] The determined sound processing parameters may be transmitted to the second information processing terminal on the far end side. The second information processing terminal on the far end side may perform sound processing based on the received sound processing parameters. Thereby, the first information processing terminal on the near end side can also control the second information processing terminal on the far end side. Further, in the second information processing terminal on the far end side, the operations of the above embodiment may be performed, and the determined sound processing parameters may be transmitted to the first information processing terminal on the near end side. In this case, the first information processing terminal on the near end side may compare the sound processing parameters determined by its own device with the sound processing parameters determined by the second information processing terminal on the far end side and determine the final sound processing parameters.

[0062] In the above embodiment, an example in which the signal processing device 20 executes the information processing method of this embodiment has been shown. However, for example, an application program of the PC 30 may execute the information processing method of this embodiment.

[0063] Also, in the above embodiment, an example of performing beamforming processing on the sound signals of a plurality of microphones has been shown. However, for example, microphones may be installed for each of a plurality of participants in the same location, and sound processing may be performed on the sound signals of the microphones of each participant. Also, when performing sound processing on the sound signal output to the speaker 12, sound processing may be performed on the sound signals received from the information processing terminals at a plurality of locations.

[0064] In the above embodiment, a remote conference has been shown as an example of an event. However, the information processing method of this embodiment can be applied to various events such as online live, online lesson, and online session. For example, in an online live, when the signal processing device 20 detects an output image with low brightness or an output image with little brightness change, it determines that it is in the first state. For example, in the case of a ballad performance or a song introduction, the output image has low brightness or little brightness change. In this case, the signal processing device 20 determines that it is in the first state, and performs equalizer processing, reverb processing, noise cancellation processing, etc. on the sound signal received from the PC 30. As a result, the sound in the remote venue becomes high-quality and easier for the listener to hear. Also, when the signal processing device 20 determines that it is in the first state, it does not transmit the sound signal received from the microphone 13 to the far-end side. As a result, the listener's voice is not reproduced at the far-end venue, and a quiet song such as a ballad is not disturbed.

[0065] On the other hand, when the signal processing device 20 detects an output image with high brightness or an output image with much brightness change, it determines that it is in the second state. For example, in the case of an exciting song with an exciting melody, the output image has high brightness or much brightness change. When the signal processing device 20 determines that it is in the second state, it performs low-latency sound processing and transmits the sound to the far-end venue. As a result, the cheer of the listener can be reproduced in real time with lower latency at the far-end venue.

[0066] Also, for example, in an online lesson, when the signal processing device 20 analyzes the output image and determines that the participant on the remote side is performing, and analyzes the video signal from the camera 11 and determines that the participant on the near side is not performing, it determines that it is in the first state. When the signal processing device 20 determines that it is in the first state, it does not transmit the sound signal received from the microphone 13 to the remote side. As a result, the sound on the near side is not reproduced at the remote side venue, and the performance on the remote side is not hindered.

[0067] Also, the signal processing device 20 may also determine that it is in the first state when it analyzes the output image and determines that the participant on the remote side is not performing, and analyzes the video signal from the camera 11 and determines that the participant on the near side is performing. However, in this case, the signal processing device 20 performs equalizer processing, reverb processing, noise cancellation processing, etc. on the sound signal received from the microphone 13. As a result, the transmitted performance sound becomes high-quality and is easier for the listeners on the remote side to hear.

[0068] On the other hand, when the signal processing device 20 analyzes the output image and determines that the participant on the remote side is performing, and analyzes the video signal from the camera 11 and determines that the participant on the near side is performing, it determines that it is in the second state (session state). When the signal processing device 20 determines that it is in the second state, it performs low-latency sound processing and transmits the performance sound to the remote side. As a result, the performer on the remote side can listen to the performance sound in real time with lower latency, and can obtain a new customer experience of enjoying the remote session more comfortably.

[0069] The description of this embodiment should be considered illustrative in all respects and not restrictive. The scope of the present invention is indicated by the claims rather than the above-described embodiments. Furthermore, the scope of the present invention includes the scope equivalent to the claims.

Explanation of Reference Numerals

[0070] 1: Information processing system, 10: Display, 11: Camera, 12: Speaker, 13: Microphone, 14: Processor, 15: Memory, 16: I / F, 20: Signal processing device

Claims

1. Receiving an audio signal from a microphone, receiving output information output from an information processing terminal, outputting a determination result indicating a first state or a second state based on the received output information, determining an audio processing parameter based on the determination result, performing audio processing on the audio signal based on the audio processing parameter, outputting the audio signal on which the audio processing has been performed, An information processing method.

2. The output information includes an output image for display on a display, The information processing method according to claim 1.

3. The information processing terminal includes a first information processing terminal and a second information processing terminal connected via a network, The output image includes image information transmitted and received from the second information processing terminal to the first information processing terminal, The information processing method according to claim 2.

4. The output information further includes audio information transmitted and received from the second information processing terminal to the first information processing terminal, The information processing method according to claim 2.

5. The output image further includes a video signal received from a camera, The information processing method according to claim 2.

6. Outputting the determination result based on the received output information and the received audio signal, The information processing method according to any one of claims 1 to 5.

7. Detecting a speaker based on the output information, and outputting the determination result based on the state of the detected speaker, The information processing method according to any one of claims 1 to 5.

8. The first state corresponds to a single speaker, The second state corresponds to a plurality of speakers, The information processing method according to any one of claims 1 to 5.

9. The audio processing includes echo cancellation processing, The information processing method according to any one of claims 1 to 5.

10. Performing the audio processing on the audio information transmitted and received from the second information processing terminal to the first information processing terminal, The information processing method according to any one of claims 1 to 5.

11. The audio processing parameter corresponding to the first state is audio processing with a higher load than the audio processing parameter corresponding to the second state, The information processing method according to any one of claims 1 to 5.

12. The audio processing parameter corresponding to the second state is audio processing with a lower delay than the audio processing parameter corresponding to the first state, The information processing method according to any one of claims 1 to 5.

13. The sound processing includes a process of amplifying voice based on the received sound signal. The information processing method according to any one of claims 1 to 5.

14. The sound processing includes a beamforming process. The information processing method according to any one of claims 1 to 5.

15. The sound processing parameters include a first sound processing parameter corresponding to the first state, a second sound processing parameter corresponding to the second state, and a third sound processing parameter corresponding to an intermediate state between the first state and the second state. The information processing method according to any one of claims 1 to 5.

16. Receiving a sound signal from a microphone, Receiving output information output from an information processing terminal, Outputting a determination result indicating a first state or a second state based on the received output information, Determining sound processing parameters based on the determination result, Performing sound processing on the sound signal based on the sound processing parameters, Outputting the sound signal on which the sound processing has been performed. An information processing system.

17. Receiving a sound signal from a microphone, Receiving output information output from an information processing terminal, Outputting a determination result indicating a first state or a second state based on the received output information, Determining sound processing parameters based on the determination result, Performing sound processing on the sound signal based on the sound processing parameters, Outputting the sound signal on which the sound processing has been performed. A program for causing an information processing apparatus to execute the process.

Citation Information

Patent Citations

  • Teleconference system

    JP2007060460A