Signal processing system, signal processing device, and signal processing method
By generating a mixed audio signal by combining audio data specified from multiple terminal devices, the problem of difficulty in understanding user reactions when transmitting streaming data in electronic conferencing systems is solved, thus enhancing the sense of presence.
Patent Information
- Application Number
- CN202080105235.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-01
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2040-10-01
AI Technical Summary
In electronic conferencing systems, it is difficult to know the exact reaction of the target user when transmitting streaming data, and using images to display applause lacks a sense of presence.
The signal processing system mixes the audio data specified by multiple terminal devices to generate a mixed audio signal, which is then sent to the communication system so that the transmission source can understand the user's reaction in real time.
It enables precise knowledge of the target user's reaction while transmitting streaming data, enhancing the user's sense of presence.
Smart Images

Figure CN116114014B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a signal processing system, a signal processing apparatus, and a signal processing method. BACKGROUND
[0002] In recent years, a communication system such as an electronic conference system using a Web application is known (for example, refer to Patent Document 1). In such a communication system, streaming data such as a speech or a presentation is transmitted to a plurality of users via a network, and communication (exchange) is performed among the plurality of users.
[0003] Patent Document 1: Japanese Patent Application Publication No. 2007-277492 SUMMARY
[0004] However, in the communication system as described above, in a case where streaming data such as a presentation is transmitted to a user without communication such as a sound signal from the user of the transmission target, it is not clear at the transmission source that the user of the transmission target has a reaction such as feedback, and it is sometimes felt that it is difficult to exchange.
[0005] Therefore, in the communication system as described above, in a case where streaming data is transmitted, it is difficult to know the reaction of the user of the transmission target accurately.
[0006] Further, as the communication system, there is a system in which a display of a picture of applause is performed by pressing a "clap" button. However, the display of the picture lacks a sense of presence.
[0007] The present application has been made to solve the above problems, and has an object to provide a signal processing system, a signal processing apparatus, and a signal processing method in which, in a case where streaming data is transmitted, the reaction of the user of the transmission target can be known accurately. It also has an object to provide a signal processing system, a signal processing apparatus, and a signal processing method in which a user feels a sense of presence.
[0008] To solve the above problems, one embodiment of the present application is a signal processing system connected to a plurality of devices including at least a first terminal device that receives streaming data and a second terminal device, and a communication system that enables communication with the plurality of devices, the signal processing system including: an accepting unit that accepts designation of first sound data obtained from the first terminal device that receives the streaming data and designation of second sound data obtained from the second terminal device that receives the streaming data; a signal processing unit that acquires the first sound data and the second sound data, generates a third sound signal in which a first sound signal based on the first sound data and a second sound signal based on the second sound data are mixed; and a transmitting unit that transmits the third sound signal to the communication system via a communication line that connects the communication system and the first terminal device.
[0009] Effects of the Invention
[0010] According to the present application, it is possible to know the reaction of a user who is a transmission target exactly in the case of transmitting streaming data. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is a block diagram showing one example of a signal processing system according to the first embodiment.
[0012] Figure 2 is a diagram showing one example of data of a sound data storage unit in the first embodiment.
[0013] Figure 3 is a diagram showing one example of data of a history information storage unit in the first embodiment.
[0014] Figure 4 is a diagram showing one example of a designation menu screen used for designation of sound data in the first embodiment.
[0015] Figure 5 is a diagram showing one example of an operation of the signal processing system according to the first embodiment.
[0016] Figure 6 is a diagram showing one example of an operation of the signal processing system according to the first embodiment in which sound data is designated from a plurality of terminal devices.
[0017] Figure 7 is a diagram showing one example of sound data that is designated in the first embodiment.
[0018] Figure 8 is a diagram showing one example of a mixed sound signal in which sound data that is designated in the first embodiment is mixed.
[0019] Figure 9is a diagram showing an example of an action of the signal processing system according to the first embodiment, in which the signal processing system according to the first embodiment specifies a plurality of sound data from one terminal device.
[0020] Figure 10 is a block diagram showing an example of the signal processing system according to the second embodiment.
[0021] Figure 11 is a diagram showing an example of an action of the signal processing system according to the second embodiment.
[0022] Figure 12 is a diagram showing an example of an action of the signal processing system according to the second embodiment, in which the signal processing system according to the second embodiment specifies sound data from a plurality of terminal devices. DETAILED DESCRIPTION
[0023] Hereinafter, a signal processing system, a signal processing device, and a signal processing method according to one embodiment of the present application will be described with reference to the accompanying drawings.
[0024] [First Embodiment]
[0025] Figure 1 is a block diagram showing an example of the signal processing system 1 according to the present embodiment.
[0026] The signal processing system 1 is a system connected to a communication system 200 and a plurality of devices including at least a terminal device 20 and a terminal device 30, and has an application host device 10.
[0027] The communication system 200 is a system that performs communication between a plurality of devices via a communication line, such as an electronic conference system or a video transmission system. The communication system 200 is connected to the signal processing system 1 via a network NW1, for example. In the present embodiment, with respect to the communication system 200, an example in which a presenter PU transmits stream data such as a presentation or a lecture to the terminal device 20 and the terminal device 30 via the network NW1 will be described.
[0028] The communication system 200 has a communication system server 210, a transmission source device 220, a microphone 230, and a speaker 240, and is capable of connecting to a plurality of devices including at least the terminal device 20 and the terminal device 30.
[0029] The communication system server 210 is a server device that manages the communication system 200, for example. The communication system server 210 receives information such as stream data, image data, and sound signals from various devices via the network NW1, and transmits (transfers) the information such as stream data to a connection target device via the network NW1.
[0030] The transmission source device 220 is a terminal device that transmits stream data such as a presentation or a lecture in real time, for example, by a presenter PU. The transmission source device 220 is connected with a microphone 230 and a speaker 240. In addition, the transmission source device 220 is built-in with a camera (not shown) for live transmission of stream data, for example. Here, the presenter PU is a user of the communication system 200 that transmits stream data, for example. The transmission source device 220 transmits stream data to the communication system server 210 via the network NW1.
[0031] The microphone 230 picks up a sound around the transmission source device 220 and outputs a sound signal to the transmission source device 220.
[0032] The speaker 240 outputs a sound signal transmitted from the communication system server 210. In addition, the speaker 240 outputs a sound signal based on sound data such as applause designated from the terminal device 20 and the terminal device 30 described later, for example.
[0033] The terminal device 20 (one example of a first terminal device) is connected to the communication system 200 through a communication line and receives stream data transmitted from the communication system 200. The terminal device 20 is a personal computer, a tablet terminal device, a smart phone, or the like, for example. The terminal device 20 receives stream data transmitted from the transmission source device 220 through the communication system 200 and outputs an image signal and a sound signal based on the stream data to the user AU1.
[0034] In addition, the terminal device 20 designates the application host device 10 for sound data by executing an application program (hereinafter, sometimes referred to as an application) for transmitting sound data such as applause to the transmission source device 220.
[0035] In addition, the terminal device 20 has an NW (network) communication section 21, an input section 22, a display section 23, a microphone 24, a speaker 25, and a terminal control section 26.
[0036] The NW communication section 21 is an interface section capable of connecting to the network NW1 and performs communication with the communication system 200 and the application host device 10 via the network NW1.
[0037] The input section 22 is an input device such as a keyboard mouse, a touch panel, or the like and accepts various inputs from the user AU1. The input section 22 is used for various operations when using the communication system 200, designation of sound data such as applause, or the like.
[0038] The display section 23 is a display device such as a liquid crystal display and displays an image or an image of various operation screens, an image of an image transmitted through the communication system 200, or the like.
[0039] The microphone 24 (one example of a sound pickup unit) picks up a sound around the terminal device 20, generates a sound signal, and outputs it to the terminal control section 26. The microphone 24 can be a built-in microphone built in the terminal device 20, or can be an external microphone.
[0040] The speaker 25 outputs a sound signal such as sound data transmitted through the communication system 200. The speaker 25 can be a built-in speaker built in the terminal device 20, or can be an external speaker.
[0041] The terminal control section 26 is a processor including, for example, a CPU (Central Processing Unit), and controls the terminal device 20. The terminal control section 26 has a communication processing section 261 and an application processing section 262.
[0042] The communication processing section 261 is connected to the communication system 200 via the NW communication section 21, and performs communication with other devices. The communication processing section 261 receives stream data transmitted from the transmission source device 220 via the NW communication section 21, causes an image signal based on the stream data to be displayed on the display section 23, and causes a sound signal based on the stream data to be output to the speaker 25.
[0043] In addition, the communication processing section 261 mixes (mixes) a sound signal (mixed sound signal) received from the application host device 10 described later via the NW communication section 21 and sound data picked up by the microphone 34, and outputs it to the communication system 200 via the NW communication section 21.
[0044] The application processing section 262 is connected to the application host device 10 via the NW communication section 21, and performs processing of designating sound data such as applause. The application processing section 262 causes an image of a designation menu screen transmitted from the application host device 10 to be displayed on the display section 23. Here, the designation menu screen is a menu screen for selecting and designating sound data transmitted to the communication system 200 from among a plurality of sound data, for example. In addition, the application processing section 262 transmits designation information of sound data designated by the user AU1 from the designation menu screen using the input section 22 to the application host device 10 via the NW communication section 21.
[0045] The terminal device 30 (one example of a second terminal device) is connected to the communication system 200 through a communication line, and receives stream data transmitted from the communication system 200. The terminal device 30 is, for example, a personal computer, a tablet terminal device, a smart phone, or the like. The terminal device 30 receives stream data transmitted from the transmission source device 220 through the communication system 200, and outputs an image signal and a sound signal based on the stream data to the user AU2.
[0046] The terminal device 30 designates the sound data to the application host device 10 by executing an application program for transmitting sound data such as applause to the transmission source device 220. The terminal device 30 has an NW communication section 31, an input section 32, a display section 33, a microphone 34, a speaker 35, and a terminal control section 36.
[0047] The NW communication section 31 has the same structure and functions as those of the NW communication section 21. The input section 32 has the same structure and functions as those of the input section 22. The display section 33 has the same structure and functions as those of the display section 23. The microphone 34 has the same structure and functions as those of the microphone 24. The speaker 35 has the same structure and functions as those of the speaker 25.
[0048] The terminal control section 36 has the same structure and functions as those of the terminal control section 26. The terminal control section 36 has a communication processing section 361 and an application processing section 362.
[0049] The communication processing section 361 has the same structure and functions as those of the communication processing section 261.
[0050] The application processing section 362 has the same structure and functions as those of the application processing section 262.
[0051] Further, in the present embodiment, the difference is that the terminal device 20 transmits the mixed sound signal generated by the application host device 10 to the communication system 200, in contrast to the terminal device 30 not transmitting the mixed sound signal to the communication system 200.
[0052] In addition, in the present embodiment, the terminal device 30 is one, but the signal processing system 1 can have a plurality of terminal devices 30. Figure 1
[0053] The application host device 10 is a server device connected to the network NW1, and is a Web server that implements a Web site for designating sound data such as applause, for example. The application host device 10 generates a mixed sound signal that mixes sound data designated by the terminal device 20 and the terminal device 30 using the Web site for designating sound data. The application host device 10 transmits the mixed sound signal to the communication system 200 through a connection line of the terminal device 20.
[0054] The application host device 10 has an NW communication section 11, a storage section 12, and an application control section 13.
[0055] The NW communication section 11 (one example of a transmission section) is an interface section capable of connecting to the network NW1, and performs communication with the terminal device 20 and the terminal device 30 via the network NW1.
[0056] The storage section 12 stores various information used by the application host device 10. The storage section 12 has a sound data storage section 121 and a history information storage section 122.
[0057] The sound data storage section 121 stores sound data that can be specified from the terminal device 20 and the terminal device 30. Here, refer to Figure 2 , a data example of the sound data storage section 121 is described.
[0058] Figure 2 is a view showing a data example of the sound data storage section 121 in the present embodiment.
[0059] The sound data storage section 121 stores a sound data name in association with sound data. Here, the sound data name is the name of sound data that can be specified in the application host device 10, and is identification information that identifies the sound data. In addition, the sound data shows the file name of the sound data. As described above, the sound data storage section 121 stores identification information of sound data in association with the sound data.
[0060] For example, in the example shown in Figure 2 , it is shown that "clap-01.wav" is stored as sound data with the sound data name "clap 01" in the sound data storage section 121. In addition, it is shown that "clap-02.wav" is stored as sound data with the sound data name "clap 02" in the sound data storage section 121.
[0061] Further, the sound data can be, for example, voice data such as "Hey—", "Can't hear the sound", and the like, sound data of a bell that indicates the end time of a presentation, and the like, in addition to sound data of different kinds of clapping.
[0062] Returning to the description of Figure 1 , the history information storage section 122 stores history information of sound data that is specified from the terminal device 20 and the terminal device 30. Here, refer to Figure 3 , a data example of the history information storage section 122 is described.
[0063] Figure 3 is a view showing a data example of the history information storage section 122 in the present embodiment.
[0064] The history information storage section 122 stores a date and time, terminal information, and a sound data name in association. Here, the date and time indicates the date and time when the sound data was specified by the terminal device 20 and the terminal device 30, and the terminal information indicates identification information of the terminal device 20 and the terminal device 30. In addition, the sound data name indicates the sound data name that was specified by the terminal device 20 and the terminal device 30.
[0065] InFigure 3 In the illustrated example, it is shown that sound data with a sound data name of "Applause 01" is designated from "Terminal A" at a date and time of "2020 / 7 / 20 10:00:00". In addition, it is shown that sound data with a sound data name of "Applause 02" is designated from "Terminal B" at a date and time of "2020 / 7 / 20 10:00:15".
[0066] Returning again to the explanation of Figure 1 , the application control section 13 is a processor such as a CPU, for example, and controls the application host device 10. The application control section 13 has a sound data accepting section 131 and a signal processing section 132.
[0067] The sound data accepting section 131 (one example of an accepting section) accepts designation of sound data from the terminal device 20 and the terminal device 30. The sound data accepting section 131 accepts designation of, for example, the 1st sound data from the terminal device 20 that received the stream data, and designation of the 2nd sound data from the terminal device 30 that received the stream data.
[0068] The sound data accepting section 131 causes a designation menu screen that designates, for example, sound data to be displayed on the terminal device 20 and the terminal device 30, and accepts designation of sound data. The sound data accepting section 131 transmits an image of the designation menu screen shown in, for example, Figure 4 to the terminal device 20 and the terminal device 30 via the NW communication section 11 to be displayed.
[0069] Figure 4 is a diagram showing one example of a designation menu screen for designation of sound data in the present embodiment.
[0070] The sound data accepting section 131 causes an image (designation menu image G1) of a designation menu screen that contains sound data information indicating sound data that can be designated, and a play bar that visually indicates a play time (play period) of the sound data to be displayed on the terminal device 20 and the terminal device 30 as shown in Figure 4 . Here, the play bar becomes a length corresponding to the play time (play period). In addition, the play bar can also be set to the same length, and the moving speed of the play bar is changed in correspondence with the play time (play period) (for example, the inverse of the play time, and the like).
[0071] The users AU1 to AU3 operate the input section 22 of the terminal device 20 or the input section 32 of the terminal device 30, and designate sound data by pressing (clicking) a portion of the sound data in the designation menu image G1.
[0072] The sound data accepting section 131, in a case where designation of sound data is accepted from the terminal device 20 and the terminal device 30, causes a designation menu screen shown in, for example,Figure 3 The date and time, terminal information, and sound data name are associated as shown and stored in the history information storage section 122.
[0073] In addition, the sound data accepting section 131 acquires the specified sound data in a case where the specification of the sound data is accepted. In this case, the sound data accepting section 131 acquires the specified sound data from the sound data storage section 121. The sound data accepting section 131 transmits the specified sound signal based on the acquired specified sound data to the terminal device of the specified source among the terminal devices 20 and 30, and causes the specified sound signal to be output to the terminal device of the specified source.
[0074] The signal processing section 132 acquires the sound data specified by the sound data accepting section 131 from the terminal devices 20 and 30 from the sound data storage section 121, and generates a mixed sound signal by mixing (mixing) sound signals based on the plurality of sound data that are specified. The signal processing section 132 acquires, for example, the first sound data specified by the terminal device 20 and the second sound data specified by the terminal device 30 from the sound data storage section 121, and generates a mixed sound signal (third sound signal) by mixing a first sound signal based on the first sound data and a second sound signal based on the second sound data.
[0075] Further, the signal processing section 132 acquires the specification of the sound data accepted by the sound data accepting section 131 by reading out the history information stored in the history information storage section 122. The signal processing section 132 acquires the sound data name of the history information stored in the history information storage section 122, and acquires the sound data corresponding to the acquired sound data name from the sound data storage section 121.
[0076] In addition, the signal processing section 132 generates a mixed sound signal (third sound signal) by mixing a plurality of sound signals based on a plurality of sound data that are specified in a case where the sound data accepting section 131 repeatedly accepts the specification of a plurality of sound data from the same terminal device among the terminal devices 20 and 30. In this case, the signal processing section 132 acquires the plurality of sound data that are specified from the sound data storage section 121, converts the plurality of sound data into a plurality of sound signals, and generates a mixed sound signal by mixing the plurality of sound signals.
[0077] The signal processing section 132 transmits the generated mixed sound signal (third sound signal) to the terminal device 20 via the NW communication section 11. The NW communication section 11 transmits the mixed sound signal (third sound signal) to the communication system 200 through the communication line connecting the communication system 200 and the terminal device 20.
[0078] Next, the signal processing system 1 according to the present embodiment will be described with reference to the drawings.
[0079] Figure 5 is a diagram illustrating one example of the operation of the signal processing system 1 according to the present embodiment. In addition, Figure 6 is a diagram illustrating one example of the operation of the signal processing system 1 according to the present embodiment to designate sound data from a plurality of terminal devices (20, 30).
[0080] One example of a case where one terminal device 20 and two terminal devices 30 (30-1, 30-2) are connected to the communication system 200 and the signal processing system 1 is shown. In addition, user AU1 represents a user of the terminal device 20 who views and listens to the stream data transmitted from the transmission source device 220, and user AU2 represents a user of the terminal device 30-1 who views and listens to the stream data transmitted from the transmission source device 220. In addition, user AU3 represents a user of the terminal device 30-2 who views and listens to the stream data transmitted from the transmission source device 220.
[0081] The terminal device 20 is connected to the communication system 200 through a communication line RT1, and the terminal device 30-1 is connected to the communication system 200 through a communication line RT2. In addition, the terminal device 30-2 is connected to the communication system 200 through a communication line RT3.
[0082] As shown in Figure 6 , first, the transmission source device 220 starts the recording of stream data (step S101). The transmission source device 220 records the stream data of the speech or presentation of the presenter PU through the connected microphone 230 and the built-in camera (not shown).
[0083] Next, the transmission source device 220 transmits the recorded stream data to the communication system server 210 (step S102). The transmission source device 220 transmits the recorded stream data to the communication system server 210 in real time.
[0084] Next, the communication system server 210 transmits the stream data to the terminal device 20 and the terminal devices 30 (30-1, 30-2) (step S103).
[0085] The terminal device 20 starts playing the stream data transmitted from the communication system server 210 (step S104). The communication processing section 261 receives the stream data through the communication line RT1, causes the image signal based on the received stream data to be displayed on the display section 23, and causes the sound signal based on the stream data to be output to the speaker 25.
[0086] In addition, the terminal device 30-1 starts playing the stream data transmitted from the communication system server 210 (step S105). The communication processing section 361 of the terminal device 30-1 receives the stream data via the communication line RT2, causes an image signal based on the received stream data to be displayed on the display section 33, and causes a sound signal based on the stream data to be output to the speaker 35.
[0087] In addition, the terminal device 30-2 starts playing the stream data transmitted from the communication system server 210 (step S106). The communication processing section 361 of the terminal device 30-2 receives the stream data via the communication line RT2, causes an image signal based on the received stream data to be displayed on the display section 33, and causes a sound signal based on the stream data to be output to the speaker 35.
[0088] Next, the terminal device 20 sends a designation of sound data to the application host device 10 (step S107). The application processing section 262 of the terminal device 20 sends a designation of sound data such as applause to the application host device 10 in correspondence with an operation of the input section 22 by the user AU1.
[0089] In addition, the terminal device 30-1 sends a designation of sound data to the application host device 10 (step S108). The application processing section 362 of the terminal device 30-1 sends a designation of sound data such as applause to the application host device 10 in correspondence with an operation of the input section 22 by the user AU2.
[0090] In addition, the terminal device 30-2 sends a designation of sound data to the application host device 10 (step S109). The application processing section 362 of the terminal device 30-2 sends a designation of sound data such as applause to the application host device 10 in correspondence with an operation of the input section 22 by the user AU3.
[0091] Next, the application host device 10 performs a reception process of sound data (step S110). The sound data reception section 131 of the application host device 10 receives sound data designated by the terminal device 20 (first sound data) and sound data designated by the terminal device 30-1 and the terminal device 30-2 (second sound data). The sound data reception section 131 associates a date and time, terminal information, and a sound data name in a case where a designation of sound data is received, and stores them in the history information storage section 122.
[0092] Next, the sound data accepting section 131 of the application host device 10 transmits the designated sound signal to the terminal device 20 (step S111). The sound data accepting section 131 acquires the sound data designated by the terminal device 20 from the sound data storage section 121, and transmits the sound signal based on the acquired sound data to the terminal device 20. Thus, the application processing section 262 of the terminal device 20 outputs the received designated sound signal to the speaker 25 (refer to Figure 5 ).
[0093] In addition, the sound data accepting section 131 of the application host device 10 transmits the designated sound signal to the terminal device 30-1 (step S112). The sound data accepting section 131 acquires the sound data designated by the terminal device 30-1 from the sound data storage section 121, and transmits the sound signal based on the acquired sound data to the terminal device 30-1. Thus, the application processing section 362 of the terminal device 30-1 outputs the received designated sound signal to the speaker 35 (refer to Figure 5 ).
[0094] In addition, the sound data accepting section 131 of the application host device 10 transmits the designated sound signal to the terminal device 30-2 (step S113). The sound data accepting section 131 acquires the sound data designated by the terminal device 30-2 from the sound data storage section 121, and transmits the sound signal based on the acquired sound data to the terminal device 30-2. Thus, the application processing section 362 of the terminal device 30-2 outputs the received designated sound signal to the speaker 35 (refer to Figure 5 ).
[0095] Next, the application host device 10 mixes all the designated sound signals (step S114). The signal processing section 132 of the application host device 10 acquires the names of all the sound data designated from the history information storage section 122, and acquires all the sound data corresponding to the acquired names of all the sound data from the sound data storage section 121. The signal processing section 132 mixes (mixes down) the sound signals based on the acquired all the sound data to generate a mixed sound signal (a 3rd sound signal).
[0096] For example, it is assumed that the terminal device 20 designates the sound data of "clap-01.wav", the terminal device 30-1 designates the sound data of "clap-02.wav", and the terminal device 30-2 designates the sound data of "clap-03.wav". In this case, the signal processing section 132 generates the sound signals of "clap-01.wav", "clap-02.wav", and "clap-03.wav" as shown in Figure 7 , and mixes the sound signals to generate a mixed sound signal as shown in Figure 8The mixed sound signal.
[0097] Further, the signal processing section 132 can generate the mixed sound signal as a mono sound signal as shown in Figure 8 Further, the signal processing section 132 can generate the mixed sound signal as a mono sound signal as shown in
[0098] Returning to the explanation of Figure 6 , the application host device 10 transmits the mixed sound signal to the terminal device 20 (step S115). The NW communication section 11 of the application host device 10 transmits the mixed sound signal generated by the signal processing section 132 to the terminal device 20 in order to make the mixed sound signal transmitted from the terminal device 20 to the communication system 200 through the communication line RT1.
[0099] Next, the terminal device 20 mixes the received mixed sound signal and the microphone input signal (input sound signal) (step S116). The communication processing section 261 of the terminal device 20 has a mix section 263 (one example of a mixing section) as shown in Figure 5 The mix section 263 mixes (mixes down) the mixed sound signal (the 3rd sound signal) received from the application host device 10 and the input sound signal (the 4th sound signal) picked up by the microphone 24, and generates a mixed sound signal (the 5th sound signal). Further, the mix section 263 outputs the mixed sound signal (the 3rd sound signal) received from the application host device 10 in the case where there is no sound signal input from the microphone 24.
[0100] Further, the terminal device 20 can be a structure which does not perform mixing with the input sound signal (the 4th sound signal) picked up by the microphone 24. For example, the terminal device 20 can be a structure which switches the sound signal input from the microphone 24 and the mixed sound signal (the 3rd sound signal) by a switch.
[0101] Next, the terminal device 20 transmits the mixed sound signal to the communication system server 210 (step S117). That is, the communication processing section 261 transmits the mixed sound signal (the 3rd sound signal or the 5th sound signal) output from the mix section 263 to the communication system 200 through the communication line RT1. Here, the communication line RT1 is a communication line which connects the communication system 200 and the terminal device 20.
[0102] Next, the communication system server 210 transmits the mixed sound signal to the transmission source device 220 (step S118). The communication system server 210 transmits the mixed sound signal (the 3rd sound signal or the 5th sound signal) received from the terminal device 20 to the transmission source device 220.
[0103] Next, the transmission source device 220 plays the received mixed sound signal (step S119). That is, the transmission source device 220 outputs the mixed sound signal such as applause designated by the users AU1 to AU3 through the terminal devices 20 and 30 (30-1, 30-2) to the speaker 240.
[0104] Thus, it is possible to convey the reactions such as applause of the users AU1 to AU3 to the transmitter PU who is transmitting a speech or a presentation in real time.
[0105] Next, the operation of the signal processing system 1 according to the present embodiment in which the plurality of sound data is designated from one terminal device will be described with reference to Figure 9 to the signal processing system 1 according to the present embodiment in which the plurality of sound data is designated from one terminal device.
[0106] Figure 9 is a view showing an example of the operation of the signal processing system 1 according to the present embodiment in which the plurality of sound data is designated from one terminal device.
[0107] In Figure 9 , the processes from step S201 to step S206 are the same as the processes from step S101 to step S106 shown in the above-described Figure 6 , and thus the description thereof is omitted here.
[0108] In step S207, the terminal device 30-1 transmits the plurality of designations of sound data to the application host device 10. The application processing section 362 of the terminal device 30-1 transmits the designations of, for example, three sound data of "Applause 01" to "Applause 03" to the application host device 10 in correspondence with the operation of the input section 22 by the user AU2.
[0109] Next, the application host device 10 performs the reception processing of sound data (step S208). The sound data reception section 131 of the application host device 10 receives the three sound data (the 2nd sound data) designated by the terminal device 30-1. The sound data reception section 131 associates the date and time, the terminal information, and the sound data names (for example, "Applause 01" to "Applause 03") in the case where the designations of the three sound data are received, and stores them in the history information storage section 122, as shown in Figure 3
[0110] Next, the sound data accepting section 131 of the application host device 10 transmits the specified three sound signals to the terminal device 20 (step S209). The sound data accepting section 131 acquires the three sound data (e.g., "Applause 01" to "Applause 03") specified by the terminal device 30-1 from the sound data storage section 121, and transmits the three specified sound signals based on the acquired three sound data to the terminal device 30-1. Thereby, the application processing section 362 of the terminal device 30-1 outputs the received three specified sound signals to the speaker 25.
[0111] Next, the application host device 10 mixes all of the specified sound signals (step S210). The signal processing section 132 of the application host device 10 acquires all of the specified three sound data names from the history information storage section 122, and acquires the three sound data corresponding to the acquired three sound data names from the sound data storage section 121. The signal processing section 132 mixes (mixes down) the three sound signals based on the acquired three sound data, and generates a mixed sound signal (third sound signal).
[0112] Next, the application host device 10 transmits the mixed sound signal to the terminal device 20 (step S211). The processes from step S211 to step S215 are the same as the processes from step S115 to step S119 shown in the above-described Figure 6 The processes from step S211 to step S215 are the same as the processes from step S115 to step S119 shown in the above-described
[0113] As described above, the signal processing system 1 according to the present embodiment is a system that is connected to a plurality of devices including at least a terminal device 20 (first terminal device) that receives stream data and a terminal device 30 (second terminal device), and a communication system 200 that enables communication between the plurality of devices, and has a sound data accepting section 131 (accepting section), a signal processing section 132, and a NW communication section 11 (transmitting section). The sound data accepting section 131 accepts a specification of first sound data from the terminal device 20 that receives stream data and a specification of second sound data from the terminal device 30 that receives stream data. The signal processing section 132 acquires the first sound data and the second sound data, and generates a third sound signal (mixed sound signal) that mixes a first sound signal based on the first sound data and a second sound signal based on the second sound data. The NW communication section 11 transmits the third sound signal to the communication system 200 through a communication line RT1 that connects the communication system 200 and the terminal device 20. Further, instead of the NW communication section 11, the transmitting section can be the communication processing section 261 or the NW communication section 21.
[0114] According to such a configuration, the signal processing system 1 according to the present embodiment transmits a third sound signal (mixed sound signal) in which a first sound signal based on the first sound data designated by the terminal device 20 and a second sound signal based on the second sound data designated by the terminal device 30 are mixed, to the communication system 200 via the communication line RT1 connected to the terminal device 20. Thus, the signal processing system 1 according to the present embodiment can aggregate sound data (applause, etc.) designated by the terminal device 20 and the terminal device 30 together, and transmit the same from one place to a transmission source of streaming data, etc., without changing the existing communication system 200. Thereby, the signal processing system 1 according to the present embodiment can convey the reaction of the user to the streaming data. Thus, the signal processing system 1 according to the present embodiment can accurately know the reaction of the user of the transmission target in the case of transmitting streaming data. The signal processing system 1 according to the present embodiment can convey, for example, the atmosphere of actually clapping to the transmission source (e.g., the transmitter PU) of the streaming data.
[0115] In addition, the signal processing system 1 according to the present embodiment mixes and transmits sound signals based on sound data designated by the sound data accepting section 131, and thus can suppress the occurrence of howling, echo, etc., unlike the input from the microphones (24, 34).
[0116] In addition, the signal processing system 1 according to the present embodiment has a mixing section 263 (mixing section) that mixes a fourth sound signal (input sound signal) picked up by the microphone 24 (sound pickup section) connected to the terminal device 20 and the third sound signal (mixed sound signal) to generate a fifth sound signal (mixed sound signal). The NW communication section 11 causes the fifth sound signal (mixed sound signal) to be transmitted to the transmission source (e.g., the transmission source device 220) of the streaming data of the communication system 200. Further, it can be configured such that the mixing section 263 is included in the signal processing section 132. That is, the signal processing section 132 mixes the fourth sound signal picked up by the microphone 24 (sound pickup section) connected to the terminal device 20 and the third sound signal to generate a fifth sound signal, and the fifth sound signal can be transmitted to the transmission source of the streaming data of the communication system 200 by the communication processing section 261 or the NW communication section 21 (one example of the transmission section).
[0117] Thereby, the signal processing system 1 according to the present embodiment can transmit the microphone input (fourth sound signal) mixed with the designated sound data (first sound signal and second sound signal) from one communication line RT1, and can convey the reaction of the user to the streaming data to the transmission source (e.g., the transmission source device 220) in real time.
[0118] In addition, in the present embodiment, the signal processing section 132 generates a third sound signal (mixed sound signal) that mixes a plurality of sound signals based on a plurality of sound data designated in the case where the sound data accepting section 131 repeatedly accepts a plurality of sound data from the same terminal device among the terminal device 20 and the terminal device 30.
[0119] Thus, the signal processing system 1 according to the present embodiment can mix a plurality of sound signals based on a plurality of sound data designated from one terminal device (the terminal device 20 or the terminal device 30) and transmit to the transmission source of the streaming data, and thus can more realistically convey the user's reaction such as applause with respect to the streaming data, and can perform a sense of presence.
[0120] Further, the signal processing system 1 according to the present embodiment has a sound data storage section 121 that stores sound data that can be designated by the sound data accepting section 131 (refer to Figure 2 ). The signal processing section 132 acquires the first sound data and the second sound data from the sound data storage section 121.
[0121] Thus, by storing various sound data in the sound data storage section 121, the signal processing system 1 according to the present embodiment can transmit sound signals based on various sound data to the transmission source of the streaming data, and can transmit a plurality of reactions of the user according to the situation to the transmission source of the streaming data.
[0122] In addition, the sound data accepting section 131 acquires the designated sound data in the case where the designation of the sound data is accepted, transmits a designated sound signal based on the acquired designated sound data to the terminal device of the designated source among the terminal device 20 and the terminal device 30, and causes the designated sound signal to be output to the terminal device of the designated source.
[0123] Thus, the signal processing system 1 according to the present embodiment can confirm the designated sound data in the form of a sound signal in the designated terminal device.
[0124] In addition, in the present embodiment, the sound data accepting section 131 causes a designation menu image G1 (image of a designation menu screen) including a sound data information indicating sound data that can be designated, and a play bar that can visually indicate a play time of the sound data, that is, a play bar of a length corresponding to the play time to be displayed on the terminal device 20 and the terminal device 30 (refer to Figure 4 ).
[0125] Therefore, the signal processing system 1 according to this embodiment enables users to select audio data based on the playback time of the audio data. Thus, the signal processing system 1 according to this embodiment improves user convenience.
[0126] [Second Implementation]
[0127] Next, the signal processing system 1a according to the second embodiment will be described. The difference between the second embodiment and the first embodiment is that the application host device 10 in the first embodiment is functionally divided into the application host device 10a and the application processing device 40.
[0128] Figure 10 This is a block diagram illustrating an example of the signal processing system 1a according to this embodiment.
[0129] The signal processing system 1a is a system connected to the communication system 200 and a plurality of devices including at least terminal device 20 and terminal device 30, and has an application host device 10a and an application processing device 40.
[0130] In addition, Figure 10 In the middle, regarding the above-mentioned Figure 1 The same structure is labeled with the same number, and its description is omitted.
[0131] Application host device 10a is a server device connected to network NW1, and is a web server that implements a website for specifying sound data, such as clapping. Application host device 10a uses the website for specifying sound data to receive specifications for sound data, such as clapping, from terminal devices 20 and 30. When the sound data receiving unit 131 receives a specification for first sound data from terminal device 20 or a specification for second sound data from terminal device 30, application host device 10a generates an instruction indicating that the first or second sound data has been specified and sends it to the application processing device 40 described below.
[0132] The application host device 10a includes an NW communication unit 11, a storage unit 12, and an application control unit 13a.
[0133] The application control unit 13a includes a processor such as a CPU, and controls the application host device 10a. The application control unit 13a has a voice data receiving unit 131 and an instruction processing unit 133.
[0134] The instruction processing section 133 generates an instruction indicating that the first sound data or the second sound data has been designated, in a case where the sound data accepting section 131 accepts designation of the first sound data from the terminal device 20 or designation of the second sound data from the terminal device 30. The instruction processing section 133 generates an instruction containing information (e.g., a sound data name, etc.) indicating the designated sound data, based on, for example, information stored in the history information storage section 122. The instruction processing section 133 transmits the generated instruction to the application processing device 40 via the NW communication section 11.
[0135] The application processing device 40 acquires the designated sound data, and generates a mixed sound signal (third signal) mixing sound signals based on the acquired sound data. The application processing device 40 acquires the designated sound data (e.g., the first sound data and the second sound data) based on the instruction generated by the instruction processing section 133 of the application host device 10a described above, mixes sound signals (e.g., the first sound signal and the second sound signal) based on the acquired sound data, and generates a mixed sound signal (third signal).
[0136] The application processing device 40 has a NW communication section 41, a sound data storage section 42, and a signal processing section 43.
[0137] The NW communication section 41 (one example of a transmission section) is an interface section capable of connecting to the network NW1, and performs communication with the application host device 10a and the terminal device 20 via the network NW1. The NW communication section 41 transmits the mixed sound signal (third sound signal) to the communication system 200 through the communication line RT1 connecting the communication system 200 and the terminal device 20.
[0138] The sound data storage section 42 stores sound data names in association with sound data, similarly to the sound data storage section 121.
[0139] The signal processing section 43 acquires sound data designated by the sound data accepting section 131 from the terminal device 20 and the terminal device 30 from the sound data storage section 42 based on an instruction received from the application host device 10a, mixes (mixes sound) sound signals based on the designated plurality of sound data, and generates a mixed sound signal. The signal processing section 43 acquires, for example, the first sound data designated by the terminal device 20 and the second sound data designated by the terminal device 30 from the sound data storage section 121, and generates a mixed sound signal (third sound signal) mixing a first sound signal based on the first sound data and a second sound signal based on the second sound data.
[0140] In addition, based on the instructions received from the application host device 10a, the signal processing unit 43 generates a mixed audio signal (the third audio signal) that mixes multiple audio signals based on the specified multiple audio data when the audio data receiving unit 131 repeatedly receives multiple audio data from the same terminal device, either the terminal device 20 or the terminal device 30.
[0141] The signal processing unit 43 transmits the generated mixed audio signal (third audio signal) to the terminal device 20 via the NW communication unit 41, and transmits the mixed audio signal (third audio signal) to the communication system 200 via the communication line RT1 connecting the communication system 200 and the terminal device 20. That is, the NW communication unit 41 transmits the mixed audio signal (third audio signal) to the communication system 200 via the communication line RT1 connecting the communication system 200 and the terminal device 20.
[0142] Next, refer to Figure 11 The operation of the signal processing system 1a according to this embodiment will be explained.
[0143] exist Figure 11 In this embodiment, the operation of the application host device 10a and the application processing device 40 differs from that in the first embodiment. The operation of other structures is also different. Figure 5 The first embodiment shown is the same.
[0144] The instruction processing unit 133 of the application host device 10a generates an instruction containing information representing the specified sound data (e.g., sound data name) based on information stored in the history information storage unit 122. The instruction processing unit 133 sends the generated instruction to the application processing device 40 via the NW communication unit 11.
[0145] Furthermore, the signal processing unit 43 of the application processing device 40, based on the instructions received from the application host device 10a, obtains the specified audio data from the audio data storage unit 42 and generates a mixed audio signal that combines audio signals based on the obtained audio data. Then, the NW communication unit 41 sends the generated mixed audio signal to the terminal device 20, and the mixed audio signal (the third audio signal) is sent to the communication system server 210 via the communication line RT1.
[0146] Figure 12 This is a diagram illustrating an example of the actions of a signal processing system 1a in specifying sound data from multiple terminal devices (20, 30).
[0147] exist Figure 12 In this process, the steps from S301 to S313 are the same as those described above. Figure 6 The processes shown from step S101 to step S113 are the same, so their description is omitted here.
[0148] Next, in step S314, the application host device 10a sends a command to mix the specified audio signals. The command processing unit 133 of the application host device 10a generates a command for mixing the specified audio signals, for example, based on information stored in the history information storage unit 122, containing a command that includes audio data names corresponding to the audio data specified by, for example, terminal device 20, terminal device 30-1, and terminal device 30-2. The command processing unit 133 sends the generated command to the application processing device 40 via the NW communication unit 11.
[0149] Next, the application processing device 40 mixes all the specified sound signals (step S315). Based on the received instructions, the signal processing unit 43 of the application processing device 40 retrieves all the sound data corresponding to all the specified sound data names from the sound data storage unit 42. The signal processing unit 43 mixes (mixes) the sound signals based on all the retrieved sound data to generate a mixed sound signal (third sound signal).
[0150] Next, the application processing device 40 sends the mixed audio signal to the terminal device 20 (step S316). The NW communication unit 41 of the application processing device 40 sends the mixed audio signal generated by the signal processing unit 43 from the terminal device 20 to the communication system 200 via the communication line RT1.
[0151] Next, the processing from step S317 to step S320 is the same as described above. Figure 6 The processes shown from step S116 to step S119 are the same, so their description is omitted here.
[0152] As explained above, the signal processing system 1a according to this embodiment includes an application host device 10a (server device) and an application processing device 40. The application host device 10a includes a voice data receiving unit 131, which generates an instruction indicating that the first or second voice data has been specified when the voice data receiving unit 131 receives a specification of first voice data or second voice data. The application processing device 40 includes a signal processing unit 43 and an NW communication unit 41 (transmitting unit). The signal processing unit 43 obtains the instruction generated by the application host device 10a, and based on the obtained instruction, obtains the first voice data specified by the terminal device 20 and the second voice data specified by the terminal device 30, generating a mixed voice signal (third voice signal) that mixes the first voice signal based on the first voice data with the second voice signal based on the second voice data. The NW communication unit 41 transmits the mixed voice signal (third voice signal) to the communication system 200 via the communication line RT1 connecting the communication system 200 and the terminal device 20.
[0153] Thus, the signal processing system 1a according to the present embodiment achieves the same effects as the signal processing system 1 according to the first embodiment described above, and can accurately know the reaction of the user of the transmission target in the case of the transmission stream data. In addition, the signal processing system 1a according to the present embodiment can reduce the processing load of the application host device 10a by dividing the processing to the application processing device 40 by generating the instruction by the application host device 10a.
[0154] Further, the present application is not limited to the embodiments described above, and can be changed within the scope of the gist of the present application.
[0155] For example, in each of the embodiments described above, an example in which the signal processing system 1 (1a) has the sound data storage section 121 (42) and acquires the sound data from the sound data storage section 121 (42) is described, but is not limited thereto, and for example, can be configured to acquire the sound data from an external file server or the like.
[0156] In addition, in the first and second embodiments described above, an example in which the terminal device 20 has the mixing section 263 is described, but is not limited thereto, and can be configured to have the mixing section 263 outside the terminal device 20. In addition, the signal processing system 1 (1a) can be configured to have the mixing section 263 as a part of the signal processing section 132 (43), or can be configured to have no mixing section 263.
[0157] In addition, in the first and second embodiments described above, an example in which the terminal device 20 that receives the stream data transmits the mixed sound signal to the communication system 200 is described, but can be configured such that, for example, a dedicated terminal that does not receive the stream data transmits the mixed sound signal to the communication system 200. In addition, can be configured such that the mixed sound signal is directly transmitted from the application host device 10 or the application processing device 40 to the communication system 200 using an account (for example, a prescribed user ID) that receives the stream data in the communication system 200.
[0158] In addition, in each of the embodiments described above, an example in which the signal processing system 1 (1a) generates the mixed sound signal in the form of a monaural sound signal when generating the mixed sound signal of the sound signal designated is described, but is not limited thereto, and can be configured to generate, for example, a multi-channel signal such as a stereophonic sound signal. In this case, the signal processing system 1 (1a) can also, for example, localize the sound image of the sound signal such as applause to each user ID. In addition, the signal processing system 1 (1a) can associate the user ID participating in the conference with the conference room ID, and adjust the sound field based on the conference room ID.
[0159] In addition, in each of the above-described embodiments, an example in which the signal processing system 1 (1a) simply mixes (mixes down) the sound signals when generating a mixed sound signal of the designated sound signals is described, but is not limited thereto, and can be configured to perform, for example, beat conversion, sound color conversion based on an equalizer, a change in volume, a change in PAN (PAN: Positioning), the application of an acoustic effect such as reverb, the replacement of a sound signal for a plurality of people that is created in advance, or the like, and generate a mixed sound signal.
[0160] In addition, in each of the above-described embodiments, an example in which the application host device 10 (10a) transmits the designated sound signal to the terminal device (20, 30) as described from step S111 to step S113 of the application host device 10 (10a) when the designation of the sound data is accepted is described, but can be configured such that the application host device 10 (10a) transmits the sound data in advance when the application of each terminal device (20, 30) is started. Figure 6
[0161] In addition, each of the above-described signal processing systems 1 (1a) has a computer system inside. Thus, the processing procedure of each of the above-described signal processing systems 1 (1a) is stored in a computer-readable storage medium in the form of a program, and the program is read out by a computer to be executed, whereby the above-described processing is performed. The computer-readable storage medium here refers to a magnetic disk, a magneto-optical disk, a CD-ROM, a DVD-ROM, a semiconductor memory, or the like. In addition, the computer program can be transmitted to a computer through a communication line, and the computer that receives the transmission can execute the program.
[0162] Explanation of Reference Signs
[0163] 1, 1a … signal processing system, 10, 10a … application host device, 11, 21, 31, 41 … NW communication section, 12 … storage section, 13, 13a … application control section, 20, 30, 30-1, 30-2, 30a, 30a-1, 30a-2 … terminal device, 22, 32 … input section, 23, 33 … display section, 24, 34, 230 … microphone, 25, 35, 240 … speaker, 26, 36 … terminal control section, 40 … application processing device, 42, 121, 121a … sound data storage section, 43, 132 … signal processing section, 122 … history information storage section, 131 … sound data acceptance section, 133 … instruction processing section, 200 … communication system, 210 … communication system server, 220 … transmission source device, 261, 361 … communication processing section, 262, 362 … application processing section, 263 … mixdown section, PU … transmitter, AU1, AU2, AU3 … user
Claims
1. A signal processing system connected to a plurality of devices including at least a first terminal device that receives streaming data and a second terminal device, and a communication system capable of communicating with the plurality of devices, the signal processing system having: an accepting section that accepts designation of first sound data from the first terminal device that receives the streaming data and designation of second sound data from the second terminal device that receives the streaming data by designating a menu image; a signal processing section that acquires the first sound data and the second sound data, and generates a third sound signal that mixes a first sound signal based on the first sound data and a second sound signal based on the second sound data; and a transmitting section that transmits the third sound signal to the communication system via a communication line that connects the communication system and the first terminal device, wherein the first terminal device transmits the third sound signal to the communication system, and, in contrast, the second terminal device does not transmit the third sound signal to the communication system.
2. The signal processing system according to claim 1, wherein the signal processing section mixes a fourth signal picked up by a pickup section connected to the first terminal device and the third sound signal, and generates a fifth sound signal, the transmitting section transmits the fifth sound signal to a transmission source of the streaming data of the communication system.
3. The signal processing system according to claim 1 or 2, wherein the signal processing section generates the third sound signal that mixes a plurality of sound signals based on a plurality of sound data designated repeatedly by the accepting section from the same one of the first terminal device and the second terminal device.
4. The signal processing system according to claim 1 or 2, wherein the signal processing system has a sound data storage section that stores sound data that can be designated by the accepting section, the signal processing section acquires the first sound data and the second sound data from the sound data storage section.
5. The signal processing system according to claim 1 or 2, wherein the accepting section, when the designation of sound data is accepted, acquires the designated sound data, transmits a designated sound signal based on the acquired designated sound data to a terminal device of a designated source among the first terminal device and the second terminal device, and causes the designated sound signal to be output to the terminal device of the designated source.
6. The signal processing system according to claim 1 or 2, wherein the signal processing system has a server device that has the accepting section, and, when the designation of the first sound data or the second sound data is accepted by the accepting section, generates an instruction indicating that the first sound data or the second sound data has been designated. The signal processing section acquires the instruction generated by the server device, acquires the first sound data and the second sound data based on the acquired instruction, and generates a third sound signal in which a first sound signal based on the first sound data and a second sound signal based on the second sound data are mixed.
7. The signal processing system according to claim 1 or 2, wherein The accepting section causes the specified menu image including a play bar and sound data information indicating sound data that can be specified to be displayed on the first terminal device and the second terminal device, the play bar being a play bar that visually indicates a play time of the sound data, or a play bar having a length corresponding to the play time.
8. A signal processing device connected to a plurality of devices including at least a first terminal device that receives stream data and a second terminal device, and a communication system that enables communication between the plurality of devices, The signal processing device has: an accepting section that accepts, through a specified menu image, a specification of first sound data from the first terminal device that receives the stream data and a specification of second sound data from the second terminal device that receives the stream data; a signal processing section that acquires the first sound data and the second sound data and generates a third sound signal in which a first sound signal based on the first sound data and a second sound signal based on the second sound data are mixed; and a transmitting section that transmits the third sound signal to the communication system through a communication line that connects the communication system and the first terminal device, wherein the first terminal device transmits the third sound signal to the communication system, and, in contrast, the second terminal device does not transmit the third sound signal to the communication system.
9. A signal processing method that is a signal processing method of a signal processing system connected to a plurality of devices including at least a first terminal device that receives stream data and a second terminal device, and a communication system that enables communication between the plurality of devices, In the signal processing method, an accepting section accepts, through a specified menu image, a specification of first sound data from the first terminal device that receives the stream data and a specification of second sound data from the second terminal device that receives the stream data, a signal processing section acquires the first sound data and the second sound data and generates a third sound signal in which a first sound signal based on the first sound data and a second sound signal based on the second sound data are mixed, a transmitting section transmits the third sound signal to the communication system through a communication line that connects the communication system and the first terminal device, wherein the first terminal device transmits the third sound signal to the communication system, and, in contrast, the second terminal device does not transmit the third sound signal to the communication system.
Citation Information
Patent Citations
Method for producing acetoacetylated polyvinyl alcohol-based resin
JP2007277492A
Server, communication method, and spectator terminal
JP2004094683A
Communication system, distribution device, and program
JP2016146544A