Media distribution system, media distribution method, and program

The media distribution system addresses the lack of realism in existing communication models by grouping users and using SFUs and MCUs to distribute group voices as background sound, improving the immersive experience and resource efficiency.

WO2025169454A1PCT designated stage Publication Date: 2025-08-14NT T INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/004513
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-09
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing communication models, such as SFU and MCU, fail to enhance the sense of realism in real-time communication by not distinguishing between user groups, leading to difficulty in creating an immersive interaction experience.

Method used

A media distribution system that divides user terminals into groups, using SFUs for intra-group communication and MCUs for inter-group communication, with a server control unit managing media distribution and attributes, synthesizing and distributing group voices as background sound.

Benefits of technology

Enhances the sense of realism and spaciousness in online communication by allowing users to hear group voices as background sound, reducing load on user terminals and server resources, and supporting high participant counts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024004513_14082025_PF_FP_ABST
    Figure JP2024004513_14082025_PF_FP_ABST
Patent Text Reader

Abstract

A media distribution system according to the present invention comprises: a plurality of user terminals (P) divided into a plurality of groups; and a control device (1) for controlling media distribution between the user terminals (P). The control device (1) includes: first media servers (121) that are provided in each of groups (G1-G3) and distribute media between each of the user terminals (P) belonging to each of the groups (G1-G3); a second media server (122) that distributes media between the groups (G1-G3); and a server control unit (13) that controls the connection of the first and second media servers (121, 122) and attributes of the media distributed by the first and second media servers (121, 122).
Need to check novelty before this filing date? Find Prior Art

Description

Media distribution system, media distribution method, and program

[0001] The present disclosure relates to a media distribution system, a media distribution method, and a program.

[0002] Online real-time communication is used in conferences and seminars attended by multiple users in remote locations. Known real-time communication models include a two-way communication model in which multiple users converse with each other, and a one-way communication model in which one user unilaterally distributes media to multiple other users.

[0003] Patent Document 1 describes a selective forwarding unit (SFU) that realizes a two-way communication model, and a multipoint control unit (MCU) that realizes a one-way communication model.

[0004] "The Next Paradigm Shift after IP Telephony: The Challenge of WebRTC," NTT Technical Journal, Vol. 27, No. 8, 2015, https: / / journal.ntt.co.jp / backnumber2 / 1508 / files / jn201508036.pdf, p. 38.

[0005] Recently, there has been a growing need for a communication model that allows users to share the sense of realism, such as the buzz and energy that arises when many people interact freely. However, in the SFU described above, all users are connected flatly, and if a small number of users form a group, it is impossible to distinguish between the voices of users in this group and those of users in other groups. This makes it difficult to enhance the sense of realism for users participating in real-time communication.

[0006] Furthermore, in the above-mentioned MCU, media is distributed in one direction, so many users cannot interact with each other. As a result, there is a problem that, like the above-mentioned SFU, it is not possible to enhance the sense of presence of users participating in real-time communication.

[0007] The present disclosure has been made in consideration of the above circumstances, and its purpose is to provide a media distribution system, a media distribution method, and a program that can enhance the sense of realism in real-time communication.

[0008] A media distribution system according to one aspect of the present disclosure comprises a plurality of user terminals divided into a plurality of groups, and a control device that controls media distribution between each of the user terminals, wherein the control device comprises a first distribution unit provided in each group that distributes media between each of the user terminals belonging to each group, a second distribution unit that distributes media between each of the groups, and a server control unit that controls the connection between the first and second distribution units and the attributes of the media distributed by the first and second distribution units.

[0009] A media distribution method according to one aspect of the present disclosure is a media distribution method for distributing media between a plurality of user terminals divided into a plurality of groups, wherein a first distribution unit performs media distribution between each user terminal belonging to each group, a second distribution unit performs media distribution between each group, and a server control unit controls the connection between the first and second distribution units and the attributes of the media distributed by the first and second distribution units.

[0010] One aspect of the present disclosure is a program for causing a computer to function as a control device of the media distribution system.

[0011] According to the present disclosure, it is possible to enhance the sense of realism in real-time communication.

[0012] FIG. 1 is a block diagram showing the configuration of a media distribution system according to an embodiment. FIG. 2 is an explanatory diagram showing media distribution within and between groups. FIG. 3 is an explanatory diagram showing the configuration of a first embodiment of a media server. FIG. 4 is a flowchart showing the processing procedure of a media server according to the first embodiment. FIG. 5 is an explanatory diagram showing the configuration of a second embodiment of a media server. FIG. 6 is a flowchart showing the processing procedure of a media server according to the second embodiment. FIG. 7 is an explanatory diagram showing the configuration of a third embodiment of a media server, showing a state in which the first arrangement is selected. FIG. 8 is an explanatory diagram showing the configuration of a third embodiment of a media server, showing a state in which the second arrangement is selected. FIG. 9 is a flowchart showing the processing procedure of a media server according to the third embodiment, showing processing for switching between the first arrangement and the second arrangement depending on the number of participating user terminals. FIG. 10 is a flowchart showing the processing procedure for switching a second media server from the first arrangement to the second arrangement according to the third embodiment. FIG. 11 is a flowchart showing the processing procedure for switching a second media server from the second arrangement to the first arrangement according to the third embodiment. FIG. 12 is a block diagram showing the hardware configuration of this embodiment.

[0013] An embodiment will be described below. In this embodiment, a plurality of users participating in online communication are divided into a plurality of groups. For example, as shown in FIG. 2, a plurality of users p are divided into four groups G1 to G4. Within each group, users p belonging to that group use user terminals to perform two-way communication using media (video, audio, etc.).

[0014] Furthermore, between groups, the speech of each user p in one group is synthesized to generate a synthetic speech, which is then distributed to the other groups as background sound. For example, the speech of each user p belonging to group G1 is synthesized and distributed to groups G2 to G4. In this way, the conversation of a group at one table is distributed as background sound to the groups at other tables, just like at a cocktail party.

[0015] Fig. 1 is a block diagram showing the configuration of a media distribution system 100 according to an embodiment. As shown in Fig. 1, the media distribution system 100 includes a plurality of user terminals P (only two are shown in the figure) operated by users, and a control device 1 connected to each user terminal P via a network line NW. The control device 1 is installed in, for example, a base station that comprehensively monitors each user terminal P.

[0016] The user terminals P are operated by users who are seated at different locations, for example. The user terminals P are mobile terminals, smartphones, PCs, etc. that can use online communication services.

[0017] The control device 1 includes a signaling server 11, a plurality of media servers 12, a server control unit 13, and a database 14.

[0018] The signaling server 11 controls sessions with user terminals P based on a preset protocol. Specifically, the signaling server 11 starts and stops communications between user terminals P, sets connection codecs (audio conversion, etc.), and negotiates participating groups (setting communication standards, etc.). The signaling server 11 outputs control commands necessary for media communication to the server control unit 13. The signaling server 11 collects various information, such as billing information for each group, as necessary.

[0019] The media server 12 distributes media (video, audio, etc.) between each user terminal P. The media server 12 appropriately mixes audio streams sent from the user terminal P that is the audio sender, and distributes this to the user terminal P on the viewing side. The media server 12 includes a first media server 121 (first distribution unit) and a second media server 122 (second distribution unit). The first media server 121 controls the distribution of media between the user terminals P that belong to each of the groups G1 to G4. The second media server 122 controls the distribution of audio between each of the groups G1 to G4.

[0020] That is, the first media server 121 is provided in each group and distributes media between the user terminals belonging to each group, and the second media server 122 distributes media between the groups.

[0021] The server control unit 13 acquires control commands output from the signaling server 11, and refers to the connection information of each user terminal P to perform various settings for deploying the media server 12 and delivering audio streams to the media server 12. Specifically, the server control unit 13 sets the synthesis method for the input audio streams, the delivery destination, etc. The connection information of the user terminal P includes information on the connected user terminals P and information on the groups in which the user terminal P is participating. In other words, the server control unit 13 controls the connections of the first and second media servers 121, 122 and the attributes of the media delivered by the first and second media servers 121, 122.

[0022] The database 14 stores connection information for each user terminal P set by the signaling server 11 or the server control unit 13. The database 14 may be installed in a housing separate from the components of the control device 1 in order to maintain the consistency of the stored data and to avoid the effects of a failure.

[0023] The media distribution system 100 according to this embodiment divides a plurality of user terminals P into a plurality of groups, and performs bidirectional media communication between the user terminals P in each group via a first media server 121. Furthermore, a second media server 122 distributes, to the user terminals P belonging to one group, speech uttered by a user terminal P belonging to another group as background sound.

[0024] A user of each user terminal P can participate in any of the multiple groups. Furthermore, each user can freely change the group they participate in. First to third embodiments of the media server 12 will be described below.

[0025] [Explanation of First Example] Fig. 3 is an explanatory diagram showing the configuration of a first example of the media server 12 in the media distribution system 100 according to the embodiment. As shown in Fig. 3, the media server 12 according to the first example has three SFUs 51 to 53 set in the first media server 121, and an MCU 54 set in the second media server 122.

[0026] That is, the server control unit 13 sets the first media server 121 as SFUs 51-53 (first method) that distributes media output from one user terminal belonging to the corresponding group to other user terminals belonging to the same group, and sets the second media server 122 as MCU 54 (second method) that combines media output from each first media server 121 and distributes it to user terminals belonging to other groups.

[0027] "SFU" is a method of distributing an upstream media stream from one user terminal P to the control device 1 to another user terminal P. "SFU" performs synthesis processing of multiple voices within the user terminal P. The number of streams that "SFU" distributes to the user terminal P is one upstream and "N-1" downstream (N is the number of user terminals). "Upstream" refers to the direction from the user terminal P to the media server 12, and "downstream" refers to the direction from the media server 12 to the user terminal P.

[0028] "MCU" is a method of multicasting the upstream media streams of user terminals P belonging to one group, and also synthesizing media streams from other user terminals P. The number of streams that "MCU" distributes to user terminals P is one for both upstream and downstream.

[0029] 3 illustrates an example in which multiple user terminals P participating in online communication are divided into three groups G1 to G3, with each group G1 to G3 having three user terminals P. That is, the first group G1 includes three user terminals P11, P12, and P13, the second group G2 includes three user terminals P21, P22, and P23, and the third group G3 includes three user terminals P31, P32, and P33. In the following, when a user terminal is specifically referred to, it will be referred to with a suffix, such as "user terminal P11," and when a user terminal is not specifically referred to or is referred to collectively, it will be referred to without a suffix, such as "user terminal P."

[0030] The number of groups is not limited to three, but may be two or four or more. The number of user terminals P belonging to each of groups G1 to G3 is not limited to three, but may be two or four or more. The number of user terminals P belonging to each group may be different. The same applies to Figures 5, 7, and 8, which will be described later.

[0031] The SFU 51 set in the first media server 121 relays media (audio, video, etc.) transmitted from one user terminal P belonging to the first group G1 and transmits it to another user terminal P belonging to the first group G1. The SFU 52 relays media transmitted from one user terminal P belonging to the second group G2 and transmits it to another user terminal P belonging to the second group G2. The SFU 53 relays media transmitted from one user terminal P belonging to the third group G3 and transmits it to another user terminal P belonging to the third group G3.

[0032] For example, the SFU 51 acquires the voice spoken by the user of the user terminal P11 in the first group G1 via line L1 and distributes it to the user terminals P12 and P13 via line L2. The SFU 51 acquires the voice spoken by the user of the user terminal P12 in the first group G1 via line L3 and transmits it to the user terminals P11 and P13 via line L4. The SFU 51 acquires the voice spoken by the user of the user terminal P13 in the first group G1 via line L5 and transmits it to the user terminals P11 and P12 via line L6. The same applies to the SFUs 52 and 53.

[0033] The MCU 54 set in the second media server 122 synthesizes the sounds transmitted from the user terminals P belonging to one group and transmits the synthesized sounds to another group.

[0034] For example, the MCU 54 acquires speech from users of three user terminals P11, P12, and P13 belonging to a first group G1 via lines L11, L12, and L13, respectively, and synthesizes the acquired speeches to generate synthetic speech R1. The MCU 54 acquires speech from users of three user terminals P21, P22, and P23 belonging to a second group G2 via lines L21, L22, and L23, respectively, and synthesizes the acquired speeches to generate synthetic speech R2. The MCU 54 acquires speech from users of three user terminals P31, P32, and P33 belonging to a third group G3 via lines L31, L32, and L33, respectively, and synthesizes the acquired speeches to generate synthetic speech R3.

[0035] The MCU 54 sets the volume of each of the synthesized voices R1, R2, and R3 to be lower than the volume of the voices distributed within each group.

[0036] The MCU 54 distributes the synthesized speech R1 and R2 to the user terminals P31, P32, and P33 of the third group G3 via line Lx3. The MCU 54 distributes the synthesized speech R1 and R3 to the user terminals P21, P22, and P23 of the second group G2 via line Lx2. The MCU 54 distributes the synthesized speech R2 and R3 to the user terminals P11, P12, and P13 of the first group G1 via line Lx1.

[0037] Therefore, a voice synthesized from the voices spoken in the second and third groups G2 and G3 is delivered at a low volume to the first group G1. A voice synthesized from the voices spoken in the first and third groups G1 and G3 is delivered at a low volume to the second group G2. A voice synthesized from the voices spoken in the first and second groups G1 and G2 is delivered at a low volume to the third group G3. In other words, in one group, the voices spoken in the other groups are heard as background sound.

[0038] Next, the procedure for media distribution by the media server 12 according to the first embodiment will be described with reference to the flowchart shown in FIG.

[0039] 4, the SFU set in the first media server 121 distributes audio and video generated at a user terminal P in the same group to other user terminals P. For example, the SFU distributes audio uttered by a user at user terminal P11 in the first group G1 to user terminals P12 and P13.

[0040] In step S12, the MCU 54 of the second media server 122 acquires the sounds generated by the user terminals P of the groups G1 to G3.

[0041] In step S13, the MCU 54 synthesizes the speech generated at each user terminal P in one group. Specifically, the MCU 54 synthesizes the speech generated at each user terminal P11 to P13 in the first group G1 shown in Fig. 3 to generate synthetic speech R1, synthesizes the speech generated at each user terminal P21 to P23 in the second group G2 to generate synthetic speech R2, and synthesizes the speech generated at each user terminal P31 to P33 in the third group G3 to generate synthetic speech R3.

[0042] In step S14, the MCU 54 reduces the volume of each of the synthesized voices R1 to R3.

[0043] In step S15, the MCU 54 distributes the synthesized speech after the volume reduction to the user terminals P of the other groups. Specifically, the MCU 54 distributes synthesized speech R1 to the second and third groups G2 and G3, the MCU 54 distributes synthesized speech R2 to the first and third groups G1 and G3, and the MCU 54 distributes synthesized speech R3 to the first and second groups G1 and G2.

[0044] Thus, the media distribution system 100 of this embodiment comprises a plurality of user terminals P divided into a plurality of groups G1 to G3, and a control device 1 that controls media distribution between each user terminal P, and the control device 1 is provided in each group G1 to G3 and comprises a first media server 121 (first distribution unit) that distributes media between each user terminal P belonging to each group G1 to G3, a second media server 122 (second distribution unit) that distributes media between each group G1 to G3, and a server control unit 13 that controls the connection between the first and second media servers 121, 122 and the attributes of the media distributed by the first and second media servers 121, 122.

[0045] Therefore, media communication can be performed between user terminals P within each group, and furthermore, each user terminal P belonging to one group can hear sounds generated in other groups as background sound. This enhances the sense of realism experienced by each user terminal P participating in online communication. It also allows users of each user terminal P to feel the spaciousness of the space.

[0046] In the first embodiment, the MCU 54 is set up in the second media server 122, and speech generated in one group is synthesized and distributed to the other group. For example, synthesized speech R1 generated in the first group G1 and synthesized speech R2 generated in the second group G2 are acquired, and then these are synthesized and distributed to the third group G3. Therefore, the number of downstream streams for the third group G3 is one (line Lx1 shown in Figure 3). This reduces the load required for speech synthesis on each user terminal P, and allows the upper limit on the number of participating user terminals P to be set high. In other words, this is extremely useful when the number of participating user terminals P is large.

[0047] [Explanation of Second Example] Next, a second example will be described. Fig. 5 is an explanatory diagram showing the configuration of a second example of the media server 12 in the media distribution system 100 according to the embodiment. As shown in Fig. 5, in the media distribution system 100 according to the second example, three SFUs 61 to 63 are set in the first media server 121, and three MCUs 64 to 66 and one SFU 67 are set in the second media server 122.

[0048] That is, the server control unit 13 sets the first media server 121 as SFUs 51 to 53 (first method) that distribute media output from one user terminal belonging to the corresponding group to other user terminals belonging to the same group, and sets the second media server 122 as MCUs 64 to 66 (third method) that combine multiple media distributed from multiple user terminals belonging to each group, and as SFU 67 (fourth method) that distributes the media combined by MCUs 64 to 66 to other groups.

[0049] The processing by the SFUs 51 to 53 set in the first media server 121 is the same as that in the first embodiment, and therefore a description thereof will be omitted.

[0050] The MCUs 64 to 66 set in the second media server 122 synthesize the voices transmitted from the user terminals P belonging to one group.

[0051] The MCU 64 acquires speech from users of the three user terminals P11, P12, and P13 belonging to the first group G1 via lines L11, L12, and L13, respectively, and generates synthetic speech R4 by synthesizing the acquired speech. The MCU 64 outputs the synthetic speech R4 to the SFU 67.

[0052] The MCU 65 acquires speech from users of the three user terminals P21, P22, and P23 belonging to the second group G2 via lines L21, L22, and L23, respectively, and generates synthetic speech R5 by synthesizing the acquired speech. The MCU 65 outputs the synthetic speech R5 to the SFU 67.

[0053] The MCU 66 acquires speech from users of the three user terminals P31, P32, and P33 belonging to the third group G3 via lines L31, L32, and L33, respectively, and generates synthetic speech R6 by synthesizing the acquired speech. The MCU 66 outputs the synthetic speech R6 to the SFU 67.

[0054] The SFU 67 set in the second media server 122 acquires the synthesized voices R4 to R6 output from the three MCUs 64 to 66. The SFU 67 sets the volume of each synthesized voice R4 to R6 to be lower than the volume of the voice distributed within each group.

[0055] The SFU 67 distributes the synthesized speech R4 output from the MCU 64 to the SFUs 52 and 53 of the first media server 121. The synthesized speech R4 is distributed as background sound to the user terminals P21 to P23 and P31 to P33 belonging to the second and third groups G2 and G3.

[0056] The SFU 67 distributes the synthesized voice R5 output from the MCU 65 to the SFUs 51 and 53 of the first media server 121. The synthesized voice R5 is distributed as background sound to the user terminals P11 to P13 and P31 to P33 belonging to the first and third groups G1 and G3.

[0057] The SFU 67 distributes the synthesized speech R6 output from the MCU 66 to the SFUs 51 and 52 of the first media server 121. The synthesized speech R6 is distributed as background sound to the user terminals P11 to P13 and P21 to P23 belonging to the first and second groups G1 and G2.

[0058] Next, a procedure for media distribution by the media server 12 according to the second embodiment will be described with reference to the flowchart shown in FIG.

[0059] 6, the SFU 51 set in the first media server 121 distributes audio and video generated at a user terminal P in the same group to other user terminals P. For example, the audio uttered by a user at user terminal P11 in the first group G1 is distributed to user terminals P12 and P13.

[0060] In step S22, each of the MCUs 64 to 66 of the second media server 122 acquires speech generated by each of the user terminals P in each of the groups G1 to G3 and synthesizes the acquired speech. Specifically, the MCU 64 synthesizes speech generated by each of the users at the user terminals P11 to P13 in the first group G1 to generate synthetic speech R4. Similarly, the MCUs 65 and 66 generate synthetic speech R5 and R6.

[0061] In step S23, the SFU 67 of the second media server 122 acquires the synthesized voices R4 to R6 synthesized by the MCUs 64 to 66 corresponding to the groups G1 to G3.

[0062] In step S24, the SFU 67 reduces the volume of each of the synthesized voices R4 to R6.

[0063] In step S25, the SFU 67 distributes the synthetic speeches R4 to R6 after reducing the volume to each user terminal P in each of the groups G1 to G3. Specifically, the SFU 67 distributes the synthetic speeches R4 and R5 to the third group G3, distributes the synthetic speeches R4 and R6 to the second group G2, and distributes the synthetic speeches R5 and R6 to the first group G1.

[0064] In the second embodiment, as in the first embodiment, media communication can be performed between user terminals P within each group, and furthermore, each user terminal P belonging to one group can hear sounds generated in other groups as background sound. This enhances the sense of realism experienced by each user terminal P participating in online communication. Furthermore, the users of each user terminal P can feel the spaciousness of the space.

[0065] In the second embodiment, the second media server 122 is configured with three MCUs 64 to 66 and one SFU 67, and the voices uttered in each of the groups G1 to G3 are synthesized by each of the MCUs 64 to 66. The SFU 67 acquires the voices synthesized by each of the MCUs 64 to 66 and distributes them to the user terminals P of each of the groups G1 to G3. This reduces the calculation load required for the synthesis process in the second media server 122 and also reduces computer resources.

[0066] [Description of Third Example] Next, a third example will be described. Figures 7 and 8 are explanatory diagrams showing the configuration of a third example of the media server 12 in the media delivery system 100 according to an embodiment. As shown in Figures 7 and 8, in the media delivery system 100 according to the third example, three SFUs 51 to 53 are set in the first media server 121.

[0067] The second media server 122 also has a first arrangement 91 including an MCU 54 and a second arrangement 92 including three MCUs 64 - 66 and one SFU 67 .

[0068] The processing by the SFUs 51 to 53 set in the first media server 121 is the same as that of the SFUs 51 to 53 shown in the first and second embodiments.

[0069] The MCU 54 of the first arrangement 91 set in the second media server 122 is the same as the MCU 54 shown in the first embodiment (FIG. 3). The MCUs 64 to 66 and the SFU 67 of the second arrangement 92 set in the second media server 122 are the same as the MCUs 64 to 66 and the SFU 67 shown in the second embodiment (FIG. 5).

[0070] In the third embodiment, the server control unit 13 selects either the first arrangement 91 or the second arrangement 92 depending on the number of user terminals P participating in online communication. In detail, the server control unit 13 counts the number of participating user terminals P, and selects the first arrangement 91 if the number of participating user terminals P is equal to or greater than a predetermined threshold, and selects the second arrangement 92 if the number is less than the threshold. In addition, the server control unit 13 executes processing to mute the audio of each user terminal P when switching the connection to the second media server 122.

[0071] 7 shows the connection state when the first arrangement 91 is selected. With this connection, the media server 12 has the same configuration as that of the first embodiment (FIG. 3).

[0072] 8 shows the connection state when the second arrangement 92 is selected. With this connection, the media server 12 has the same configuration as that of the second embodiment (FIG. 5).

[0073] The reason for the switching operation will be explained below. In the media distribution by the media server 12 shown in the first embodiment (FIG. 3) described above, the number of streams H1 to be combined at the user terminal P is "H1 = (N-1) + 1" (where "N" is the number of user terminals in the group). Therefore, the number of streams H1 does not depend on the number of groups.

[0074] On the other hand, in the media distribution by the media server 12 shown in the second embodiment (FIG. 5) described above, the number of streams H2 to be synthesized at the user terminal P is "H2 = (N-1) + (M-1)" (where "M" is the number of groups). Therefore, the number of streams H2 depends on the number of groups.

[0075] Therefore, since the processing capacity of the audio stream in the user terminal P is limited, when the number of participating user terminals P is large (above a predetermined threshold), it is better to set it to the first arrangement 91 (corresponding to the first embodiment), which has a small number of streams to be synthesized.

[0076] On the other hand, as explained in the second embodiment above, by setting the second media server 122 to the second arrangement 92 (corresponding to the second embodiment), there is an advantage in that the computer resources required for synthesis processing in the second media server 122 can be reduced.

[0077] Therefore, in the third embodiment, the server control unit 13 controls the selection of either the first arrangement 91 or the second arrangement depending on the increase or decrease in the number of participating user terminals P. Specifically, if the number of participating user terminals P is equal to or greater than a predetermined threshold, the first arrangement 91 is selected as shown in Fig. 7, and if it is less than the threshold, the second arrangement 92 is selected as shown in Fig. 8.

[0078] That is, the server control unit 13 configures the first media server 121 (first distribution unit) as SFUs 51 to 53 (first method) that distribute media output from one user terminal belonging to a corresponding group to other user terminals belonging to this group, and configures the second media server 122 (second distribution unit) as MCU 54 (second method) that synthesizes media output from each first media server and distributes it to user terminals belonging to other groups. In this case, it is possible to set either one of the following arrangements: a first arrangement 91 in which the first media server 121 is set to SFUs 51 to 53 (first method) that distribute media to other user terminals belonging to this group, and a second media server 122 is set to MCUs 64 to 66 (third method) that combine multiple media distributed from multiple user terminals belonging to each group, and an SFU 67 (fourth method) that distributes the media combined by MCUs 64 to 66 to other groups; and a second arrangement 92 in which the first media server 121 is set to SFUs 51 to 53 (first method) that distribute media distributed from multiple user terminals belonging to each group to other user terminals belonging to this group, and a second media server 122 is set to MCUs 64 to 66 (third method) that combine multiple media distributed from multiple user terminals belonging to each group, and an SFU 67 (fourth method) that distributes the media combined by MCUs 64 to 66 to other groups. When the number of participants of user terminals P is equal to or greater than a predetermined threshold number, the first arrangement 91 is used, and when the number is less than the threshold number, the second arrangement 92 is used.

[0079] Next, a processing procedure for media distribution by the media server 12 according to the third embodiment will be described with reference to the flowcharts shown in Figures 9 to 11. Figure 9 is a flowchart showing the processing for setting the media server 12 to either a first arrangement 91 or a second arrangement 92.

[0080] 9, the server control unit 13 determines whether the number of participating user terminals P is equal to or greater than a predetermined threshold. If it is equal to or greater than the threshold (S1; YES), the process proceeds to step S2; if it is not equal to or greater than the threshold (S1; NO), the process proceeds to step S3.

[0081] In step S2, the server control unit 13 sets the second media server 122 in the first arrangement 91.

[0082] In step S3, the server control unit 13 sets the second media server 122 in the second arrangement 92. After that, this process ends.

[0083] FIG. 10 is a flowchart showing the processing procedure when the second media server 122 is set in the first arrangement 91 and the second arrangement is selected.

[0084] 10, the server control unit 13 determines whether or not there is a switching input from the first arrangement 91 to the second arrangement 92. If there is a switching input (S31; YES), the process proceeds to step S32.

[0085] In step S32, the server control unit 13 deploys a container / VM of the MCU specification used in the second deployment 92 and installs the software.

[0086] In step S33, the server control unit 13 sets information about the participating user terminal P in the MCU 67.

[0087] In step S34, the server control unit 13 mutes the audio of all user terminals P. That is, it temporarily stops accepting audio input.

[0088] In step S35, the server control unit 13 notifies the user of each user terminal P that the audio has been muted.

[0089] In step S36, the server control unit 13 switches the second media server 122 from the first arrangement 91 to the second arrangement 92. That is, the connection shown in FIG. 7 is switched to the connection shown in FIG.

[0090] In step S37, the server control unit 13 unmutes all the user terminals P.

[0091] In step S38, the server control unit 13 deletes the container / VM of the MCU 54 in the first arrangement 91. Then, this process ends. In this way, when the second media server 122 is set to the first arrangement 91, if the number of participating user terminals P falls below the threshold, the second media server 122 is switched to the second arrangement 92.

[0092] Next, a description will be given of a procedure for switching from the second arrangement 92 to the first arrangement 91. Fig. 11 is a flowchart showing a processing procedure when the first arrangement 91 is selected when the media server 12 is set to the second arrangement 92.

[0093] 11, the server control unit 13 determines whether or not there is a switching input from the second arrangement 92 to the first arrangement 91. If there is a switching input (S41; YES), the process proceeds to step S42.

[0094] In step S42, the server control unit 13 deploys a container / VM of the MCU specification used in the first deployment 91 and installs the software.

[0095] In step S43, the server control unit 13 sets information about the participating user terminal P in the MCU 54.

[0096] In step S44, the server control unit 13 mutes the audio of all user terminals P. That is, it temporarily stops accepting audio input.

[0097] In step S45, the server control unit 13 notifies each user terminal P that the audio has been muted.

[0098] In step S46, the server control unit 13 switches the second media server 122 from the second arrangement 92 to the first arrangement 91. That is, the connection shown in FIG. 8 is switched to the connection shown in FIG.

[0099] In step S47, the server control unit 13 unmutes all the user terminals P.

[0100] In step S48, the server control unit 13 deletes the containers / VMs of the MCUs 64 to 66 and the SFU 67 in the second arrangement 92. Then, this process ends. In this way, when the second media server 122 is set in the second arrangement 92 and the number of participating user terminals P exceeds the threshold, the second media server 122 is switched to the first arrangement 91.

[0101] In the third embodiment, as in the first and second embodiments, media communication can be performed between user terminals P within each group, and furthermore, each user terminal P belonging to one group can hear sounds generated in other groups as background sound. This enhances the sense of realism experienced by each user terminal P participating in online communication. Furthermore, the users of each user terminal P can feel the spaciousness of the space.

[0102] In the third embodiment, when the number of participating user terminals P is equal to or greater than the threshold, the first arrangement 91 is selected, which makes it possible to reduce the load required for voice synthesis processing at each user terminal P, as shown in the first embodiment described above. Also, when the number of participating user terminals P is less than the threshold, the second arrangement 92 is selected, which makes it possible to reduce the calculation load on the second media server 122 and save computer resources, as shown in the second embodiment described above.

[0103] Furthermore, when switching between the first arrangement 91 and the second arrangement 92, the audio of each user terminal P is muted, thereby preventing unpleasant audio generated during switching from being distributed to each user terminal P.

[0104] Furthermore, the volume of the audio distributed from other groups is set lower than the audio distributed within one's own group, so that audio generated at a user terminal P of another group can be heard as background sound within one's own group, enhancing the sense of realism of online communication and allowing each user to feel the spaciousness of the space.

[0105] The control device 1 of the media distribution system 100 of the present embodiment described above can be, for example, a general-purpose computer system including a CPU (Central Processing Unit, processor) 901, a memory 902, a storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), a communication device 904, an input device 905, and an output device 906, as shown in Fig. 12. The memory 902 and the storage 903 are storage devices. In this computer system, the CPU 901 executes a predetermined program loaded on the memory 902, thereby realizing each function of the control device 1.

[0106] The control device 1 may be implemented by one computer or by multiple computers, or may be a virtual machine implemented on a computer.

[0107] The program for the control device 1 can be stored in a computer-readable recording medium such as a HDD, SSD, USB (Universal Serial Bus) memory, CD (Compact Disc), or DVD (Digital Versatile Disc), or can be distributed via a network. The computer-readable recording medium is, for example, a non-transitory recording medium.

[0108] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the present disclosure.

[0109] REFERENCE SIGNS LIST 1 control device 11 signaling server 12 media server 13 server control unit 14 database 91 first arrangement 92 second arrangement 100 media distribution system 121 first media server 122 second media server G1 to G3 first to third groups NW network line P user terminal

Claims

1. A media distribution system comprising: a plurality of user terminals divided into a plurality of groups; and a control device that controls media distribution between each of the user terminals, wherein the control device is provided in each group and comprises: a first distribution unit that distributes media between each of the user terminals belonging to each group; a second distribution unit that distributes media between each of the groups; and a server control unit that controls the connection between the first and second distribution units and the attributes of the media distributed by the first and second distribution units.

2. The media distribution system of claim 1, wherein the server control unit sets the first distribution unit to a first method of distributing media output from one user terminal belonging to a corresponding group to other user terminals belonging to the group, and sets the second distribution unit to a second method of synthesizing media output from each first distribution unit and distributing the combined media to user terminals belonging to other groups.

3. The media distribution system of claim 1, wherein the server control unit sets the first distribution unit to a first method of distributing media output from one user terminal belonging to a corresponding group to other user terminals belonging to the group, and sets the second distribution unit to a third method of synthesizing multiple media distributed from multiple user terminals belonging to each group, and a fourth method of distributing the media synthesized by the third method to other groups.

4. The server control unit is capable of setting either one of the following arrangements: a first arrangement in which the first distribution unit is set to a first method of distributing media output from one user terminal belonging to a corresponding group to other user terminals belonging to the group, and the second distribution unit is set to a second method of synthesizing media output from each first distribution unit and distributing the media to user terminals belonging to other groups; or a second arrangement in which the first distribution unit is set to a first method of distributing media output from one user terminal belonging to a corresponding group to other user terminals belonging to the group, and the second distribution unit is set to a third method of synthesizing multiple media distributed from multiple user terminals belonging to each group, and a fourth method of distributing media synthesized by the third method to other groups; and the media distribution system described in claim 1, wherein the first arrangement is used when the number of participants in the user terminals is equal to or greater than a predetermined threshold number, and the second arrangement is used when the number of participants is less than the threshold number.

5. A media distribution method for distributing media between multiple user terminals divided into multiple groups, wherein a first distribution unit executes media distribution between each user terminal belonging to each group, a second distribution unit executes media distribution between each group, and a server control unit controls the connection between the first and second distribution units and the attributes of the media distributed by the first and second distribution units.

6. A program that causes a computer to function as a control device for the media distribution system according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Conference control method, system, and program

    WO2008078555A1

  • Video display system and video display method

    WO2022220308A1