Information processing system, information processing method, and program
The system facilitates private conversations in web conferences by managing voice distribution based on user requests, addressing the lack of personal communication in existing online meeting technologies.
Patent Information
- Application Number
- JP2025176449
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-01-08
AI Technical Summary
Existing online communication technologies do not allow for private conversations between specific individuals during web conferences, limiting the ability to share personal thoughts with intended recipients.
An information processing system that includes a request receiving mechanism for whispered conversations, a selection receiving mechanism for response methods, and a control mechanism to manage voice distribution between participants, enabling private conversations during web conferences.
Enables confidential discussions with specific individuals during online meetings by controlling voice distribution, allowing participants to share personal thoughts privately.
Smart Images

Figure 2026002928000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]
[0002] Recently, as the trend towards promoting remote work has accelerated, meetings are increasingly being held online (web conferences) rather than in real-life face-to-face spaces. While this type of remote communication has the advantage of eliminating the need to secure a physical conference room, it can also lead to inconveniences that would not occur in face-to-face meetings, as participants cannot share the same space and converse face-to-face.
[0003] Patent Document 1 discloses a technology that solves one of the issues of remote communication, "collisions of speech." Specifically, the technology describes a technology that reduces collisions of speech by predicting the next speaker and highlighting that speaker. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-111643 DISCLOSURE OF THE INVENTION [Problem to be solved by the invention]
[0005] The technology described in Patent Document 1 is based on the premise that comments are shared with all conference participants, but in conversations, what people want to say is often something they want to convey only to specific individuals.
[0006] Therefore, an object of the present invention is to provide a mechanism that enables a conversation with a specific individual in an online conference or the like. [Means for solving the problem]
[0007] The information processing system of the present invention is characterized by comprising a request receiving means for receiving a request for a whispered conversation from a first user participating in a Web conference to a second user, a selection receiving means for receiving a selection of a response method for the whispered conversation from the second user who is the recipient of the request for the whispered conversation received by the request receiving means, and a control means for controlling the distribution of the voices of the first user and the second user in accordance with the response method selected and received by the selection receiving means. [Effects of the Invention]
[0008] According to the present invention, it is possible to have a conversation with a specific individual in an online conference or the like. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of an information processing system to which the present invention can be applied. [Figure 2] A block diagram showing an example of the hardware configuration of an information processing device used as a client terminal 101 or a server device 102. [Figure 3] Flowchart showing the processing contents of the present invention [Figure 4] Flowchart showing details of the process in step S302 [Figure 5] Flowchart showing details of the process in step S304 [Figure 6] Flowchart showing details of the process in step S308 [Figure 7] Flowchart showing details of the process in step S608 [Figure 8] An example of a screen displayed on the client terminal 101 [Figure 9] An example of a screen displayed on the client terminal 101 [Figure 10] An example of a screen displayed on the client terminal 101 [Figure 11] An example of a screen displayed on the client terminal 101 [Figure 12] FIG. 1 is a diagram illustrating an outline of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0011] FIG. 1 is a diagram showing an example of the configuration of an information processing system to which the present invention can be applied.
[0012] As shown in FIG. 1, a client terminal 101 and a server device 102 are connected to each other via a network 110 so that they can communicate with each other.
[0013] The client terminal 101 is a terminal used and operated by a user participating in a web conference or the like, and is a terminal on which the screens shown in Figs. 8 to 11 are displayed. As shown in Fig. 1, a plurality of client terminals 101 are connected via a network 110. This configuration is based on the premise that each user participating in the web conference will use one client terminal. Specific examples of the client terminal 101 include, but are not limited to, personal computers (such as notebook PCs or desktop PCs), tablet terminals, and smartphones.
[0014] The server device 102 is a device that controls the exchange of data between client terminals in the Web conference building. It has the function of receiving user voice data and image data from the client terminal based on user operations on the client terminal 101, and distributing them to other client terminals. Although only one server device 102 is shown in FIG. 1, it may also be a system in which multiple server devices cooperate with each other, or a so-called cloud system.
[0015] FIG. 2 is a block diagram showing an example of the hardware configuration of an information processing device used as the client terminal 101 or the server device 102 of the present invention.
[0016] As shown in FIG. 2, the information processing device is connected to a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, a storage device 204, an input controller 205, an audio controller 206, a video controller 207, a memory controller 208, and a communication I / F controller 209 via a system bus 200.
[0017] The CPU 201 controls all devices and controllers connected to the system bus 200 .
[0018] ROM202 or external memory 213 stores the BIOS (Basic Input / Output System) and OS (Operating System), which are control programs executed by CPU201, computer-readable and executable programs for realizing this information processing method, and various necessary data (including data tables).
[0019] The RAM 203 functions as a main memory, a work area, etc. for the CPU 201. The CPU 201 loads programs and the like required for executing processing from the ROM 202 or the external memory 213 into the RAM 203, and executes the loaded programs to realize various operations.
[0020] The input controller 205 controls input from input devices such as a keyboard 210 and a pointing device such as a mouse (not shown). If the input device is a touch panel, the user can issue various instructions by pressing (touching with a finger or the like) icons, cursors, or buttons displayed on the touch panel.
[0021] The touch panel may also be a touch panel capable of detecting positions touched by multiple fingers, such as a multi-touch screen.
[0022] The video controller 207 controls the display on an external output device such as a display 212. The display also includes the display of a notebook computer integrated with the main body. Note that the external output device is not limited to a display, and may be, for example, a projector. In addition, for devices capable of receiving the above-mentioned touch operation, an input device is also provided.
[0023] The video controller 207 can control a video memory (VRAM) for display control, and can use part of the RAM 203 as a video memory area, or can provide a separate dedicated video memory.
[0024] The memory controller 208 controls access to the external memory 213. The external memory may be an external storage device (hard disk) that stores a boot program, various applications, font data, user files, edited files, and various data, a flexible disk (FD), or a CompactFlash (registered trademark) memory connected to a PCMCIA card slot via an adapter.
[0025] The communication I / F controller 209 connects and communicates with external devices via a network, and executes communication control processing on the network. For example, communication using TCP / IP, telephone lines such as ISDN, and 4G and 5G mobile phone lines are possible.
[0026] The CPU 201 enables display on the display 212 by, for example, executing a process of expanding (rasterizing) an outline font into a display information area in the RAM 203. The CPU 201 also enables user instructions using a mouse cursor (not shown) on the display 212.
[0027] Next, an outline of the present invention will be described with reference to FIG.
[0028] When a Web conference starts, a screen including images of the conference participants (images taken by an in-camera, etc.) is displayed on the client terminal 101 of each conference participant, as shown in Fig. 12(1). Fig. 12(1) shows that four users, A, B, C, and D, are participating in the conference. It is not necessary to display images of all participants participating in the conference, and a configuration may be adopted in which only a predetermined number of participants are displayed.
[0029] Here, user A selects user B's area by a predetermined operation (such as clicking or tapping continuously for a certain period of time) to request a whisper with user B. When a whisper request is made, the screen displays user A's image in the area of user B's image, as shown in Figure 12(2), and information identifying user A (such as name) is displayed in user A's area.
[0030] When a whisper is requested, User B selects a response method that corresponds to the request via the response selection interface shown in Figure 11. If a response that allows the whisper is selected, the screen transitions to Figure 12(3), and the whisper conversation begins.
[0031] When a whispered conversation begins, the voices of user A and user B are controlled so that they can only be heard by the person they whispered to (in this case, user A's voice is heard only by user B, and user B's voice is heard only by user A). Note that the voices of the other conference participants can also be heard by users A and B, but while a specific operation (for example, an operation to select the area displaying the image of the user they are whispering to) is being performed, the volume of the voices of the other conference participants is controlled to be lowered. This makes it easier to hear the voice of the person they whispered to.
[0032] When the whispered conversation ends, the screen returns to the screen shown in Figure 12(1). The audio will also be transmitted to all conference participants.
[0033] As described above, by using the whisper function of the present invention, it becomes possible to talk only with desired users during a conference.
[0034] Next, the processing contents of the present invention will be described with reference to the flowcharts of FIGS.
[0035] FIG. 3 shows a process in which the CPU 201 of the client terminal 101 reads and executes a predetermined control program.
[0036] In step S301, it is determined whether a whisper response request has been received from another user. A whisper response request is a request to confirm whether another user can respond to a whisper (a function that delivers voice to the person who whispered among Web conference participants but not to other participants), and is a request issued by another user operating the client terminal 101 and received via the server device 102.
[0037] If a whisper response request has been received (step S301: YES), the process proceeds to step S302; if not, the process proceeds to step S303.
[0038] Fig. 9 shows an example of a screen displayed when user B receives a whisper response request from user A. As shown in Fig. 9, a small image of user A is displayed in the area for user B. Note that this display screen is just an example, and any form of display may be used as long as it is possible to identify which user sent the whisper response request to which user.
[0039] In step S302, a response process is performed in response to the whisper response request. Details of the response process will be described with reference to the flowchart in FIG.
[0040] In step S303, it is determined whether a whisper conversation request has been received. A whisper conversation request is a request to have a conversation using the whisper function, and is a request issued from a client terminal 101 used by another user and received via the server device 102.
[0041] In step S304, processing for whispered conversation (listener) is executed, the details of which will be explained using the flowchart in FIG.
[0042] In step S305, it is determined whether a predetermined selection operation (for example, an operation of clicking or tapping continuously for a predetermined period of time or more on an area where other users are displayed on the screen, or an operation of displaying an icon indicating a whisper function and accepting the selection of that icon) has been accepted. If the operation has been accepted (step S305; YES), the process proceeds to step S306, and if the operation has not been accepted (step S305: NO), the process of this flowchart ends.
[0043] In step S306, it is determined whether the user (client terminal used by the user) associated with the area determined in step S305 to have had a selection operation has turned on (enabled) the whisper function. Information on whether each user (client terminal used by the user) has turned on the whisper function is collected and managed by the server device 102 from each client terminal and distributed to each client terminal.
[0044] If the user to whom you are whispering has the whisper function turned on (step S306: YES), the process proceeds to step S307. If the user does not have the whisper function turned on (step S306: NO), the process returns to step S301.
[0045] In step S307, a whisper response request is sent to the whispered user.
[0046] In step S308, a process of waiting for a whispered response is performed. The details of the process of waiting for a whispered response will be explained using the flowchart of FIG.
[0047] Next, the response process shown in step S302 will be described in detail with reference to the flowchart of FIG.
[0048] First, the screen displayed when a whisper response request is received will be described.
[0049] When the client terminal 101 receives a whisper response request, it displays a screen, an example of which is shown in Fig. 9. Fig. 9 is an example of a screen that is displayed when user A sends a whisper response request to user B and user B receives the whisper response request. When this screen is displayed and user B's client terminal accepts a predetermined selection operation in the displayed area for user B (user B), a response selection interface, an example of which is shown in Fig. 11, is displayed. User B's client terminal accepts the selection of a response method from user B via this response selection interface, and responds to user A's request.
[0050] The response selection interface in Figure 11 displays four types of reception areas: "Voice," "Text," "Yes," and "No." A response to a whispered response request is selected by accepting the selection of one of the reception areas. The circle in the center of the reception area indicates the location where a click or tap operation was performed. The selection of a reception area can be accepted by performing a flick operation or the like from the location of the circle in the direction of the reception area, but the method of accepting the selection is not limited to a flick operation and any method can be used.
[0051] The response process will be described below with reference to FIG.
[0052] In step S401, it is determined whether a selection of "Sentence" has been received from the receiving unit shown in Fig. 11. If a selection has been received (step S401: YES), the process proceeds to step S402, and if a selection has not been received (step S402: NO), the process proceeds to step S403.
[0053] In step S402, a text message indicating a text response is sent to the sender of the whisper response request. Note that a text response is a response indicating that communication (e.g., chat) will be conducted with the user who sent the whisper response request by text rather than by voice. If a text response is selected, a chat screen (not shown) will be displayed for communication between the sender and destination users of the whisper request.
[0054] In step S403, it is determined whether the selection of "voice" has been accepted by the accepting unit. If the selection of "voice" has been accepted (step S403: YES), the process proceeds to step S404, and if it has not been accepted (step S403: NO), the process proceeds to step S405. Note that "voice" is a response indicating that communication will be conducted by voice.
[0055] In step S404, a whisper response is sent to the sender of the whisper response request, and this flowchart then ends.
[0056] In step S405, it is determined whether the selection of "o" has been accepted from the accepting section. If a selection has been accepted (step S405: YES), the process proceeds to step S406, and if not (step S405: NO), the process proceeds to step S407.
[0057] Note that "〇" indicates a response indicating that the person is willing to whisper.
[0058] In step S406, a text message indicating that the request will be accepted is sent to the sender of the whisper response request.
[0059] When "Sentence," "Voice," and "◯" are selected, a screen such as the one shown in FIG. 10 is displayed on the client terminal 101, indicating that the whisper function is being used in User AB's building.
[0060] In step S407, it is determined whether the selection of "x" has been received from the reception unit. If the selection of "x" has been received (step S407: YES), the process proceeds to step S408; if not (step S407: NO), the process returns to step S401.
[0061] Note that "x" is a response indicating that the person cannot respond to the whisper.
[0062] In step S408, a text message indicating that the request cannot be complied with is sent to the sender of the whisper response request.
[0063] In step S409, a simple response is sent to the sender of the whisper response request.
[0064] Through the above processing, a user who receives a whispered response request can inform the other user how they will / can respond.
[0065] Next, the process of whispered conversation (listener) will be described with reference to the flowchart of FIG.
[0066] Although details will be described later, when the flowchart of Fig. 5 is executed, the screen shown in Fig. 10 is displayed on the display unit of the client terminal 101. That is, the image of the user who is whispering (user A and user B in Fig. 10) is displayed (evenly divided) in the display area of user B, and information identifying user A (such as name) is displayed in the display area of user A.
[0067] In step S501, it is determined whether the user to whom you are whispering (speaker) has performed a predetermined operation such as a click operation or a tap operation on the displayed area.
[0068] If the predetermined operation has been performed (step S501: YES), the process proceeds to step S501, and if the predetermined operation has not been performed (step S501: NO), the process proceeds to step S503.
[0069] In step S502, the volume of the voices of users (conference participants) other than the user who is the whispering partner is lowered to a set value (set to a predetermined value). Then, the process returns to step S501. That is, while the predetermined operation is being performed, the volume of the voices of users (conference participants) other than the user who is the whispering partner is lowered to a set value.
[0070] In step S503, the voice volumes of users (conference participants) other than the user who is the whispering partner are set to default values (values set in each client terminal).
[0071] In step S504, it is determined whether an end quest has been received from the whispered user. If it has been received (step S504: YES), this flowchart ends. If it has not been received (step S504: NO), the process proceeds to step S505.
[0072] In step S505, it is determined whether the user to whom the whisper is to be sent has performed a predetermined operation (operation to end the whisper), such as a double-click operation or a double-tap operation, on the displayed area.
[0073] If a predetermined operation has been performed (step S505; YES), the process proceeds to step S506, and if not (step S505: NO), the process returns to step S501.
[0074] In step S506, a request to terminate the whisper function is sent to the user who is the whisper destination.
[0075] Next, the response waiting process in step S308 will be described with reference to the flowchart of FIG.
[0076] In step S601, the image of the user who has sent the whisper response request is reduced in size and displayed superimposed on the image of the user who is the whisper destination.
[0077] In step S602, the name of the user who sent the whisper response request (information for identifying the user, such as a name) is displayed in the area of the user.
[0078] As a result of the processing in steps S601 and S602, a screen example of which is shown in FIG. 9 is displayed.
[0079] In step S603, the voice of the user who sent the whisper response request is set to be audible only to the user to whom the whisper was made. Note that this process is a process of controlling the voice so that it cannot be heard by conference participants other than the user to whom the whisper was made, and is not limited to controlling the voice to be completely blocked, but also includes a process of reducing the volume to a level that is inaudible to the human ear.
[0080] In step S604, it is determined whether a whisper response has been received from the user who is the whisper destination. If a whisper response has been received (step S604; YES), the process proceeds to step S607, and if not (step S604: NO), the process proceeds to step S605.
[0081] In step S605, it is determined whether a simple reply has been received from the whispered user. If a simple reply has been received (step S605: YES), the process proceeds to step S609. If a simple reply has not been received (step S605: NO), the process proceeds to step S606.
[0082] In step S606, it is determined whether a predetermined operation (such as a double-click or double-tap) has been received in the whispered user's area. If the predetermined operation has been received (step S606: YES), the process proceeds to step S609. If the predetermined operation has not been received (step S606: NO), the process returns to step S604.
[0083] In step S607, a whisper conversation request is sent to the user who is the whisper destination.
[0084] In step S608, a whisper conversation (speaker) process is executed. Details of this process will be explained using the flowchart in FIG.
[0085] In step S609, the reduced display executed in step S601 is cancelled, and in step S610, the user's own image is displayed in the user's own area. Through this processing, the screen shown in Fig. 8 is displayed on the client terminal.
[0086] In step S611, control is performed so that the human voice can be heard by all the conference participants.
[0087] By the above process, when the whispering function is started, the voice of the speaking user can be heard only by the person being whispered to. Also, the conference screen displays a display that makes it possible to identify which users are whispering to each other.
[0088] Next, the whispered conversation (speaker) processing in step S608 will be described with reference to the flowchart of FIG.
[0089] In step S701, images of both the listener and speaker are displayed in the listener user's area as shown in Fig. 10. At this time, by using a display format such as equally dividing the listener user's area as shown in Fig. 10, it is possible to distinguish this state from the response waiting state shown in Fig. 9.
[0090] In step S702, the user sets the voice of the user who is whispering to him (listener user) so that only he can hear it.
[0091] In step S703, it is determined whether a selection by a predetermined operation (click operation, tap operation, etc.) on the area of the user to whom the whisper is to be made on the screen is accepted. If it is accepted (step S703: YES), the process proceeds to step S704, and if it is not accepted (step S703: NO), the process proceeds to step S705.
[0092] In step S704, the volume of the voices of users (conference participants) other than the user who is the whispering partner is lowered to a set value (set to a predetermined value). Then, the process returns to step S703. That is, while the predetermined operation is being performed, the volume of the voices of users (conference participants) other than the user who is the whispering partner is lowered to a set value.
[0093] In step S705, the voice volumes of users (conference participants) other than the user who is the whispering partner are set to default values (values set in each client terminal).
[0094] In step S706, it is determined whether a hunting request has been received from the whispered user. If a hunting request has been received (step S706: YES), the process proceeds to step S709. If a hunting request has not been received (step S706: NO), the process proceeds to step S707.
[0095] In step S707, it is determined whether a selection by a predetermined operation (such as a double-click operation or a double-tap operation) on the area of the user to whom the whispering is to be made on the screen has been accepted, that is, whether an operation to send an end request has been accepted.
[0096] If it has been accepted (step S707: YES), the process proceeds to step S708, and if it has not been accepted (step S707: NO), the process proceeds to step S703.
[0097] In step S708, an end request is sent to the whispered user.
[0098] In step S709, the screen display returns to the screen shown in FIG.
[0099] In step S710, the voice of the user who is the whispering partner is set to be heard by all the participants in the conference.
[0100] By using the whisper function described above, it becomes possible to talk only with desired users during a web conference.
[0101] The present invention can be embodied, for example, as a system, an apparatus, a method, a program, or a recording medium, etc. Specifically, the present invention may be applied to a system consisting of multiple devices, or may be applied to an apparatus consisting of a single device.
[0102] Furthermore, the program of the present invention is a program that enables a computer to execute the processing methods of the flowcharts shown in Figures 3 to 7, and the storage medium of the present invention stores a program that enables a computer to execute the processing methods of Figures 3 to 7. Note that the program of the present invention may be a program for each processing method of each device in Figures 3 to 7.
[0103] As described above, it goes without saying that the object of the present invention can also be achieved by supplying a recording medium on which a program that realizes the functions of the above-mentioned embodiments is recorded to a system or device, and having the computer (or CPU or MPU) of that system or device read and execute the program stored on the recording medium.
[0104] In this case, the program itself read from the recording medium will realize the novel functions of the present invention, and the recording medium on which the program is recorded will constitute the present invention.
[0105] Examples of recording media for supplying the program include flexible disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, DVD-ROMs, magnetic tapes, non-volatile memory cards, ROMs, EEPROMs, and silicon disks.
[0106] Furthermore, it goes without saying that not only are the functions of the above-mentioned embodiments realized by the computer executing a program it has read, but also cases are included in which an OS (operating system) running on the computer performs some or all of the actual processing based on the instructions of the program, and the functions of the above-mentioned embodiments are realized through that processing.
[0107] Furthermore, it goes without saying that this also includes cases where a program read from a recording medium is written into a memory provided on a function expansion board inserted into a computer or a function expansion unit connected to the computer, and then a CPU or the like provided on the function expansion board or function expansion unit performs some or all of the actual processing based on the instructions of the program code, thereby realizing the functions of the above-mentioned embodiments.
[0108] Furthermore, the present invention may be applied to a system consisting of multiple devices, or to a device consisting of a single device. It goes without saying that the present invention can also be applied to a case where the present invention is achieved by supplying a program to a system or device. In this case, the system or device can enjoy the effects of the present invention by reading a recording medium containing a program for achieving the present invention into the system or device.
[0109] Furthermore, by downloading and reading a program for achieving the present invention from a server, database, etc. on a network using a communication program, the system or device can enjoy the effects of the present invention. Note that the present invention also includes configurations that combine the above-mentioned embodiments and their modified examples. [Explanation of symbols]
[0110] 101 client terminals 102 Server device 110 Network
Claims
1. a request receiving means for receiving a request for whisper conversation from a first user participating in a Web conference with a second user; a selection receiving means for receiving a selection of a method of responding to the whispered conversation from the second user who is the request recipient of the request for the whispered conversation received by the request receiving means; a control means for controlling distribution of the voices of the first user and the second user in accordance with the response method selected by the selection receiving means; An information processing system comprising:
2. 2. The information processing system according to claim 1, wherein the selection receiving means receives a selection from among response methods including a voice response and a text response.
3. 3. The information processing system according to claim 2, wherein the control means controls distribution of the voices of the first user and the second user when a voice response is received by the selection receiving means.
4. 4. The information processing system according to claim 1, wherein the control means controls so that the voices of the first user and the second user are not distributed to conference participants other than the first user and the second user.
5. The information processing system according to any one of claims 1 to 4, characterized in that the control means, when receiving a predetermined operation from the first user and the second user, controls the distribution of voices of conference participants other than the first user and the second user to the first user and the second user.
6. The information processing system according to any one of claims 1 to 5, characterized in that, when the control means receives a predetermined operation from the first user and the second user, it lowers the volume of the voices of conference participants other than the first user and the second user and distributes the voices to the first user and the second user.
7. a display control means for controlling the display of images of the Web conference participants on the information processing devices used by each of the participants; 7. The information processing system according to claim 1, wherein the display control means controls the display so that the user who is having the whispered conversation is identifiable.
8. The information processing system according to claim 7, characterized in that the display control means controls to display a chat screen with the first user and the second user when the selection receiving means receives a selection of a text response.
9. a request receiving step in which a request receiving means of the information processing system receives a request for a whisper conversation from a first user participating in the Web conference with a second user; a selection receiving step in which a selection receiving means of the information processing system receives, from the second user who is the request recipient of the request for whispered conversation received in the request receiving step, a selection of a method of responding to the whispered conversation; a control step in which a control means of the information processing system controls distribution of the voices of the first user and the second user in accordance with the response method selected in the selection receiving step; An information processing method comprising:
10. A program for causing a computer to function as each of the means according to claims 1 to 8.
Citation Information
Patent Citations
Video conference system
JP2015041885A
Server device and video conferencing system
JP2021064833A
Online conference system
JP2021184189A
Web conference system, information processing method, and program
JP2017111643A