Screen compositing method using a web conferencing system
The screen composition method for web conferences enhances participant visibility and engagement by processing and displaying key participants, addressing the limitations of existing systems to extract and highlight relevant interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- INTERACTIVE SOLUTIONS CORP
- Filing Date
- 2023-07-26
- Publication Date
- 2026-05-22
AI Technical Summary
Existing web conference systems lack the ability to effectively extract and display specific participants during a conference, limiting the clarity and engagement of the communication process.
A screen composition method for web conferences that includes document display, participant image reception, selection, and display steps, along with speaker identification and image processing to highlight relevant participants, using a predetermined display pattern and storing the display images for later retrieval.
Enables clear extraction and display of specific participants, enhancing communication clarity and engagement by focusing on key interactions, while allowing for image and information correction and storage.
Smart Images

Figure 0007863887000001 
Figure 0007863887000002 
Figure 0007863887000003
Abstract
Description
Technical Field
[0001] This invention relates to a screen composition method using a web conference system and the like.
Background Art
[0002] Japanese Patent No. 7062126 describes a web conference system using an avatar.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of this invention is to provide a screen composition method using a web conference system that can extract and display specific participants in a web conference system.
Means for Solving the Problems
[0005] The first invention relates to a screen composition method using a web conference system 1. This method includes a document display step (S110), a participant image reception step (S120), a participant selection step (S130), and a selected participant display step (S140). The document display step (S110) is a step for the web conference system 1 to display the document used in the web conference on the display unit 3. The participant image reception step (S120) is a step for the web conference system 1 to receive images of each of a plurality of participants participating in the web conference, where each participant is photographed. The participant selection step (S130) is a step for the web conference system 1 to receive that two or more participants among the participants received in the participant image reception step are selected. The selected participant display step (S140) is a step in which the web conferencing system 1 extracts the image region of each participant from images taken of the participants for two or more participants selected in the participant selection step, and displays the image based on the extracted image region of each participant, along with the materials, on the display unit 3. The selected participant display step displays an image based on the participant's image area on the display unit 3, based on a predetermined display pattern.
[0006] A preferred example of this screen synthesis method further includes a speaker identification step (S131). The speaker identification step (S131) is a step in which the web conferencing system 1 identifies the person who made the statement among the participants in the web conference. Then, the participant selection step (S130) uses the information about the identified speaker to select two or more participants from among the participants.
[0007] A preferred example of this screen composition method is that when the web conferencing system 1 determines that a conversation has taken place between the first participant and the second participant, the selected participant display step (S130) displays an image based on the image area of the first participant and an image based on the image area of the second participant on the display unit 3.
[0008] A preferred example of this screen synthesis method is that when the web conferencing system 1 determines that a questioner has asked a question regarding the presenter's explanation, the selected participant display step (S130) displays an image based on the presenter's video area and an image based on the questioner's video area on the display unit 3.
[0009] A preferred example of this screen synthesis method further includes a display image storage step (S150). The display image storage step (S150) is a step in which the web conferencing system 1 stores a display image that includes the materials and images based on the image areas of the participants that were displayed on the display unit 3 in the selected participant display step (S140).
[0010] A preferred example of this screen synthesis method further includes a recording information correction step (S160). The recorded information correction step (S160) is a step in which, if the data has been corrected, the image related to the data from the display images stored in the display image storage step (S150) is corrected and stored.
[0011] An example of a screen composition method is that the web conferencing system 1 includes a code information display step (S210). The code information display step (S210) is a step of displaying code information on the display unit 3. In the code information display step (S210), for example, the code information corresponding to each page of the document is displayed on the display unit 3. A preferred example of this step is that as the web conference progresses, the web conferencing system 1 updates the code information to obtain updated code information, and the web conferencing system 1 displays the updated code information on the display unit 3.
[0012] The following invention relates to a program and a non-temporary information recording medium that stores the program. This program is a program that causes a computer to execute one of the screen composition methods described above. [Effects of the Invention]
[0013] This invention provides a screen composition method using a web conferencing system that allows for the extraction and display of specific participants within the web conferencing system. [Brief explanation of the drawing]
[0014] [Figure 1] Figure 1 is a flowchart illustrating the screen composition method. [Figure 2] Figure 2 is a schematic diagram illustrating a web conferencing system. [Figure 3] Figure 3 is a block diagram of a web conferencing system. [Figure 4] Figure 4 is a conceptual diagram showing an example where documents and participant images are displayed on the administrator screen. [Figure 5] Figure 5 shows how the administrator selects participants. [Figure 6] Figure 6 shows an example of an image displayed based on the participant's image region. [Figure 7] FIG. 7 is a different view showing an example in which an image based on the image area of the participant is displayed. **Embodiments for Carrying Out the Invention**
[0015] Hereinafter, embodiments for carrying out the present invention will be described with reference to the drawings. The present invention is not limited to the embodiments described below, and also includes those appropriately modified by those skilled in the art within an obvious range from the following embodiments.
[0016] FIG. 1 is a flowchart for explaining a screen composition method. As shown in FIG. 1, this method includes a material display step (S110), a participant image reception step (S120), a participant selection step (S130), and a selected participant display step (S140). Also, as shown in FIG. 1, this method may further include any one or more of a speaker identification step (S131), a display image storage step (S150), and a recording information correction step (S160). Note that the example in FIG. 1 is an example of steps, and the order may be changed as appropriate, or each step may be performed simultaneously. For example, the participant image reception step (S120) may be performed after the material display step (S110), or the material display step (S110) may be performed after the participant image reception step (S120). Also, as shown in FIG. 1, this method may include a code information display step (S210). Note that not limited to this method, the code information display step (S210) may be used to store and read out the display image displayed on the display unit. This method is executed by a computer or a processor.
[0017] Computers and processors have an input unit, an output unit, a control unit, an arithmetic unit, and a storage unit, and each element is connected by a bus or the like so that information can be exchanged. For example, a program may be stored in the storage unit, or various types of information may be stored. When predetermined information is input from the input unit, the control unit reads out the program stored in the storage unit. Then, the control unit appropriately reads out the information stored in the storage unit and transmits it to the arithmetic unit. Also, the control unit appropriately transmits the input information to the arithmetic unit. The arithmetic unit performs arithmetic processing using the various types of received information and stores it in the storage unit. The control unit reads out the arithmetic result stored in the storage unit and outputs it from the output unit. In this way, various processes and steps are executed. Each unit and each means execute these various processes. A computer may have a processor, and the processor may realize various functions and various steps.
[0018] Figure 2 is a schematic diagram showing a web conference system. In the example shown in Figure 2, a server 25 and a plurality of terminals (clients) 27 are connected via a network (intranet or Internet). The web conference system 1 is a system for connecting a plurality of terminals and holding a web conference. A program for the web conference system may be installed on each terminal, or a program for the web conference system may be installed on the server.
[0019] Figure 3 is a block diagram showing a web conference system. As shown in Figure 3, the web conference system 1 includes a document display unit 5, a participant image reception unit 7, a participant selection unit 9, and a selected participant display unit 11. Also, as shown in Figure 3, this system 1 may further include any one or more of a speaker identification unit 13, a display image storage unit 15, and a recorded information correction unit 17. Also, as shown in Figure 3, this method may have a code information display unit 21. The document display unit 5 is an element for displaying documents used in a web conference on the display unit 3. The participant image receiving unit 7 is an element that receives images taken of each participant from multiple participants participating in a web conference. The participant selection unit 9 is an element that receives that two or more participants have been selected from the participants received by the participant image receiving unit. The selected participant display unit 11 is an element in which the web conferencing system 1 extracts the image area of each participant from the images taken of the participants for two or more participants selected by the participant selection unit, and displays an image based on the extracted image area of the participants along with the documents on the display unit 3. The selected participant display unit 11 displays an image based on the image area of the participants on the display unit 3 based on a predetermined display pattern. Each unit may be read as each means, and each unit performs the corresponding process.
[0020] The document display step (S110) is a step in which the web conferencing system 1 displays the documents to be used for the web conference on the display unit 3. The display unit 3 may be the display unit of the server (monitor, etc.) or the display unit of each terminal (monitor, etc.). The server's display unit and each terminal's display unit may display images based on the same information, or they may display images based on different information. A web conferencing system is a system that allows multiple terminals to meet simultaneously via a network (Internet or intranet). Web conferencing systems themselves are publicly known. Examples of web conferencing systems include ZOOM®, Teams®, Meets®, WebEX®, Skype®, and LINE® Meeting. For example, suppose several people are participating in a web conference. Then, a speaker (presenter) wants to share a document. For example, the speaker's client computer receives a command to read the document and reads the document from its storage unit. The client may be a personal computer (PC) or a mobile terminal such as a smartphone. The client then outputs the retrieved document to the system along with a sharing instruction. The system, having received the sharing instruction and the document, displays the document on its display unit. The system also outputs information to the display units of the client of each participant in the web conference, providing them with the necessary information to display the document. Examples of documents include presentation materials and various documents to be shared with participants. The code information display step (S210), described later, may be performed simultaneously with or after the document display step (S110).
[0021] The participant image receiving step (S120) is the step in which the web conferencing system 1 receives images taken of each participant from multiple participants participating in the web conference. It is not necessary for this step to receive images of all participants from multiple participants in the web conference. That is, some participants may keep their cameras off during the web conference, so images of such participants are not required. For example, each participant's client camera 23 may take a picture of each participant and send the captured image to the system 1. In this way, the web conferencing system 1 can receive images taken of each participant from multiple participants participating in the web conference. In a web conference, when a camera is turned on, the camera takes a picture of the participant and sends it to the web conferencing system. The web conferencing system then shares the video (including a series of images) of those whose cameras are on. Therefore, a typical web conferencing system can receive images taken of each participant from multiple participants participating in the web conference. Alternatively, the system 1 may store the participants' images in advance in its storage unit, and when the system 1 receives information about the participants, it may read the participants' images from the storage unit.
[0022] Figure 4 is a conceptual diagram showing an example where documents and participant images are displayed on the administrator screen. In this example, the administrator screen 31 displays an image 33 related to the documents, as well as a photograph 35 of the participant. In this example, code information 37 is also displayed on the administrator screen 31. This code information 37 may also be displayed on the display units of participants other than the administrator. System 1 stores, for example, the displayed documents and participant information in relation to this code information 37. Then, by using this code information, participants can restore the documents, participant information, or the displayed screen.
[0023] The participant selection step (S130) is the step in which the web conferencing system 1 receives notification that two or more participants have been selected from among the participants whose images were received in the participant image reception step. For example, a speaker or administrator may specify two or more participants from among those whose images are displayed on the display unit. The system may then receive notification that one or more participants have been selected from among the participants whose images were received in the participant image reception step. This step may be performed automatically by the system or based on input from a server or terminal. An example of this being performed automatically by the server is to randomly select two or more participants from among the participants, as will be described later. Note that in one example, two or more participants were selected. However, it may also be the case that only one participant is selected.
[0024] Figure 5 illustrates how an administrator selects participants. In this example, the administrator's terminal display shows the participant management screen, where multiple participants are displayed. When the administrator selects a participant's image with their finger, the system processes the input and registers the participant as selected. The system can also select participants using a mouse or voice. For example, the memory unit stores participant names in association with participant information (e.g., participant IDs). The system analyzes the input voice, identifies names, and stores them in the memory unit as appropriate. The system then reads the stored names (names included in the input voice) and the stored participant names and performs a comparison. As a result, the system finds participant names that match the stored names. The system then uses the information of the matched participant name to read the information about that participant stored in the memory unit and select the participant. In this way, for example, a moderator can simply call out a participant's name, and the called participant will be selected, and an image associated with that participant will be displayed on the screen.
[0025] The participant selection step (S130) may include a speaker identification step (S131). The speaker identification step (S131) is a step in which the web conferencing system 1 identifies the person who made a statement among the participants in the web conference. For example, suppose two or more participants make statements consecutively. In this case, the audio is input to the participants' clients. The audio input to the clients is then transmitted to the system along with client information (participant information). The system then receives the audio input to the clients and the client information (participant information). The system stores the received audio input to the clients and the client information (participant information) in its memory unit. In this way, the system can identify the person who made a statement among the participants in the web conference. The participant selection step (S130) then selects participants using the information about the identified speaker. The system reads the client information (participant information) stored in the memory unit and uses the read client information (participant information) to select multiple participants. In this case, for example, two or more participants who made statements consecutively may be selected.
[0026] A preferred example of this screen synthesis method is that when the web conferencing system 1 determines that a conversation has taken place between a first participant and a second participant, the selected participant display step (S130) may display an image based on the image area of the first participant and an image based on the image area of the second participant on the display unit 3. The system 1 analyzes the terminals from which audio is input and determines that a conversation has taken place between a first participant and a second participant when audio from a predetermined number of people is input within a certain period of time. In this way, the system may display an image based on the image area of the first participant and an image based on the image area of the second participant on the display unit 3.
[0027] A preferred example of this screen synthesis method is that when the web conferencing system 1 determines that a questioner has asked a question regarding the presenter's explanation, the web conferencing system 1 displays an image based on the presenter's video area and an image based on the questioner's video area on the display unit 3 during the selected participant display step (S130).
[0028] The Selected Participant Display Step (S140) is a step in which the web conferencing system 1 extracts the image regions of the participants from images taken of the participants for two or more participants selected in the participant selection step, and displays the images based on the extracted image regions of the participants, along with the materials, on the display unit 3. In this case, the materials may be materials displayed on the display unit (a page of a certain document or a part of a certain document). The Selected Participant Display Step displays the images based on the image regions of the participants on the display unit 3 based on a predetermined display pattern. An example of a predetermined display pattern is to display the captured image as is. Another example of a predetermined display pattern is to display only the images based on the image regions of the participants on the display unit. In a normal web conference, the captured image of the participant is displayed as is (or an image with the background adjusted) on the display unit. In this example, only the images based on the image regions of the participants are displayed on the display unit. This makes it possible to display the conference as if the discussion is being conducted by the selected participants.
[0029] An example of an image based on the participant's image region may be an image obtained by removing the background from a photograph taken by the participant's device. An example of an image based on the participant's image region may be an image obtained by extracting the participant's face from a photograph taken by the participant's device. Alternatively, the participant's image region may be extracted from a photograph taken by the participant's device, and then a predetermined processing may be applied to it. Examples of predetermined processing include making the participant's face region larger than the body region, controlling the opening and closing of the participant's mouth at a predetermined frequency or in accordance with speech, or opening and closing the participant's eyes at a predetermined frequency. Such an image can be obtained by storing a program for image processing in the memory unit, reading the program by command from the control unit, and having the calculation unit perform a predetermined calculation.
[0030] Figure 6 shows an example of how images based on the image regions of the participants are displayed. In Figure 6, the top and middle participants displayed on the administrator screen are selected, and through image processing, images based on the image regions of the participants (images with the background removed) 39 are extracted and displayed on the left and right sides of the screen. In this way, it is possible to display the image 33 related to the document as if the two participants were discussing or debating it.
[0031] Another aspect of the selected participant display step (S140) is that the web conferencing system 1 displays images related to the selected participants on the display unit 3 for one or more participants selected in the participant selection step. For example, the system 1 stores avatars associated with the participants in a memory unit. Then, based on the information about the one or more participants selected in the participant selection step, it reads the avatars corresponding to the participants from the memory unit. The system 1 may then manipulate the read avatars as appropriate to display them as if the participants' avatars were speaking or conversing.
[0032] Figure 7 shows an example of an image being displayed based on the image area of a participant. In this example, the hair color of the middle participant, which was displayed on the administrator screen, has been changed, and an exaggerated image 41 with an enlarged face is displayed. In this example, the avatar information of the top participant, which was displayed on the administrator screen, has been read from the memory unit, and an avatar 43 corresponding to that participant is displayed on the screen. In this way, this system may display participants in an exaggerated manner or as avatars. By employing such display techniques, it is possible to attract the viewer's attention.
[0033] A preferred example of this screen synthesis method further includes a display image storage step (S150). The display image storage step (S150) is a step in which the web conferencing system 1 stores a display image that includes the materials and images based on the image areas of the participants that were displayed on the display unit 3 in the selected participant display step (S140). Because there is a display image storage step, a participant or a third party can obtain the image that was displayed on the display unit at a later date. At this time, the audio that was present when the predetermined image was displayed may also be stored in the storage unit. A participant or a third party can reproduce the display image and audio that were displayed on the display unit. At this time, code information, which will be described later, may be displayed on the display unit, and the display screen may be reproduced based on that code information.
[0034] A preferred example of this screen synthesis method further includes a recording information correction step (S160). The recorded information correction step (S160) is a step in which, if the material has been corrected, the display image stored in the display image storage step (S150) that relates to the material is corrected and stored. The storage unit of System 1 stores information about the pages and parts of the material that were displayed on the display unit in relation to the display image. When System 1 receives information that the pages or parts of the displayed material have been corrected, it replaces the information about the pages or parts of the material in the display image stored in the storage unit with the corrected information and stores it in the storage unit as the corrected display image. Then, if a participant or a third party tries to obtain the display image at a later date, they will be able to obtain the corrected image. For example, if something incorrect was said in a lecture, or if the information has changed due to a legal revision, or if the system has changed, it will be possible to provide a display image with the latest information. Furthermore, if the audio data stored in the storage unit is also corrected at the same time, the speech itself will also be updated and provided to the participant or a third party. Furthermore, if the user wishes to remove any of the participants displayed on the screen at a later date, they can use information about other participants (for example, the identification information of another participant) to retrieve the image or avatar of that other participant and replace the image of the participant they wish to remove. This allows the user to continue providing available video information without compromising the content itself.
[0035] An example of a screen composition method is a web conferencing system 1 that includes a code information display step (S210). The code information display step (S210) is a step of displaying code information on the display unit 3. In the code information display step (S210), for example, for each page of the document, the code information corresponding to the page is displayed on the display unit 3. A preferred example of this step is that as the web conference progresses, the web conferencing system 1 updates the code information to obtain the updated code information, and the web conferencing system 1 displays the updated code information on the display unit 3. This code information is stored in the storage unit of the system 1 in relation to the document displayed on the display unit. Therefore, participants or third parties can obtain the document by referring to this code information. Also, as explained earlier, in addition to the page or part of the document, the storage unit also stores audio data from when that page or part was displayed in relation to the code information. In this example, by referring to the code information, not only the page or part of the document but also the audio from when it was displayed can be obtained.
[0036] The following invention relates to a non-temporary information recording medium that can be read by a program or a computer that stores a program. This program is a program that causes a computer to execute one of the screen composition methods described above. [Industrial applicability]
[0037] This method can be used in web conferencing systems, etc. [Explanation of Symbols]
[0038] 1. Web conferencing system 5 Material display area 7. Participant Image Receipt Department 9. Participant Selection Section 11. Selected Participant Display Section 13. Speaker Identification Unit 15 Display Image Storage Unit 17. Record Information Correction Section 21 Code Information Display Unit
Claims
1. A method for combining screens using a web conferencing system, The aforementioned web conferencing system includes a document display step in which the documents to be used for the web conference are displayed on a display unit, The web conferencing system includes a participant image receiving step, in which it receives images taken of each participant from multiple participants participating in the web conferencing, The web conferencing system includes a participant selection step in which it receives that two or more participants have been selected from among the participants who received the images in the participant image receiving step, The web conferencing system includes a selected participant display step which involves extracting the image region of two or more participants selected in the participant selection step from images taken of the participants, and displaying an image based on the extracted image region of the participants together with the materials on the display unit, The image based on the participant's image region is an image obtained by extracting the participant's face portion from an image of the participant taken. The aforementioned selected participant display step is a screen composition method using a web conferencing system, which displays an image based on the image area of the participant on the display unit based on a predetermined display pattern, The participants in the aforementioned web conference include presenters and questioners, The web conferencing system receives information about the audio input to the questioner's client and client information about the questioner's client, the web conferencing system identifies the questioner based on the client information about the questioner's client, and the web conferencing system determines, using the information about the audio input to the questioner's client, that the questioner asked a question about the presenter's explanation, and the selected participant display step displays an image based on the presenter's image area and an image based on the questioner's image area on the display unit, The web conferencing system further includes a code information display step in which code information is displayed on the display unit, The aforementioned document includes multiple pages, and the code information display step is a method of displaying code information corresponding to each page on the display unit.
2. A screen composition method using the web conferencing system described in claim 1, The web conferencing system further includes a speaker identification step that identifies the person who made the statement among the participants in the web conferencing, The participant selection step further includes the step of selecting two or more participants from among the participants using information about the identified person who made the statement.
3. A screen composition method using the web conferencing system described in claim 1, A method comprising a web conferencing system and a display image storage step of storing a display image that includes the materials displayed on the display unit in the selected participant display step and an image based on the image area of the participant.
4. A screen composition method using the web conferencing system described in claim 3, A method further comprising a recording information modification step, in which, if the aforementioned material is modified, the image related to the aforementioned material among the display images stored in the display image storage step is modified and stored.
5. A screen composition method using the web conferencing system described in claim 1, A method comprising the following steps: as the web conference progresses, the web conferencing system updates the code information to obtain updated code information, and the web conferencing system displays the updated code information on the display unit.
6. On the computer, A method for combining screens using a web conferencing system, The aforementioned web conferencing system includes a document display step in which the documents to be used for the web conference are displayed on a display unit, The web conferencing system includes a participant image receiving step, in which it receives images taken of each participant from multiple participants participating in the web conferencing, The web conferencing system includes a participant selection step in which it receives that two or more participants have been selected from among the participants whose images were received in the participant image receiving step, The web conferencing system includes a selected participant display step which involves extracting the image region of two or more participants selected in the participant selection step from images taken of the participants, and displaying an image based on the extracted image region of the participants together with the materials on the display unit, The image based on the participant's image region is an image obtained by extracting the participant's face portion from an image of the participant taken. The aforementioned selected participant display step is a screen composition method using a web conferencing system, which displays an image based on the image area of the participant on the display unit based on a predetermined display pattern, The participants in the aforementioned web conference include presenters and questioners, The web conferencing system receives information regarding the audio input from the questioner's client and client information regarding the questioner's client, and the web conferencing system identifies the questioner based on the client information regarding the questioner's client. If the web conferencing system determines, using audio information input to the questioner's client, that the questioner has asked a question about the presenter's explanation, the selected participant display step is a method of displaying an image based on the presenter's image area and an image based on the questioner's image area on the display unit, The web conferencing system further includes a code information display step in which code information is displayed on the display unit, The aforementioned document includes multiple pages, and the code information display step is a method of displaying code information corresponding to each page on the display unit for each page displayed on the display unit. A program to execute.
7. A non-temporary information recording medium that can be read by a computer storing the program described in Claim 6.