Information processing device, information processing method and program

The information processing device facilitates understanding web conference content by storing and modifying display modes based on participant actions, addressing the challenge of multiple windows obscuring content during playback.

JP2025146189APending Publication Date: 2025-10-03JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024046831
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

When recording a web conference, the conference state is captured on a single screen, resulting in multiple windows showing shared materials and participants, making it difficult for users to understand the content during playback.

Method used

An information processing device that stores the web conference state as multiple different types of content, allows selective display on a user's terminal device, and modifies the display mode based on detected participant actions using tags.

Benefits of technology

Enables proper understanding of web conference content from recorded data by allowing selective display and mode changes based on participant actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025146189000001_ABST
    Figure 2025146189000001_ABST
Patent Text Reader

Abstract

To appropriately grasp contents of a web conference from recording data of the web conference in an appropriate manner.SOLUTION: An information processing device includes a storage control part for storing a state of a web conference as a plurality of different kinds of contents in a storage part, a reproduction part for displaying at least a portion of the plurality of contents stored in the storage part in a display part of a terminal device of a user relating to the web conference, a change part for changing display modes of the contents reproduced by the reproduction part, and a tag creation part for creating a tag obtained by associating prescribed action with timing at which the prescribed action is detected and attaching it to the contents in the case of detecting the prescribed action of a participant on the basis of a state of the web conference of the participant of the web conference. The change part changes display modes of the contents on the basis of the tag.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] There are known technologies for recording the proceedings of web conferences. For example, Patent Document 1 discloses a technology that records the circumstances under which the minutes of a conference are discussed, and allows the recorded information to be searched and reproduced. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2002-101398 Summary of the Invention [Problem to be solved by the invention]

[0004] Typically, when recording a web conference, the conference state is recorded on a single screen and a video file is generated. In this case, when the video file is played back, multiple windows showing the shared materials and participants of the web conference are displayed, making it difficult for users to see the content of the web conference.

[0005] The present disclosure aims to provide an information processing device, an information processing method, and a program that enable the content of a web conference to be appropriately understood from recorded data of the web conference. [Means for solving the problem]

[0006] The information processing device of the present disclosure includes a memory control unit that stores the state of a web conference as multiple different types of content in a memory unit; a playback unit that displays at least some of the multiple contents stored in the memory unit on a display unit of a terminal device of a user related to the web conference; a modification unit that changes the display mode of the content being played back by the playback unit; and a tag creation unit that, when a predetermined action of a participant of the web conference is detected based on the state of the participant during the web conference, creates a tag that associates the predetermined action with the timing at which the predetermined action was detected and assigns the tag to the content, and the modification unit changes the display mode of the content based on the tag.

[0007] The information processing method disclosed herein includes a step in which a computer stores the state of a web conference as multiple different types of content in a memory unit, a step in which at least some of the multiple contents stored in the memory unit are displayed on a display unit of a terminal device of a user associated with the web conference, a step in which the computer changes the display mode of the content being played, a step in which, if a predetermined action of a participant of the web conference is detected based on the state of the participant during the web conference, the computer creates a tag that associates the predetermined action with the timing at which the predetermined action was detected and assigns it to the content, and a step in which the computer changes the display mode of the content based on the tag.

[0008] The program disclosed herein causes a computer to perform the following steps: storing the state of a web conference as multiple different types of content in a memory unit; displaying at least some of the multiple contents stored in the memory unit on a display unit of a terminal device of a user associated with the web conference; changing the display mode of the content being played; when a predetermined action of a participant in the web conference is detected based on the state of the participant during the web conference, creating a tag that associates the predetermined action with the timing at which the predetermined action was detected and assigning it to the content; and changing the display mode of the content based on the tag. [Effects of the Invention]

[0009] According to the present disclosure, the content of a web conference can be properly understood from the recorded data of the web conference. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a conference system according to the first embodiment. [Figure 2] FIG. 2 is a block diagram showing an example of the configuration of a terminal device according to the first embodiment. [Figure 3] FIG. 3 is a diagram illustrating content according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the configuration of the information processing device according to the first embodiment. [Figure 5] FIG. 5 is a flowchart showing the flow of the recording process according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing an example of a content setting screen during recording according to the first embodiment. [Figure 7] FIG. 7 is a diagram for explaining a content storage method according to the first embodiment. [Figure 8] FIG. 8 is a diagram for explaining a method for recording a web conference according to the first embodiment. [Figure 9] FIG. 9 is a flowchart showing the flow of the playback process according to the first embodiment. [Figure 10] FIG. 10 is a diagram showing an example of a content setting screen during playback according to the first embodiment. [Figure 11] FIG. 11 is a diagram for explaining a content playback method according to the first embodiment. [Figure 12] FIG. 12 is a diagram for explaining a method for changing the display mode of the display screen for playing back recorded data according to the first embodiment. [Figure 13] FIG. 13 is a diagram for explaining a method for changing the display mode according to the first embodiment. [Figure 14]FIG. 14 is a flowchart showing the flow of the recording process according to the second embodiment. [Figure 15] FIG. 15 is a flowchart showing the flow of the reproduction process according to the first example of the second embodiment. [Figure 16] FIG. 16 is a diagram showing a method for displaying tags according to the second embodiment. [Figure 17] FIG. 17 is a flowchart showing the flow of a reproduction process according to a modified example of the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that the present disclosure is not limited to these embodiments, and in the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0012] [First embodiment] (Conference system) The conference system according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the configuration of the conference system according to the first embodiment.

[0013] The conference system 1 includes multiple terminal devices 10 and an information processing device 100. The terminal devices 10 and the information processing devices 100 are communicatively connected via a wired or wireless network N. The conference system 1 may include any number of terminal devices 10 and information processing devices 100. The conference system 1 is a web conference system in which multiple users can hold a web conference using each of the terminal devices 10. The conference system 1 is also a web conference system in which multiple users can play and display the contents of the web conference using each of the terminal devices 10. Here, the multiple users who view the contents of the web conference using the terminal devices 10 include not only users who participated in the web conference but also users who were absent from the web conference, i.e., users related to the web conference. The conference system 1 records the web conference during the web conference. The conference system 1 allows users to appropriately understand the contents of the web conference from the recorded data of the web conference.

[0014] (Terminal Device) An example of the configuration of a terminal device according to the first embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the configuration of a terminal device according to the first embodiment.

[0015] The terminal device 10 is a device used by a user to participate in a web conference. The terminal device 10 is, for example, a notebook PC (Personal Computer), a desktop PC, a tablet terminal, a smartphone, an HMD (Head Mounted Display), an HUD (Head Up Display), etc., but is not limited to these.

[0016] As shown in FIG. 2, the terminal device 10 includes a camera 12, an operation unit 14, an audio input unit 16, a display unit 18, an audio output unit 20, a communication unit 22, a storage unit 24, and a control unit 26.

[0017] The camera 12 is a camera that captures an image of a user using the terminal device 10. The camera 12 captures, for example, the face, body, facial expressions, and movements of the user using the terminal device 10.

[0018] The operation unit 14 accepts various operations for the terminal device 10. The operation unit 14 is realized by various input devices such as a keyboard, a mouse, a switch, a button, and a touch panel.

[0019] The voice input unit 16 detects the voice of the user using the terminal device 10. The voice input unit 16 converts the detected voice into a voice signal. The voice input unit 16 is realized by a microphone.

[0020] The display unit 18 has a display surface that displays various types of information. The display unit 18 is realized by a display including, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display. The display unit 18 displays, for example, a display screen that shows information about a web conference. If the operation unit 14 is a touch panel, the operation unit 14 and the display unit 18 are configured as an integrated unit. Here, if the terminal device 10 is a device that displays image information by projecting or projecting it, such as a HUD, the display unit 18 does not need to have a display surface.

[0021] The audio output unit 20 is a speaker that outputs various types of audio. For example, the audio output unit 20 outputs the user's own voice and the voices of the users of the other terminal devices 10 participating in the web conference.

[0022] The communication unit 22 is a communication interface that performs communication between the terminal device 10 and an external device. The communication unit 22 executes communication between, for example, the terminal device 10 and the information processing device 100. The communication unit 22 is realized by, for example, a wireless LAN (Local Area Network), a wired LAN, Wi-Fi (registered trademark), or the like.

[0023] The storage unit 24 stores various types of information. The storage unit 24 stores information such as the contents of calculations performed by the control unit 26 and programs. The storage unit 24 includes at least one of a RAM (Random Access Memory), a main storage device such as a ROM (Read Only Memory), and an external storage device such as an HDD (Hard Disk Drive).

[0024] The control unit 26 controls each unit of the terminal device 10. The control unit 26 has, for example, an information processing device such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and a storage device such as a RAM or a ROM. The control unit 26 executes a program that controls the operation of the terminal device 10 according to the present invention. The control unit 26 may be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 26 may be realized by a combination of hardware and software.

[0025] The control unit 26, for example, acquires an image or video of the user captured by the camera 12. The control unit 26 transmits the image or video acquired from the camera 12 to the information processing device 100, for example, via the communication unit 22. The control unit 26, for example, performs processing to display information related to the web conference received from the information processing device 100 via the communication unit 22 on the display unit 18.

[0026] (Information processing device) The information processing device 100 generates and plays, for example, web conference information. The web conference information is information for reproducing the contents of the web conference after the web conference has ended for users involved in the web conference. The information processing device 100 is realized, for example, by a server device. The information processing device 100 may be realized, for example, by a terminal device 10 of a user participating in the web conference. The information processing device 100 may be realized, for example, by both the server device and the terminal device 10.

[0027] When recording a web conference, the information processing device 100 records each window displaying the shared screen and participants displayed on the display screen during the web conference. For example, the information processing device 100 generates as content the same number of video files and audio files as the number of windows displayed on the display unit 18 during the web conference. However, the number of content files generated by the information processing device 100 is not limited to this. For example, if it is difficult to display all participants at once on the display unit 18, such as when 100 participants are participating in a web conference, the information processing device 100 may generate as content more video files and audio files than the number of windows displayed on the display unit 18.

[0028] When playing back a recorded file of a web conference, the information processing device 100 allows the user to select the desired content and displays the selected content on a single screen. The user can zoom in, zoom out, and mute the content being played back, just like during an actual web conference. In other words, when playing back content, the user can select and delete information that they deem unnecessary.

[0029] The content according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram for explaining the content according to the first embodiment.

[0030] As shown in FIG. 3, table TB1 includes items such as "Timing" and "Content."

[0031] "Timing" indicates the timing of adding content. Recording is the timing when a user selects content to record when recording a web conference. Playback is the timing when a user selects content to play when playing back a recorded web conference.

[0032] "Content" refers to the type of content that can be added during recording or playback. Examples of content that can be selected during recording include, but are not limited to, a participant's shared video, a participant's camera video, and a designated screen video of the participant who is not sharing their screen. For example, if a specific participant's camera video and shared video are selected as the content to be recorded, the video of the window displaying the selected specific participant and the video of the window displaying the shared image are recorded as separate files.

[0033] Examples of content that can be selected during playback include, but are not limited to, specified images or videos, and website content (e.g., website URLs (Uniform Resource Locators)). For example, if a camera image of a specific participant and a URL of a video provided by a video distribution service provider are selected as content to be played, the video of the selected specific participant and the selected video will be played on one screen in separate windows.

[0034] An example of the configuration of the information processing device according to the first embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of the configuration of the information processing device according to the first embodiment.

[0035] As shown in FIG. 4, the information processing device 100 includes a communication unit 102, a storage unit 104, and a control unit .

[0036] The communication unit 102 is a communication interface that performs communication between the information processing device 100 and an external device. The communication unit 102 executes communication between the information processing device 100 and the terminal device 10, for example. The communication unit 102 is realized by, for example, a wireless LAN or a wired LAN.

[0037] The storage unit 104 stores various types of information. The storage unit 104 stores information such as the contents of calculations performed by the control unit 106 and programs. The storage unit 104 includes at least one of, for example, a RAM, a main storage device such as a ROM, and an external storage device such as an HDD.

[0038] The storage unit 104 stores information for reviewing the contents of the web conference, such as video recording data of the participants participating in the web conference, video recording data of the screen of the shared materials shared in the web conference, and audio recording data of the remarks of the participants in the web conference. Note that the information for reviewing the contents of the web conference may be stored in the storage unit 104 of the information processing device 100 as described above, or may be stored in the storage unit 24 of the terminal device 10 of the participant. Details of the video recording data stored in the storage unit 104 will be described later.

[0039] The control unit 106 controls each unit of the information processing device 100. The control unit 106 has, for example, an information processing device such as a CPU or an MPU, and a storage device such as a RAM or a ROM. The control unit 106 executes a program that controls the operation of the information processing device 100 according to the present invention. The control unit 106 may be realized by, for example, an integrated circuit such as an ASIC or an FPGA. The control unit 106 may be realized by a combination of hardware and software.

[0040] The control unit 106 includes, as functional blocks realized by processing by the control unit 106, a display control unit 110, a storage control unit 112, a playback unit 114, a change unit 116, a tag creation unit 118, and a determination unit 120. Each of these functional blocks may be realized by the control unit 26 of the terminal device 10, or may be realized by both the control unit 106 and the control unit 26 of the terminal device 10.

[0041] The display control unit 110 displays the state of the web conference on the display unit 18. The display control unit 110 displays a display screen of the web conference on the display unit 18. More specifically, the display control unit 110 transmits information for reproducing the web conference to the terminal device 10 via the communication unit 102, and the display unit 18 of the terminal device 10 displays an image that reproduces the display screen of the web conference.

[0042] The storage control unit 112 stores the state of the web conference as multiple different types of content in the storage unit 104 of the information processing device 100, or in the storage unit 24 of the terminal device 10 via the communication unit 102. For example, the storage control unit 112 stores content specified by a user in separate files in the storage unit 104 or the storage unit 24. For example, the storage control unit 112 generates a video recording file and an audio file for each participant of the web conference and stores them in the storage unit 104. More specifically, the storage control unit 112 stores the video recording file and audio file of each participant in the storage unit 24 of the terminal device 10 of each participant. Details of the storage control unit 112 will be described later.

[0043] The playback unit 114 outputs at least some of the content stored in the storage unit 104 or the storage unit 24 of the terminal device 10 to the display control unit 110, thereby displaying the content on the display unit 18 of the terminal device 10. For example, the playback unit 114 outputs content selected by the user of the terminal device 10 from among the plurality of content stored in the storage unit 104 or the storage unit 24 to the display control unit 110, thereby displaying and playing the content on a single screen displayed on the display unit 18 of the terminal device 10. Here, the playback unit 114 does not need to play all of the selected content when, for example, the number of content selected by the user is greater than the number that can be displayed on the display unit 18 at one time. For example, the playback unit 114 outputs content selected by the user of the terminal device 10 from among the plurality of content stored in the storage unit 104 to the display control unit 110, thereby displaying and playing the content side by side on a single screen displayed on the display unit 18 of the terminal device 10. The playback unit 114 may select content to display based on tags, which will be described later, from a plurality of pieces of content stored in the storage unit 104 or the storage unit 24. Details of the playback unit 114 will be described later.

[0044] The change unit 116 changes the display mode of the content being played back by the playback unit 114. The change unit 116 changes the display mode of the content being played back by the playback unit 114 in accordance with a user instruction. If a tag is assigned to the content being played back by the playback unit 114, the change unit 116 changes the display mode of the content in accordance with the content of the tag. The change unit 116 changes the display mode of the content based on a tag created by a tag creation unit 118 (described later). More specifically, the change unit 116 changes the display mode of the content based on the priority of the tag determined by the determination unit 120. Details of a method for changing the display mode of the content will be described later.

[0045] The tag creation unit 118 determines whether or not a predetermined action of a participant has been detected based on the appearance of the participant in the web conference displayed on the display unit 18. The tag creation unit 118 detects the predetermined action of the participant, for example, by performing a voice recognition process on the participant's video recording data and a voice recognition process on the participant's voice data. When the tag creation unit 118 detects the predetermined action of the participant, it creates a tag by associating the detected predetermined action with the timing at which the predetermined action was detected. The tag creation unit 118 assigns the created tag to the timing at which the predetermined action was detected in the content.

[0046] The determination unit 120 determines whether or not multiple different types of tags are attached at the same time to the content played back by the playback unit 114. If multiple different types of tags are attached at the same time, the determination unit 120 determines the priority of the multiple tags. Details of the priority determination process will be described later.

[0047] (Recording process) The flow of the recording process according to the first embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the flow of the recording process according to the first embodiment.

[0048] Before the web conference starts, the storage control unit 112 selects the type of content to be recorded by having the user set it (step S10). Here, the users who set the type of content to be recorded include users who are involved in the web conference. Specifically, the storage control unit 112 sets the type of content to be recorded according to the user's selection. FIG. 6 is a diagram showing an example of a content setting screen during recording according to the first embodiment. A content setting screen 200 is displayed on the display unit 18 of the user's terminal device 10. The content setting screen 200 includes a content setting area 210.

[0049] The content setting area 210 includes, for example, a content setting area 210a, a content setting area 210b, and a content setting area 210c. The content setting area 210a is a display for setting content related to user A, who is a participant in the web conference. The content setting area 210b is an area for setting content related to user B, who is a participant in the web conference. The content setting area 210c is an area for setting content related to user C, who is a participant in the web conference. The content setting screen 200 displays a content setting area 210 for each user participating in the web conference.

[0050] Content setting area 210a includes camera video selection section 211a, shared video selection section 212a, and audio selection section 213a. Content setting area 210b includes camera video selection section 211b, shared video selection section 212b, and audio selection section 213b. Content setting area 210c includes camera video selection section 211c, shared video selection section 212c, and audio selection section 213c.

[0051] Camera image selection units 211a to 211c display check boxes for specifying whether or not to record video of users participating in a web conference. If a user wants to record video of a user participating in a web conference, the user simply checks the check box.

[0052] Shared video selection units 212a to 212c display check boxes for specifying whether or not to record the shared video shared by users during a web conference. If a user wants to record the shared video shared by the user during a web conference, the user simply checks the check box.

[0053] The audio selection sections 213a to 213c display checkboxes for specifying whether or not to record the audio of users during a web conference. If a user wants to record the audio spoken by the user during a web conference, the user simply checks the checkbox.

[0054] In the content setting area 210a, the checkboxes for the camera image selection section 211a and the audio selection section 213a are checked. In this case, the storage control section 112 sets the camera image of user A and the audio of user A as the content to be recorded or sounded related to user A.

[0055] In content setting area 210b, the checkboxes for camera video selection section 211b, shared video selection section 212b, and audio selection section 213b are checked. In this case, storage control section 112 sets the camera video of user B, the shared video shared by user B, and the audio of user B as video or audio content related to user B.

[0056] In the content setting area 210c, the check box of the voice selection section 213c is checked. In this case, the storage control section 112 sets the voice of the user C as the content related to the user C to be recorded.

[0057] Returning to FIG. 5, when the web conference starts, the storage control unit 112 starts storing each set content (step S12). FIG. 7 is a diagram for explaining a content storage method according to the first embodiment. The display screen 300 shown in FIG. 7 is a screen displayed on the display unit 18 of the terminal device 10 during the web conference. The display screen 300 includes a camera button 301, a microphone button 302, an exit button 303, a first display area 310, a second display area 311, a third display area 312, a fourth display area 313, and a fifth display area 314.

[0058] The camera button 301 is a button for switching on and off the camera 12 that captures the user of the terminal device 10. The microphone button 302 is a button for switching on and off the voice input unit 16 that detects the voice of the user of the terminal device 10. The exit button 303 is a button for exiting the web conference.

[0059] The first display area 310 displays the shared video shared by user B. The second display area 311 displays the camera video taken of user A. The third display area 312 displays the camera video taken of user B. The fourth display area 313 is an area where the camera video taken of user C is displayed, but is not displayed in the example shown in FIG. 7. The fifth display area 314 is an area where the camera video taken of user C is displayed, but is not displayed in the example shown in FIG. 7.

[0060] In the example shown in Fig. 7, the storage control unit 112 performs recording so as to create a recording file for each video displayed in the first display area 310, the second display area 311, and the third display area 312, in accordance with the content set on the content setting screen 200 shown in Fig. 6. The storage control unit 112 also performs recording so as to create an audio file for each of the audio of user A, user B, and user C.

[0061] Returning to FIG. 5, the control unit 106 determines whether the web conference has ended (step S14). If it is determined that the web conference has ended (step S14; Yes), the process proceeds to step S16. If it is determined that the web conference has not ended (step S14; No), the process proceeds to step S12.

[0062] If the determination in step S14 is Yes, the storage control unit 112 ends storing the files for each set content in the storage unit 104 (step S16). FIG. 8 is a diagram for explaining a method for recording a web conference according to the first embodiment. As shown in FIG. 8, files 401 to 406 are stored in the storage unit 104. File 401 is an audio file of user A. File 402 is a video file of camera footage of user A. File 403 is an audio file of user B. File 404 is a video file of a shared screen shared by user B during the web conference. File 405 is a video file of camera footage of user B. File 406 is an audio file of user C. In this way, the storage control unit 112 generates a file for each set content and stores it in the storage unit 104.

[0063] (Recycling) The flow of the playback process according to the first embodiment will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the flow of the playback process according to the first embodiment.

[0064] The playback unit 114 selects the type of content to be played back by having the user set it (step S20). Specifically, the playback unit 114 sets the content to be played back in accordance with the user's selection. FIG. 10 is a diagram showing an example of a content setting screen during playback according to the first embodiment. A content setting screen 500 is displayed on the display unit 18 of the user's terminal device 10. The content setting screen 500 includes a content setting area 510 and an additional data setting area 520.

[0065] Content setting area 510 includes, for example, content setting area 510a, content setting area 510b, and content setting area 510c. Content setting area 510a is a display for setting content related to user A, who is a participant in the web conference. Content setting area 510b is an area for setting content related to user B, who is a participant in the web conference. Content setting area 510c is an area for setting content related to user C, who is a participant in the web conference. Content setting screen 500 displays content setting area 510 for each user participating in the web conference.

[0066] Content setting area 510a includes camera image selection section 511a and audio selection section 512a. Content setting area 510b includes camera image selection section 511b, shared image selection section 512b, and audio selection section 513b. Content setting area 510c includes audio selection section 511c.

[0067] Camera image selection section 511a and camera image selection section 511b display checkboxes for specifying whether or not to play back images of users who are participating in a web conference. If a user wants to play back images of users who are participating in a web conference, the user simply checks the checkbox.

[0068] Shared video selection section 512b displays a check box for specifying whether or not to play back the shared video shared by users during a web conference. If a user wants to play back the shared video shared by the user during a web conference, the user simply checks the check box.

[0069] The audio selection sections 512a, 513b, and 511c display check boxes for specifying whether or not to play back the audio of users during a web conference. If a user wants to play back the audio spoken by the user during a web conference, the user simply checks the check box.

[0070] The additional data setting area 520 includes an additional image data selection section 520a and an additional video data selection section 520b.

[0071] The additional image data selection section 520a displays check boxes for specifying whether or not to play specified image data, comments written in chat, image data posted in chat, etc. The additional video data selection section 520b displays check boxes for specifying whether or not to play specified video data and various website content, etc.

[0072] In content setting area 510a, check boxes for camera video selection section 511a and audio selection section 513a are checked. In this case, playback section 114 sets the camera video of user A and the audio of user A as content to be played back related to user A.

[0073] In content setting area 510b, the checkboxes for shared video selection section 512b and audio selection section 513b are checked. In this case, playback section 114 sets the shared video shared by user B and the audio of user B as content to be played back related to user B.

[0074] In the content setting area 510c, the check box of the audio selection section 513 is checked. In this case, the playback section 114 sets the audio of the user C as the content related to the user C to be played back.

[0075] The check box in the additional image data selection section 520a is checked in the additional data setting area 520. In this case, the playback section 114 sets any image data designated by the user as the content to be played back.

[0076] Returning to FIG. 9, the playback unit 114 plays the set content (step S22). FIG. 11 is a diagram for explaining the content playback method according to the first embodiment. The display screen 600 shown in FIG. 11 is a screen displayed on the display unit 18 of the terminal device 10 during a web conference. The display screen 600 includes a first display area 601, a second display area 602, a third display area 603, a fourth display area 604, and a fifth display area 605.

[0077] In the first display area 601, camera footage of user A is played. In the second display area 602, camera footage of user B is played. In the third display area 603, the shared screen shared by user B is played. In the fourth display area 604, information indicating that the audio of user C is being played is displayed. In the fifth display area 605, additional image data is displayed. As shown in FIG. 11, the playback unit 114 plays the set content for each display area. In the example shown in FIG. 11, the third display area 603 is displayed with the largest size.

[0078] In step S22, if the number of set contents is greater than the number that can be displayed on the display unit 18 at one time, the playback unit 114 may play back only the number of contents that can be displayed on the display unit 18.

[0079] Returning to FIG. 9, the change unit 116 determines whether or not the user has issued an instruction to change the display mode of the display screen 600 (step S24). FIG. 12 is a diagram for explaining a method for changing the display mode of the display screen for playing back recorded data according to the first embodiment. The user can change the display mode of a display area, for example, by positioning the cursor 610 on the display area whose display mode the user wishes to change. For example, when the user positions the cursor 610 on the first display area 601, a pop-up screen 620 is displayed. The pop-up screen 620 includes instructions for changing the display mode of the first display area 601. The pop-up screen 620 includes, but is not limited to, instructions such as "maximize window," "minimize window," "mute," "unmute," and "delete from content." The user may be able to change the position or size of each display area, for example, by a drag operation. That is, the user can arbitrarily change the sizes of the first display area 601 to the fifth display area 605. If it is determined that the display mode is to be changed (step S24; Yes), the process proceeds to step S26. If it is not determined that the display mode should be changed (step S24; No), the process proceeds to step S28.

[0080] If the determination in step S24 is Yes, the change unit 116 changes the display mode in accordance with an instruction from the user (step S26). FIG. 13 is a diagram for explaining a method for changing the display mode according to the first embodiment. FIG. 13 shows a display screen 600A including a first display area 601 to a fifth display area 605. For example, assume that the user instructs the first display area 601 to be made the largest in the example shown in FIG. 12. In this case, as shown in FIG. 13, the first display area 601 is displayed as the largest on the display screen 600A. In this way, the change unit 116 can change the display mode of each display area included in the display screen in accordance with an instruction from the user.

[0081] Furthermore, the change unit 116 may change the mode of content playback between users who have participated in the web conference and users who have not participated in the web conference.

[0082] For example, if a user viewing web conference information played back by the information processing device 100 is a user who participated in the web conference, the modification unit 116 may modify the display mode of the display area to display content that is likely to be overlooked during the web conference. In this case, the modification unit 116 may maximize the display area for displaying the contents of the shared materials or camera footage at the time when a URL or the like was pasted in the chat field during the web conference. This is because, when a URL or the like is pasted in the chat field during the web conference, the participants of the web conference are likely to browse the website based on the attached URL and overlook the screen of the shared materials or the contents of the camera footage during the web conference.

[0083] In another example, for example, in the case of a user participating in a web conference, since it is assumed that the content of the shared materials has already been grasped during the web conference, the playback unit 114 may maximize the display area showing the camera footage of the participant who is speaking rather than the content of the shared materials.

[0084] For example, in the case of a user who is not participating in a web conference, the change unit 116 may display the material displayed on the shared screen in the largest size possible so that the user can properly understand the contents of the conference.

[0085] Returning to Fig. 9, the playback unit 114 determines whether or not to end the process (step S28). Specifically, the playback unit 114 determines to end the process when an operation to end the playback is received or when playback of the selected content is completed. If it is determined to end the process (step S28; Yes), the process of Fig. 9 ends. If it is not determined to end the process (step S28; No), the process proceeds to step S22.

[0086] In the first embodiment, the user is supposed to check all check boxes on the setting screens shown in Figures 6 and 10, but this is not limiting. For example, each setting may have a default value. Also, the default setting value may be different depending on the user's attributes, for example, between users who have participated in a web conference and users who have not participated in a web conference.

[0087] As described above, in the first embodiment, the state of a web conference is recorded for each set content. In the first embodiment, when playing back recorded data of a web conference, the selected content is displayed side by side on one screen, and the recorded data is played back. In this way, in the first embodiment, only the information that the user considers necessary can be played back, allowing the user to properly understand the content of the web conference.

[0088] [Second embodiment] In the second embodiment, when a web conference is recorded, if a participant in the web conference performs a predetermined action, a tag indicating the content of the predetermined action is added at the timing when the predetermined action was performed.

[0089] (Recording process) The recording process according to the second embodiment will be described with reference to Fig. 14. Fig. 14 is a flowchart showing the flow of the recording process according to the second embodiment.

[0090] The processing of steps S40 and S42 is the same as the processing of steps S10 and S12 shown in Fig. 5, respectively, and therefore description thereof will be omitted. However, in the second embodiment, step S40, that is, processing similar to step S10 shown in Fig. 5, may be omitted. When step S40 is omitted, the storage control unit 112 stores, for example, all of the content that is displayed in the content setting area 210 in Fig. 6 and that can be selected by the user, in separate files in the storage unit 104 or the storage unit 24 (step S42).

[0091] The tag creation unit 118 determines whether a predetermined action has been detected during recording of the web conference (step S44). The tag creation unit 118 detects the predetermined action, for example, by performing image recognition processing on camera images of each participant to be recorded. The tag creation unit 118 detects the predetermined action, for example, by performing voice recognition processing on the voice of each participant to be recorded. The tag creation unit 118 detects the predetermined action, for example, by acquiring input information input to the operation unit 14 of the terminal device 10.

[0092] The predetermined action detected by the tag creation unit 118 may be set in advance, for example, before the start of a web conference. Examples of the predetermined action include, but are not limited to, the start or end of a speech, a mute on / off operation, an operation to start or end screen sharing, a participant to be recorded joining a web conference midway, a participant to be recorded leaving a web conference midway, and a user manually assigning a tag. For example, when detecting the start or end of a speech as the predetermined action, the tag creation unit 118 executes the aforementioned voice recognition process to detect the predetermined action. For example, when detecting joining or leaving a web conference midway as the predetermined action, the tag creation unit 118 executes the aforementioned image recognition process to detect the predetermined action. For example, when detecting the mute on / off operation, an operation to start or end screen sharing, a user's tag assignment operation, or the like as the predetermined action, the tag creation unit 118 detects the predetermined action by acquiring the aforementioned input information.

[0093] If it is determined that the predetermined action has been detected (step S44; Yes), the process proceeds to step S46. If it is determined that the predetermined action has not been detected (step S44; No), the process proceeds to step S50.

[0094] If the determination in step S44 is Yes, tag creation unit 118 creates a tag that associates the predetermined action with the timing at which the predetermined action was detected (step S46). Here, tag creation unit 118 may combine detailed information of the tag, including information on the content in which the predetermined action was detected, time information on the timing at which the predetermined action was detected, and information on the type of the detected predetermined action, into tag information, and store the combined information in storage unit 104, for example.

[0095] The tag creating unit 118 assigns the tag created in step S46 to the content in which the predetermined action is detected at the timing when the predetermined action is detected (step S48).

[0096] The processes in steps S50 and S52 are the same as those in steps S14 and S16 shown in FIG. 5, respectively, and therefore will not be described further.

[0097] (Recycling) (First example) The flow of the playback process according to the first example of the second embodiment will be described with reference to Fig. 15. Fig. 15 is a flowchart showing the flow of the playback process according to the first example of the second embodiment.

[0098] The processes of steps S60 and S62 are the same as the processes of steps S20 and S22 shown in FIG. 9, respectively, and therefore description thereof will be omitted. However, in the first example of the second embodiment, step S60, that is, the process similar to step S20 shown in FIG. 9, may be omitted. When the process of step S60 is omitted, if the total number of contents is greater than the number that can be displayed at one time on the display unit 18, the playback unit 114 selects as many contents as can be displayed on the display unit 18 and plays the contents (step S62). Here, the playback unit 114 can select as many contents as can be displayed based on tags, for example. The playback unit 114 selects content related to the assigned tags.

[0099] The playback unit 114 determines whether or not a tag has been assigned to the content to be played back (step S64). The playback unit 114 determines whether or not a tag has been assigned to the content to be played back for all times from the start time to the end of the web conference to be played back, for example, based on tag information stored in the storage unit 104. If it is determined that a tag has been assigned (step S64; Yes), the playback unit 114 proceeds to step S66. If it is determined that at least one tag has been assigned to any of the content to be played back, the playback unit 114 determines that a tag has been assigned. If it is not determined that a tag has been assigned (step S64; No), the playback unit 114 proceeds to step S76.

[0100] If the determination in step S64 is Yes, the playback unit 114 displays the tags on the display screen (step S66). FIG. 16 is a diagram showing a method for displaying tags according to the second embodiment. As shown in FIG. 16, the playback unit 114 displays tags on, for example, a seek bar 701 displayed on a display screen 700. The left end of the seek bar 701 represents the start time of the web conference, and the right end represents the end time of the web conference. In other words, the tags displayed on the seek bar 701 represent tags at later times as they move to the right.

[0101] 16, tags 702a, 702b, 702c, 702d, 702e, and 702f are displayed. The user can check the details of a tag by hovering cursor 710 over the tag. For example, when cursor 710 is hovered over tag 702b, a pop-up screen 720 is displayed. The pop-up screen 720 displays the timing at which the tag was added and the details of the tag. The details of the tag include the type of predetermined action detected by tag creation unit 118.

[0102] The pop-up screen 720 shows that tags such as "User B's comment" and "User B's screen sharing is ON" have been added 10 minutes and 12 seconds after the content started to be played. For example, when a user clicks on a tag with the cursor 710, the playback unit 114 may skip the content to the point of the clicked tag. For example, when the playback unit 114 receives an operation by the user with the cursor 710 indicating that the tag is unnecessary, the playback unit 114 may delete the display of the tag from the display screen 700. Note that even if the display of the tag is deleted from the display screen 700, the tag added to the content, i.e., the tag information stored in the storage unit 104, is not deleted.

[0103] Returning to FIG. 15, the playback unit 114 determines whether multiple tags are attached at the same timing (step S68). If it is determined that multiple tags are attached at the same timing (step S68; Yes), the process proceeds to step S70. Here, the playback unit 114 determines whether multiple tags are attached by comparing the time information of the timing at which a predetermined action was detected, which is included in the multiple pieces of tag information. If the time information of the timing at which the predetermined action was detected is the same, or if the difference in the time information is within one second, for example, the playback unit 114 determines whether multiple tags are attached at the same timing. If it is determined that multiple tags are attached at the same timing (step S68; No), the process proceeds to step S74.

[0104] If step S68 returns Yes, the determination unit 120 determines the priority of the multiple tags (step S70). Specifically, in the example shown in FIG. 16, the determination unit 120 compares the priorities of two tags, "User B speaks" and "User B turns on screen sharing," which are assigned at the time 10 minutes and 12 seconds have elapsed. The determination unit 120 may determine the priority of the multiple tags based on, for example, a priority determined in advance according to the type of predetermined action detected from the tag information. For example, priorities are determined in advance for all predetermined actions that can be detected. For example, the predetermined actions are ranked from highest to lowest priority, such as starting sharing > starting to speak > joining a web conference midway. For example, if the types of predetermined actions detected from the multiple tag information are voice-based, such as starting to speak, the determination unit 120 may perform voice recognition processing to determine the priority based on the content of the voice. For example, if the type of detected predetermined action is the start of a speech, the judgment unit 120 performs speech recognition processing on the audio content from the time information at which the start of the speech is detected to the time information at which the end of the speech is detected, and determines that the higher the importance of the speech content, the higher the priority.

[0105] The change unit 116 changes the display mode to one corresponding to the tag with the highest priority among the multiple tags (step S72). For example, if the tag with the highest priority is "user B turns on screen sharing," the change unit 116 changes the display mode so that the display area shared by user B is maximized.

[0106] If the determination in step S68 is No, the change unit 116 changes the display mode according to the assigned tag (step S74). For example, if the tag assigned at a certain point in time is "User B turns on screen sharing," the change unit 116 changes the display mode so that the display area of ​​the screen shared by user B is maximized. For example, if the tag assigned at a certain point in time is "User B makes a statement," the change unit 116 changes the display mode so that the display area of ​​the camera image capturing user B is maximized. In other words, the change unit 116 changes the display mode so that the display area related to the assigned tag is maximized.

[0107] The process of step S76 is the same as the process of step S28 shown in FIG. 9, and therefore a description thereof will be omitted.

[0108] As described above, in the second embodiment, tags are added to content when recording a web conference. In the second embodiment, a user can properly understand the content by referring to the tags. Furthermore, in the second embodiment, tags are added at the timing when a predetermined action is detected, so that the user can properly understand the points to be checked in the web conference.

[0109] In addition, in the second embodiment, the display mode is changed according to the tag that has been added, so that in the second embodiment, the mode of the conference can be properly grasped according to the tag.

[0110] In the second embodiment, when multiple tags are attached at the same time, the display mode is changed to that corresponding to the tag with the highest priority among the multiple tags. As a result, in the second embodiment, the state of the conference can be properly grasped according to the tag.

[0111] [Modification of the second embodiment] (Recycling) The flow of the playback process according to the modified example of the second embodiment will be described with reference to Fig. 17. Fig. 17 is a flowchart showing the flow of the playback process according to the modified example of the second embodiment.

[0112] The processes from step S80 to step S88 are the same as the processes from step S60 to step S68 shown in FIG. 15, respectively, and therefore will not be described here.

[0113] If the determination in step S88 is Yes, the change unit 116 changes the display form in order according to the tags (step S90). Specifically, after 10 minutes and 12 seconds have passed, the playback unit 114 plays the content whose display form has been changed by the change unit 116 to correspond to "User B makes a statement" up to the time when the tag "User B finishes speaking" was added, and then plays the content whose display form has been changed by the change unit 116 to correspond to "User B turns on screen sharing" from the time when the tag "User B turns on screen sharing" was added, that is, from 10 minutes and 12 seconds. In other words, the playback unit 114 replays video and the like before and after the time when multiple tags are added.

[0114] The processes in steps S92 and S94 are the same as those in steps S74 and S76 shown in FIG. 15, respectively, and therefore will not be described here.

[0115] In the second embodiment, content with multiple tags attached is replayed. As a result, in the modified example of the second embodiment, video footage with multiple tags attached that is assumed to be an important point is replayed, allowing the user to more accurately understand the content of the web conference.

[0116] The components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads and usage conditions. This distribution and integration configuration may also be performed dynamically.

[0117] Although the embodiments of the present disclosure have been described above, the present disclosure is not limited to the contents of these embodiments. Furthermore, the above-described components include those that can be easily imagined by a person skilled in the art, those that are substantially the same, and those that are within the so-called equivalent range. Furthermore, the above-described components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the spirit of the above-described embodiments. [Explanation of symbols]

[0118] 1. Conference System 10 Terminal Equipment 12 Camera 14 Control section 16 Audio input section 18 Display 20 Audio output section 22,102 Communications Department 24,104 storage section 26,106 Control Unit 100 Information processing device 110 Display control unit 112 Memory control unit 114 Playback Department 116 Changes 118 Tag Creation Department 120 Judgment section

Claims

1. a storage control unit that stores the state of the web conference as a plurality of different types of content in a storage unit; a playback unit that displays at least some of the plurality of contents stored in the storage unit on a display unit of a terminal device of a user associated with the web conference; a change unit that changes the display mode of the content being played back by the playback unit; a tag creation unit that, when detecting a predetermined action of a participant of the web conference based on the state of the participant during the web conference, creates a tag that associates the predetermined action with a timing at which the predetermined action was detected, and assigns the tag to the content; Equipped with the change unit changes the display mode of the content based on the tag. Information processing device.

2. a determination unit that determines a priority based on the contents of the tags when a plurality of different types of tags are assigned to the content at the same time; the change unit changes the display mode of the content in accordance with the priority. The information processing device according to claim 1 .

3. when a plurality of different types of tags are assigned to the content at the same timing, the playback unit replays the content at the timing when the tags are assigned in accordance with the number of the tags assigned; the change unit changes the display mode of the content in accordance with the content of the tag each time the content at the timing when the tag is assigned is replayed.

3. The information processing device according to claim 1.

4. The computer storing the state of the web conference as a plurality of different types of content in a storage unit; displaying at least some of the plurality of contents stored in the storage unit on a display unit of a terminal device of a user associated with the web conference; changing the display mode of the content being played; When a predetermined action of a participant of the web conference is detected based on the state of the participant during the web conference, creating a tag that associates the predetermined action with the timing at which the predetermined action was detected, and assigning the tag to the content; changing a display mode of the content based on the tag; An information processing method that performs the above.

5. storing the state of the web conference as a plurality of different types of content in a storage unit; displaying at least some of the plurality of contents stored in the storage unit on a display unit of a terminal device of a user associated with the web conference; changing a display mode of the content being played; When a predetermined action of a participant of the web conference is detected based on the state of the participant during the web conference, creating a tag that associates the predetermined action with the timing at which the predetermined action was detected, and assigning the tag to the content; changing a display mode of the content based on the tag; A program that causes a computer to execute the following.

Citation Information

Patent Citations

  • Electronic conference system

    JP2002101398A