Lecture support system and method

The system addresses the challenge of monitoring both face-to-face and online students in hybrid lectures by analyzing and synthesizing remote student images with face-to-face footage, enabling lecturers to efficiently check the concentration levels of all students.

JP7722120B2Active Publication Date: 2025-08-13エフサステクノロジーズ株式会社
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021168959
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-14
Publication Date
2025-08-13
Estimated Expiration
2041-10-14

AI Technical Summary

Technical Problem

Lecturers face a significant burden in conducting hybrid lectures that combine face-to-face and online participation, as they need to check the concentration and reactions of both types of students, which is challenging due to the different viewing methods for each group.

Method used

A system that acquires and analyzes both face-to-face and remote student videos, synthesizes remote student images as avatars based on concentration levels, and displays the analysis results alongside face-to-face students on a single screen, allowing lecturers to monitor all students efficiently.

Benefits of technology

Facilitates effective monitoring of both face-to-face and online students' concentration levels, reducing the lecturer's burden and enhancing the hybrid lecture experience by providing comprehensive student status visibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007722120000001
    Figure 0007722120000001
  • Figure 0007722120000002
    Figure 0007722120000002
  • Figure 0007722120000003
    Figure 0007722120000003
Patent Text Reader

Abstract

To support confirmation of a state of a trainee by a lecturer who performs a hybrid type lecture combined between a confrontation lecture and an on-line lecture.SOLUTION: A lecture support system analyzes a concentration degree of each of a confrontation student and a remote student by performing video analysis of each of a confrontation video imaging an inside of a lecture room including confrontation students who receive a lecture in confrontation, and a remote video imaging remote students who receive a lecture on-line; synthesizes an avatar corresponding to the concentration degree of the remote student with a vacant seat in the confrontation video to create a seat video; creates video data of a display video that includes the seat video and an analysis result video representing an analysis result of the concentration degree of the student in a screen; and outputs the created video data to a display device 32 that can be visually recognized by a teacher.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed technology relates to a lecture support system and a lecture support method. [Background technology]

[0002] Conventionally, there have been technologies that allow participants to participate in online meetings or classes from their own locations, such as online conferences or online lectures. In such online conferences or online classes, captured images of participants or processed images of the captured images are displayed in a list format or the like on an information processing terminal such as a personal computer used by each participant.

[0003] For example, a method for implementing computer-mediated communication has been proposed, in which a computer analyzes an image captured by a digital camera used by a participant to recognize the state of the participant, and a computer used by one participant displays a display corresponding to the recognized state of the other participant instead of displaying an image captured by a digital camera used by another participant.

[0004] Also, for example, a video conference system has been proposed in which an information processing device holds a video conference with a communication destination information processing device. In this system, the information processing device sets a group including multiple participants in the video conference and creates image data of one avatar with a standard emotion type corresponding to all participants in the group. The information processing device then determines emotion information of all participants in the video conference and reflects the determined emotion type in the created image data of the avatar.

[0005] In recent years, hybrid lecture formats that combine face-to-face lectures, where participants meet face-to-face with a lecturer, and online lectures, where participants attend online, have become increasingly common in classes at educational institutions such as schools, cram schools, corporate training sessions, etc. In a typical hybrid lecture, the lecture for face-to-face participants is filmed with a camera and the filmed video is streamed in real time to the information processing devices used by the online participants, allowing all participants to receive the lecture in the same way. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Patent No. 6872066 [Patent Document 2] Patent Publication No. 2021-114642 Summary of the Invention [Problem to be solved by the invention]

[0007] Lecturers need to check the concentration and other status of students and their reactions as they lecture. In the case of hybrid lectures that combine face-to-face and online lectures, lecturers check the actual status of students attending face-to-face. On the other hand, for students attending online, lecturers display footage of online students on their own personal computers and check the status of students by looking at the screen. In this way, hybrid lectures require lecturers to conduct the lecture while checking the status of both face-to-face and online students, which creates a significant burden and a difficult situation for lecturers.

[0008] In one aspect, the disclosed technology aims to support instructors who give hybrid lectures that combine face-to-face lectures and online lectures in checking the status of students. [Means for solving the problem]

[0009] In one aspect, the disclosed technology includes an acquisition unit that acquires a first video captured of a venue including a first student attending a lecture and a second video captured of a second student attending the lecture remotely via a network. The disclosed technology also includes an analysis unit that analyzes the first video and the second video to analyze the status of each of the first student and the second student. The disclosed technology also includes a synthesis unit that generates a composite video by synthesizing each of the second videos or each of videos obtained by processing the second video into an area of the first video where the first student is not present. The disclosed technology also includes an output unit that generates video data for a display video including the composite video and a video showing the analysis results by the analysis unit on a single screen, and outputs the generated video data to a display device. [Effects of the Invention]

[0010] One of the advantages of this technology is that it can help instructors who give hybrid lectures that combine face-to-face and online lectures to check the status of students. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing a schematic configuration of a lecture support system. [Figure 2] FIG. 10 is a diagram showing an example of a face-to-face video. [Figure 3] FIG. 1 is a schematic diagram of a remote student and remote video. [Figure 4] FIG. 2 is a functional block diagram of the lecture support device. [Figure 5] FIG. 10 is a diagram showing an example of a seat DB. [Figure 6] FIG. 10 is a diagram illustrating an example of an analysis result DB. [Figure 7] FIG. 10 is a diagram showing an example of avatars according to concentration levels. [Figure 8] FIG. 10 is a diagram for explaining generation of a seat image. [Figure 9]FIG. 10 is a diagram illustrating an example of a display image. [Figure 10] FIG. 1 is a block diagram showing a schematic configuration of a computer that functions as a lecture support device. [Figure 11] 10 is a flowchart illustrating an example of face-to-face video processing. [Figure 12] 10 is a flowchart illustrating an example of remote video processing. [Figure 13] 10 is a flowchart illustrating an example of a display process. [Figure 14] FIG. 10 is a diagram showing another example of a seat image. DETAILED DESCRIPTION OF THE INVENTION

[0012] An example of an embodiment of the disclosed technology will be described below with reference to the drawings. Note that in the following embodiment, a case where the disclosed technology is applied to a class at a school will be described as an example of a lecture. Also, in the following embodiment, only parts related to the disclosed technology in a system for realizing a hybrid class will be described, and parts not directly related to the disclosed technology, such as distribution of video of a teacher or blackboard writing, will be omitted.

[0013] As shown in Fig. 1, a lecture support system 100 according to this embodiment includes a lecture support device 10, a camera 30, a display device 32, and a remote student terminal 40. The lecture support device 10, the camera 30, and the display device 32 are connected to each other via a network such as an intranet. The lecture support device 10 and the remote student terminal 40 are connected to each other via a network such as the Internet. The numbers of cameras 30, display devices 32, and remote student terminals 40 included in the lecture support system 100 are not limited to the example shown in Fig. 1.

[0014] Camera 30 captures images of the classroom, including students (hereinafter referred to as "face-to-face students") attending classes face-to-face with the teacher, and outputs the captured images. Camera 30 is installed in a position that allows it to capture the upper body, including the face, of each face-to-face student from the eye level of the teacher standing at the podium, i.e., from the front to the back of the classroom. Figure 2 shows an example of an image captured by camera 30 (hereinafter referred to as "face-to-face image"). Note that face-to-face students are an example of a "first student" in the disclosed technology, and face-to-face image is an example of a "first image" in the disclosed technology. In addition, in this specification, "image" includes both one frame of an image and multiple frames of an image. In other words, "image" in this specification can also be referred to as "image."

[0015] The display device 32 displays a display image (details of which will be described later) generated by the lecture support device 10. The display device 32 may be realized, for example, as a head-mounted display worn by the teacher, a freestanding display, etc. In the case of a freestanding display, it may be a small display placed in a position visible only to the teacher, such as on a teacher's desk, or a large display placed in a position visible to the teacher and the students facing it, such as in a corner of the classroom.

[0016] The remote student terminal 40 is an information processing terminal used by a student (hereinafter referred to as a "remote student") who takes online classes remotely, such as at home. The remote student terminal 40 is realized by a personal computer, tablet terminal, smartphone, or the like equipped with a camera, microphone, display device, network connection function, etc. FIG. 3 shows a schematic diagram of a remote student taking a class using the remote student terminal 40. The camera 40A of the remote student terminal 40 captures an image of the remote student and outputs the captured image. The camera 40A is installed, for example, in a position where it can capture the upper body, including the face, of the remote student. The upper view of FIG. 3 is an example of an image captured by the camera 40A (hereinafter referred to as a "remote image"). The remote student terminal 40 acquires the remote image output from the camera 40A and transmits it to the lecture support device 10 via a network together with the remote student's identification information (hereinafter referred to as a "remote student ID"). The remote student is an example of a "second student" in the disclosed technology, and the remote image is an example of a "second image" in the disclosed technology.

[0017] 4, the lecture support device 10 functionally includes an acquisition unit 11, an analysis unit 12, a synthesis unit 13, and an output unit 14. In addition, a predetermined storage area 20 of the lecture support device 10 stores a video DB (Database) 21, a seat DB 22, and an analysis result DB 23.

[0018] The acquisition unit 11 acquires the face-to-face video output from the camera 30 and stores it in the video DB 21. The acquisition unit 11 also acquires the remote video and remote student ID transmitted from the remote student terminal 40 and stores it in the video DB 21. Time information is assigned to each frame of the face-to-face video and the remote video, and the face-to-face video and the remote video can be synchronized based on the time information.

[0019] The analysis unit 12 analyzes the status of each of the face-to-face students and the remote students by analyzing the face-to-face video and the remote video. Specifically, the analysis unit 12 acquires the face-to-face video from the video DB 21 and analyzes the status of each face-to-face student for each seat based on the seat DB 22. FIG. 5 shows an example of the seat DB 22. In the example of FIG. 5, row numbers are assigned from the front of the classroom (the teacher's desk side) in order, such as row A, row B, row C, etc., and array numbers are assigned from left to right in each row, such as 1, 2, 3, etc., and the seat number is represented by a combination of the row number and the array number. Hereinafter, a seat number with row number i and array number j will be represented as "ij." The analysis unit 12 identifies the area of each seat in the face-to-face video based on the layout of each seat defined in the seat DB 22 and the angle of view determined from the installation position and installation angle of the camera 30. For each identified seat area, the analysis unit 12 analyzes the status of the face-to-face students present in that area. If no facing student is present in the area of a seat, the analysis unit 12 identifies the seat as an empty seat.

[0020] For example, the analysis unit 12 analyzes the face-to-face student's level of concentration in class as an example of the state of the face-to-face student. Specifically, the analysis unit 12 analyzes the level of concentration based on at least one of the face-to-face student's facial expression, head orientation, and body movement. More specifically, the analysis unit 12 estimates the type of emotion of the face-to-face student from the face-to-face student's facial expression, and analyzes the face-to-face student's level of concentration based on a predefined relationship between the type of emotion and the level of concentration and the estimated type of emotion. Note that a conventionally known method may be used to estimate the type of emotion from the facial expression. Furthermore, the analysis unit 12 may analyze that the level of concentration is high when the face-to-face student's head is facing forward, and low when the head is facing sideways or downward. Furthermore, the analysis unit 12 may analyze that the level of concentration is low when the magnitude of the face-to-face student's body movement is equal to or greater than a predetermined value. Note that the magnitude of the body movement may be obtained using face-to-face video for the past few frames including the frame to be processed.

[0021] The analysis unit 12 stores the analyzed concentration level of the face-to-face student for each seat in the analysis result DB 23, in association with the seat number of the seat and time information of the face-to-face video. FIG. 6 shows an example of the analysis result DB 23. In the example of FIG. 6, each row corresponds to one seat, and the following information is stored in association with each seat: "Seat Number," "Presence or Absence of Face-to-Face Student," "Remote Student ID," and "Analysis Result." "Presence or Absence of Face-to-Face Student" is information indicating whether or not a face-to-face student is present at that seat. In the example of FIG. 6, a seat with a face-to-face student is represented by "Present," and a seat without a face-to-face student is represented by "Absent." The "Analysis Result" stores the analysis results in chronological order, in association with time information (t1, t2, ... in the example of FIG. 6). In this embodiment, the concentration level is classified into four levels, and classification symbols S1, S2, S3, and S4 are assigned to each level in descending order of concentration level. FIG. 6 shows an example in which these classification symbols are stored as analysis results. The "Remote Student ID" will be described later.

[0022] The analysis unit 12 also analyzes the concentration levels of the remote students from the remote video. The analysis unit 12 stores the remote student IDs in association with vacant seats in the analysis result DB 23, and stores the analysis results of the remote students indicated by the remote student IDs in association with the time information of the remote video. Specifically, as shown in FIG. 6, the analysis unit 12 stores the remote student ID in the "Remote Student ID" field of the row in the analysis result DB 23 where the "Whether Face-to-Face Student Is There" field is "None." Then, the analysis unit 12 stores the analysis results of the remote student indicated by the remote student ID in the corresponding time information field of the "Analysis Results" for that row.

[0023] The synthesis unit 13 generates seating images by synthesizing each of the remote images or each of the images obtained by processing the remote images in areas of the face-to-face image where no face-to-face students are present. In this embodiment, a case will be described in which avatars that resemble remote students are used as images obtained by processing the remote images. The seating images are an example of the "synthetic images" of the disclosed technology, and the avatars are an example of the "human-shaped images" of the disclosed technology. The image data of the avatars is stored in the image DB 21.

[0024] Specifically, the synthesis unit 13 identifies avatars with different appearances depending on the remote student's concentration level based on the analysis results of the remote student stored in the analysis result DB 23. For example, as shown in FIG. 7, avatars are defined according to classification symbols representing the level of concentration. The "analysis result display information" in FIG. 7 will be described later. The synthesis unit 13 determines an avatar corresponding to the classification symbol indicating the analysis result stored in the analysis result DB 23, and reads video data of the determined avatar from the video DB 21. The synthesis unit 13 references the analysis result DB 23 to identify the seat area indicated by the seat number assigned to the remote student in the face-to-face video, and synthesizes the read avatar with the identified seat area to generate a seat video. For example, as shown in FIG. 8, it is assumed that a remote student with a remote student ID of R2 is assigned to vacant seat A3, and an analysis result S1 is obtained from the remote video at time information t7. In this case, the synthesis unit 13 reads out the video data of the avatar corresponding to the classification symbol S1 from the video DB 21 and synthesizes it with the area of the seat A3 in the face-to-face video for the time information t7, thereby generating a seat video for the time information t7. The synthesis unit 13 stores the generated seat video in the video DB 21.

[0025] The output unit 14 generates an analysis result video showing the analysis results based on the analysis results stored in the analysis result DB 23. For example, the output unit 14 generates an analysis result video showing at least one of the aggregated results of the students' concentration levels and the time-series changes in the concentration levels. Specifically, the output unit 14 may generate an analysis result video showing aggregated results, such as "S1: 10 people, S2: 10 people, S3: 5 people, S4: 2 people," in text or a table based on the analysis results associated with the relevant time information. Also, as shown in FIG. 7, for example, different colors may be defined as analysis result display information for each classification symbol representing the analysis results. The output unit 14 may then generate an analysis result video in meter format, in which each analysis result associated with the relevant time information is represented by a cell, the number of cells corresponding to the total number of students is arranged in order of concentration level, and each cell is colored in a color indicated by the analysis result display information. Also, for example, the output unit 14 may generate an analysis result video showing the time-series changes in concentration levels as a line graph or by color changes indicated by the analysis result display information. In this case, an analysis result video may be generated for each student, or an analysis result video may be generated based on the average of the analysis results for all students or for each predetermined group such as each row.

[0026] The output unit 14 generates video data of a display video including the generated analysis result video and a seating video generated by the synthesis unit 13 based on the face-to-face video and the remote video of the corresponding time information on a single screen. The output unit 14 stores the generated video data in the video DB 21 and outputs it to the display device 32. FIG. 9 shows an example of a display video displayed on the display device 32. The example of FIG. 9 shows an example in which the display video is displayed on a head-mounted display worn by the teacher. In addition, the example of FIG. 9 shows an example in which the upper part of the display video is a seating video and the lower part is an analysis result video, and the analysis result video is shown in the meter format described above. Note that the output unit 14 may switch between different types of analysis result video and output them at predetermined time intervals, or may generate multiple types of analysis result video and switch between them by operation of the teacher.

[0027] The lecture support device 10 may be realized by, for example, a computer 50 shown in FIG. 10 . The computer 50 includes a CPU (Central Processing Unit) 51, a memory 52 as a temporary storage area, and a non-volatile storage unit 53. The computer 50 also includes an input / output I / F (Interface) 54 for connecting to external devices including the camera 30 and the display device 32, and an R / W (Read / Write) unit 55 for controlling reading and writing of data from and to a storage medium 59. The computer 50 also includes a communication I / F 56 for connecting to a network such as the Internet. The computer 50 also includes a GPU (Graphics Processing Unit) 58 for performing video analysis. The CPU 51, memory 52, storage unit 53, input / output I / F 54, R / W unit 55, communication I / F 56, and GPU 58 are connected to one another via a bus 57.

[0028] The storage unit 53 may be realized by a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage unit 53 as a storage medium stores a lecture support program 60 for causing the computer 50 to function as the lecture support device 10 by executing lecture support processing including face-to-face video processing, remote video processing, and display processing, which will be described later. The lecture support program 60 includes an acquisition process 61, an analysis process 62, a synthesis process 63, and an output process 64. The storage unit 53 also includes an information storage area 70 in which information constituting each of the video DB 21, the seat DB 22, and the analysis result DB 23 is stored.

[0029] The CPU 51 reads the lecture assistance program 60 from the storage unit 53, loads it into the memory 52, and sequentially executes the processes of the lecture assistance program 60. The CPU 51 executes the acquisition process 61 to operate as the acquisition unit 11 shown in FIG. 4. The CPU 51 executes the analysis process 62 to operate as the analysis unit 12 shown in FIG. 4. The CPU 51 executes the synthesis process 63 to operate as the synthesis unit 13 shown in FIG. 4. The CPU 51 executes the output process 64 to operate as the output unit 14 shown in FIG. 4. The CPU 51 reads information from the information storage area 70 and loads the video DB 21, seat DB 22, and analysis result DB 23 into the memory 52. As a result, the computer 50 that has executed the lecture assistance program 60 functions as the lecture assistance device 10. The CPU 51 that executes the program is hardware.

[0030] The functions realized by the lecture assistance program 60 can also be realized by, for example, a semiconductor integrated circuit, more specifically, an ASIC (Application Specific Integrated Circuit).

[0031] Furthermore, the lecture support device 10 is placed near the display device so that the delay time from when the face-to-face video and the remote video are acquired until the display video is displayed on the display device is within a predetermined time.

[0032] Next, the operation of the lecture support device 10 according to this embodiment will be described. As shown in FIG. 11, when the camera 30 starts capturing images and outputting the captured face-to-face images (step S10), face-to-face image processing is performed in the lecture support device 10. As shown in FIG. 12, when the camera 40A of the remote student terminal 40 starts capturing images and the remote student terminal 40 starts transmitting the remote images and the remote student ID (step S30), the lecture support device 10 starts remote image processing. Furthermore, as shown in FIG. 13, the lecture support device 10 starts display processing. The lecture support processing including the face-to-face image processing, the remote image processing, and the display processing is an example of a lecture support method of the disclosed technology. Each of the face-to-face image processing, the remote image processing, and the display processing may be performed for each frame of the face-to-face image and the remote image, or may be performed for each predetermined frame. Each of the face-to-face image processing, the remote image processing, and the display processing will be described in detail below.

[0033] First, face-to-face video processing will be described with reference to FIG.

[0034] In step S21, the acquisition unit 11 determines whether or not the acquisition unit 11 has acquired a face-to-face video output from the camera 30. If a face-to-face video has been acquired, the process proceeds to step S22; if not, the face-to-face video processing ends.

[0035] In step S22, the analysis unit 12 identifies the area of each seat in the face-to-face video acquired in step S21 above, based on the arrangement of each seat defined in the seat DB 22 and the angle of view determined from the installation position and installation angle of the camera 30.

[0036] Next, in step S23, the analysis unit 12 performs video analysis for each of the identified seat areas to analyze the concentration levels of the face-to-face students present in that area. The analysis unit 12 then associates the analyzed concentration levels of the face-to-face students for each seat with the seat number of that seat and the time information of the face-to-face video, stores the association results in the analysis result DB 23, and returns to step S21.

[0037] Next, remote video processing will be described with reference to FIG.

[0038] In step S41, the acquisition unit 11 determines whether or not it has acquired the remote video and the remote student ID transmitted from the remote student terminal 40. If it has acquired the remote video and the remote student ID, it proceeds to step S42, and if it has not acquired the remote video and the remote student ID, it ends the remote video processing.

[0039] In step S42, the analysis unit 12 performs video analysis of the remote video to analyze the concentration levels of the remote students. The analysis unit 12 then stores the remote student IDs in association with vacant seats in the analysis result DB 23, and stores the analysis results of the remote students indicated by the remote student IDs in association with the time information of the remote video.

[0040] Next, in step S43, the synthesis unit 13 determines an avatar that corresponds to the level of concentration indicated by the analysis results of the remote student from among avatars whose appearances differ depending on the level of concentration, based on the analysis results of the remote student stored in the analysis result DB23, and returns to step S41.

[0041] Next, the display process will be described with reference to FIG.

[0042] In step S51, the synthesis unit 13 determines whether or not a face-to-face video corresponding to the time information to be processed is stored by referring to the video DB 21. If a face-to-face video is stored, the process proceeds to step S52, and if not, the process proceeds to step S55.

[0043] In step S52, the synthesis unit 13 determines whether or not a remote video corresponding to the time information to be processed is stored by referring to the video DB 21. If a remote video is stored, the process proceeds to step S53, and if not, the process proceeds to step S54.

[0044] In step S53, the synthesis unit 13 reads out the face-to-face video corresponding to the time information to be processed from the video DB 21, and identifies the seat area indicated by the seat number assigned to the remote student on the face-to-face video by referring to the analysis result DB 23. Then, the synthesis unit 13 reads out the video data of the avatar determined in step S43 of the remote video processing from the video DB 21 based on the remote video corresponding to the time information to be processed, and synthesizes it with the identified seat area to generate a seat video. Meanwhile, in step S54, the synthesis unit 13 uses the face-to-face video as the seat video as is.

[0045] In step S55, the synthesis unit 13 determines whether or not a remote video corresponding to the time information to be processed is stored by referring to the video DB 21. If a remote video is stored, the process proceeds to step S56, and if not, the process returns to step S51.

[0046] In step S56, the synthesis unit 13 reads, based on the remote video corresponding to the time information to be processed, the video data of the avatars determined in step S43 of the remote video processing from the video DB 21. Then, the synthesis unit 13 generates a seating video in which the avatars are arranged.

[0047] Next, in step S57, output unit 14 generates an analysis result video showing the analysis result of the concentration level based on the analysis result stored in analysis result DB 23. Next, in step S58, output unit 14 generates video data of a display video including the generated analysis result video and the seat video generated in step S53, S54, or S56 on one screen, and outputs the video data to display device 32. As a result, the display video is displayed on display device 32 (step S70).

[0048] As described above, in the lecture support system according to this embodiment, the lecture support device acquires face-to-face video captured inside a classroom, including face-to-face students, and remote video captured inside a classroom, including remote students. The lecture support device then analyzes the states of the face-to-face students and the remote students, such as their levels of concentration, by performing video analysis on the acquired face-to-face video and remote video. The lecture support device also generates seating videos by combining the remote video or processed remote video with areas of the face-to-face video where no face-to-face students are present, such as empty seats. The lecture support device then generates video data for display video including the seating video and video showing the results of the state analysis on a single screen, and outputs the generated video data to a display device. This helps teachers who teach hybrid classes that combine face-to-face and online classes to check the state of their students.

[0049] Furthermore, in the lecture support device according to the above embodiment, avatars are used as images obtained by processing the remote video according to the concentration level of the remote students. This allows consideration to be given to the privacy of students participating from home, etc. Furthermore, even when it is difficult to see the facial expressions and reactions of students in the remote video due to issues such as image quality or screen size, the use of avatars makes it easier to check the status of the students. Note that the generation of seating images is not limited to combining the avatars of the remote students with the face-to-face video, and the remote video may be combined directly.

[0050] Furthermore, the lecture support device according to the above embodiment can quantitatively grasp the state of students by including the analysis results of the state of students in the displayed image in addition to the images of the face-to-face students and the remote students.

[0051] In the above embodiment, an avatar is superimposed on an empty seat where no face-to-face student is present, but the present invention is not limited to this. For example, as shown in Fig. 14, a seating image may be generated by superimposing a panel-like image in which each remote image or each image of an avatar created by processing the remote image is arranged in an area such as a corner of a classroom with the face-to-face image.

[0052] In the above embodiment, the concentration level is analyzed as the state of the students, but the present invention is not limited to this. For example, other indicators that can grasp the students' reactions to the lesson, such as the level of alertness or satisfaction, may be analyzed, or the type of emotion may be analyzed from the students' facial expressions.

[0053] In the above embodiment, the application of the disclosed technology to a school class has been described as an example of a lecture, but the present invention is not limited to this. The disclosed technology can be applied to various hybrid lectures that combine face-to-face and online lectures, such as classes at other educational institutions or cram schools, corporate training sessions, and various other lectures.

[0054] In the above embodiment, the display image generated by the lecture support device is transmitted to the display device for display, but the present invention is not limited to this. The functions of the lecture support device of the above embodiment may be installed as an application on the display device, and the display image may be generated and displayed on the display device.

[0055] In the above embodiment, the lecture assistance program is pre-stored (installed) in the storage unit, but the present invention is not limited to this. The program according to the disclosed technology can also be provided in a form stored on a storage medium such as a CD-ROM, a DVD-ROM, or a USB memory.

[0056] The following additional notes are provided regarding the above-described embodiments.

[0057] (Appendix 1) an acquisition unit that acquires a first video of a venue including a first student attending a lecture and a second video of a second student attending the lecture remotely via a network; an analysis unit that analyzes the state of each of the first student and the second student by performing video analysis on each of the first video and the second video; a synthesis unit that generates a synthesized video by synthesizing each of the second videos or each of the videos obtained by processing the second videos in an area of the first video where the first student does not exist; an output unit that generates video data of a display video including the composite video and a video showing the analysis result by the analysis unit on one screen, and outputs the generated video data to a display device; A lecture support system including:

[0058] (Appendix 2) The lecture support system described in Appendix 1, wherein the synthesis unit synthesizes a human-shaped image imitating the second student, the human-shaped image having a different appearance depending on the state of the second student analyzed by the analysis unit, into the first image as a processed image of the second image.

[0059] (Appendix 3) the analysis unit identifies an empty seat where the first student is not present based on the seating information of the venue and the first video; The synthesizing unit synthesizes each of the second images or each of the images obtained by processing the second images onto the vacant seats in the first image. A lecture support system according to appendix 1 or appendix 2.

[0060] (Appendix 4) The lecture support system described in Appendix 1 or Appendix 2, wherein the synthesis unit synthesizes a panel-like image in which each of the second images or each of images processed from the second images are arranged in an area of the first image where the first student is not present.

[0061] (Appendix 5) The lecture support system according to any one of Supplementary Notes 1 to 4, wherein the analysis unit analyzes the concentration levels of the first and second students as the states of the first and second students based on at least one of the facial expressions, head orientations, and body movements of the first and second students in the first and second videos, respectively.

[0062] (Appendix 6) the analysis unit stores the states of the first student and the second student in a storage unit in a chronological order; The output unit generates an image showing the analysis result based on the information stored in the storage unit. A lecture support system according to any one of Supplementary Note 1 to Supplementary Note 5.

[0063] (Appendix 7) The lecture support system described in Appendix 6, wherein the output unit generates a video showing the analysis results that represent at least one of the summary results of the status of each of the first student and the second student and the time series changes in the status based on the information stored in the memory unit.

[0064] (Appendix 8) A lecture support system described in any one of Appendix 1 to Appendix 7, wherein the output unit outputs the video data to at least one of a display device installed in the venue and a head-mounted display worn by the lecturer of the lecture.

[0065] (Appendix 9) A lecture support system as described in any one of Supplementary Notes 1 to 8, wherein a lecture support device including the acquisition unit, the analysis unit, the synthesis unit, and the output unit is arranged in the vicinity of the display device so that the delay time from acquisition of the first image and the second image to display of the display image on the display device is within a predetermined time.

[0066] (Appendix 10) Acquire a first video of a venue including a first student attending a lecture and a second video of a second student attending the lecture remotely via a network; analyzing the first video and the second video to analyze the states of the first student and the second student; generating a composite video by combining each of the second videos or each of the videos obtained by processing the second videos with an area of the first video where the first student is not present; generating image data of a display image including the composite image and an image showing the state analysis result on one screen, and outputting the generated image data to a display device; A lecture support method in which a computer executes a process including the steps of:

[0067] (Appendix 11) A lecture support method as described in Appendix 10, in which a human-shaped image modeled after the second student, the human-shaped image having different appearances depending on the analyzed state of the second student, is synthesized into a processed image of the second image and then combined with the first image.

[0068] (Appendix 12) Identifying an empty seat where the first student does not exist based on the seat information of the venue and the first video; Each of the second images or each of the images obtained by processing the second images is superimposed on the vacant seats in the first image. A lecture support method according to Appendix 10 or Appendix 11.

[0069] (Appendix 13) A lecture support method as described in Appendix 10 or Appendix 11, in which a panel-like image in which each of the second images or each of images processed from the second images is arranged is synthesized in an area of the first image in which the first student is not present.

[0070] (Appendix 14) A lecture support method according to any one of Supplementary Note 10 to Supplementary Note 13, in which the concentration levels of the first and second students are analyzed as the states of the first and second students based on at least one of the facial expressions, head orientations, and body movements of the first and second students in the first and second videos, respectively.

[0071] (Appendix 15) storing the states of the first student and the second student in a storage unit in a chronological order; generating an image showing the analysis result based on the information stored in the storage unit; A lecture support method according to any one of Supplementary Note 10 to Supplementary Note 14.

[0072] (Appendix 16) A lecture support method as described in Appendix 15, which generates an image showing the analysis results that represent at least one of the aggregated results of the status of each of the first student and the second student and the time-series changes in the status based on the information stored in the memory unit.

[0073] (Appendix 17) The lecture support method according to any one of Supplementary Note 10 to Supplementary Note 16, wherein the video data is output to at least one of a display device installed in the venue and a head-mounted display worn by the lecturer of the lecture.

[0074] (Appendix 18) Acquire a first video of a venue including a first student attending a lecture and a second video of a second student attending the lecture remotely via a network; analyzing the first video and the second video to analyze the states of the first student and the second student; generating a composite video by combining each of the second videos or each of the videos obtained by processing the second videos with an area of the first video where the first student is not present; generating image data of a display image including the composite image and an image showing the state analysis result on one screen, and outputting the generated image data to a display device; A lecture support program that causes a computer to execute processes including the above.

[0075] (Appendix 19) A lecture support program as described in Appendix 18, which synthesizes a human-shaped image of the second student, the human-shaped image having different appearances depending on the analyzed state of the second student, into the first image as a processed image of the second image.

[0076] (Appendix 20) Acquire a first video of a venue including a first student attending a lecture and a second video of a second student attending the lecture remotely via a network; analyzing the first video and the second video to analyze the states of the first student and the second student; generating a composite video by combining each of the second videos or each of the videos obtained by processing the second videos with an area of the first video where the first student is not present; generating image data of a display image including the composite image and an image showing the state analysis result on one screen, and outputting the generated image data to a display device; A non-transitory storage medium storing a lecture support program for causing a computer to execute a process including the above. [Explanation of symbols]

[0077] 10 Lecture support equipment 11 Acquisition Department 12 Analysis Department 13 Synthesis section 14 Output section 20 storage area 21 Video DB 22 Seat DB 23 Analysis result DB 30 Camera 32 Display device 40 Remote Student Devices 40A Camera 50 Computers 51 CPU 52 memory 53 Storage section 54 Input / Output Interface 55 R / W section 56 Communication I / F 57 Bus 58 GPU 59 Storage medium 60 Lecture Support Program 100 Lecture Support System

Claims

1. an acquisition unit that acquires a first video of a venue including a first student attending a lecture and a second video of a second student attending the lecture remotely via a network; an analysis unit that analyzes the first video and the second video to analyze the states of the first student and the second student; a synthesis unit that generates a synthesized video by synthesizing each of the second videos or each of the videos obtained by processing the second videos in an area of the first video where the first student does not exist; an output unit that generates video data of a display video including the composite video and a video showing the analysis result by the analysis unit on one screen, and outputs the generated video data to a display device; A lecture support system including:

2. The lecture support system of claim 1, wherein the synthesis unit synthesizes a human-shaped image that imitates the second student, the human-shaped image having a different appearance depending on the state of the second student analyzed by the analysis unit, into the first image as a processed image of the second image.

3. the analysis unit identifies an empty seat where the first student is not present based on the seating information of the venue and the first video; The synthesizing unit synthesizes each of the second images or each of the images obtained by processing the second images onto the vacant seats in the first image.

3. The lecture support system according to claim 1 or 2.

4. The lecture support system described in claim 1 or claim 2, wherein the synthesis unit synthesizes a panel-shaped image in which each of the second images or each of images processed from the second images are arranged in an area of the first image where the first student is not present.

5. The lecture support system described in any one of claims 1 to 4, wherein the analysis unit analyzes the concentration levels of the first and second students as the states of the first and second students based on at least one of the facial expressions, head orientation, and body movements of the first and second students in the first and second videos, respectively.

6. the analysis unit stores the states of the first student and the second student in a storage unit in a chronological order; The output unit generates an image showing the analysis result based on the information stored in the storage unit. The lecture support system according to any one of claims 1 to 5.

7. The lecture support system of claim 6, wherein the output unit generates a video showing the analysis results that represent at least one of the summary results of the states of the first student and the second student and the time series changes in the states based on the information stored in the memory unit.

8. The lecture support system according to any one of claims 1 to 7, wherein the output unit outputs the video data to at least one of a display device installed in the venue and a head-mounted display worn by the lecturer of the lecture.

9. The lecture support system according to any one of claims 1 to 8, wherein the lecture support device including the acquisition unit, the analysis unit, the synthesis unit, and the output unit is arranged near the display device so that the delay time from acquisition of the first image and the second image to display of the display image on the display device is within a predetermined time.

10. Acquire a first video of a venue including a first student attending a lecture and a second video of a second student attending the lecture remotely via a network; analyzing the first video and the second video to analyze the states of the first student and the second student; generating a composite video by combining each of the second videos or each of the videos obtained by processing the second videos with an area of the first video where the first student is not present; generating image data for a display image including the composite image and an image showing the state analysis result on one screen, and outputting the generated image data to a display device; A lecture support method in which a computer executes a process including the steps of:

Citation Information

Patent Citations

  • Attendance status discriminating device, method, and program

    JP2006330464A

  • Information processing system, information processing device and program

    JP2018060375A

  • Video conference device, video conference system, and program

    JP2021114642A

  • System, method and program for implementing computer-mediated communication

    JP6872066B1

  • JPP6872066B