Educational support system, educational support method, program, and program recording medium
The educational support system effectively categorizes and labels postures in educational videos, enhancing review processes by providing chronological displays and interactive features, thus improving lesson quality.
Patent Information
- Application Number
- JP2025092585
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Existing systems fail to provide an easy and effective way to grasp and evaluate the status of educational videos for review purposes, particularly in classrooms or similar settings.
An educational support system that includes an acquisition unit, a situation estimation unit, and a display unit to analyze and label the posture of individuals in videos, displaying these labels chronologically and allowing for comments and code recognition to enhance review capabilities.
Enables easy review and understanding of educational videos by categorizing and labeling postures, facilitating improved lesson evaluation and content enhancement.
Smart Images

Figure 0007802996000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an education support system, an education support method, a program, and a program recording medium. [Background technology]
[0002] In recent years, a method has been implemented in which educators such as teachers and lecturers film the lessons they give to students, and then review the videos later to improve the lessons.
[0003] Patent Document 1 discloses a feedback information processing system that can quickly provide feedback on the content of expressions of intent at a meeting held for a large number of participants. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-48381 Summary of the Invention [Problem to be solved by the invention]
[0005] However, with the technology described in Patent Document 1, even if feedback can be provided by someone other than the presenter, it is not possible to easily grasp the situation of the target video and easily evaluate or review it.
[0006] In view of the above-mentioned circumstances, an object of the present invention is to provide a novel technique that allows a user to easily grasp the status of a target video for review purposes. [Means for solving the problem]
[0007] [1] An educational support system, An acquisition unit, a situation estimation unit, and a display unit, The acquisition unit acquires a target video, which is a video capturing one or more people present at a certain location; the situation estimation unit assigns a situation label to each time segment of the target video based on a posture of one or more people estimated based on the target video; the display unit displays a video playback screen including a video playback area for playing the target video and the status label for each category; Educational support system. [2] The education support system further includes one or more image capture devices, The said photographing device is a device that is placed in one or more classrooms, and photographs the said target video, which is a video photograph of one or more students participating in the class; the situation estimation unit assigns the situation label based on the posture of the one or more students; [1] The educational support system described in [1]. [3] The situation labels are displayed in sections along a timeline corresponding to the time of the target video. [1] The educational support system described in [1]. [4] The education support system further includes a posture estimation unit, the posture estimation unit estimates a body flexion degree relating to a degree of flexion of the upper body as a posture of one or more persons based on the target video; the situation estimation unit assigns the situation label based on the body flexure degree. An educational support system according to any one of [1] to [3]. [5] The display unit further displays a graph showing the body undulation degree in time series. [4] The educational support system described in [4]. [6] The education support system further includes a reception unit and a photographing device, the reception unit receives schedule information regarding a class schedule including a classroom for the class, a start time and an end time of the class; the imaging device performs imaging based on the schedule information, and stores the acquired target video in association with the schedule information; the display unit causes the imaging device to display the schedule information associated with the target video; An educational support system according to any one of [1] to [5]. [7] The education support system further includes a reception unit, When the receiving unit receives an input of a comment, the receiving unit identifies a playback time of the target video at the time the comment was input, and stores the comment and the playback time in association with each other; the display unit displays a comment display area showing the comment on the video playback screen; When the comment is pressed, the display unit plays the target video at the time when the comment was associated in the video playback area. An educational support system according to any one of [1] to [6]. [8] The educational support system further includes a code acquisition unit, the code acquisition unit recognizes the code based on the target video that captures a code that can identify a document and / or a page of the document, identifies the time when the code was displayed, and acquires code information from the code; the display unit displays material information related to the material corresponding to the code information based on the acquired code information, in correspondence with the time when the code is displayed. An educational support system according to any one of [1] to [7]. [9] A computer-implemented educational support method, comprising: The method includes an acquisition step, a situation determination step, and a display step, In the acquisition step, a target video is acquired, which is a video of one or more people present at a certain location; In the situation determination step, a situation label is assigned to each time segment of the target video based on a posture of one or more people estimated based on the target video; In the display step, a video playback screen including a video playback area for playing the target video and the status label for each category is displayed. Educational support methods.
[10] A program that causes the computer to execute the method described in [9].
[11] A recording medium storing a program that causes the computer to execute the method described in [9].
[0008] The invention of [1] makes it possible to easily review the target video by classifying it according to the situation that can be grasped from the target video.
[0009] The invention according to [2] makes it possible to assign situation labels based on the posture of students participating in class.
[0010] The invention according to [3] makes it possible to display situation labels in chronological order.
[0011] [4] The invention of the present invention allows the situation of the location captured in the video to be estimated based on the degree of undulation of the upper body, and allows for a review.
[0012] According to the invention of [5], the degree of upper body undulation, which indicates the degree of upper body undulation, can be easily confirmed along with the video.
[0013] According to the invention of [6], the target video can be shot by the shooting device based on the registered schedule, and information about the schedule can be confirmed together with the target video.
[0014] [7] The invention allows users to check comments and, by pressing a comment, play the part of the video associated with the comment.
[0015] [8] The invention of the present invention allows for the automatic setting of a timestamp from a code captured in a video, and the corresponding part of the video can be easily played back. [Effects of the Invention]
[0016] The present invention can provide a novel technique that allows for easy understanding of the status of a target video for review purposes. [Brief explanation of the drawings]
[0017] [Figure 1] Image of the system during shooting [Figure 2] System configuration block diagram [Figure 3] Hardware configuration diagram [Figure 4] Data structure diagram [Figure 5] Flowchart showing the process flow [Figure 6] A diagram showing an example of a screen display [Figure 7] A diagram showing an example of a screen display [Figure 8] A diagram showing an example of a screen display DETAILED DESCRIPTION OF THE INVENTION
[0018] The present embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which preferred embodiments are shown, which may, however, be embodied in many different forms and are not limited to the embodiments set forth herein.
[0019] For example, in the present embodiment, an educational support system and the like will be described, but similar effects can be achieved by a method, an apparatus, a computer program, a computer recording medium, etc. The program may be provided as a non-transitory computer-readable recording medium, or may be provided so as to be downloadable from an external server, etc.
[0020] In the embodiment described below, as shown in Fig. 1, a classroom in which a lecturer is teaching students is filmed using a filming device 3 such as a camera, and a user such as a lecturer reviews the video while reviewing it, thereby improving the content of the lesson and their own abilities. Note that in this embodiment, an example is described in which a lesson is filmed and reviewed in order to improve the quality of education at educational institutions such as cram schools. However, the subject of review is not limited to lessons; for example, a seminar, presentation, event, or the like may be filmed and reviewed as long as reviewing the video can lead to improvement.
[0021] FIG. 1 is a diagram illustrating an example of a system according to an embodiment capturing a classroom in which a class is being held. The camera device 3 is arranged so as to be able to capture one or more people in the classroom and is configured to be able to communicate with the education support device 1 and the database DB via a network NW. As shown in FIG. 1, the classroom contains students, a lecturer, and a board displaying the lesson content, such as a blackboard. Multiple camera devices 3 are provided, including a forward-facing camera device 3 that captures the students attending the class and a rear-facing camera device 3 that captures the lecturer and a board, such as a blackboard. In this embodiment, the rear-facing camera device 3 that captures the students attending the class captures images with an angle of view that captures only the students, while the forward-facing camera device 3 that captures the lecturer capturing the lesson captures images with an angle of view that captures only the lecturer, not the students. In this embodiment, the board displaying the lesson content is a blackboard, but it may also be a whiteboard, a screen displaying materials, a monitor, or the like. In this embodiment, the education support device 1 performs processing using video (target video) captured by the camera device 3 arranged in each classroom, as shown in FIG. 1.
[0022] Fig. 2 is a block diagram showing the configuration of a system according to an embodiment. As shown in Fig. 2, the education support system 0 includes an education support device 1, a user terminal 2, and an image capture device 3, and is configured to be connectable via a network NW. The education support system 0 also includes a database DB that stores various data used in the processing described below, and is connectable via the network NW.
[0023] The education support device 1 is a device such as a server device that executes an education support method, and is configured to realize the functional configuration described below. As the education support device 1, one or more general-purpose server devices or computer devices such as personal computers can be used.
[0024] The user terminal 2 is a terminal device operated by a user such as an instructor who uses the system. Terminal devices such as a personal computer, smartphone, or tablet terminal can be used as the user terminal 2. The user terminal 2 connects to the education support device 1 by using various programs (education support device usage programs) such as a browser application or a client application, and executes the processes related to education support described below. The education support device usage program may be a browser application pre-installed or downloaded in advance to the user terminal 2, or may be a client application provided by a program providing device (not shown) and downloaded. Furthermore, there may be multiple user terminals 2.
[0025] The image capturing device 3 is equipped with an imaging element such as a camera and captures the target video. It transmits data related to the target video to the education support device 1 in real time or in a recorded format via the network NW. In this embodiment, the image capturing device 3 is connected to the network NW wirelessly or via a wired connection to facilitate acquisition of the target video. However, it does not need to be connected to the network NW if the target video can be acquired using a portable recording medium or the like. In this embodiment, the image capturing device 3 is a fixed camera installed on the ceiling or wall of a classroom and captures video from a fixed viewing angle all the time or only during class hours. However, it may also be a device equipped with a movable imaging element, such as a smartphone or tablet terminal. In this embodiment, the image capturing device 3 captures the target video of a class based on a class schedule (schedule information) registered in advance by the instructor or the system administrator. However, the image capturing may be performed by pressing a capture button or by being instructed by a timer or the like. The image capturing device 3 may also be operated by instructions received wirelessly or via a wired connection using a controller or the like.
[0026] The network NW is an IP (Internet Protocol) network, but there are no restrictions on the type of communication protocol, the type of network, etc.
[0027] The database DB stores various data necessary for processing related to educational support, which will be described later. In this embodiment, the database DB is configured by a database server accessible via a network NW including an IP network or the like, but may also be configured using, for example, the processing unit 101 and the storage unit 102 that configure the educational support device 1. In this embodiment, the database DB is configured by one or more computer devices such as server devices.
[0028] Hereinafter, the hardware of the education support device 1 and the terminal device 20 (user terminal 2, photographing device 3) will be described with reference to FIG.
[0029] Fig. 3(a) is a hardware configuration diagram of the education support device 1. As shown in Fig. 3(a), the education support device 1 includes a processing unit 101, a storage unit 102, and a communication unit 103, which are used to perform the functions of each unit and each process.
[0030] The processing unit 101 has a processor such as a CPU that can execute an instruction set, and executes an OS, an educational support program, and the like. The storage unit 102 includes a volatile memory such as a RAM capable of storing an instruction set, and a non-volatile recording medium such as an HDD or SSD capable of recording an OS, an educational support program, and the like. The communication unit 103 has an interface for connecting to the network NW, and controls communication with the network NW to input and output information.
[0031] Fig. 3(b) is a hardware configuration diagram of the terminal device 20 (user terminal 2, photographing device 3). As shown in Fig. 3(b), the terminal device 20 (user terminal 2, photographing device 3) has a processing unit 201, a storage unit 202, a communication unit 203, and an output unit 205, which are used to exert the effects of each unit and each process.
[0032] The processing unit 201 has a processor such as a CPU that can execute an instruction set, and executes programs such as an OS and a program for using the educational support device. The storage unit 202 has a volatile memory such as RAM capable of storing an instruction set, and a non-volatile storage medium such as an HDD or SSD capable of recording an OS and an application program that can use the education support system (such as an education support device usage program). In this embodiment, the storage unit 202 of the image capture device 3 stores the target video for a certain period of time, but may also be configured to send the target video to a database DB and delete it immediately. The communication unit 203 has an interface for connecting to the network NW, and controls communication with the network NW to input and output information. Note that if the target video can be provided using a portable recording medium, the image capturing device 3 does not need to be equipped with the communication unit 203 that can be connected to the network NW. The input unit 204 has input devices such as an operation input device capable of input processing, such as a touch panel or keyboard, and an image input device, such as a camera, capable of image input. The output unit 205 has an output device such as a display device capable of display processing such as a display etc. Note that the image capturing device 3 does not necessarily have to include the output unit 205.
[0033] <Data structure in Educational Support System 0> An example of the configuration of data used in the education support system 0 in this embodiment will be described below with reference to Fig. 4. In this embodiment, the database DB stores various necessary data, which will be described later, but the storage unit 102 of the education support device 1 may store the necessary data.
[0034] 4(a) shows an example of the camera information. The camera information is information about the camera 3 that is connected to the network NW and captures images of a place (classroom) where lessons or the like are held, and in this embodiment, includes identification information (device ID) of the camera 3 and the installation position of the camera 3 (in front of or behind the classroom). In this embodiment, the camera information is stored in association with location information about the place where the camera 3 is installed (a classroom in this embodiment) using identification information such as a location ID.
[0035] 4(b) shows an example of location information. The location information is information about the location where the image capture device 3 is located, and in this embodiment, includes location identification information (location ID), the name of the classroom, the school building in which the classroom is located, and the district in which the school building is located.
[0036] In this embodiment, the location information and the photographing device information are registered in advance by the system administrator and stored in the database DB.
[0037] An example of user information is shown in Figure 4(c). The user information is information about users, such as instructors, who use the system, and includes the user's identification information (user ID), username, and email address. In this embodiment, the user logs in in advance using an email address, password, etc., so that the terminal operated by a specific user can be identified.
[0038] FIG. 4(d) shows an example of schedule information. The schedule information is information about the schedule of an event in a classroom (in this embodiment, a class in a classroom), and includes a start time and an end time. In this embodiment, the image capture device 3 captures an event such as a class held in a classroom in accordance with the schedule information, and stores the video data (target video) in a database DB. In this embodiment, processing related to posture estimation and calculation of upper body flexion degrees is performed for people appearing in the target video. The schedule information is also stored in association with location information and user information using identification information including identification information (such as a location ID) of the location where the image capture device 3 is located and identification information (such as a user ID) of the user. Although not shown in FIG. 4(d), the schedule information in this embodiment includes the grade, class category, subject, and unit of the class of students taking the class.
[0039] The database DB also stores target videos. The target videos are videos of an event (a class in this embodiment) that occurred at the location where the camera device 3 was installed, captured using the camera device 3. In this embodiment, the target videos are videos of one or more people present at the location where the event occurred, such as students and instructors attending a class in a classroom. Processing such as posture estimation is performed on the people captured in the target videos. In this embodiment, a target video captured for a certain class is stored indirectly associated with schedule information for the class via location information and camera device information. The target video also includes images, audio, and video playback time for each frame. In the processing related to posture estimation and situation label assignment, which will be described later, information related to the estimated posture (such as position coordinates for each body part and posture label), upper body flexion degree, and situation label are stored in association with the playback time of the frame used for estimation. This playback time indicates the time information of the frame relative to the playback time of the entire video, for example, "10:10:00 / 20:00:00", and is used to accurately refer to the person's posture at a specific point in time in subsequent processing and display.
[0040] In this embodiment, a status indicating the progress of the posture estimation process by the posture estimation unit 13 and the assignment of a situation label by the situation estimation unit 14 is attached to the target video. At this time, a status such as "analysis completed" is attached to a target video for which processing has been completed, and "analyzing" is attached to a target video for which processing is in progress, but the content of the status is not limited to those described here. Note that the education support device 1 can identify the area of the face of a person in a target image acquired by the image capture device 3 installed in a classroom facing backward to capture an image of a student, and perform processing to conceal the face so that it cannot be recognized, such as by applying mosaic or blurring to that area.
[0041] <Functional configuration> The functional configuration of the education support device 1 in this embodiment will be described below. As shown in Fig. 2, the education support device 1 includes a reception unit 11, an acquisition unit 12, a posture estimation unit 13, a situation estimation unit 14, a display unit 15, and a code acquisition unit 16. Note that the functional configuration may be realized by executing part of the processing described below in another computer device such as a user terminal 2 or a server device.
[0042] <Reception Section 11> The reception unit 11 receives various necessary information, such as data input to the user terminal 2 and stores it in the database DB, or performs processing to pass it on to other functional components. In this embodiment, the reception unit 11 receives input from the user terminal 2, such as schedule information, and stores it in the database DB. The reception unit 11 may also receive imaging device information, location information, and the like in advance and store it in the database DB.
[0043] The reception unit 11 receives a comment and a comment type for a target video input by a user, and stores the comment in the database DB in association with the target video. The comment type is a type of comment, and in this embodiment, includes "Like" for praising a good point, "Get better" for giving advice, "Note" for leaving a marker, and "Help" for asking other users for advice. In this embodiment, the reception unit 11 associates the comment received from the user terminal 2 with information about the actual date and time when the comment was made (comment date and time information), user information about the user who made the comment, and video time designation information, and stores them in the database DB. The video time designation information is information for designating a certain point in the target video, such as information indicating how far the target video has been played (e.g., 00:20:35 / 00:44:59), and is information about the location in the target video where a timestamp related to the comment is placed. In this embodiment, the video time designation information is information about the part of the target video that is displayed on the video playback screen when the comment is entered, such as information specifying the 20-minute point in the video if the comment is made when the video is being played at the 20-minute point, but it may also be information about the part of the target video that is specified by the user.
[0044] It should be noted that the timestamp referred to here is a mark for specifying a certain point in a video, and is displayed on the screen, and when pressed, the video at the specified point can be played. In this embodiment, the receiving unit 11 registers the part of the target video being played at the time the comment is input as the part specified by the timestamp in association with the comment, but it may also receive a specification of the part of the target video from the user.
[0045] <Acquisition part 12> The acquisition unit 12 performs a process of acquiring target moving images from the camera device 3 and storing them in the database DB. In this embodiment, the acquisition unit 12 acquires target moving images from the camera device 3 placed in a location (classroom) specified in the schedule information so as to capture moving images from a start time to an end time based on registered schedule information. At this time, the acquisition unit 12 identifies the target camera device 3 by referring to the camera device information and location information, and acquires the target moving images. Furthermore, the acquisition unit 12 associates the schedule information used when issuing the instruction with the acquired target moving images and stores them in the database DB; however, if the target moving images are stored in association with the camera device information and indirectly associated via the location information and the camera device information, they do not need to be stored in association.
[0046] <Posture estimation unit 13> The posture estimation unit 13 uses a posture estimation model to estimate the postures (poses) of one or more people appearing in the target video captured by the imaging device 3. The posture estimation unit 13 estimates the people and their postures appearing in the target video using an image recognition algorithm such as a convolutional neural network (CNN) and outputs the estimation result. In this embodiment, the posture estimation result by the posture estimation unit 13 is expressed as information including two-dimensional position coordinates of body parts and posture labels. However, it may also be expressed by one or more feature quantities such as vectors or labels. As a posture estimation process, the posture estimation unit 13 identifies the positions of each body part of one or more people appearing in the target video and assigns a posture label based on the identified positions. Note that in this embodiment, to estimate the posture of a sitting person, the posture estimation unit 13 performs posture estimation process only on the upper body. However, posture estimation process including identifying the position and assigning a posture label may also be performed on the entire body or parts of the body other than the upper body. In this embodiment, the posture estimation unit 13 identifies the positions of body parts as two-dimensional position coordinates along the x and y axes. However, the positions of body parts may be identified as three-dimensional position coordinates along the x, y, and z axes. The posture estimation unit 13 may perform a process using a skeleton, as described below, to identify each body part in the image. Alternatively, the posture estimation unit 13 may perform a process for detecting each body part in a pixel unit, such as semantic segmentation, or a process for detecting each region, such as a process for detecting an object using a bounding box. The posture estimation unit 13 performs a process for detecting a person appearing in the target video in order to calculate the upper body undulation, as described below. In this case, the posture estimation unit 13 may perform a process for detecting a person using a bounding box or the like before the process for identifying the person's parts, or may detect a person by appropriately connecting the identified body parts.
[0047] In this embodiment, the posture estimation unit 13 estimates the posture of a person appearing in a target video by using a posture estimation model that has learned the process of estimating key points of a person from images such as videos or still images and generating a skeleton by connecting the key points, as a process of recognizing images and estimating postures. The posture estimation model is a model configured using an image recognition algorithm such as CNN. In this embodiment, the posture estimation model and / or parameters of the posture estimation model are stored in the storage unit 102 of the education support device 1, but may also be stored in a separate server device or the like.
[0048] Key points estimated using a posture estimation model are important points for estimating a person's posture, such as joints such as elbows and knees, the corners of the eyes, and the start and end points of the nose. In this embodiment, to estimate the posture of a seated person (particularly a student), key points of the upper body (in this embodiment, the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, and neck) are estimated. In this embodiment, the posture estimation unit 13 outputs a set of position coordinates for each key point as the key point estimation result (for example, a set of position coordinates such as nose (x, y), left eye (x, y), ...). By connecting these key points, a skeletal structure (skeleton) that can represent the posture of a person, such as arms, legs, and eyes, is generated. A skeleton is a skeletal structure formed by line segments connecting key points and represents the posture of the human body. Note that, as long as the posture estimation unit 13 can connect appropriate key points to represent the posture of a human body, such as connecting key points indicating elbows and key points indicating shoulders, the posture estimation unit 13 may first identify the area in which each person appears in the target video for each person, and then estimate and connect key points for each person present in the identified area to estimate a skeleton. Alternatively, the posture estimation unit 13 may estimate key points appearing in the target video, and then connect appropriate key points representing the structure of the human body to estimate the skeleton of each person and identify the person. Furthermore, in this embodiment, the posture estimation unit 13 counts the number of detected people (students) to calculate the upper body undulation degree described below. Furthermore, in order to continue tracking the movement of the same person over time, the posture estimation results of each person (such as the position coordinates of body parts) output by the posture estimation unit 13 may be accompanied by a person label as identification information for identifying each person.
[0049] In this embodiment, the target video is a video captured during a class, as shown in FIG. 1 , and therefore, multiple people (students, etc.) often appear in the video. Furthermore, during a class, if students' upper bodies are often raised, it is highly likely that the teacher is giving a lecture, and if students' upper bodies are often lowered, it is highly likely that they are practicing. Therefore, the posture estimation unit 13 uses a posture estimation model to output a skeleton for each person appearing in the target video, and assigns a posture label indicating whether the upper body is raised or lowered. In this embodiment, the posture estimation unit 13 assigns a posture label to each person based on the person's skeleton obtained using the posture estimation model, assuming that the upper body is raised if the person's eyes are above the shoulders, and that the upper body is lowered otherwise. The posture labels may be assigned using a model trained using a skeleton connecting keypoints as input data and the assigned labels (posture labels) as output data, or may be assigned when the positional relationship of the skeletons satisfies a specific condition (for example, the skeleton representing the eyes is positioned above the skeleton representing the shoulders).
[0050] The posture estimation unit 13 calculates the upper body undulation degree based on the assigned posture label. The upper body undulation degree is an index showing the proportion of people (students in this embodiment) in the classroom who have their upper bodies raised or lowered, and in this embodiment, if all people in the classroom have their upper bodies raised, the index is 100 (100%). Note that if only one person is shown, the posture estimation unit 13 may estimate values such as the angle of the upper body and the distance between the eyes for that person and use them as the upper body undulation degree, or if multiple people are shown, the upper body undulation degree may be the average of the values of the degree of upper body undulation calculated for each person.
[0051] In this embodiment, posture estimation unit 13 finally outputs the upper body flexure as information related to the estimated posture of the person. However, it may output a skeleton connecting key points, or may output further information obtained using a skeleton connecting key points such as the tilt angle of the person's upper body or a posture label related to the posture (upper body raised, upper body lowered, sitting, standing, etc.). Furthermore, posture estimation by posture estimation unit 13 is not limited to a method of estimating a skeleton that indicates the bones of the human body. For example, the posture estimation model may be a model capable of recognizing a person and their face appearing in a target image, and posture estimation unit 13 may recognize the person appearing in the target image and determine whether the person's face is raised (whether the face from the eyes to the mouth can be recognized), and if the person can be recognized, it may label the person's situation as having their upper body raised. If some or all of the face features cannot be recognized, it may label the person's situation as having their face lowered, and calculate the upper body flexure.
[0052] <Situation Estimation Unit 14> The situation estimation unit 14 assigns a situation label to each time segment of the target video based on the posture of the person estimated from the target video. In this embodiment, the situation estimation unit 14 divides the target video into sections based on the upper body flexure estimated as a posture by the processing in the posture estimation unit 13, dividing the time segment of the target video in which the frequency of periods with high upper body flexure is equal to or greater than a threshold into "lecture" and "exercise" sections in which the frequency of periods with low upper body flexure is equal to or greater than a threshold, and assigns a situation label to each section. Furthermore, the situation estimation unit 14 may assign a situation label to each time segment based on the posture label assigned as a posture, provided that a predetermined number or a predetermined percentage or more of people with a specific posture label are present. Furthermore, the situation estimation unit 14 may determine the time segment by assigning a situation label to each frame, or may assign a situation label to each predetermined time unit (e.g., 10 seconds).
[0053] In this embodiment, the posture estimation unit 13 performs the above-described posture estimation process for each frame of the target video or every predetermined time (e.g., 1 second). At the same time, the situation estimation unit 14 also performs a process for assigning a situation label for each frame or every predetermined time. The posture estimation results (body part position coordinates, posture label, and upper body flexion degree) output by the posture estimation unit 13 and the situation label determined by the situation estimation unit 14 are stored in association with information indicating a time point in the target video (e.g., the playback time of the video). By using the estimation results and situation labels associated with the time point in the target video, the display unit 15 displays the situation labels, upper body flexion degrees, etc. in chronological order. The situation estimation unit 14 determines the temporal division for each situation label by connecting frames or predetermined time units assigned with the same situation label as divisions assigned with the same situation label and dividing frames or predetermined time units assigned with different situation labels into different divisions. After completing the process of estimating the posture for the entire target video, the situation estimation unit 14 assigns a situation label to the target video for each time segment based on the estimated posture. In addition, in this embodiment, the posture estimation unit 13 and the situation estimation unit 14 perform processing using the target video that is a target video of students and that was captured by a camera device 3 installed facing backward in the classroom, but they may also perform processing using the target video that was captured by a camera device 3 installed facing forward in the classroom.
[0054] In the above-described embodiment, the posture is estimated using a target video captured from a single angle. However, to achieve more accurate posture estimation, the posture may be estimated using multiple target videos captured by capturing a person from different angles. In this case, the posture estimation unit 13 may improve the accuracy of identifying body parts such as a skeleton by using target videos captured from multiple angles. Furthermore, the situation estimation unit 14 may estimate the upper body flexure degree for each of the target videos captured from multiple angles and estimate the upper body flexure degree by averaging the estimated values.
[0055] Moreover, posture estimation unit 13 repeats the process related to posture estimation at predetermined timing intervals (every 10 frames in this embodiment) to estimate the posture until the end of the target moving image.
[0056] <Display section 15> The display unit 15 performs processing to display a screen based on various necessary data such as data stored in the database DB and output from the posture estimation unit 13 and the situation estimation unit 14.
[0057] <Code Acquisition Section 16> In this embodiment, a board placed in a classroom displays a code indicating the materials (teaching materials, pages thereof, etc.) currently being used in the lesson, along with notes and materials indicating the lesson content. The code may be a character string, a one-dimensional code, or a two-dimensional code such as a QR code (registered trademark). The code may be printed on paper and posted on the board, which may be photographed by the camera device 3. Alternatively, if the board is a screen or monitor, the code may be displayed on the board and photographed by the camera device 3. The code acquisition unit 16 recognizes the code displayed in the target video, identifies the playback time of the target video in which the code is displayed, and acquires code information from the code. The code acquisition unit 16 identifies the target material and acquires corresponding material information based on the acquired code information. The code is a two-dimensional code, such as a QR code (registered trademark), printed or displayed for each material or page of material presented in the lesson, and is a code for identifying the material and / or its page currently being presented. The code information is information acquired from a code displayed on the board and photographed, and is information for identifying the currently displayed material and / or page of the material. In this embodiment, the code information includes identification information and the page number indicating the material, but may also include identification information indicating the material and / or the page of the material, or the name of the material, as long as the material and / or the page of the material can be identified. The material information is information about materials used in classes, and includes the material name and the page number, and is stored in the database DB. The material information may also include information about the subject or unit of the class in which the material is used, the class in which the material is used (e.g., grade, class category (highest class), etc.). If the code information includes information about the material name, the subject or unit of the class, etc., and is configured so that all information about the material can be obtained from the code, the database DB does not need to store the material information.
[0058] The code acquisition unit 16 acquires code information including identification information of the material and / or the page of the material from the code displayed in the target video, and identifies the material or its page that is being presented at a certain point in time in the target video (the point at which the code is displayed or the display of the code starts). In this way, the code acquisition unit 16 recognizes the code and identifies the material and the point in time that is being displayed, so that a timestamp corresponding to the material presented by the teacher can be placed in the target video at a point corresponding to the point at which the material was presented.
[0059] In addition, if the unit, subject, etc. identified from the information regarding the material acquired by the code acquisition unit 16 from the code is different from the unit or subject previously registered in the calendar information, the code acquisition unit 16 may change the unit, subject, etc. registered in the calendar information based on the information acquired from the code.
[0060] <Processing flow of educational support method> The processing flow of the education support method will be described below with reference to Fig. 5. Note that the processing flow described below is merely an example, and the processing may be realized by a different flow.
[0061] Before a class is held, the reception unit 11 receives schedule information entered via the schedule registration screen W1 and stores it in the database DB. Below, an example of the schedule registration screen W1 displayed by the display unit 15 when registering schedule information will be described with reference to FIG. 6.
[0062] The schedule registration screen W1 shown in FIG. 6 is a screen for registering a schedule and can accept schedule information including the school building, classroom, start date and time, class time (photography time), grade, class category, subject, instructor, and unit. In this embodiment, when a school building is selected, a list of classrooms included in the school building is displayed, allowing a classroom to be selected from the list. To determine the time period during which photography will be performed by the photography device 3, the schedule registration screen W1 may be configured to allow the user to input an end date and time instead of the class time. In this embodiment, the schedule registration screen W1 is also configured to allow the user to input whether or not repetition is required. When a repetition setting is entered, schedule information that is the same except for the date, such as the time period and classroom, is registered, with only the date changed. The date in the schedule registered by the repetition setting may be the same day of the week in the week following the date entered during the initial registration, or the same day in a different month.
[0063] In this embodiment, the camera devices 3 installed in each location (classroom) capture target videos based on registered schedule information. The acquisition unit 12 acquires target videos from the camera devices 3 and stores them in the database DB in association with the schedule information (S101). The camera devices 3 may capture target videos only during the time period registered in the schedule information and provide the target videos, or may capture videos at all times and provide only target videos during the time period registered in the schedule information. Furthermore, when multiple camera devices 3 are installed in the same location (classrooms in this embodiment), such as before and after a classroom, target videos are captured and acquired by the multiple camera devices 3 based on one schedule information. In this case, multiple target videos, such as a target video captured by a camera device 3 installed facing the front of the classroom and a target video captured by a camera device 3 installed facing the back of the classroom, are stored in association with each other in one schedule information.
[0064] The posture estimation unit 13 estimates a posture using a posture estimation model based on the acquired target video (S102). In this embodiment, the posture estimation unit 13 performs the following processes to estimate the posture of a person appearing in the target video: outputting a skeleton indicating the posture of the person using the posture estimation model, assigning a posture label to each person based on the skeleton, and determining a body flexion degree based on the posture label. The situation estimation unit 14 assigns a situation label to each time segment of the target video based on the posture estimated by the posture estimation unit 13 (in this embodiment, the body flexion degree) (S103). When the assignment of situation labels for each time segment as described above has been completed for the entire target video, the user terminal 2 sends a request to display a video playback screen (S201), and the display unit 15 displays the video playback screen (S104).
[0065] An example of the video playback screen W2 displayed in S104 will be described below with reference to FIG.
[0066] The video playback screen W2 is a screen for playing the target video, and includes a video information display area W21, a video playback area W22, a timeline display area W23, a comment type selection area W24, a comment display area W25, and a comment input area W26.
[0067] The video information display area W21 is an area that displays information about the target video being played, and in this embodiment, displays the name of the teacher teaching the lesson, the district, the classroom name, the class category, the number of viewers, and the synchronization rate value related to the degree of synchronization of the postures of the students in the class. In this embodiment, the display unit 15 displays information about the target video shown in the video information display area based on schedule information and the like that is stored in the database DB in association with the target video.
[0068] The video playback area W22 is an area in which the target video is played. When target videos are shot in multiple directions in the same location, such as when multiple camera devices 3 are installed in the same classroom, including a rear-facing camera device 3 that shoots the students and a forward-facing camera device 3 that shoots the teacher, the video playback area W22 displays the multiple target videos by displaying one of the target videos in a smaller size, as shown in FIG. 7. When the switch button displayed in the upper right corner of the video playback area W22 is pressed, the target video displayed in a larger size is switched to the target video displayed by another camera device 3 (for example, from the target video that shot the back of the classroom to the target video that shot the front of the classroom). Furthermore, when multiple target videos are displayed, using a seek bar or the like to operate the playback point of a target video, multiple target videos that were shot at the same time (the playback time of the target videos is the same, or the registered shooting date and time is the same) are played in the video playback area W22.
[0069] The timeline display area W23 is an area for displaying various information about the target video in chronological order, and in this embodiment, it is an area that displays, along the time of the target video, a seek bar indicating the playback point of the target video being played in the video playback area W22, a graph of the upper body flexion degree, status labels for each time division of the target video that are superimposed on the graph of the upper body flexion degree, and a graph of the synchronization rate indicating the degree of synchronization of the students' posture.
[0070] The seek bar is a screen display element for controlling the playback position of the target video and can be moved left and right. The upper body flexibility graph is a graph that shows the upper body flexibility values over the time of the target video, with the higher the value, the closer it is to 100 (100%). In this embodiment, the situation labels for each time segment that are displayed superimposed on the upper body flexibility graph include "narration," "exercise," etc. In this embodiment, the situation label is assigned to a time segment with a high frequency of high upper body flexibility values, and to a time segment with a low upper body flexibility value, as "exercise." In this embodiment, the situation labels for each segment are displayed superimposed on the graph showing the upper body flexibility values and the segment. The segments are distinguished by color.
[0071] The timeline display area W23 also displays a timestamp indicating the portion of the material to be displayed (e.g., page 2 of the material), based on information acquired by the code acquisition unit 16 from the code displayed in the target video. The timestamp is placed in the timeline display area W23 at a position corresponding to the playback time specified by the timestamp (e.g., if the timestamp is for specifying 21 minutes, the position indicating the playback time is 21 minutes). Furthermore, a timestamp indicating the portion where the comment was entered, which is generated when a comment is entered, is also displayed in the timeline display area W23. The timestamp displayed based on the comment is displayed as an icon corresponding to the comment type (in this embodiment, "◎", "△", "◯", or "?") at a position corresponding to the specified playback time. When these timestamps are pressed, the display unit 15 plays the target video for the playback time specified by the timestamp in the video playback area W22.
[0072] The comment type selection area W24 is an area for selecting the type of comment to be displayed in the comment display area W25, and in this embodiment, is an area that displays check boxes for inputting whether or not to display comments for each comment type, an icon for each comment type, and the number of comments for each comment type. Comments of the comment type selected in the comment type selection area W24 are displayed in the comment display area W25. Furthermore, only the timestamps of the comment type selected in the comment type selection area W24 may be displayed in the timeline display area W23.
[0073] The comment display area W25 is an area for displaying comments received from the user terminal 2, and in this embodiment, displays an icon indicating the comment type, the playback time of the target video specified at the time the comment was entered, the name of the user who made the comment, and the comment itself. Furthermore, the comment display area W25 displays only comments of the comment type selected in the comment type selection area W24.
[0074] The comment input area W26 is an area for accepting input of a comment, and allows input of the comment, the range of users who can view the comment, and the comment type. The accepting unit 11 accepts the comment, the range of users who can view the comment, and the comment type input in the comment input area W26. Furthermore, based on the viewing range set in the comment input area W26, comments corresponding to the comment display area W25 of users who fall within the viewing range are displayed.
[0075] In this embodiment, the education support device 1 further includes a search unit that searches for target videos. The search unit searches for target videos based on search criteria entered on the video search screen W3 and extracts target videos that meet the criteria. At this time, the search unit identifies target videos that meet the search criteria using the target videos, schedule information associated with the target videos, location information directly or indirectly associated with the schedule information, etc. The display unit 15 displays a list of target videos that meet the search criteria based on the target videos that meet the search criteria extracted by the search unit.
[0076] An example of the screen display of the video search screen W3 will be described with reference to Fig. 8. The video search screen W3 is a screen for searching for a target video, and includes a search condition input area W31 and a video list display area W32.
[0077] The search condition input area W31 is an area for inputting search conditions. In this embodiment, the search condition input area W31 allows for inputting conditions related to items such as date, instructor, grade, class category, subject, unit, and location (district, school building, classroom), which correspond to the items input in the schedule information. In addition, in this embodiment, the search condition input area W31 is configured to allow for input of search conditions including items such as the number of comments and likes, camera location, status (analysis in progress or analysis completed), whether or not transcription is included, and keywords. After the search conditions are input, when the "Search button" displayed in the search condition input area W31 is pressed, the search unit identifies and extracts target videos that match the search conditions.
[0078] The video list display area W32 is an area for displaying a list of information related to the target video. In this embodiment, the area displays the school building, district, grade, class category, subject, unit, instructor, number of viewers, number of likes, index value for lesson evaluation (synchronization rate), recording date and time, number of comments by comment type, status, and camera position based on schedule information and comment information associated with the target video. In this embodiment, information related to the target video that matches the search criteria extracted by the search unit is displayed in a list. However, if no search criteria are entered in the search criteria input area W31 or a search for the target video is performed, the video list display area W32 may simply display a list of the target video. In this embodiment, the information items displayed in the video list display area W32 correspond to the items that can be entered in the search criteria input area W31, but they may also be different items.
[0079] Furthermore, the faces of each user (instructor) may be registered in advance. In this case, the education support device 1 may further include a person identification unit, and may identify the user (instructor) appearing in the target video by using a person identification model. The person identification model is a model that has learned to identify a person by identifying facial features of a person from a still image or video. Furthermore, if the user identified as appearing in the target video is different from the user registered in the calendar information as the instructor who will teach the class, the person identification unit may modify the calendar information by setting the user identified as appearing in the target video as the instructor. [Explanation of symbols]
[0080] 0 Educational Support System 1 Educational support equipment 11 Reception 12 Acquisition Department 13 Posture estimation section 14 Situation Estimation Unit 15 Display 16 Code Acquisition Section 2. User terminal 3. Imaging equipment NW Network DB Database
Claims
1. An educational support system, The device includes an acquisition unit, a situation estimation unit, a display unit, and a posture estimation unit, The acquisition unit acquires a target video, which is a video capturing one or more people present at a certain location; the posture estimation unit estimates a degree of upper body flexion as a posture of one or more persons based on the target video; the situation estimation unit assigns a situation label to each time segment of the target video based on the upper body flexure degrees associated with the postures of one or more people, the upper body flexure degrees being estimated based on the target video; the display unit displays a video playback screen including a video playback area for playing the target video and the status label for each category; Educational support system.
2. An educational support system, The apparatus includes an acquisition unit, a situation estimation unit, a display unit, a reception unit, and an imaging device, the reception unit receives schedule information regarding a class schedule including a classroom for the class, a start time and an end time of the class; the imaging device performs imaging based on the schedule information, and stores the acquired target video in association with the schedule information; The acquisition unit acquires the target video, which is a video capturing one or more people present at a certain location; the situation estimation unit assigns a situation label to each time segment of the target video based on a posture of one or more people estimated based on the target video; the display unit displays a video playback screen including a video playback area for playing the target video and the status label for each category, and the schedule information associated with the target video; Educational support system.
3. An educational support system, The device includes an acquisition unit, a situation estimation unit, a display unit, and a reception unit, The acquisition unit acquires a target video, which is a video capturing one or more people present at a certain location; the situation estimation unit assigns a situation label to each time segment of the target video based on a posture of one or more people estimated based on the target video; the display unit displays a video playback screen including a video playback area for playing the target video and the status label for each category; When the receiving unit receives an input of a comment, the receiving unit identifies a playback time of the target video at the time the comment was input, and stores the comment and the playback time in association with each other; the display unit displays a comment display area showing the comment on the video playback screen; When the comment is pressed, the display unit plays the target video at the time when the comment was associated in the video playback area. Educational support system.
4. An educational support system, The device includes an acquisition unit, a situation estimation unit, a display unit, and a code acquisition unit, The acquisition unit acquires a target video, which is a video capturing one or more people present at a certain location; the situation estimation unit assigns a situation label to each time segment of the target video based on a posture of one or more people estimated based on the target video; the display unit displays a video playback screen including a video playback area for playing the target video and the status label for each category; the code acquisition unit recognizes the code based on the target video in which a code capable of identifying a document and / or a page of the document is captured, identifies the time when the code was displayed, and acquires code information from the code; the display unit displays material information related to the material corresponding to the code information based on the acquired code information, in correspondence with the time when the code is displayed. Educational support system.
5. The education support system further includes one or more image capture devices, The imaging device is one or more devices placed in a classroom, and captures the target video, which is a video of one or more students participating in the class; the situation estimation unit assigns the situation label based on the posture of the one or more students; The education support system according to claim 1 .
6. the situation labels are displayed in sections along a timeline corresponding to the time of the target video; The education support system according to claim 1 .
7. The display unit further displays a graph showing the body undulation degree in time series. The education support system according to claim 1 .
8. A computer-implemented educational support method, comprising: The method includes an acquisition step, a situation determination step, a display step, and a posture estimation step, In the acquiring step, a target video is acquired, which is a video of one or more people present at a certain location; In the posture estimation step, a body flexion degree relating to the degree of flexion of the upper body is estimated as a posture of one or more persons based on the target moving image; In the situation determination step, a situation label is assigned to each time segment of the target video based on the upper body flexion degrees associated with the postures of one or more people, which are estimated based on the target video; In the display step, a video playback screen including a video playback area for playing the target video and the status label for each category is displayed. Educational support methods.
9. A computer-implemented educational support method, comprising: The method includes an acquisition step, a status determination step, a display step, and a reception step, In the receiving step, schedule information regarding a class schedule including a classroom, a start time, and an end time of the class is received; the imaging device performs imaging based on the schedule information, and stores the acquired target video in association with the schedule information; In the acquisition step, the target video is acquired, which is a video capturing one or more people present at a certain location; In the situation determination step, a situation label is assigned to each time segment of the target video based on the postures of one or more people estimated based on the target video; In the display step, a video playback screen including a video playback area for playing the target video and the status label for each category, and the schedule information associated with the target video are displayed. Educational support methods.
10. A computer-implemented educational support method, comprising: The method includes an acquisition step, a status determination step, a display step, and a reception step, In the acquiring step, a target video is acquired, which is a video of one or more people present at a certain location; In the situation determination step, a situation label is assigned to each time segment of the target video based on the postures of one or more people estimated based on the target video; In the display step, a video playback screen including a video playback area for playing the target video and the status label for each section is displayed; In the display step, a comment display area showing a comment is displayed on the video playback screen, In the receiving step, when an input of a comment is received, a playback time of the target video at the time the comment was input is identified, and the comment and the playback time are associated and stored; In the display step, when the comment is pressed, the target video at the time when the comment was associated is played in the video playback area. Educational support methods.
11. A computer-implemented educational support method, comprising: The method includes an acquisition step, a situation determination step, a display step, and a code acquisition step, In the acquiring step, a target video is acquired, which is a video of one or more people present at a certain location; In the situation determination step, a situation label is assigned to each time segment of the target video based on the postures of one or more people estimated based on the target video; In the display step, a video playback screen including a video playback area for playing the target video and the status label for each section is displayed; In the code acquisition step, the code is recognized based on the target video in which a code capable of identifying a document and / or a page of the document is captured, and the time point at which the code is displayed is identified, and code information is acquired from the code; In the display step, material information relating to the material corresponding to the code information is displayed based on the acquired code information, in correspondence with the time when the code is displayed. Educational support methods.
12. A program for causing the computer to execute the method according to any one of claims 8 to 11.
13. A recording medium storing a program for causing the computer to execute the method according to any one of claims 8 to 11.
Citation Information
Patent Citations
Information processor and information processing method
JP2008090570A
Device, system, server device and terminal device for content evaluation
JP2016218658A
Educational learning activity support system
JP2018032276A
Behavior estimation apparatus and behavior estimation program
JP2019053647A
Education support method, education support program, and education system
JP2019159152A