Information processing system, information processing method, and computer program
The system automates childcare quality assessment by analyzing video footage to detect gaze relationships, addressing behavioral changes and subjectivity in traditional methods, ensuring accurate and continuous evaluation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2025-09-15
- Publication Date
- 2026-05-15
AI Technical Summary
Existing childcare quality assessment methods face challenges such as changes in behavior due to evaluator presence, lack of specific evaluation situations, and individual evaluator subjectivity, making it difficult to accurately measure eye contact and joint attention between caregivers and children.
An information processing system that automatically detects gaze relationships like eye contact and joint attention by analyzing video footage using an eye-tracking network, integrating person estimation and gaze estimation to evaluate childcare quality.
Automates the assessment of childcare quality, reducing behavioral changes and subjectivity, and providing accurate, continuous evaluation of gaze relationships.
Smart Images

Figure JP2025032431_15052026_PF_FP_ABST
Abstract
Description
Information Processing System, Information Processing Method, and Computer Program
[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to an information processing system, an information processing method, and a computer program that perform processing for evaluating the quality of human communication (for example, the quality of childcare practice of childcare workers communicating with children).
[0002] In order to support children's development and improve childcare, it is essential to evaluate the quality of childcare practice. As "childcare quality assessment scales", for example, ECERS (Early Childhood Environment Rating Scale) and CLASS (Classroom Assessment Scoring System) are known. Among these, ECERS is a scale developed in the United States in 1980 for comprehensively measuring the quality of childcare. The original is a scale for measuring the quality of group childcare for children aged 3 and above, and it has 6 sub-scales and 35 items such as "indoor space" and "health and hygiene". These evaluation methods are generally carried out manually by trained evaluators visiting the site and observing. However, there are problems such as changes in the activities of children and childcare workers due to the entry of evaluators into the site, the absence of specific scenes for evaluation during observation, and individual differences among evaluators.
[0003] Further, there has been proposed a person-in-charge terminal device that includes a generation unit that generates a childcare image by photographing a childcare scene with an imaging device, a server processing activation unit that transmits the childcare image to a server and causes the server to process and form other person-in-charge opinion information obtained from other person-in-charge terminals corresponding to the childcare image, an acquisition unit that acquires the childcare image and the other person-in-charge opinion information corresponding to the childcare image from the server, and an output unit that presents the childcare image and the other person-in-charge opinion information corresponding to the childcare image to the person-in-charge, and is used by the person-in-charge of childcare to improve the evaluation ability for childcare and the like (see Patent Document 1).
[0004] Japanese Patent Application Laid-Open No. 2023-169233, Japanese Patent Application Laid-Open No. 2014-171586
[0005] Eunji et al. "Detecting Attended Visual Targets in Video"(arXiv:2003.02501v2 [cs.CV] 30 Mar 2020)
[0006] The purpose of this disclosure is to provide an information processing system, an information processing method, and a computer program that automate the acquisition of information that serves as a clue to evaluating the quality of human communication.
[0007] This disclosure has been made in consideration of the above-mentioned problems, and its first aspect is an information processing system comprising: a person estimation unit that estimates the person appearing in each frame of a video; a gaze estimation unit that estimates the gaze of the person appearing in each frame of a video; and a detection unit that combines the estimation results from the person estimation unit and the gaze estimation unit to detect a specific gaze relationship between people that occurred in each frame of a video.
[0008] However, the term "system" as used here refers to a logical collection of multiple devices (or functional modules that perform specific functions), without regard to whether each device or functional module resides within a single enclosure. In other words, both a single device consisting of multiple components or functional modules, and a collection of multiple devices, qualify as a "system."
[0009] The gaze estimation unit estimates gaze using a gaze estimation network. The detection unit then detects that at least one of the following gaze relationships has occurred: eye contact where two people make eye contact within a frame, joint attention where two people focus their attention on the same object, gaze sharing where joint attention occurs within a certain time after eye contact, eye region gaze where one person looks into the other's eyes, and body region gaze where one person looks into the other's body.
[0010] The information processing system relating to the first aspect further includes an evaluation unit that integrates specific gaze relationships between people detected by the detection unit in each frame of the video in the temporal direction and evaluates the people.
[0011] The information processing system relating to the first aspect further comprises a presentation unit that presents the detection results from the detection unit or the evaluation results from the evaluation unit. The presentation unit displays the gaze information corresponding to the specific gaze relationship superimposed on the video playback screen. The presentation unit also presents the evaluation results from the evaluation unit for a person specified on the video playback screen, or presents the temporal progression of the evaluation results from the evaluation unit.
[0012] Furthermore, a second aspect of this disclosure is an information processing method comprising: a person estimation step for estimating the person appearing in each frame of a video; a gaze estimation step for estimating the gaze of the person appearing in each frame of a video; and a detection step for detecting a specific gaze relationship between people that occurred in each frame of a video by combining the estimation results from the person estimation step and the gaze estimation step.
[0013] Furthermore, a third aspect of this disclosure is a computer program written in a computer-readable format to cause a computer to function as: a person estimation unit that estimates the person appearing in each frame of a video; a gaze estimation unit that estimates the gaze of the person appearing in each frame of a video; and a detection unit that combines the estimation results from the person estimation unit and the gaze estimation unit to detect a specific gaze relationship between people that occurred in each frame of a video.
[0014] The computer program relating to the third aspect of this disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes via a storage medium, communication medium, such as an optical disk, magnetic disk, semiconductor memory, or a network, in a computer-readable format. By installing the computer program relating to the third aspect of this disclosure onto a computer via any of these media, collaborative effects can be achieved on the computer, similar to the effects of the information processing system relating to the first aspect of this disclosure.
[0015] Figure 1 is a diagram showing an example of the functional configuration of the childcare quality evaluation system 100. Figure 2 is a flowchart showing the operation of the childcare quality evaluation system 100. Figure 3 is a diagram showing the results of person estimation within a frame. Figure 4 is a diagram showing the results of gaze estimation within the same frame as in Figure 3. Figure 5 is a diagram showing the eye contact judgment output obtained by integrating the estimation results shown in Figures 3 and 4. Figure 6 is a diagram showing the results of person estimation within a frame. Figure 7 is a diagram showing the results of gaze estimation within the same frame as in Figure 6. Figure 8 is a diagram showing the joint attention judgment output obtained by integrating the estimation results shown in Figures 6 and 7. Figure 9 is a diagram showing the LookEye judgment output. Figure 10 is a diagram showing the LookBody judgment output. Figure 11 is a diagram to explain the method of determining pairs of people from within a frame (when there are 3 subjects). Figure 12 is a diagram to explain the method of determining pairs of people from within a frame (when there are 4 subjects). Figure 13 shows an example of the configuration of a UI screen (video selection screen) for viewing information related to the quality assessment of childcare. Figure 14 shows an example of the configuration of a UI screen (showing a video selected on the video selection screen) for viewing information related to the quality assessment of childcare. Figure 15 shows an example of the configuration of a UI screen (video playback screen) for viewing information related to the quality assessment of childcare. Figure 16 shows an example of the configuration of a UI screen (video playback screen with superimposed joint attention detection results) for viewing information related to the quality assessment of childcare. Figure 17 shows an example of the configuration of a UI screen (video playback screen with superimposed eye contact detection results) for viewing information related to the quality assessment of childcare. Figure 18 shows an example of the configuration of a UI screen (video playback screen with superimposed eye area gaze detection results) for viewing information related to the quality assessment of childcare. Figure 19 shows an example of the configuration of a UI screen (screen for selecting a childcare worker) for viewing the results of the quality assessment of childcare. Figure 20 shows an example of the configuration of a UI screen (screen displaying the childcare quality score value of the selected childcare worker) for viewing the results of the quality assessment of childcare. Figure 21 shows an example of the configuration of a UI screen for viewing the evaluation results of childcare quality (a screen showing the temporal changes in childcare quality). Figure 22 shows an example of the hardware configuration of the information processing device.
[0016] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.
[0017] A. Overview A-1. Background and Challenges A-3. Application of Eye-Tracking Networks A-3. Application of Eye-Tracking Networks B. System Configuration B-1. Overall Configuration B-2. Detection of Eye-Tracking Relationships Between Individuals B-3. Quality Assessment of Childcare C. System Operation D. How to View Video Analysis Results and Childcare Quality Assessment Results E. Technical Points F. Hardware Configuration of Information Processing Device
[0018] A. Overview A-1. Background and Challenges In daycare centers, it is common practice to review the day's work, discuss the positive aspects and areas for improvement of the childcare workers' actions, and evaluate the quality of childcare. Generally, when evaluating the quality of human communication, eye contact is considered an important clue. In evaluating the quality of childcare, the extent to which childcare workers made eye contact with children and engaged in joint attention (paying attention to the same object as the child) can be incorporated as evaluation criteria.
[0019] For example, CLASS PRE-K (Preschool Kindergarten) is a quality assessment standard for childcare consisting of three domains: emotional support, classroom organization, and instructional support. In the emotional support domain, the quality of trust and support between caregivers and children is evaluated, and an important clue is the occurrence of eye contact and joint attention between caregivers and children.
[0020] Eye contact and joint attention are behaviors that indicate a person's interest and emotions, and are frequently used as behavioral indicators in psychology. Eye contact is when two people make eye contact with each other. When caregivers make eye contact with children, they can build a sense of security and trust. In other words, eye contact can be said to be an important element in evaluating the emotional connection and quality of communication between caregivers and children. Joint attention is when two people direct their attention to the same object. When caregivers and children focus their attention on the same object, it contributes to deeper learning and improved understanding. For example, when a caregiver reads a picture book to children, the scene of them looking at the book together and discussing it can be evaluated as an indicator of the quality of care. Also, since children observe where adults are looking and learn from it, it is important that joint attention occurs after eye contact. This is called gaze following and is an important indicator in the development of language and social skills. Thus, not only is it important whether eye contact or joint attention occurs at a certain time, but the temporal relationship between them (i.e., whether joint attention occurs within a certain time after eye contact) is also an important indicator of the quality of care.
[0021] Here, it is extremely difficult for childcare workers to accurately remember every single instance of eye contact with a child. Furthermore, there are several problems with the method by which evaluators count how often eye contact, joint attention, and other forms of eye contact occurred between childcare workers and children throughout the day. For example, the presence of the evaluator may influence the behavior of both the child and the childcare worker, the eye contact relationships to be evaluated may not appear during the evaluator's observation, and the evaluation may be influenced by the evaluator's subjectivity, oversights, and individual differences.
[0022] Installing cameras in the nursery school and reviewing the footage of childcare workers' actions during their work from the beginning to identify where eye contact, joint attention, and shared gaze occurred during their long workdays is a manual process, making it costly and difficult to implement on a continuous basis.
[0023] While a method of measuring gaze information by attaching a gaze detection device equipped with an eye camera to the head is known (see, for example, Patent Document 2), wearing such a wearable device may affect the behavior of children and childcare workers, and prolonged wear can be psychologically and physically burdensome. Furthermore, the cost of the device makes continuous implementation difficult. Moreover, such gaze detection devices themselves output independent gaze information for each wearer and do not directly output gaze relationships between people, such as eye contact or joint attention.
[0024] A-2. System Overview This disclosure addresses the challenges mentioned in Section A-1 above and is a technology for automatically detecting gaze relationships related to the evaluation of childcare quality from video footage captured by a camera. Specifically, the system to which this disclosure is applied is configured to automatically detect specific gaze relationships between people, such as eye contact and joint attention, by extracting gaze information from people appearing in videos of a nursery school and combining gaze information between different people.
[0025] Specifically, in a system to which this disclosure applies, a mobile device such as a smartphone is first installed as a fixed-point camera in the nursery school to acquire video footage of the actions of childcare workers and children during their work. Then, the acquired video data is analyzed to identify the people in the video and extract gaze information, such as where each person is looking. By combining the gaze information obtained for each person, it is then possible to detect whether a specific gaze relationship, such as eye contact or joint attention, occurred between those people.
[0026] In a system that applies this disclosure, the function for extracting gaze information from videos, as described above, can be combined with a function for identifying people within the video to accurately count who made eye contact or shared attention with whom and how many times. This information, counted for each childcare worker, can then be used to evaluate the quality of childcare.
[0027] A-3. Application of Eye-Tracking Networks The system described herein uses an eye-tracking network to extract eye-tracking information from a person. An eye-tracking network is a deep learning technique that estimates the starting point and point of focus of a gaze from images and videos. Basically, it performs eye-tracking by first detecting the face region in the image, then detecting the position of the eyes on the detected face, and then estimating the direction of the gaze from the position of the eyes and the orientation of the face. For example, an example of eye-tracking is performed using Attention Target Detection (ATN), a model based on a transformer architecture (see, for example, Non-Patent Document 1).
[0028] The gaze estimation network itself outputs independent gaze information for each frame and each person, and does not directly output gaze relationships between people such as eye contact or joint attention. Therefore, in a system to which this disclosure is applied, in each frame, the gaze information of each person extracted by the gaze estimation network and the person information (profile) identified by the person identification function are processed in the spatial direction to detect gaze relationships between people such as eye contact and joint attention. In other words, when multiple gaze information is provided within a single frame by the gaze estimation network, the spatial distance between the information of the start and end points (points of attention) of those gazes is taken into consideration to detect each gaze relationship such as eye contact, joint attention, gaze sharing, eye region gaze, and body region gaze.
[0029] Furthermore, in the system to which this disclosure is applied, the gaze relationships between individuals detected on a frame-by-frame basis are aggregated along the time axis and used for calculating the quality of childcare. In other words, the system to which this disclosure is applied enables the automatic calculation of the quality of childcare by aggregating the information output independently by the gaze estimation network for each frame and each individual in the spatial and temporal directions.
[0030] In short, according to this disclosure, by automating quality assessment in childcare using deep learning technology, it is possible to solve problems in existing childcare quality assessment methods, such as changes in the activities of children and childcare workers when evaluators are present on-site, the lack of specific situations for assessment during observation, and individual differences among evaluators.
[0031] The effects described herein are illustrative only, and the effects brought about by this disclosure are not limited to those described herein. Furthermore, this disclosure may produce additional effects beyond those described above. Further objectives, features, and advantages of this disclosure will become apparent from the more detailed descriptions based on the embodiments and accompanying drawings described later.
[0032] B. System Configuration B-1. Overall Configuration Diagram 1 shows an example of the functional configuration of the childcare quality evaluation system 100 to which this disclosure applies. The illustrated childcare quality evaluation system 100 includes an imaging device 110, an image analysis device 120, a quality evaluation device 130, and an information presentation device 140.
[0033] In the operational configuration of the childcare quality evaluation system 100 shown in Figure 1, the imaging device 110 and the information display device 140 are installed on the user's side, while the video analysis device 120 and the quality evaluation device 130 are built as a cloud server to provide video analysis and childcare quality evaluation services to multiple users. For example, the imaging device 110 is installed as a fixed-point camera in the nursery school. The information display device 140 may be installed in the nursery school director's office or in the room of the person in charge of managing or supervising the childcare workers, or the child's guardian.
[0034] In Figure 1, for the sake of simplification, only one each of the imaging device 110, video analysis device 120, quality evaluation device 130, and information display device 140 is depicted. However, in reality, a cloud server consisting of one set of video analysis device 120 and quality evaluation device 130 provides video analysis and childcare quality evaluation services to multiple imaging devices 110 (one for each fixed-point camera) and multiple information display devices 140 (one for each viewer). Of course, it is also conceivable that the imaging device 110, video analysis device 120, quality evaluation device 130, and information display device 140 are all installed in one location, and the childcare quality evaluation system 100 is configured as a standalone device.
[0035] The imaging device 110 is intended to utilize a camera mounted on a mobile phone such as a smartphone. Specifically, the smartphone camera is used as a fixed-point camera installed in a nursery room or similar location to acquire video footage of the actions of childcare workers and children during work hours, and upload it to a cloud server. The imaging device 110 transfers the captured video files to the cloud server using FTP (File Transfer Protocol), for example, but this disclosure is not limited to any specific transfer method. On the cloud server side, the videos uploaded from the imaging device 110 are stored in a video database 170.
[0036] The information display device 140 is a device for users to view information such as the results of analyzing videos uploaded from the imaging device 110 by the video analysis device 120, and the quality score of childcare calculated by the quality evaluation device 130 based on the video analysis results (detection results of eye contact and joint attention). The information display device 140 is, for example, an information terminal such as a smartphone, tablet, or personal computer (PC) used by the user. The information display device 140 presents information to the user via a web front-end system such as an internet browser, including detection results of gaze relationships such as eye contact, joint attention, gaze sharing, eye area gaze, and body area gaze, as well as the quality score of childcare, which are acquired from a cloud server.
[0037] The video analysis device 120 retrieves the video uploaded from the imaging device 110 from the video database 170, analyzes the video frame by frame of the video capturing the actions of the childcare workers and children during their work, detects the gaze relationships between people, including eye contact, joint attention, eye-area gaze, and body-area gaze, and saves them in the frame analysis history database 150. The video analysis device 120 includes a face recognition unit 121, a gaze estimation unit 122, an object estimation unit 123, a person estimation unit 124, and a person-to-person gaze relationship detection unit 125.
[0038] The face recognition unit 121 detects human faces in each frame of the video and performs face recognition on the detected faces. The face recognition results from the face recognition unit 121 are sent to the gaze estimation unit 122. The face recognition unit 121 may be configured using existing libraries that are stored, shared, and made publicly available through source code management services or the like.
[0039] The gaze estimation unit 122 receives the face recognition results from the face recognition unit 121 and estimates where each person in the video frame is looking. In this embodiment, the gaze estimation unit 122 is assumed to use a gaze estimation network, and detects the eye positions of the detected faces detected by the face recognition unit 121, and further estimates the gaze direction (start point and point of focus of the gaze) from the eye positions and face orientation. The gaze estimation unit 122 then sends the gaze estimation results, with gaze information attached to each detected face in the frame, to the person-to-person gaze relationship detection unit 125.
[0040] The gaze estimation network used by the gaze estimation unit 122 estimates the starting point and point of interest of the gaze from an RGB image, for example, Attention Target Detection (described above). The gaze estimation unit 122 may be configured using existing libraries that are stored, shared, and made publicly available through source code management services, etc.
[0041] The object estimation unit 123 has the function of estimating what the objects shown in each frame of the video are and makes a determination as to whether the object is a person or an object. The object estimation unit 123 estimates the object using existing algorithms that use deep networks, such as YOLO (You Only Look Once). The object estimation result from the object estimation unit 123 is sent to the person estimation unit 124. The object estimation unit 123 may be configured using existing libraries that are stored, shared, and made public in source code management services, etc.
[0042] The person identification unit 124 inputs the object estimation result by the object estimation unit 123 and estimates who the person appearing in each frame of the video is. The candidates for the person estimated by the person identification unit 124 are children and childcare workers registered in advance, assuming that they are registered in the person database 160 in advance. Also, the person identification unit 124 assumes that for unregistered persons, they will be registered in the person database 127 as guests. The person identification unit 124 may have a function of estimating whether the guests are the same person. Then, the person identification unit 124 sends the person identification result with person information (profile) attached to each person area in the frame to the inter-person line-of-sight relationship detection unit 125.
[0043] The inter-person line-of-sight relationship detection unit 125 integrates the line-of-sight estimation result of each person for each frame sent from the line-of-sight estimation unit 122 and the person information (profile) for each frame sent from the person identification unit 124 in the spatial direction, detects the line-of-sight relationship between persons occurring in each frame, and stores it in the frame analysis history database 150.
[0044] In this embodiment, the inter-person line-of-sight relationship detection unit 125 includes an eye contact detection unit 125-1, a joint attention detection unit 125-2, a line-of-sight sharing detection unit 125-3, an eye area fixation (LookEye) detection unit 125-4, and a body area fixation (LookBody) detection unit 125-5. That is, the inter-person line-of-sight relationship detection unit 125 is configured to detect five types of line-of-sight relationships between persons, namely eye contact, joint attention, line-of-sight sharing, eye area fixation, and body area fixation, as the line-of-sight relationship between persons. However, what kind of line-of-sight relationship the inter-person line-of-sight relationship detection unit 125 actually detects is arbitrary, the above five types are not essential, and the inter-person line-of-sight relationship detection unit 125 may be configured to further detect line-of-sight relationships other than these. However, when the present disclosure is applied to the quality evaluation of childcare based on CLASS, it is essential for the inter-person line-of-sight relationship detection unit 125 to detect at least eye contact and joint attention.
[0045] The person database 160 is a database for storing information for identifying individuals. Information on nursery teachers who conduct nursery operations and information on children who are the subjects of nursery care are input and stored in the person database 160 in advance.
[0046] The quality evaluation device 130 acquires data on the line-of-sight relationship between persons detected by the video analysis device 120 for each frame of the video from the frame analysis history database 150, aggregates the line-of-sight relationship between persons for each frame in the time axis direction, calculates information such as the score of the quality of nursery care, and provides it to the information presentation device 140.
[0047] The quality evaluation device 130 takes statistics on the line-of-sight relationship between persons such as eye contact and joint attention by the nursery teacher obtained by analyzing the video captured by the fixed-point camera (imaging device 110) by the video analysis device 120, and calculates the score of the quality of nursery care of the nursery teacher based on the statistics. Specifically, the quality evaluation device 130 integrates the detection results of the line-of-sight relationship between persons such as eye contact, joint attention, eye area fixation, and body area fixation in each frame of the video in the time direction, and calculates the relationship between the persons shown in each frame of the video, in other words, the score of the quality of nursery care of the nursery teacher taking care of the child. The calculated score of the quality of nursery care is stored in the video database 170.
[0048] The video database 170 is a database for storing the video captured by the fixed-point camera (imaging device 110) and storing the data combining the analysis result quality of the video analysis device 120 and the result of the quality of nursery care obtained by the evaluation calculation device 130. In the video analysis device 120, the line-of-sight relationship between persons such as eye contact and joint attention is acquired in units of frames. Also, the score of the quality of nursery care calculated by the quality evaluation device 130 varies from moment to moment. In the video database 170, the outputs from the video analysis device 120 and the quality evaluation device 130 are stored in synchronization with the video frames. The user can view the video stored in the video database 170, the video analysis result for the video, and the evaluation result of the quality of nursery care via a WEB front-end system (information presentation device 140) such as an Internet browser.
[0049] B-2. Person-to-Person Gaze Relationship Detection In this embodiment, the person-to-person gaze relationship detection unit 125 includes an eye contact detection unit 125-1, a joint attention detection unit 125-2, a gaze sharing detection unit 125-3, an eye-area gaze detection unit 125-4, and a body-area gaze detection unit 125-5. The person-to-person gaze relationship detection unit 125 is capable of implementing any detection module. Eye contact, joint attention, gaze sharing, eye-area gaze, and body-area gaze are all important gaze relationships for conveying nonverbal messages in communication and human relationships, and serve as a way to show interest in and empathy for the other person. However, this disclosure does not require the detection of all five types of gaze relationships, and may be configured to further detect other types of gaze relationships. However, if this disclosure is applied to quality assessment of childcare based on CLASS, the person-to-person gaze relationship detection unit 125 is required to detect at least eye contact and joint attention.
[0050] The eye contact detection unit 125-1 integrates the person estimation results from the person estimation unit 124 and the gaze estimation results from the gaze estimation unit 122 in the spatial direction to detect eye contact that occurred between people in each frame of the video. Specifically, the eye contact detection unit 125-1 obtains person information (profile) of people shown in the video frame from the person estimation unit 124 and gaze information (starting point of gaze and point of focus) of people shown in the video frame from the gaze estimation unit 122. For each pair of people shown in the video frame, the eye contact detection unit 125-1 determines whether eye contact has occurred by checking whether the vectors from the estimated starting point of gaze to the point of focus are in approximately opposite directions. The eye contact detection unit 125-1 then stores the eye contact detection results for each frame of the video in the frame analysis history database 150.
[0051] For example, the eye contact detection unit 125-1 receives frames from the person estimation unit 124 in which person information (profiles) such as "Teacher A" and "Girl B" are labeled, as shown in Figure 3, and also receives the estimated gaze information for each person for the same frame from the gaze estimation unit 122, as shown in Figure 4. The eye contact detection unit 125-1 then integrates the person estimation result shown in Figure 3 and the gaze estimation result shown in Figure 4 in the spatial direction and detects that eye contact occurred between "Teacher A" and "Girl B" in this frame because the distance between the starting point of "Teacher A's" gaze and the point of interest of "Girl B's" gaze, and the distance between the starting point of "Girl B's" gaze and the point of interest of "Teacher A's" gaze are both smaller than a predetermined threshold. The eye contact detection unit 125-1 then outputs a determination result that eye contact occurred between "Teacher A" and "Girl B" in this frame, as shown in Figure 5.
[0052] The joint attention detection unit 125-2 integrates the person estimation results from the person estimation unit 124 and the gaze estimation results from the gaze estimation unit 122 in the spatial direction to detect joint attention occurring between people in each frame of the video. Specifically, the joint attention detection unit 125-2 receives person information (profiles) of people shown in the video frame from the person estimation unit 124 and obtains gaze information (start point and point of focus of the gaze) of people shown in the video frame from the gaze estimation unit 122. It then determines whether joint attention is occurring based on whether each pair of people shown in the video frame is not making eye contact and whether the distance between the points of focus of their gazes is smaller than a predetermined threshold. The joint attention detection unit 125-2 then saves the joint attention detection results for each frame of the video in the frame analysis history database 150.
[0053] For example, the joint attention detection unit 125-2 receives frames from the person estimation unit 124 in which person information (profiles) such as "Teacher C" and "D-chan" are labeled, as shown in Figure 6, and also receives the estimation results of the gaze information of each person for the same frame from the gaze estimation unit 122, as shown in Figure 7. The joint attention detection unit 125-2 then integrates the person estimation results shown in Figure 6 and the gaze estimation results shown in Figure 7 in the spatial direction and determines that eye contact has not occurred because the distance between the starting point of "Teacher C's" gaze and the point of focus of "D-chan's" gaze, and the distance between the starting point of "D-chan's" gaze and the point of focus of "Teacher C's" gaze are both greater than a predetermined threshold. Furthermore, since the distance between the point of focus of "Teacher C's" gaze and the point of focus of "D-chan's" gaze is smaller than a predetermined threshold, it detects joint attention that occurred between "Teacher C" and "D-chan" in this frame. The joint attention detection unit 125-2 then outputs a determination result that joint attention occurred between "Teacher C" and "D-chan" in this frame, as shown in Figure 8.
[0054] The gaze sharing detection unit 125-3 integrates the detection results of the eye contact detection unit 125-1 and the joint attention detection unit 125-2 to detect gaze sharing that occurs between people in each frame of the video. Specifically, the gaze sharing detection unit 125-3 determines gaze sharing if the joint attention detection unit 125-2 detects joint attention and it has been within a predetermined time since the eye contact detection unit 125-1 detected eye contact. The gaze sharing detection unit 125-3 then stores the gaze sharing detection results for each frame of the video in the frame analysis history database 150.
[0055] The eye-area gaze detection unit 125-4 integrates the person estimation results from the person estimation unit 124 and the gaze estimation results from the gaze estimation unit 122 in the spatial direction to detect LookEyes that occur between people in each frame of the video. A LookEy is the act of looking into the eyes of another person and is a gaze relationship that indicates interest or attention to that person. For example, a LookEy can show interest or empathy towards someone by looking into their eyes during a conversation. The eye-area gaze detection unit 125-4 receives person information (profile) of people shown in the video frame from the person estimation unit 124 and obtains gaze information (start point of gaze and point of focus) of people shown in the video frame from the gaze estimation unit 122. It then determines whether a LookEy has occurred based on whether the point of focus of one person in each pair of people shown in the video frame is within the eye area of the other person or is located very close to the eye area. The eye gaze detection unit 125-4 then stores the eye gaze detection results for each frame of the video in the frame analysis history database 150.
[0056] LookEye is a gaze-line relationship that evaluates how much attention and interest a caregiver is showing a child. The caregiver's gaze information is an important evaluation target, while the child's gaze information is not. Therefore, when the eye-area gaze detection unit 125-4 receives frames from the person estimation unit 124 in which each person is labeled with person information (profile), such as "Teacher A" and "Child B," it evaluates the gaze information of "Teacher A" received from the gaze estimation unit 122 and determines whether a LookEye has occurred between "Teacher A" and "Child B" based on whether the point of focus of that gaze is within or very close to Child B's eye area. If the eye-area gaze detection unit 125-4 can detect a LookEye, it outputs a determination result that a LookEye has occurred between "Teacher A" and "Child B" in this frame, as shown in Figure 9.
[0057] The body region gaze detection unit 125-5 integrates the person estimation results from the person estimation unit 124 and the gaze estimation results from the gaze estimation unit 122 in the spatial direction to detect LookBody occurring between people in each frame of the video. LookBody is the act of looking at the entire body of another person and is a gaze relationship that indicates interest or attention to that person. For example, by looking at the posture and movements of another person, one can read the other person's emotions and intentions. The body region gaze detection unit 125-5 receives person information (profile) of people shown in the video frame from the person estimation unit 124 and obtains gaze information (starting point of gaze and point of focus) of people shown in the video frame from the gaze estimation unit 122, and determines whether LookBody is occurring based on whether the point of focus of one person in each pair of people shown in the video frame is within the body region of the other person or is located very close to the body region. The body gaze detection unit 125-5 then stores the body gaze detection results for each frame of the video in the frame analysis history database 150.
[0058] Like LookEye, LookBody is a gaze relationship that evaluates how much attention and interest a caregiver is paying to a child. The caregiver's gaze information is an important evaluation target, while the child's gaze information is not. Therefore, when the body region gaze detection unit 125-5 receives frames from the person estimation unit 124 in which each person is labeled with person information (profile), such as "Teacher A" and "Child B," it evaluates the gaze information of "Teacher A" received from the gaze estimation unit 122 and determines whether LookBody has occurred between "Teacher A" and "Child B" based on whether the point of focus of that gaze is within or very close to Child B's body region. If LookEye is detected, the body region gaze detection unit 125-5 outputs a determination result that LookBody has occurred between "Teacher A" and "Child B" in this frame, as shown in Figure 10.
[0059] B-3. Quality Assessment of Childcare The quality assessment device 130 integrates the results of detecting eye-gaze relationships between individuals for each frame of a video, stored in the frame analysis history database 150, in the time axis direction, performs statistical processing, and calculates a score for the quality of childcare provided by childcare workers based on these statistics. Section B-3 describes how the quality assessment device 130 integrates the results of detecting eye-gaze relationships between individuals for each frame of a video, performs statistical processing, and calculates a score for the quality of childcare provided by childcare workers based on these statistics. The following describes a method for calculating the quality of childcare for the entire childcare room, but a method for calculating the quality of childcare for each childcare worker is also possible. Furthermore, the calculation method described below is just one example of quality assessment of childcare, and this disclosure is not limited to this.
[0060] Definition of Constants and Variables First, the variables used in calculating the quality score for childcare are defined as follows.
[0061] T: A certain time interval. A score for the quality of childcare is calculated for each time interval T. F: The number of frames included in the time interval T. N: The number of childcare workers (or persons being evaluated) that appear in the video frames during the time interval T. n: The number of children that appear in the video frames during the time interval. Ai (i ≤ N): The i-th childcare worker (i is the childcare worker's serial number). Cj (j ≤ n): The j-th child (j is the child's serial number). Wj (j ≤ n): The level of supervision needed for the j-th child. Assume ΣW = 1. Assume that a level of supervision needed Wj is predetermined for each child. Simply put (i.e., if all children are equal), Wj = 1 / n.
[0062] Function Definitions Next, the functions S(f,p), G(f,a,c), G'(f,a,c), E(f,p,q), and J(f,p,q), which are used to calculate the score for the quality of childcare, are defined as follows.
[0063] S(f,p): A function that indicates whether a person p (whether a childcare worker or a child) is recognized (i.e., whether they are visible in the frame) in the f-th frame. Here, f is the sequential number of the frame (frame ID within the time interval T), and f ≤ F. S(f,p) = 1 if person p is visible in frame f, and S(f,p) = 0 if they are not visible.
[0064] G(f, a, c): An eye-area gaze (LookEye) function that indicates whether a person a is looking at the eye area of another person c in the f-th frame. If a is looking at c's eye area in frame f, G(f, a, c) = 1; otherwise, G(f, a, c) = 0. a and c are not interchangeable (G(f, a, c) ≠ G(f, c, a)). The eye-area gaze detection unit 125-4 outputs a determination result as shown in Figure 9 for each frame of the video, making it possible to use the function G(f, a, c). The eye-area gaze detection unit 125-4 may also output the detection result of eye-area gaze for each frame in the form of this eye-area gaze function G(f, a, c).
[0065] G'(f, a, c): A LookBody function that indicates whether a person a is looking at the body region of another person c in the f-th frame. If a is looking at c's body region in frame f, G'(f, a, c) = 1; otherwise, G'(f, a, c) = 0. a and c are inexchangeable (G'(f, a, c) ≠ G'(f, c, a)). The LookBody detection unit 125-5 outputs a determination result as shown in Figure 10 for each frame of the video, making it possible to use the function G'(f, a, c). The LookBody detection unit 125-5 may also output the LookBody detection result for each frame in the form of this LookBody function G'(f, a, c).
[0066] Persons a and c can be either a childcare worker or a child, but since this is for visual scanning (described later) where a childcare worker looks around and watches over the children, in the above functions G(f, a, c) and G'(f, a, c), a is the childcare worker and c is the child.
[0067] E(f, p, q): An eye contact function that indicates whether a person p and a person q made eye contact in the f-th frame. If eye contact occurred between people p and q in frame f, E(f, p, q) = 1; otherwise, E(f, p, q) = 0. p and q are interchangeable, and p is looking at q's eye area while q is also looking at p's eye area. That is, E(f, p, q) = E(f, q, p) = G(f, p, q) × G(f, q, p) holds true. The eye contact detection unit 125-1 outputs a determination result as shown in Figure 5 for each frame of the video, making it possible to use the function E(f, p, q). The eye contact detection unit 125-1 may also output the eye contact detection result for each frame in the form of this eye contact function G(f, p, q).
[0068] J(f, p, q): A joint attention function that indicates whether two people p and q shared joint attention in the f-th frame. If joint attention occurs between people p and q in frame f (when the endpoints of their gazes almost coincide), J(f, p, q) = 1; otherwise, J(f, p, q) = 0. p and q are interchangeable, and J(f, p, q) = J(f, q, p) holds true. The joint attention detection unit 125-2 outputs a determination result as shown in Figure 8 for each frame of the video, making it possible to use the function J(f, p, q). The joint attention detection unit 125-2 may also output the joint attention detection result for each frame in the form of this eye contact function J(f, p, q).
[0069] Example of calculating a childcare worker quality score based on eye contact and joint attention First, childcare worker Ai is fixed. In a certain frame f, whether childcare worker Ai is making eye contact with any of the children Cj is calculated using the following formula (1).
[0070]
[0071] Since it is not possible to make eye contact with more than one person at the same time, even if we add up the eye contact functions with all the children Cj, the value of equation (1) above will be either 1 or 0.
[0072] On the other hand, in the case of joint attention, childcare worker Ai can simultaneously engage in joint attention with multiple children. Therefore, if childcare worker Ai simultaneously engages in joint attention with one or more children, the joint attention function J(f, Ai, Cj) is rounded to 1 using equation (2) below. That is, equation (2) below is either 0 or 1.
[0073]
[0074] Then, by summing the values calculated using equation (1) above for all frames F within the time interval T, we can obtain the number of frames in which eye contact occurred within the time interval T. There are also frames in which the person Ai did not appear. The number of frames in which childcare worker Ai appeared within the time interval T is expressed by equation (3) below.
[0075]
[0076] Within a time interval T, the total number of frames in which childcare worker Ai made eye contact with any child Cj, divided by the number of frames in which childcare worker Ai appeared, yields a value between 0 and 1. Averaging this value across all childcare workers Ai allows us to calculate the EyeContact score, which represents the quality of childcare based on eye contact, as shown in formula (4) below.
[0077]
[0078] Furthermore, within a time interval T, the total number of frames in which childcare worker Ai jointly pays attention to any child Cj is divided by the number of frames in which childcare worker Ai appears, resulting in a value between 0 and 1. Averaging this value across all childcare workers Ai allows us to calculate the JointAttention score, which represents the quality of childcare based on joint attention, as shown in formula (5) below.
[0079]
[0080] However, equations (4) and (5) above have the problem that they are heavily influenced by childcare workers with an extremely small number of appearances (i.e., an extremely small function S). Therefore, it is possible to modify equations (4) and (5) above into calculation formulas that use a weighted average according to the number of appearances, or to calculate equations (4) and (5) above by excluding childcare workers with an extremely small number of appearances (i.e., an extremely small function S).
[0081] This specification will explain how to formulate the formula for the GazeFollowing score, which measures the quality of childcare based on eye contact. For example, a gaze-sharing function J'(f, p, q) is defined to indicate whether joint attention occurred between two individuals p and q in the f-th frame (provided that the frame is within a predetermined time after the first eye contact occurred). The formula for the GazeFollowing score can then be formulated using this gaze-sharing function J'(f, p, q).
[0082] Visual Scan Calculation Example: In a nursery room, each caregiver looks around and monitors each child (Visual Scan). The premise is that it is sufficient for a child to be observed by one caregiver. Each child has a set level of need for supervision Wj (j ≤ n, 0 ≤ Wj, ΣW = 1).
[0083] We interpret that children with a small W value can be left unsupervised to some extent, and formulate the Visually Scan score calculation formula accordingly. For example, for the j-th child with a Wj of 0.5, we assume that it is safe to leave them unsupervised for half of the time interval.
[0084] First, we fix the position of child Cj. In a given frame f, whether child Cj is being observed by at least one caregiver can be calculated using the body domain gaze function in equation (6) below. In equation (6) below, if two or more caregivers are looking at child Cj simultaneously, the body domain gaze function G'(f, Ai, Cj) is rounded to 1. That is, equation (6) below is either 0 or 1.
[0085]
[0086] The number of frames in which child Cj was seen by the caregiver within a time interval T can be determined by summing up equation (6) above for all frames F within the time interval T. Then, in order to normalize to [0-1], equation (7) below divides by the number of frames in which child Cj appeared.
[0087]
[0088] If a child Cj has an extremely small value in equation (7) above, it means that the child is not receiving any attention from any of the childcare workers.
[0089] The value of equation (7) above must exceed the level of need for supervision Wj. If supervision exceeds the level of need, the value is discarded. The average of the values calculated using equation (7) above for all children Cj is shown in equation (8) below. Equation (8) below is adopted as the score value for Visually Scan.
[0090]
[0091] In equation (8) above, if there are children with an extremely small number of appearance frames, the error will be large, so it is necessary to take measures such as removing such cases.
[0092] Ideally, if there are multiple childcare workers in a class, they should cooperate to ensure that they do not take their eyes off any child for too long. This point has not yet been taken into consideration in equation (8) above. Also, even when we say "taking your eyes off a child," the quality of supervision differs greatly depending on whether you take your eyes off the child for just one second every six seconds or for ten consecutive seconds every minute. Not only the proportion of time spent taking your eyes off the child, but also the length of time spent taking your eyes off the child should be considered.
[0093] Example of Seek Out Teachers Calculation Seek Out Teachers is a score value that observes and evaluates a child's attempts to communicate with a caregiver when tackling new things or when they are having trouble. To calculate this score value, it is necessary to understand the child's situation. The difficulty in interpreting a child's situation is a challenge in formulating this score value. In this specification, we will formulate a mathematical formula to represent the extent to which a child is looking at a caregiver (or their eyes) when they are not receiving direct attention from the caregiver.
[0094] First, in frame f, we focus on a specific child Cj and calculate using equation (9) below whether any of the childcare workers Ai are looking at child Cj, and whether child Cj is looking at any of the childcare workers Ai.
[0095]
[0096] In equation (9) above, the first term indicates whether any of the childcare workers Ai are looking at the child Cj, and the second term indicates whether the child Cj is looking at any of the childcare workers Ai. Equation (9) has a value of either 1 or 0.
[0097] Next, the percentage of frames within a time interval T in which no childcare worker Ai is looking at a specific child Cj is calculated using the following formula (10).
[0098]
[0099] The average of the values calculated using equation (10) above for all children Cj is shown in equation (11) below. Equation (11) below will be adopted as the SeekTeachers score.
[0100]
[0101] In Section B-3, four types of score values for childcare quality calculated by the quality evaluation device 130 are listed: Eye Contact, Joint Attention, Visually Scan, and Seek out teachers. However, this disclosure is not limited to the calculation of these four types of score values, and it is conceivable that there may be cases where at least one of these four types of score values is not calculated, or where other score values are calculated. Furthermore, although the calculation methods for each of these four types of score values have been described above, this disclosure is not limited to the calculation formulas for specific score values, and at least one of these four types of score values may be formulated in a way other than those described above.
[0102] The quality evaluation device 130 may also utilize the functions of external software via API (Application Programming Interface) to calculate a score value for a desired gaze relationship.
[0103] C. System Operation Diagram 2 shows the operation procedure of the childcare quality evaluation system 100 to which this disclosure applies in the form of a flowchart. Here, it is assumed that the video uploaded from the imaging device 110 is already stored in the video database 170.
[0104] First, at each time interval T of the video, steps S201 to S219 are performed to detect the gaze relationship between people from each frame.
[0105] The video analysis device 120 reads the video to be processed from the video database 170 in predetermined time intervals T (step S201), and then decomposes the video within this time interval T into individual frames (step S202). Then, it takes out the unprocessed frames one by one from the beginning of the video frames within the time interval T (step S203), and starts video analysis and interpersonal gaze relationship detection on a frame-by-frame basis.
[0106] The face recognition unit 121 detects the faces of people in the frame (step S204). The gaze estimation unit 122 then uses an existing gaze estimation network, such as Attention Target Detection, to estimate gaze information consisting of the starting point and point of focus of the gaze of each detected face in the frame image (step S205). Specifically, the gaze estimation unit 122 detects the eye positions of the detected faces detected by the face recognition unit 121, and further estimates the gaze direction (starting point and point of focus of the gaze) from the eye positions and the orientation of the face. The gaze estimation unit 122 sends the gaze estimation results, with gaze information attached to each detected face in the frame, to the interpersonal gaze relationship detection unit 125.
[0107] Furthermore, the object estimation unit 123 estimates the objects in the frame using existing algorithms that utilize deep networks, such as YOLO (step S206). The person estimation unit 124 then receives the object estimation results from the object estimation unit 123 and estimates who the people in each frame of the video are (step S207). The candidate people estimated by the person estimation unit 124 are pre-registered children and childcare workers, and it is assumed that they are registered in the person database 160 in advance. The person estimation unit 124 may also have a function to estimate whether the guest is the same person or not. The person estimation unit 124 sends the person estimation results, with person information (profile) attached to each person area in the frame, to the person-to-person gaze relationship detection unit 125.
[0108] The person-to-person gaze relationship detection unit 125 integrates the gaze estimation results for each person for each frame sent from the gaze estimation unit 122 and the person information (profile) for each frame sent from the person estimation unit 124 in the spatial direction to detect the gaze relationships between people that occurred in each frame and saves them in the frame analysis history database 150.
[0109] The person-to-person gaze relationship detection unit 125 determines pairs of all people within the frame whose gaze relationships are to be detected (step S208).
[0110] For example, as shown in Figure 11, if three people A, B, and C are in the frame, three possible pairs of people are determined: (A, B), (B, C), and (C, A). Also, as shown in Figure 12, if four people A, B, C, and D are in the frame, four possible pairs of people are determined: (A, B), (A, C), (A, D), (B, C), (B, D), and (C, D). However, the pairs of people for which gaze relationships such as eye contact and joint attention should be detected are basically the combination of a caregiver and a child. Therefore, the person estimation results in step S207 may be used to exclude pairs of people other than caregiver-child pairs (for example, pairs of caregivers, pairs of children).
[0111] Then, the person-to-person gaze relationship detection unit 125 takes one unprocessed pair of people at a time from the determined pair of people (step S209) and sequentially checks whether gaze relationships such as eye contact, joint attention, lookeye, and lookbody are occurring within the frame.
[0112] The eye contact detection unit 125-1 spatially integrates the gaze estimation result in step S205 and the person estimation result in step S207 to detect whether eye contact occurred between the target pair of people within the frame (step S210). The eye contact detection unit 125-1 determines whether eye contact has occurred by checking whether the vectors from the starting point of each person's gaze to the point of focus are in approximately opposite directions (see, for example, Figure 4). If eye contact is detected (Yes in step S210), the eye contact detection unit 125-1 saves the eye contact detection result (see, for example, Figure 5) in the frame analysis history database 150 (step S211).
[0113] Next, the joint attention detection unit 125-2 spatially integrates the gaze estimation result from step S205 and the person estimation result from step S207 to detect whether joint attention has occurred between the target pair of people within the frame (step S212). The joint attention detection unit 125-2 determines whether joint attention has occurred based on whether the distance between the points of interest of each person's gaze is smaller than a predetermined threshold (see, for example, Figure 7). If joint attention is detected (Yes in step S212), the joint attention detection unit 125-2 saves the joint attention detection result (see, for example, Figure 8) to the frame analysis history database 150 (step S213). At this time, the gaze sharing detection unit 125-3 also determines whether gaze attention has occurred based on whether a predetermined time has elapsed since the last time eye contact was detected by the eye contact detection unit 125-1, and if gaze sharing is detected, it also saves the detection result to the frame analysis history database 150.
[0114] Next, the eye gaze detection unit 125-4 spatially integrates the gaze estimation result from step S205 and the person estimation result from step S207 to detect whether a LookEye occurred in the target pair of people within the frame (step S214). The eye gaze detection unit 125-4 determines whether a LookEye has occurred based on whether the point of attention of one person in the target pair is within the eye area of the other person or very close to the eye area. If a LookEye is detected (Yes in step S214), the eye gaze detection unit 125-4 saves the LookEye detection result (see, for example, Figure 9) in the frame analysis history database 150 (step S215).
[0115] Next, the body region gaze detection unit 125-5 spatially integrates the gaze estimation result from step S205 and the person estimation result from step S207 to detect whether a LookBody has occurred in the frame for the target person pair (step S216). The body region gaze detection unit 125-5 determines whether a LookBody has occurred based on whether the point of attention of one person in the target person pair is within the body region of the other person or is located very close to the body region. If a LookBody is detected (Yes in step S216), the body region gaze detection unit 125-5 saves the LookBody detection result (see, for example, Figure 10) in the frame analysis history database 150 (step S217).
[0116] As described above, once the person relationship detection process is completed for the person pair extracted in step S209, it is checked whether there are any unprocessed person pairs remaining in the currently processed frame (step S218). If there are any unprocessed person pairs remaining in the currently processed frame (Yes in step S218), the process returns to step S209, one unprocessed person pair is extracted, and the person relationship detection process is repeatedly executed for that person pair.
[0117] Furthermore, if the process of detecting eye-gaze relationships between all pairs of people is completed for the frame currently being processed (No in step S218), the system checks whether there are any unprocessed frames remaining in the video within the time interval T to be processed (step S219). If there are any unprocessed frames remaining in the video within the time interval T (Yes in step S219), the system returns to step S203, extracts the next frame from the video frames within the time interval T, and repeatedly performs video analysis and eye-gaze relationship detection on that frame.
[0118] On the other hand, once video analysis and detection of interpersonal gaze relationships are completed for all frames within the time interval T to be processed (No. in step S219), the detection results of interpersonal gaze relationships for the time interval T of the video are stored in the frame analysis history database 150. Furthermore, by repeating the above steps S201 to S219 for each time interval T of the video, the detection results of interpersonal gaze relationships for the entire video are stored in the frame analysis history database 150.
[0119] The quality evaluation device 130 calculates a score for the quality of childcare based on the detection results of interpersonal gaze relationships at each time interval T of the video, which are stored in the frame analysis history database 150 (step S220).
[0120] In step S220, the quality evaluation device 130 obtains data on the gaze relationships between people detected frame by frame for a time interval T minutes in the video from the frame analysis history database 150, aggregates the gaze relationships between people for each frame along the time axis, and calculates information such as a score for the quality of childcare.
[0121] Specifically, the quality evaluation device 130 can obtain the following functions from the frame analysis history database 150 for each frame f within a time interval T: eye contact function E(f, p, q), joint attention function J(f, p, q), gaze sharing function J'(f, p, q), eye region gaze (LookEye) function G(f, a, c), and body region gaze (LookBody) function G'(f, a, c). The quality evaluation device 130 then uses these function values for each frame f within the time interval T to calculate a score value for childcare quality using eye contact as an indicator according to equation (4) above, a score value for childcare quality using joint attention as an indicator according to equation (5) above, a score value for childcare quality using eye-gaze sharing as an indicator based on the eye-gaze sharing function J'(f, p, q) above, a score value for childcare quality using Visually Scan as an indicator according to equation (8) above, and a score value for childcare quality using Seek out teacher as an indicator according to equation (11) above. The quality evaluation device 130 may also utilize APIs in calculating these score values.
[0122] The quality evaluation device 130 then saves the score values for the quality of childcare calculated at each time interval T in the video database 170, synchronized with the video frames (step S221).
[0123] Subsequently, users (such as the director's office of the nursery school, the person in charge of managing or supervising the childcare workers, or the parents of the children) can access the video database 170 via an internet browser (information display device 140) and view the video of the nursery room captured by the imaging device 110 along with the childcare quality score for each time interval T (step S222). Users can view the video analysis results and the childcare quality evaluation results, for example, via the UI screens shown in Figures 15 to 21 (described later).
[0124] D. Method for Viewing Video Analysis Results and Childcare Quality Evaluation Results In the childcare quality evaluation system 100 related to this disclosure, users can use the information display device 140 to view information such as the results of analyzing videos uploaded from the imaging device 110 with the video analysis device 120, and the childcare quality score calculated by the quality evaluation device 130 based on the video analysis results (detection results of eye contact and joint attention). The information display device 140 is, for example, an information terminal such as a smartphone, tablet, or PC used by the user. The information display device 140 presents information such as the detection results of gaze relationships such as eye contact and joint attention, and the childcare quality score, obtained from a cloud server, to the user via a web front-end system such as an internet browser. Section D explains how to view information such as video analysis results and childcare quality scores, using an example of a UI (User Interface) screen configuration displayed on the screen of the information display device 140. Each UI screen introduced below is provided as, for example, a web front-end system such as an internet browser.
[0125] After logging into the cloud server, the user specifies the location they wish to view ("○○ Nursery School") and the date and time they wish to view ("△△ Month □□ Day"). As shown in Figure 13, the UI screen displays a timeline of the schedule for ○○ Nursery School for △△ Month □□ Day. Each schedule item displayed in the schedule list is a selection menu. As shown in Figure 14, the user selects the desired schedule item, "11:30 - Lunch" by moving the cursor over it (it will be highlighted as shown in Figure 14), and then presses the "Select Video" button in the lower right corner of the UI screen to confirm the selection. The video for the schedule item "11:30 - Lunch" is then downloaded from the video database 170 to the information display device 140, and as shown in Figure 15, the playback video for "11:30 - Lunch" is displayed on the UI screen.
[0126] First, let's explain the UI screen used to view the video analysis results.
[0127] As shown in Figure 15, buttons for "Joint Attention," "Eye Contact," and "Look Eye" are located near the bottom of the video playback screen. Pressing any of these buttons displays information regarding the detection result of the corresponding gaze relationship in the playback video. For the sake of simplification of the diagram, buttons for selecting other gaze relationships, such as body gaze, have been omitted from the illustration, but the video playback screen may also include buttons for selecting other gaze relationships.
[0128] Figure 16 shows what happens when the "Joint Attention" button is selected. In this case, the gaze information of each person is superimposed on the original video playback screen, but the gaze information corresponding to joint attention is highlighted (in the example shown in the figure, the gaze information corresponding to joint attention is displayed with a thick, dark line, while the gaze information not corresponding to joint attention is displayed with a thin, light line).
[0129] Figure 17 shows what happens when the "Eye Contact" button is selected. In this case as well, the gaze information of each person is superimposed on the original video playback screen, but the gaze information corresponding to eye contact is highlighted (in the example shown in the figure, the gaze information corresponding to eye contact is displayed with a thick, dark line, while the gaze information that does not correspond to eye contact is displayed with a thin, light line).
[0130] Figure 18 shows what happens when the "Look Eye" button is selected. In this case as well, the gaze information of each person is superimposed on the original video playback screen, but the gaze information corresponding to eye gaze is highlighted (in the example shown in the figure, the gaze information corresponding to eye gaze is displayed with a thick, dark line, while the gaze information that does not correspond to eye gaze is displayed with a thin, light line).
[0131] Next, I will explain the UI screen for viewing the evaluation results of the quality of childcare.
[0132] In the video playback screen shown in Figure 19, select the desired childcare worker by clicking (or touching) near the display area (the area enclosed by a thick line in the figure). Then, the UI screen will switch to the one showing the evaluation results of the quality of care provided by the selected childcare worker, as shown in Figure 20. In the example shown in Figure 20, the video of the childcare worker in question is reduced in size and displayed in the timeline display area 1901 on the left side of the UI screen. In addition, the score display area 1902 on the right side of this UI screen displays the Eye Contact, Joint Attention, Visually Scan, and Seek out teachers scores calculated by the quality evaluation device 130 for the childcare worker, along with the track ID and track length of the video being played.
[0133] The UI screen shown in Figure 21 presents the evaluation results of the quality of childcare using a video playback area indicated by reference numeral 2101, a timeline display area for the quality of childcare indicated by reference numeral 2102, and a ranking display area indicated by reference numeral 2103. In the video playback area 2101, the gaze information of each person (childcare worker and child) is superimposed on the video being played at the current playback position. In addition, the timeline display area for the quality of childcare indicated by reference numeral 2102 shows a line graph of the temporal changes in the score values of Eye Contact, Joint Attention, Visually Scan, and Seek out teachers for the entire childcare room, calculated by the quality evaluation device 130, over the video playback time. Furthermore, the ranking display area 2103 includes a list of childcare workers sorted in descending order of score, a list of children sorted in descending order of score (unsupervised), and a list of recommended scenes from the video for review (for example, a list of scenes sorted in descending order of the overall score value of the childcare room).
[0134] E. Technical Points: The technical points of the childcare quality evaluation system 100 to which this disclosure applies are summarized below.
[0135] (1) The childcare quality evaluation system 100 is configured to automatically calculate the quality of childcare provided by childcare workers from videos of childcare work. Therefore, when reviewing the day's work at a nursery school, the quality of childcare provided by childcare workers can be evaluated based on accurate information recorded by the childcare quality evaluation system 100, without relying on vague human memory.
[0136] (2) The childcare quality evaluation system 100 is configured to use the gaze information of people in the video to detect whether childcare workers were able to make eye contact with children and provide joint attention. Therefore, the childcare quality evaluation system 100 can automatically detect to what extent childcare workers were able to make eye contact and provide joint attention, which are very important elements for evaluating the quality of childcare.
[0137] F. Hardware Configuration Diagram 22 of the Information Processing Device shows an example of the hardware configuration of the information processing device 2000 applicable to this disclosure. The cloud server shown in Figure 1 can be constructed using the information processing device 2000. Alternatively, the cloud server may be constructed using multiple information processing devices 2000. Furthermore, the video analysis device 120, the quality evaluation device 130, and the information presentation device 140 can each be configured using individual information processing devices 2000.
[0138] This information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013. The information processing device 2000 is configured as, for example, a PC, but some functions may be configured as information terminals such as tablets and smartphones.
[0139] CPU 2001 controls the overall operation of the information processing unit 2000 according to various programs. When performing computationally intensive processes such as AI (Artificial Intelligence) model training on the information processing unit 2000, it is desirable that CPU 2001 be a multi-core CPU (e.g., Apple M1 Max), and that the information processing unit 2000 also be equipped with a multi-core processor such as a GPU (Graphics Processing Unit) or GPGPU (General-purpose computing on graphics processing units) (e.g., NVIDIA's "RTX A6000"). However, for convenience, these will be collectively referred to simply as CPU 2001 below.
[0140] ROM2002 non-volatilely stores programs (such as the basic input / output system) and arithmetic parameters used by CPU2001. RAM2003 is used to load programs to be executed by CPU2001 and to temporarily store parameters such as work data that change as needed during program execution. Programs loaded into RAM2003 and executed by CPU2001 include, for example, various application programs and operating systems (OS).
[0141] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which consists of a CPU bus and other components. Through the collaborative operation of ROM 2002 and RAM 2003, the CPU 2001 can execute various application programs within the execution environment provided by the OS, thereby realizing a variety of functions and services. If the information processing device 2000 is a PC, the OS is, for example, Microsoft's Windows®, Unix®, or its successor OS. For example, a program for performing at least one of the following processes—video analysis of moving images and evaluation of childcare quality based on the video analysis results—is executed on the information processing device 2000. Note that the application program, or some modules within the application program, may utilize existing libraries stored, shared, and made publicly available through, for example, a source code management service, or the processing may be implemented via an API.
[0142] The host bus 2004 is connected to the expansion bus 2006 via the bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not need to be configured in a way that isolates its circuit components by the host bus 2004, the bridge 2005, and the expansion bus 2006; it may be an implementation in which almost all circuit components are interconnected by a single bus (not shown).
[0143] The interface unit 2007 connects peripheral devices such as the input unit 2008, output unit 2009, storage unit 2010, drive 2011, and communication unit 2013 in accordance with the expansion bus 2006 standard. However, not all peripheral devices shown in Figure 22 are necessarily required, and the information processing device 2000 may include additional peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some peripheral devices may be externally connected to the main body of the information processing device 2000.
[0144] The input unit 2008 consists of an input control circuit that generates an input signal based on user input and outputs it to the CPU 2001. If the information processing device 2000 is a PC, the input unit 2008 may include a keyboard, mouse, touch panel, camera, and microphone. The output unit 2009 may include, for example, a liquid crystal display (LCD) device, an organic EL (Electro-Luminescence) display device, and an LED (Light Emitting Diode) display device, as well as an audio output device such as a speaker. If the information processing device 2000 is used as an information presentation device 140, the output unit 2009 is used to display a GUI screen (see Figures 15 to 21).
[0145] The storage unit 2010 stores files such as programs (applications, OS, etc.) and various data executed by the CPU 2001. The storage unit 2010 is composed of, for example, a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device.
[0146] The removable storage medium 2012 is a storage medium configured in a cartridge format, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 2012. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or storage unit 2010, and writes data on the RAM 2003 or storage unit 2010 to the removable storage medium 2012.
[0147] The communication unit 2013 is a device that performs wireless communication such as Wi-Fi®, Bluetooth®, and cellular communication networks such as 4G and 5G. The communication unit 2013 may also be equipped with terminals such as USB (Universal Serial Bus) and HDMI® (High-Definition Multimedia Interface), and may further have the function of performing HDMI® communication with USB devices such as scanners and printers, and displays. Programs executed on the information processing device 2000 are installed externally, for example, through the communication unit 2013. Furthermore, when the information processing device 2000 is used as a quality evaluation device 130, it may also make API calls to external devices, for example, through the communication unit 2013. In addition, the information processing device 2000 may also be a device that receives and processes API requests from the quality evaluation device 130.
[0148] The present disclosure has been described in detail above with reference to specific embodiments. However, this disclosure should not be construed as being limited to the embodiments described above, and it will be obvious that those skilled in the art can modify or substitute these embodiments without departing from the gist of the disclosure. Furthermore, the effects described herein are merely illustrative, and the effects brought about by this disclosure are not limited, and there may be additional effects not described herein.
[0149] This specification has primarily described embodiments of this disclosure applied to the field of childcare, but the gist of this disclosure is not limited to this. The gaze relationships between people automatically acquired by this disclosure are important clues for evaluating the quality of human-to-human communication. This disclosure can be applied not only to early childhood education but also to various educational fields such as primary education, secondary education, higher education, skills training, and driving instruction. Furthermore, this disclosure can be applied to sports schools for various sports such as soccer, baseball, tennis, and swimming, as well as to settings for extracurricular activities such as tea ceremony, flower arrangement, and calligraphy. According to this disclosure, in each of these application fields, it is possible to automate the evaluation of the quality of education and learning based on video data and realize quality evaluation through data linkage that combines it with profile information, which can be used to improve daily education, instruction, practical training, and instruction. Furthermore, video data, quality evaluation results, and posted text can be shared through chat systems, etc. This disclosure can be broadly applied to various services related to human-to-human communication.
[0150] In short, this disclosure has been explained in the form of examples, and the contents of this specification should not be interpreted restrictively. The claims should be considered in order to determine the gist of this disclosure.
[0151] The series of processes described herein can be executed by hardware, software, or a combination of hardware and software. When executing the processes by software, a program recording the processing sequence related to the implementation of this disclosure is installed and executed in the memory of a computer embedded in dedicated hardware. It is also possible to install the program on a general-purpose computer capable of executing various processes and execute the processes related to the implementation of this disclosure.
[0152] The program can be pre-stored on a recording medium installed in a computer, such as an HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable recording medium such as a flexible disk, CD-ROM (Compact Disc Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), BD (Blu-Ray Disc®), magnetic disk, or USB (Universal Serial Bus) memory. Using such a removable recording medium, the program related to the implementation of this disclosure can be provided as so-called packaged software.
[0153] Alternatively, the program may be transferred to a computer wirelessly or via a wired connection from a download site to a network such as a cellular network (WAN, Wide Area Network), a LAN (Local Area Network), or the Internet. The computer can then receive the program transferred in this way and install it into a large-capacity storage device such as an HDD or SSD.
[0154] Furthermore, this disclosure may also take the following form.
[0155] (1) An information processing system comprising: a person estimation unit that estimates the person appearing in each frame of a video; a gaze estimation unit that estimates the gaze of the person appearing in each frame of a video; and a detection unit that combines the estimation results from the person estimation unit and the gaze estimation unit to detect a specific gaze relationship between people that occurred in each frame of a video.
[0156] (2) The information processing system according to (1) above, wherein the detection unit detects that at least one of the following gaze relationships has occurred: eye contact in which two people make eye contact with each other within a frame, joint attention in which two people direct their attention to the same object, gaze sharing in which joint attention occurs within a certain period of time after eye contact, eye-domain gaze in which one looks at the other person's eyes, and body-domain gaze in which one looks at the other person's body.
[0157] (3) The information processing system described in (2) above, wherein the detection unit determines whether or not eye contact has occurred for each pair of people pictured in the video frame by checking whether the vectors from the estimated viewpoint of the gaze to the point of interest are in approximately opposite directions.
[0158] (4) The information processing system according to either (2) or (3) above, wherein the detection unit determines whether joint attention is occurring based on whether each pair of people shown in the video frame is not making eye contact and the distance between their points of focus is less than a predetermined threshold, and determines whether gaze sharing occurs based on whether joint attention occurs within a certain period of time after detecting the occurrence of eye contact.
[0159] (5) The information processing system according to any one of (2) to (4) above, wherein the detection unit determines whether eye fixation is occurring based on whether the point of attention of one of the pairs of people pictured in the video frame is within the eye area of the other person or is in a position very close to the eye area.
[0160] (6) The information processing system according to any one of (2) to (5) above, wherein the detection unit determines whether body region gaze is occurring based on whether the point of attention of one of the pairs of people shown in the video frame is within the body region of the other person or is located very close to the body region.
[0161] (7) The information processing system according to any one of (1) to (6) above, wherein the gaze estimation unit estimates gaze using a gaze estimation network.
[0162] (8) The information processing system according to any one of (1) to (7) above, further comprising an evaluation unit that integrates specific gaze relationships between persons detected by the detection unit in each frame of the video in the time direction and evaluates the persons.
[0163] (9) The information processing system described in (8) above, wherein the evaluation unit calculates at least one of the following: a score value based on the detection result of eye contact by the detection unit, a score value based on the detection result of joint attention by the detection unit, a score value based on the detection result of gaze sharing by the detection unit, a score value based on the detection result of eye region gaze by the detection unit, and a score value based on the detection result of body region gaze by the detection unit.
[0164] (10) The information processing system according to either (8) or (9) above, further comprising a presentation unit for presenting the detection result by the detection unit or the evaluation result by the evaluation unit.
[0165] (11) The information processing system according to (10) above, wherein the display unit displays gaze information corresponding to the specific gaze relationship superimposed on the video playback screen.
[0166] (12) The information processing system according to either (10) or (11) above, wherein the display unit displays the evaluation results from the evaluation unit for a person specified on the video playback screen.
[0167] (13) The information processing system according to any one of (10) to (12) above, wherein the presentation unit presents the temporal progression of the evaluation results by the evaluation unit.
[0168] (14) The information processing system according to any one of (10) to (13) above, wherein the presentation unit presents a list sorted by the persons appearing in the video based on the evaluation results of the evaluation unit.
[0169] (15) The information processing system according to any one of (10) to (14) above, wherein the presentation unit presents a list of scenes selected from the video based on the evaluation results by the evaluation unit.
[0170] (16) An information processing method comprising: a person estimation step of estimating the person appearing in each frame of a video; a gaze estimation step of estimating the gaze of the person appearing in each frame of a video; and a detection step of detecting a specific gaze relationship between people that occurred in each frame of a video by combining the estimation results in the person estimation step and the gaze estimation step.
[0171] (17) A computer program written in a computer-readable format to cause a computer to function as: a person estimation unit that estimates the person appearing in each frame of a video; a gaze estimation unit that estimates the gaze of the person appearing in each frame of a video; and a detection unit that combines the estimation results from the person estimation unit and the gaze estimation unit to detect a specific gaze relationship between people that occurred in each frame of a video.
[0172] 100...Childcare quality evaluation system, 110...Imaging device, 120...Video analysis device, 121...Face recognition unit, 122...Eye gaze estimation unit, 123...Object estimation unit, 124...Person estimation unit, 125...Person-to-person eye gaze relationship detection unit, 125-1...Eye contact detection unit, 125-2...Joint attention detection unit, 125-3...Eye gaze sharing detection unit, 125-4...Eye region gaze detection unit, 125-5...Body region gaze detection unit, 130...Quality evaluation device, 140...Information presentation device, 150...Frame analysis history database, 160...Person database, 170...Video database, 2000...Information processing device, 2001...CPU, 2002...ROM, 2003...RAM, 2004...Host bus, 2005...Bridge, 2006...Expansion bus, 2007...Interface unit 2008...Input section, 2009...Output section, 2010...Storage section, 2011...Drive, 2012...Removable recording medium, 2013...Communication section
Claims
1. An information processing system comprising: a person estimation unit that estimates the person appearing in each frame of a video; a gaze estimation unit that estimates the gaze of the person appearing in each frame of a video; and a detection unit that combines the estimation results from the person estimation unit and the gaze estimation unit to detect a specific gaze relationship between people that occurred in each frame of a video.
2. The information processing system according to claim 1, wherein the detection unit detects that at least one of the following gaze relationships has occurred: eye contact in which two people make eye contact with each other within a frame; joint attention in which two people direct their attention to the same object; gaze sharing in which joint attention occurs within a certain period of time after eye contact; eye-domain gaze in which one looks at the other person's eyes; and body-domain gaze in which one looks at the other person's body.
3. The information processing system according to claim 2, wherein the detection unit determines whether or not eye contact has occurred for each pair of people pictured in the video frame by checking whether the vectors from the estimated viewpoint of the gaze to the point of interest are in approximately opposite directions.
4. The information processing system according to claim 2, wherein the detection unit determines whether joint attention is occurring based on whether each pair of people shown in the video frame is not making eye contact and the distance between their points of focus is less than a predetermined threshold, and determines whether gaze sharing occurred within a certain period of time after detecting the occurrence of eye contact.
5. The information processing system according to claim 2, wherein the detection unit determines whether eye fixation is occurring based on whether the point of focus of one person in each pair of people shown in the video frame is within the eye area of the other person or is located very close to the eye area.
6. The information processing system according to claim 2, wherein the detection unit determines whether body region gaze is occurring based on whether the point of attention of one of the pairs of people shown in the video frame is within or very close to the body region of the other person.
7. The information processing system according to claim 1, wherein the gaze estimation unit estimates gaze using a gaze estimation network.
8. The information processing system according to claim 1, further comprising an evaluation unit that integrates specific gaze relationships between persons detected by the detection unit in each frame of a video in the temporal direction and evaluates the persons.
9. The information processing system according to claim 8, wherein the evaluation unit calculates at least one of the following: a score value based on the detection result of eye contact by the detection unit; a score value based on the detection result of joint attention by the detection unit; a score value based on the detection result of gaze sharing by the detection unit; a score value based on the detection result of eye region gaze by the detection unit; and a score value based on the detection result of body region gaze by the detection unit.
10. The information processing system according to claim 8, further comprising a presentation unit for presenting the detection result from the detection unit or the evaluation result from the evaluation unit.
11. The information processing system according to claim 10, wherein the display unit overlays and displays gaze information corresponding to the specific gaze relationship onto the video playback screen.
12. The information processing system according to claim 10, wherein the display unit displays the evaluation results from the evaluation unit for a person specified on the video playback screen.
13. The information processing system according to claim 10, wherein the display unit displays the temporal progression of the evaluation results by the evaluation unit.
14. The information processing system according to claim 10, wherein the presentation unit presents a list sorted by the persons appearing in the video based on the evaluation results from the evaluation unit.
15. The information processing system according to claim 10, wherein the presentation unit presents a list of scenes selected from the video based on the evaluation results by the evaluation unit.
16. An information processing method comprising: a person estimation step for estimating the person appearing in each frame of a video; a gaze estimation step for estimating the gaze of the person appearing in each frame of a video; and a detection step for detecting a specific gaze relationship between people that occurred in each frame of a video by combining the estimation results from the person estimation step and the gaze estimation step.
17. A computer program written in a computer-readable format to cause a computer to function as: a person estimation unit that estimates the person appearing in each frame of a video; a gaze estimation unit that estimates the gaze of the person appearing in each frame of a video; and a detection unit that combines the estimation results from the person estimation unit and the gaze estimation unit to detect a specific gaze relationship between people that occurred in each frame of a video.