Interaction method for AI online education

By requesting face detection and emotion analysis from the teacher's end to collaborate with the student's end, the problem of excessive server load in remote education was solved, and video stream analysis was enabled on the student's end with insufficient hardware configuration, thus improving teaching efficiency.

CN120852115AActive Publication Date: 2025-10-28HANGZHOU DIANZI UNIV +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511010824.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-28
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

In distance education, excessive server load can prevent face detection and emotion analysis from being completed when students' hardware configurations are insufficient, thus affecting the quality of teaching.

Method used

The teacher requests the student's side to perform sentiment annotation, and the student's side performs face detection and sentiment analysis, which reduces the server load and establishes a mapping table locally for collaborative processing.

Benefits of technology

It enables face detection and emotion analysis of video streams for students with insufficient hardware configuration, reducing server load and improving teaching efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852115A_ABST
    Figure CN120852115A_ABST
Patent Text Reader

Abstract

The invention discloses an interactive method for AI online education, which is applied to the situation that a student terminal sends own video streams to a teacher terminal through a server for playing and displaying, and the student terminal is insufficient in own capability and cannot analyze video pictures in the own video streams, and when a teacher plays and displays student video pictures through the teacher terminal, the student terminal can play and display the student video pictures through the server. And when the students are questioned, the teacher clicks a first button at the teacher end to obtain the emotion labels of the students. According to the interaction method for AI online education, through cooperation of the cooperative student terminal with strong capability and the student terminal needing cooperation with insufficient capability, face detection, extraction and sentiment analysis of a video stream picture of the student terminal needing cooperation are realized, so that the pressure of a server is reduced, each student terminal supports own face detection, extraction and sentiment analysis, and the interaction efficiency of the AI online education is improved. The problem that hardware configuration of a student end with multiple students is insufficient and cannot be completed in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of online education technology, and specifically relates to an interactive method for AI-powered online education. Background Technology

[0002] In a distance education scenario, teachers and students are located in different places and conduct teaching activities via remote video conferencing through an online education platform. After the remote class begins, the teacher's end and multiple student ends connect to the online classroom and each send its own video stream (the teacher's video stream is the content being explained by the teacher, captured by the teacher's camera; the student's video stream is the video view of the student in front of them, captured by the student's camera) unicast to the server. The server encodes and processes the video stream data and then multicasts it to the teacher's and student ends. The teacher's end receives and displays the video feeds from each student end.

[0003] For student devices, there are situations where one camera on a student device covers one or more students (e.g., five students are seated in front of one student device, and the camera on that device covers all five students and is connected to the remote education platform). There are also multiple student devices connected to the online education platform. When teachers need to call on students, they want to specify a particular student (e.g., someone whose expression shows understanding or confusion). To avoid teachers manually reviewing each student's facial expressions in their own video feed, which is time-consuming and laborious, teachers want to perform facial expression (emotional) analysis and annotation on each student's video feed on their own device to improve the efficiency of calling on students.

[0004] The current method involves the server performing face detection, name matching, and sentiment analysis on each student's video feed, followed by unified sentiment annotation. However, this approach often results in the server being overloaded, increasing its workload and impacting the quality of distance education. Therefore, a newer method involves the server instructing each student to perform face detection, name matching, and sentiment analysis on their respective video feeds, followed by unified sentiment annotation. While this method is suitable for single-student devices, it becomes unsuitable for devices with multiple students due to insufficient hardware capabilities. Summary of the Invention

[0005] The purpose of this invention is to address the problems raised in the background art by proposing an interactive method for AI-based online education.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] This invention proposes an interactive method for AI-powered online education. This method is applied when students send their video streams to teachers via a server for playback, but the students themselves lack the capability to analyze the video content within their streams. Specifically:

[0008] When a teacher plays and displays student videos on the teacher's end and asks questions to students, the teacher obtains student sentiment annotations by clicking the first button on the teacher's end. Then, the teacher's end sends a first request message for student sentiment annotation to the server, and the server then sends a second request message for sentiment self-annotation to all student ends.

[0009] All student terminals return a first response message to the server indicating whether they accept or reject the sentiment self-labeling request. The student terminal corresponding to the first response message that returns acceptance is designated as a collaborating student terminal, and the student terminal corresponding to the first response message that returns rejection is designated as a collaborating student terminal.

[0010] If a first response message returns a rejection, the server unicasts a notification message for collaborative sentiment analysis to all student terminals.

[0011] After all student terminals receive the notification message, all collaborating student terminals perform face detection, extraction, and sentiment analysis on the current video stream of each collaborating student terminal, respectively obtaining the coordinates of the face detection box, the feature vector of the face, and the sentiment category in the current video stream of each collaborating student terminal;

[0012] Each student terminal sends the coordinates of the face detection box, the feature vector of the face, and the emotion category in its current video stream to the server. The server obtains the student's name corresponding to the face based on the face feature vector and marks the corresponding student's name and emotion category above the coordinates of the face detection box in the video stream of each student terminal.

[0013] The server then sends the video stream, labeled with the student's basic information and emotional category, to the teacher's end.

[0014] Preferably, after all student terminals receive the second request message, each student terminal determines whether its own capabilities meet the preset conditions: each student terminal determines whether its CPU core count, memory capacity and video memory capacity are greater than its own preset configuration, and determines whether its CPU utilization rate, memory utilization rate and video memory utilization rate are less than its own preset values. When all conditions are met, the student terminal is designated as the cooperating student terminal and is called the first cooperating student terminal; otherwise, it is designated as the student terminal that needs to be cooperated with.

[0015] The notification message includes a message ID, a first multicast group address M for sending the first image frame of the current video stream of each student terminal that needs to cooperate, a list of multicast group addresses for sending the subsequent video streams after the first image frame of the current video stream of each student terminal that needs to cooperate, and a list of student terminal IDs that need to cooperate. The list of multicast group addresses contains a second multicast group address that corresponds one-to-one with each student terminal in the list of student terminal IDs that needs to cooperate.

[0016] All student terminals establish a mapping table for collaborative sentiment analysis locally. The mapping table includes records corresponding to the student terminal IDs that need to collaborate. These records include the student terminal IDs that need to collaborate, the first multicast group address M, the second multicast group address, a list of student terminal IDs that need to collaborate, a list of face detection box coordinates, a list of face feature vectors, and a list of face sentiment categories. It also includes a face detection timestamp that corresponds one-to-one with each face sentiment category in the face sentiment category list. These face detection timestamps form a face detection timestamp list. Initially, the list of student terminal IDs, the list of face detection box coordinates, the list of face feature vectors, the list of face sentiment categories, and the list of face detection timestamps are all empty, and the number of faces that need to collaborate is zero initially.

[0017] Preferably, when each first cooperating student terminal performs face detection, extraction and emotion analysis on the first image frame of the current video stream of the student terminal to be cooperated with, after all student terminals receive the notification message, they send an IGMP member report message to the router directly connected to their local area to join the first multicast group address M in the notification message;

[0018] For each student terminal that needs to collaborate, the student terminal that needs to collaborate sends a third request message containing the first image frame of the current video stream via the first multicast group address M in a multicast manner. After receiving the first image frame in the third request message, each first collaborative student terminal that has joined the first multicast group address M starts a random countdown timer with a preset number of seconds on its local machine, and the timing of the random countdown timer is different between each first collaborative student terminal.

[0019] Each student terminal performs face detection, extraction, and sentiment analysis on the faces in the first image frame according to the order in which the random countdown timer expires;

[0020] For the first collaborating student terminal when the random countdown timer expires, first retrieve the feature vectors in the feature vector list of the face in the record corresponding to the student terminal ID to be collaborated in the local mapping table. The number of feature vectors in the feature vector list of the face is 'a'. Then, obtain the coordinates of all face detection boxes in the first image frame after being sorted in a preset order.

[0021] The first collaborative student terminal extracts and performs sentiment analysis on the face corresponding to the coordinates of the (a+1)th face detection box, and obtains the feature vector and sentiment category of the face corresponding to the coordinates of the (a+1)th face detection box, that is, the feature vector and sentiment category of the (a+1)th face.

[0022] The first collaborative student terminal updates its local mapping table. Specifically, the first collaborative student terminal adds its ID to the collaborative student terminal ID list corresponding to the collaborative student terminal ID in the local mapping table, adds the feature vector of the (a+1)th face to the face feature vector list, and adds the timestamp of detecting the (a+1)th face to the face detection timestamp list.

[0023] The first collaborative student terminal sends a second response message for the collaborative sentiment analysis request via multicast through the first multicast group address M. The second response message includes a message ID, an associated request ID, the received response status, the current first collaborative student terminal ID, the student terminal ID that needs to collaborate, the number of faces that need to collaborate, the number of the currently detected (a+1)th face, the feature vector of the (a+1)th face, the sentiment category of the (a+1)th face, and the timestamp when the (a+1)th face was detected.

[0024] The cooperating student terminal and other first cooperating student terminals that have joined the first multicast group M update their respective local mapping tables according to the second response message sent by the current first cooperating student terminal. That is, they add the current first cooperating student terminal ID to the list of cooperating student terminal IDs of the record corresponding to the cooperating student terminal ID in their respective local mapping tables, add the feature vector of the (a+1)th face to the list of face feature vectors, and add the timestamp of detecting the (a+1)th face to the list of face detection timestamps. When the cooperating student terminal updates its local mapping table, it also adds the emotion category of the (a+1)th face to the list of face emotion categories in the mapping table.

[0025] Preferably, for the first collaborative student terminal when the random countdown timer expires, a face detection algorithm is used to detect all faces in the first image frame, obtain the detection box coordinates of each face, and count the number of faces as b. Then the actual number of first collaborative student terminals required is b.

[0026] When the first collaborating student updates its local mapping table, it also adds the count of faces b to the number of faces that need to be collaborated with the student in its local mapping table.

[0027] The second response message sent by the first collaborating student terminal at the first time via the first multicast group address M in a multicast manner also includes the coordinates of the detection boxes of all faces in the first image frame after being sorted in a preset order, i.e., the list of face detection box coordinates.

[0028] When the student client that needs to collaborate and other first collaborative student clients that have joined the first multicast group M update their local mapping tables according to the second response message sent by the first first collaborative student client that arrives, they also add the number of faces b to the number of faces of the student client that needs to collaborate in the record corresponding to the ID of the student client that needs to collaborate in their local mapping tables, and add the detection box coordinates of all faces sorted in a preset order to the list of face detection box coordinates.

[0029] Preferably, when the student terminal requiring collaboration and other first collaborative student terminals joining the first multicast group M receive the second response message about the corresponding student terminal ID requiring collaboration and update their respective local mapping tables, they determine whether the number of faces of the student terminal requiring collaboration corresponding to the corresponding student terminal ID in their respective local mapping tables is consistent with the number in the list of collaborative student terminal IDs. If they are consistent, it means that the face detection task corresponding to the third request message of the corresponding student terminal ID requiring collaboration has been completed, and other first collaborative student terminals outside the list of collaborative student terminal IDs stop counting down the random countdown timer for the corresponding student terminal ID requiring collaboration.

[0030] The student client needs to retrieve the local mapping table and send the analysis information to the server. The analysis information includes the student client ID, the feature vector of each face in the feature vector list, the detection box coordinates of each face in the detection box coordinate list, and the emotion category of each face in the emotion category list.

[0031] After receiving the analysis information, the server retrieves the feature vector of each face from the pre-saved student information table to obtain the student name corresponding to the feature vector of each face. The student information table includes the student ID, name, and the feature vector of the face.

[0032] The server overlays the student's name and emotion category onto the detection box coordinates of each face in the current video stream on the student's end, and then sends the overlaid current video stream to the teacher's end. The teacher then asks questions based on the student's name and emotion category.

[0033] Preferably, the student terminal requiring collaboration also sends the subsequent video stream of the first image frame via the second multicast group address in a multicast manner. After sending a second response message about the corresponding student terminal ID, each first collaborative student terminal sends an IGMP member report message to its locally directly connected router to join the second multicast group address of the corresponding student terminal ID.

[0034] Each first collaborative student terminal that has joined the second multicast group address receives the subsequent video stream of the first image frame through the second multicast group address. It then extracts a new frame from the subsequent video stream according to a preset first cycle and performs the following processing steps: Face detection and extraction are performed on all faces in the newly extracted frame to obtain the coordinates of the corresponding face detection bounding boxes and feature vectors. Simultaneously, the feature vectors of the faces it is responsible for, stored in its local mapping table, are compared one by one with the feature vectors of all faces in the newly extracted frame.

[0035] If the comparison is unsuccessful, the test will be terminated.

[0036] If the comparison is successful, each first collaborative student terminal that has been added to the second multicast group address retrieves a new frame to obtain the coordinates of the face detection box of the successfully compared face, performs sentiment analysis to obtain the sentiment category of the successfully compared face, and then sends a collaborative sentiment analysis update message to the corresponding collaborative student terminal ID. The update message includes message ID, first collaborative student terminal ID, number of faces to be compared by the student terminal, feature vector of the face currently handled by the first collaborative student terminal, coordinates of the face detection box currently handled by the first collaborative student terminal, sentiment category of the face currently handled by the first collaborative student terminal, timestamp of the face detection currently handled by the first collaborative student terminal, and a list of feature vectors of other faces besides those currently handled by the first collaborative student terminal.

[0037] After receiving the update message, the collaborating student terminal needs to update its local mapping table, that is, update the face detection box coordinates, face emotion category, and face detection timestamp of the first collaborating student terminal in the face detection box coordinate list of the local mapping table.

[0038] Simultaneously, the student client needs to compare the feature vector list of faces other than those currently handled by the first collaborative student client in the update message with the feature vector list of faces in the local mapping table. If the feature vectors of the faces in the former are all in the list of the latter, no processing is required. If the feature vectors of the faces in the former are not in the list of the latter, it means that a new student has been added in the newly retrieved frame. At this time, the student client needs to send a fourth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fourth request message includes a message ID, the ID of the student client that needs to collaborate, the second multicast group address, the first number of newly added faces c of the student client that needs to collaborate, and the feature vector list of the first newly added faces.

[0039] Preferably, after receiving the fourth request message, the collaborative student terminal that has joined the first multicast group address M determines whether the preset conditions are met. The collaborative student terminal that meets the preset conditions is called the second collaborative student terminal. Each second collaborative student terminal starts a random countdown timer of preset number of seconds locally. For the second collaborative student terminal whose countdown timer expires first, a third response message for adding collaborative sentiment analysis request is sent via multicast through the first multicast group address M. The third response message includes message ID, associated request ID, current second collaborative student terminal ID, received response status, the first number of new faces c to be added by the collaborative student terminal, and the feature vector list of the first new faces.

[0040] After receiving the third response message, the other second collaborative student terminals stop the countdown timer. After receiving the third response message, the collaborative student terminals need to update their local mapping table, that is, add the current second collaborative student terminal ID c times to the collaborative student terminal ID list, update the number of faces in the collaborative student terminal to b+c, and add the feature vectors of each newly added face to the feature vector list of faces.

[0041] If the second collaborative student terminal has not joined the second multicast group address, after sending the third response message, it sends an IGMP member report message to its directly connected router to join the second multicast group address of the corresponding student terminal that needs to collaborate. Then, it receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the coordinates of the face detection box corresponding to the feature vector in the feature vector list of the first newly added face, and performs sentiment analysis to obtain the sentiment category of the corresponding face.

[0042] At the same time, the current second collaborative student terminal updates its local mapping table. Specifically, the current second collaborative student terminal adds the ID of the current second collaborative student terminal c times to the list of collaborative student terminal IDs corresponding to the required collaborative student terminal ID in the local mapping table, adds the feature vectors of each newly added face to the feature vector list of faces, adds the emotion category of each newly added face to the emotion category list of faces, and adds the timestamp of each newly added face when it was detected to the face detection timestamp list.

[0043] Then, the current second collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal.

[0044] Preferably, the collaborative student terminal needs to retrieve the local mapping table according to the preset second cycle, and determine whether the difference between each face detection timestamp and the current timestamp in the face detection timestamp list is greater than the threshold. For each face detection timestamp, if it is less than the threshold, no processing is required. If it is greater than the threshold, it indicates that the current face detection timestamp has exceeded the threshold and has not been updated. It is also determined whether the collaborative student terminal corresponding to the current face detection timestamp is offline due to a fault.

[0045] If there is no offline fault, it means that the current face has left. At this time, the student terminal needs to delete the information about the current face from the mapping table. That is, delete the student terminal ID corresponding to the current face, the detection box coordinates of the current face, the feature vector of the current face, the emotion category of the current face, and the detection timestamp of the current face in the mapping table, and update the number of faces that need to be coordinated with the student terminal in the mapping table to the number of faces remaining after deletion.

[0046] If there is a failure and the student terminal is offline, the collaborative student terminal needs to be replaced. In this case, the collaborative student terminal needs to send a fifth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fifth request message includes a message ID, the ID of the student terminal to be collaborated with, the second multicast group address, the second number of newly added faces of the student terminal to be collaborated with, and the list of facial feature vectors of the second newly added faces. The list of facial feature vectors of the second newly added faces is the facial feature vectors of the collaborative student terminal that is offline due to the failure.

[0047] Preferably, after the other collaborative student terminals that have joined the first multicast group address M and are not offline due to a fault receive the fifth request message, they determine whether the preset conditions are met. The collaborative student terminals that meet the preset conditions are called the third collaborative student terminals. Each third collaborative student terminal starts a random countdown timer for a preset number of seconds locally. For the third collaborative student terminal whose countdown timer expires first, a fourth response message for adding collaborative sentiment analysis request is sent via multicast through the first multicast group address M. The fourth response message includes message ID, associated request ID, current third collaborative student terminal ID, received response status, the number of second newly added faces to be added by the collaborative student terminal, and the feature vector list of the second newly added faces.

[0048] After receiving the fourth response message, other third collaborative student terminals stop the countdown timer. After receiving the fourth response message, the collaborative student terminal needs to update its local mapping table, that is, replace the collaborative student terminal ID that is offline with the third collaborative student terminal ID whose countdown timer expires first in the collaborative student terminal ID list.

[0049] If the current third collaborative student terminal has not joined the second multicast group address, after sending the fourth response message, it sends an IGMP member report message to its directly connected router to join the second multicast group address of the corresponding student terminal that needs to collaborate. Then, it receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the coordinates of the face detection box corresponding to the feature vector in the feature vector list of the second newly added face, and performs sentiment analysis to obtain the sentiment category of the second newly added face.

[0050] At the same time, the current third collaborative student terminal updates its local mapping table. Specifically, the current third collaborative student terminal replaces the collaborative student terminal ID that is offline with the current third collaborative student terminal ID in the collaborative student terminal ID list corresponding to the collaborative student terminal ID in the local mapping table; replaces the face emotion category of the collaborative student terminal ID that is offline with the emotion category of the second newly added face in the face emotion category list; and replaces the face detection timestamp of the collaborative student terminal ID that is offline with the timestamp when the second newly added face was detected in the face detection timestamp list.

[0051] Then, the current third collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal.

[0052] Preferably, when all face detection timestamps in the face detection timestamp list of the local mapping table of the student client needing to collaborate are updated, the student client needs to collaborate to retrieve the local mapping table and send analysis information to the server. The analysis information includes the ID of the student client needing to collaborate, the feature vector of each face in the feature vector list, the detection box coordinates of each face in the face detection box coordinate list, and the emotion category of each face in the emotion category list.

[0053] After receiving the analysis information, the server retrieves the feature vector of each face from the pre-saved student information table to obtain the student name corresponding to the feature vector of each face. The student information table includes the student ID, name, and the feature vector of the face.

[0054] The server overlays the student's name and emotion category onto the detection box coordinates of each face in the current video stream on the student's end, and then sends the overlaid current video stream to the teacher's end. The teacher then asks questions based on the student's name and emotion category.

[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0056] This AI-powered online education interactive method enables high-capability collaborative student terminals to collaborate with low-capability student terminals that require collaboration. This allows for face detection, extraction, and sentiment analysis of the video stream images from the student terminals requiring collaboration, thereby reducing server load. Simultaneously, it enables each student terminal to support its own face detection, extraction, and sentiment analysis, solving the problem in existing technologies where student terminals with multiple students often lack sufficient hardware configuration (capacity) to complete the task. Attached Figure Description

[0057] Figure 1 This is a flowchart of the interactive method for AI online education according to the present invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0059] In one embodiment, such as Figure 1 As shown, an interactive method for AI-powered online education is provided. This method is applied when students send their video streams to teachers via a server for playback, and the students themselves lack the capability to analyze the video footage in their streams (this application omits cases where students with strong capabilities can perform face detection, extraction, and sentiment analysis on their own video streams, as these are existing technologies). Specifically:

[0060] Step 1: When the teacher plays and displays student videos on the teacher's end and asks questions to the students, the teacher obtains student emotion annotations by clicking the first button on the teacher's end. Then, the teacher's end sends a first request message for student emotion annotation to the server (including request ID, request type (student emotion annotation request), teacher's end ID, and course information). The server then sends a second request message for emotion self-annotation to all student ends (including request ID, request type (emotion self-annotation request), and request server IP).

[0061] Step 2: All student terminals return a first response message to the server indicating whether they accept or reject the sentiment self-labeling request. The student terminal corresponding to the first response message that returns "accept" is designated as a collaborating student terminal; the student terminal corresponding to the first response message that returns "reject" is designated as a collaborating student terminal.

[0062] After receiving the second request message, each student terminal determines whether its own capabilities meet the preset conditions: each student terminal determines whether its CPU core count, memory capacity, and video memory capacity are greater than its preset configuration (e.g., in this embodiment, the preset configuration is 4 CPU cores, 16GB memory, and 4GB video memory), and whether its CPU utilization rate, memory utilization rate, and video memory utilization rate are less than its preset values ​​(e.g., in this embodiment, the preset values ​​are 70% CPU utilization rate, 75% memory utilization rate, and 80% video memory utilization rate). If all conditions are met, the student terminal is designated as the collaborating student terminal and is called the first collaborating student terminal, and returns an accepted first response message to the server; otherwise, it is designated as the student terminal requiring collaboration and returns a rejected first response message to the server.

[0063] Step 3: When a first response message returns a rejection, the server unicasts a collaborative sentiment analysis notification message to all student terminals. The notification message includes a message ID, a first multicast group address M for sending the first image frame of the current video stream to each student terminal that needs to collaborate (where the first image frame of the current video stream is the first image frame of the video stream that the student terminal that needs to collaborate needs to send at the current moment), a list of multicast group addresses for sending the video streams following the first image frame of the current video stream to each student terminal that needs to collaborate, and a list of student terminal IDs that need to collaborate. The list of multicast group addresses contains a second multicast group address that corresponds one-to-one with each student terminal in the list of student terminal IDs that needs to collaborate.

[0064] Step 4: After all student terminals receive the notification message, all collaborating student terminals perform face detection, extraction, and sentiment analysis on the current video stream of each collaborating student terminal. This yields the coordinates of the face detection box, the feature vector of the face, and the sentiment category in the current video stream of each collaborating student terminal. (The face detection, extraction, and sentiment analysis processes are all existing technologies. For example, face detection uses face detection algorithms (such as YOLO networks), obtaining the coordinates of the face detection box (i.e., the four corner coordinates of the face detection box) for each face in the image frame, extracting the feature vector of each face, and using an expression recognition algorithm (such as a CNN model) to perform expression recognition on the detected faces to analyze the emotional state of the students' learning comprehension (emotions can be categorized, such as: understanding, confusion, neutral, other)).

[0065] Step 4.1: When each of the first cooperating student terminals performs face detection, extraction, and sentiment analysis on the first image frame of the current video stream of the student terminal requiring cooperation, all student terminals, after receiving the notification message, send an IGMP member report message to their locally directly connected router to add the first multicast group address M in the notification message (for cooperating student terminals, adding the first multicast group address M is to receive the first image frame; for student terminals requiring cooperation, adding the first multicast group address M is to receive subsequent information sent through the first multicast group address M. After receiving the IGMP report, the student terminal requiring cooperation and each directly connected router of the cooperating student terminal update their IGMP group table (including the first multicast group address M, interface information, group member status, and timers, etc.), recording that a member has added the first multicast group address M for image frame forwarding on the corresponding interface of the router. After updating the IGMP group table, the directly connected router communicates with other routers through the multicast routing protocol to establish and maintain multicast forwarding table entries).

[0066] Specifically, all student terminals establish a local mapping table for collaborative sentiment analysis. This table includes records corresponding to the student terminal IDs requiring collaboration. Each record includes the student terminal ID, the first multicast group address M, the second multicast group address, a list of collaborative student terminal IDs (adding all collaborative student terminals corresponding to the required student terminal to this list), the number of faces in the required student terminal's video stream, and a list of face detection box coordinates (adding the coordinates of all face detection boxes in the required student terminal's video stream). The table contains coordinates for each face detection bounding box, including the coordinates of the four corners of the bounding box. The number of face detection bounding box coordinates in the table matches the number of members in the collaborative student ID list. Furthermore, the face detection bounding box coordinates in the table correspond one-to-one with the collaborative student IDs in the collaborative student ID list (the same applies below). It also includes a list of face feature vectors (i.e., all face feature vectors in the collaborative student video stream are added to this list, and the number of face feature vectors in this list matches the number of members in the collaborative student ID list). The list includes a list of facial feature vectors that correspond one-to-one with the collaborative student IDs in the collaborative student ID list, a list of facial emotion categories (i.e., the emotion categories of all faces in the collaborative student video stream are added to this list, and the number of facial emotion categories in the list matches the number of members in the collaborative student ID list, and the facial emotion categories in the list correspond one-to-one with the collaborative student IDs in the collaborative student ID list), and a face detection timestamp that corresponds one-to-one with each facial emotion category in the facial emotion category list. The timestamps constitute a list of face detection timestamps (the same applies below, i.e., all face detection timestamp lists mentioned in this application are like this, and will not be described repeatedly). Among them, the list of collaborating student IDs, the list of face detection box coordinates, the list of face feature vectors, the list of face emotion categories, and the list of face detection timestamps are all empty when the mapping table is initially established, and the number of faces that need to be collaborated with the student is zero when the mapping table is initially established (it should be noted that the mapping table of the student that needs to be collaborated with has only one record about itself, while the mapping table of the collaborating student has multiple records that correspond one-to-one with all the students that need to be collaborated with).

[0067] Step 4.2: For each student terminal requiring collaboration, the student terminal sending a third request message containing the first image frame of the current video stream via multicast through the first multicast group address M (each router in the network copies and forwards the third request message to all collaborative student terminals that have joined the first multicast group address M according to the multicast routing table entry and multicast forwarding table entry). After receiving the first image frame in the third request message, each first collaborative student terminal that has joined the first multicast group address M starts a random countdown timer with a preset number of seconds on its local machine (such as a 5-second random countdown timer, i.e., randomly reaching the time within 0-5 seconds), and the timing of the random countdown timer is different between the local machines of each first collaborative student terminal.

[0068] Each first collaborative student terminal performs face detection, extraction, and sentiment analysis on the face in the first image frame according to the order in which the random countdown timer expires (sorted according to the size of the upper left corner coordinates (X, Y) of each detection box, i.e., first compare the upper left corner coordinate X, and the face detection box with the smaller X is placed first; if X is equal, then compare Y, and the face detection box with the smaller Y is placed first).

[0069] For the current first collaborative student terminal when the random countdown timer expires (i.e., for each first collaborative student terminal), first retrieve the feature vector list of the face in the record corresponding to the student terminal ID to be collaborated with in the local mapping table. The number of feature vectors in the feature vector list of the face is a. Then, obtain the coordinates of all face detection boxes in the first image frame after being sorted in a preset order.

[0070] The first collaborative student terminal extracts and performs sentiment analysis on the face corresponding to the coordinates of the (a+1)th face detection box, and obtains the feature vector and sentiment category of the face corresponding to the coordinates of the (a+1)th face detection box, that is, the feature vector and sentiment category of the (a+1)th face.

[0071] The first collaborative student terminal updates its local mapping table. Specifically, the first collaborative student terminal adds its ID to the collaborative student terminal ID list corresponding to the collaborative student terminal ID in the local mapping table, adds the feature vector of the (a+1)th face to the face feature vector list, and adds the timestamp of detecting the (a+1)th face to the face detection timestamp list.

[0072] The first collaborative student terminal sends a second response message for the collaborative sentiment analysis request via multicast through the first multicast group address M. The second response message includes a message ID, an associated request ID (i.e., the associated third request message ID), the received response status, the current first collaborative student terminal ID, the student terminal ID that needs to collaborate, the number of faces that need to collaborate, the number of the currently detected (a+1)th face, the feature vector of the (a+1)th face, the sentiment category of the (a+1)th face, and the timestamp when the (a+1)th face was detected.

[0073] The coordinating student client and other first coordinating student clients joining the first multicast group M update their local mapping tables based on the second response message sent by the current first coordinating student client. Specifically, they add the current first coordinating student client ID to the coordinating student client ID list corresponding to the coordinating student client ID in their local mapping table, add the feature vector of the (a+1)th face to the face feature vector list, and add the timestamp of detecting the (a+1)th face to the face detection timestamp list. Furthermore, when updating their local mapping tables, the coordinating student client also adds the emotion category of the (a+1)th face to the emotion category list of faces in the mapping table. It should be noted that each first coordinating student client performs the above operations, but for the first coordinating student client whose random countdown timer expires, there are additional operations besides the above:

[0074] Among them, for the first collaborative student terminal when the random countdown timer expires, a face detection algorithm is used to detect all faces in the first image frame, obtain the detection box coordinates of each face, and count the number of faces as b. Then the actual number of first collaborative student terminals required is b.

[0075] When the first collaborating student updates its local mapping table, it also adds the count of faces b to the number of faces that need to be collaborated with the student in its local mapping table.

[0076] The second response message sent by the first collaborating student terminal at the first time via the first multicast group address M in a multicast manner also includes the coordinates of the detection boxes of all faces in the first image frame after being sorted in a preset order, i.e., the list of face detection box coordinates.

[0077] When the student client that needs to collaborate and other first collaborative student clients that have joined the first multicast group M update their local mapping tables according to the second response message sent by the first first collaborative student client that arrives, they also add the number of faces b to the number of faces of the student client that needs to collaborate in the record corresponding to the ID of the student client that needs to collaborate in their local mapping tables, and add the detection box coordinates of all faces sorted in a preset order to the list of face detection box coordinates.

[0078] Step 4.3: When the student client requiring collaboration and other first collaborative student clients joining the first multicast group M receive the second response message regarding the corresponding student client ID, and update their respective local mapping tables, they determine whether the number of faces of the student client requiring collaboration corresponding to the student client ID in their respective local mapping tables is consistent with the number in the list of collaborative student client IDs. If they are consistent, it means that the face detection task corresponding to the third request message of the corresponding student client ID has been completed. Then, other first collaborative student clients other than those in the list of collaborative student client IDs stop counting down the random countdown timer for the corresponding student client ID (other first collaborative student clients still save the mapping table locally and do not delete it).

[0079] The student client needs to retrieve the local mapping table and send the analysis information to the server. The analysis information includes the student client ID, the feature vector of each face in the feature vector list, the detection box coordinates of each face in the detection box coordinate list, and the emotion category of each face in the emotion category list.

[0080] After receiving the analysis information, the server retrieves the feature vector of each face from the pre-saved student information table to obtain the student name corresponding to the feature vector of each face. The student information table includes the student ID, name, and the feature vector of the face.

[0081] Based on the coordinates of the face detection boxes, the server overlays the student's name and the emotion category of each face onto the detection box coordinates of each face in the current video stream (i.e., each frame of the current video stream) on the student's end. The server then sends the overlaid current video stream to the teacher's end, where the teacher asks questions based on the student's name and emotion category.

[0082] The above describes the process of detecting faces in each collaborating student terminal for the first image frame, and the collaborating student terminal sending the analysis information about the first image frame overlaid in the video stream to the server.

[0083] The following is the subsequent video stream for the first image frame:

[0084] Step 4.4: Step 4.4.1, the student terminal that needs to cooperate also sends the subsequent video stream of the first image frame in multicast mode through the second multicast group address. After sending the second response message about the corresponding student terminal ID, each first student terminal sends an IGMP membership report message to its locally directly connected router to join the second multicast group address of the corresponding student terminal ID (obtained from its respective mapping table).

[0085] Step 4.4.2: After receiving the subsequent video stream of the first image frame through the second multicast group address, each first collaborative student terminal that has joined the second multicast group address extracts a new frame from the subsequent video stream according to a preset first period (e.g., 10 seconds). (Since the detection time of the first video frame is different for each first collaborative student terminal, the newly extracted frame for each first collaborative student terminal is also different.) The following processing steps are performed: face detection and extraction are performed on all faces in the newly extracted frame to obtain the coordinates of the corresponding face detection box and feature vector. At the same time, the feature vectors of the faces they are responsible for, stored in their local mapping tables, are compared one by one with the feature vectors of all faces in the newly extracted frame.

[0086] Step 4.4.3: If the comparison is unsuccessful (for example, the face (i.e., the student) being handled by the first collaborative student terminal has left the frame, so the feature vector of the face is not detected), then the current detection is terminated; proceed to the next cycle, return to step 4.4.2 and the first collaborative student terminal continues to extract one frame for loop operation;

[0087] Step 4.4.4: If the comparison is successful, each first collaborative student terminal that has joined the second multicast group address retrieves a new frame to obtain the coordinates of the face detection boxes of the successfully compared faces, performs sentiment analysis to obtain the sentiment category of the successfully compared faces, and then sends a collaborative sentiment analysis update message to the corresponding collaborative student terminal ID. The update message includes the message ID, the first collaborative student terminal ID, the number of faces to be compared by the student terminal, the face feature vector currently handled by the first collaborative student terminal, the coordinates of the face detection boxes currently handled by the first collaborative student terminal, the sentiment category of the face currently handled by the first collaborative student terminal, the face detection timestamp currently handled by the first collaborative student terminal, and a list of feature vectors of other faces besides those handled by the current first collaborative student terminal (i.e., other face feature vectors identified from the image besides those handled by the current first collaborative student terminal itself).

[0088] Step 4.4.5: After receiving the update message, the collaborating student terminal needs to update the local mapping table, that is, update the face detection box coordinates, face emotion category, and face detection timestamp of the first collaborating student terminal in the face detection box coordinate list of the local mapping table.

[0089] Step 4.4.6: Simultaneously, the collaborative student terminal needs to compare the feature vector list of faces other than those currently handled by the first collaborative student terminal in the update message with the feature vector list of faces in the local mapping table. If the feature vectors of the faces in the former are all in the list of the latter (indicating that no new faces have appeared in the image), no processing is required; proceed to the next preset first cycle, return to step 4.4.2 and the first collaborative student terminal continues to extract one frame for loop operation;

[0090] Step 4.4.6: If the facial feature vector of the former is not in the list of the latter, it means that a new student has been added in the newly extracted frame. At this time, the student terminal needs to send the fourth request message for new collaborative sentiment analysis in multicast mode through the first multicast group address M. The fourth request message includes message ID, ID of the student terminal to be collaborated with, second multicast group address, the first number of newly added faces c of the student terminal to be collaborated with, and the feature vector list of the first newly added faces.

[0091] Step 4.5: Step 4.5.1 After receiving the fourth request message, the collaborative student terminal that has joined the first multicast group address M determines whether it meets the preset conditions (i.e., whether its own capabilities meet the preset conditions, see step 2). The collaborative student terminal that meets the preset conditions is called the second collaborative student terminal. Each second collaborative student terminal starts a random countdown timer for a preset number of seconds locally. For the second collaborative student terminal whose countdown timer expires first, a third response message for adding collaborative sentiment analysis request is sent via multicast through the first multicast group address M. The third response message includes message ID, associated request ID (i.e., associated fourth request message), current second collaborative student terminal ID, received response status, the first number of new faces c to be added by the collaborative student terminal, and the feature vector list of the first new faces.

[0092] Step 4.5.2: After receiving the third response message, the other second collaborative student terminals stop the countdown timer. After receiving the third response message, the collaborative student terminals need to update their local mapping table, that is, add the current second collaborative student terminal ID c times to the collaborative student terminal ID list, update the number of faces in the collaborative student terminal to b+c, and add the feature vectors of each newly added face to the feature vector list of faces.

[0093] Step 4.5.3: If the current second collaborative student terminal has not joined the second multicast group address, after sending the third response message, it sends an IGMP member report message to its directly connected router to join the second multicast group address of the corresponding student terminal that needs to collaborate. Then, it receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the coordinates of the face detection box corresponding to the feature vector in the feature vector list of the first newly added face, and performs sentiment analysis to obtain the sentiment category of the corresponding face.

[0094] Step 4.5.4: Simultaneously, the current second collaborative student terminal updates its local mapping table. That is, the current second collaborative student terminal adds c times the ID of the current second collaborative student terminal to the list of collaborative student terminal IDs corresponding to the required collaborative student terminal IDs in the local mapping table, adds the feature vectors of each newly added face to the feature vector list of faces, adds the emotion category of each newly added face to the emotion category list of faces, and adds the timestamp of detecting each newly added face to the face detection timestamp list.

[0095] Step 4.5.6: Then, the current second collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal (that is, the second collaborative student terminal operates in the manner of steps 4.4.2-4.5 of the first collaborative student terminal).

[0096] Step 4.6: Step 4.6.1, the collaborative student terminal needs to retrieve the local mapping table according to the preset second cycle (e.g., every 10 seconds) and determine whether the difference between each face detection timestamp and the current timestamp in the face detection timestamp list is greater than the threshold (e.g., 20 seconds). For each face detection timestamp, if it is less than the threshold, no processing is required (then enter the next preset first cycle and return to step 4.4.2 for the first collaborative student terminal to continue to retrieve a frame for loop operation); if it is greater than the threshold, it indicates that the current face detection timestamp has exceeded the threshold and has not been updated, and it is determined whether the collaborative student terminal corresponding to the current face detection timestamp has a fault and is offline (the collaborative student terminal needs to send an ICMP Echo request (i.e., Ping test) to the collaborative student terminal. If the collaborative student terminal responds, there is no fault and offline; if there is no response, it indicates that there is a fault and offline).

[0097] Step 4.6.2: If there is no offline fault, it means that the current face has left. At this time, the student terminal needs to delete the information about the current face from the mapping table, that is, delete the student terminal ID corresponding to the current face, the detection box coordinates of the current face, the feature vector of the current face, the emotion category of the current face, and the detection timestamp of the current face in the mapping table, and update the number of faces that need to be coordinated with the student terminal in the mapping table to the number of faces remaining after deletion; then execute step (5);

[0098] Step 4.6.3: If there is a fault and the student terminal is offline, the collaborative student terminal needs to be replaced. At this time, the collaborative student terminal needs to send a fifth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fifth request message includes message ID, ID of the student terminal to be collaborated with, second multicast group address, second number of newly added faces of the student terminal to be collaborated with, and a list of facial feature vectors of the second newly added faces. The list of facial feature vectors of the second newly added faces is the facial feature vectors of the collaborative student terminal that is offline due to the fault.

[0099] Step 4.7: Step 4.7.1 After the other collaborative student terminals that have joined the first multicast group address M and are offline except for those that have failed, receive the fifth request message, they determine whether the preset conditions are met. Collaborative student terminals that meet the preset conditions are called third collaborative student terminals. Each third collaborative student terminal starts a random countdown timer for a preset number of seconds locally. For the third collaborative student terminal whose countdown timer expires first, a fourth response message for adding collaborative sentiment analysis request is sent via multicast through the first multicast group address M. The fourth response message includes message ID, associated request ID (i.e., associated fifth request message), current third collaborative student terminal ID, received response status, the number of second newly added faces to be added by the collaborative student terminal, and the feature vector list of the second newly added faces.

[0100] Step 4.7.2: After receiving the fourth response message, other third collaborative student terminals stop the countdown timer. After receiving the fourth response message, the collaborative student terminal needs to update its local mapping table, that is, replace the collaborative student terminal ID that is offline with the third collaborative student terminal ID whose countdown timer expires first in the collaborative student terminal ID list.

[0101] Step 4.7.3: If the current third collaborative student terminal has not joined the second multicast group address, after sending the fourth response message, it sends an IGMP member report message to its directly connected router to join the second multicast group address of the corresponding student terminal that needs to collaborate. Then, it receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the coordinates of the face detection box corresponding to the feature vector in the feature vector list of the second newly added face, and performs sentiment analysis to obtain the sentiment category of the second newly added face.

[0102] Step 4.7.4: Simultaneously, the current third collaborative student terminal updates its local mapping table. That is, the current third collaborative student terminal replaces the collaborative student terminal ID with the current third collaborative student terminal ID in the collaborative student terminal ID list corresponding to the collaborative student terminal ID in the local mapping table; replaces the face emotion category with the emotion category of the second newly added face in the face emotion category list; and replaces the face detection timestamp with the timestamp when detecting the second newly added face in the face detection timestamp list.

[0103] Step 4.7.5: Then, the current third collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal (that is, the third collaborative student terminal operates in the manner of steps 4.4.2-4.5 of the first collaborative student terminal).

[0104] Enter the next preset second cycle, return to step 4.6.1, and repeat the operation.

[0105] Step 5: When all face detection timestamps in the face detection timestamp list of the local mapping table of the student client needing to collaborate are updated, the student client needs to collaborate to retrieve the local mapping table and send the analysis information to the server. The analysis information includes the ID of the student client needing to collaborate, the feature vector of each face in the feature vector list, the detection box coordinates of each face in the face detection box coordinate list, and the emotion category of each face in the emotion category list.

[0106] After receiving the analysis information, the server retrieves the feature vector of each face from the pre-saved student information table to obtain the student name corresponding to the feature vector of each face. The student information table includes the student ID, name, and the feature vector of the face.

[0107] The server overlays the student's name and emotion category onto the detection box coordinates of each face in the current video stream on the student's end, and sends the overlaid current video stream to the teacher's end. The teacher then asks questions based on the student's name and emotion category.

[0108] Enter the next preset first cycle, return to step 4.4.2 and continue to extract one frame for loop operation from the first collaborative student terminal.

[0109] In another embodiment, this application also includes an interactive device for AI online education, including a processor and a memory storing a plurality of computer instructions, wherein the computer instructions, when executed by the processor, implement the steps of the method in steps 1-5. Specific limitations regarding the interactive device for AI online education can be found in the limitations of the interactive method for AI online education described above, and will not be repeated here.

[0110] This AI-powered online education interactive method enables high-capability collaborative student terminals to collaborate with low-capability student terminals that require collaboration. This allows for face detection, extraction, and sentiment analysis of the video stream images from the student terminals requiring collaboration, thereby reducing server load. Simultaneously, it enables each student terminal to support its own face detection, extraction, and sentiment analysis, solving the problem in existing technologies where student terminals with multiple students often lack sufficient hardware configuration (capacity) to complete the task.

[0111] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0112] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. An interactive method for AI-powered online education, characterized by: The method is applied when students send their video streams to teachers via a server for playback, but the students themselves lack the capability to analyze the video frames within their own streams. When a teacher plays and displays student videos on the teacher's end and asks questions to students, the teacher obtains student sentiment annotations by clicking the first button on the teacher's end. Then, the teacher's end sends a first request message for student sentiment annotation to the server, and the server then sends a second request message for sentiment self-annotation to all student ends. All student terminals return a first response message to the server indicating whether they accept or reject the sentiment self-labeling request. The student terminal corresponding to the first response message that returns acceptance is designated as a collaborating student terminal, and the student terminal corresponding to the first response message that returns rejection is designated as a collaborating student terminal. If a first response message returns a rejection, the server unicasts a notification message for collaborative sentiment analysis to all student terminals. After all student terminals receive the notification message, all collaborating student terminals perform face detection, extraction, and sentiment analysis on the current video stream of each collaborating student terminal, respectively obtaining the coordinates of the face detection box, the feature vector of the face, and the sentiment category in the current video stream of each collaborating student terminal; Each student terminal sends the coordinates of the face detection box, the feature vector of the face, and the emotion category in its current video stream to the server. The server obtains the student's name corresponding to the face based on the face feature vector and marks the corresponding student's name and emotion category above the coordinates of the face detection box in the video stream of each student terminal. The server then sends the video stream, labeled with the student's basic information and emotional category, to the teacher's end.

2. The interactive method for AI online education as described in claim 1, characterized in that: After all student terminals receive the second request message, each student terminal determines whether its own capabilities meet the preset conditions: each student terminal determines whether its CPU core count, memory capacity, and video memory capacity are greater than its own preset configuration, and whether its CPU utilization rate, memory utilization rate, and video memory utilization rate are less than its own preset value. If all conditions are met, the student terminal is designated as the collaborating student terminal and is called the first collaborating student terminal; otherwise, it is designated as the collaborating student terminal. The notification message includes a message ID, a first multicast group address M for sending the first image frame of the current video stream of each student terminal that needs to cooperate, a list of multicast group addresses for sending the subsequent video streams after the first image frame of the current video stream of each student terminal that needs to cooperate, and a list of student terminal IDs that need to cooperate. The list of multicast group addresses contains a second multicast group address that corresponds one-to-one with each student terminal in the list of student terminal IDs that needs to cooperate. All student terminals establish a mapping table for collaborative sentiment analysis locally. The mapping table includes records corresponding to the student terminal IDs that need to collaborate. These records include the student terminal IDs that need to collaborate, the first multicast group address M, the second multicast group address, a list of student terminal IDs that need to collaborate, a list of face detection box coordinates, a list of face feature vectors, and a list of face sentiment categories. It also includes a face detection timestamp that corresponds one-to-one with each face sentiment category in the face sentiment category list. These face detection timestamps form a face detection timestamp list. Initially, the list of student terminal IDs, the list of face detection box coordinates, the list of face feature vectors, the list of face sentiment categories, and the list of face detection timestamps are all empty, and the number of faces that need to collaborate is zero initially.

3. The interactive method for AI online education as described in claim 1, characterized in that: When each of the first collaborative student terminals performs face detection, extraction, and sentiment analysis on the first image frame of the current video stream of the student terminal to be collaborated with, after receiving the notification message, all student terminals send an IGMP member report message to the router directly connected to their local area to join the first multicast group address M in the notification message; For each student terminal that needs to collaborate, the student terminal that needs to collaborate sends a third request message containing the first image frame of the current video stream via the first multicast group address M in a multicast manner. After receiving the first image frame in the third request message, each first collaborative student terminal that has joined the first multicast group address M starts a random countdown timer with a preset number of seconds on its local machine, and the timing of the random countdown timer is different between each first collaborative student terminal. Each student terminal performs face detection, extraction, and sentiment analysis on the faces in the first image frame according to the order in which the random countdown timer expires; For the first collaborating student terminal when the random countdown timer expires, first retrieve the feature vectors in the feature vector list of the face in the record corresponding to the student terminal ID to be collaborated in the local mapping table. The number of feature vectors in the feature vector list of the face is 'a'. Then, obtain the coordinates of all face detection boxes in the first image frame after being sorted in a preset order. The first collaborative student terminal extracts and performs sentiment analysis on the face corresponding to the coordinates of the (a+1)th face detection box, and obtains the feature vector and sentiment category of the face corresponding to the coordinates of the (a+1)th face detection box, that is, the feature vector and sentiment category of the (a+1)th face. The first collaborative student terminal updates its local mapping table. Specifically, the first collaborative student terminal adds its ID to the collaborative student terminal ID list corresponding to the collaborative student terminal ID in the local mapping table, adds the feature vector of the (a+1)th face to the face feature vector list, and adds the timestamp of detecting the (a+1)th face to the face detection timestamp list. The first collaborative student terminal sends a second response message for the collaborative sentiment analysis request via multicast through the first multicast group address M. The second response message includes a message ID, an associated request ID, the received response status, the current first collaborative student terminal ID, the student terminal ID that needs to collaborate, the number of faces that need to collaborate, the number of the currently detected (a+1)th face, the feature vector of the (a+1)th face, the sentiment category of the (a+1)th face, and the timestamp when the (a+1)th face was detected. The cooperating student terminal and other first cooperating student terminals that have joined the first multicast group M update their respective local mapping tables according to the second response message sent by the current first cooperating student terminal. That is, they add the current first cooperating student terminal ID to the list of cooperating student terminal IDs of the record corresponding to the cooperating student terminal ID in their respective local mapping tables, add the feature vector of the (a+1)th face to the list of face feature vectors, and add the timestamp of detecting the (a+1)th face to the list of face detection timestamps. When the cooperating student terminal updates its local mapping table, it also adds the emotion category of the (a+1)th face to the list of face emotion categories in the mapping table.

4. The interactive method for AI online education as described in claim 3, characterized in that: For the first collaborative student terminal whose random countdown timer expires, a face detection algorithm is used to detect all faces in the first image frame, obtain the detection box coordinates of each face, and count the number of faces as b. Then the actual number of first collaborative student terminals required is b. When the first collaborating student updates its local mapping table, it also adds the count of faces b to the number of faces that need to be collaborated with the student in its local mapping table. The second response message sent by the first collaborating student terminal at the first time via the first multicast group address M in a multicast manner also includes the coordinates of the detection boxes of all faces in the first image frame after being sorted in a preset order, i.e., the list of face detection box coordinates. When the student client that needs to collaborate and other first collaborative student clients that have joined the first multicast group M update their local mapping tables according to the second response message sent by the first first collaborative student client that arrives, they also add the number of faces b to the number of faces of the student client that needs to collaborate in the record corresponding to the ID of the student client that needs to collaborate in their local mapping tables, and add the detection box coordinates of all faces sorted in a preset order to the list of face detection box coordinates.

5. The interactive method for AI online education as described in claim 4, characterized in that: When the student client requiring collaboration and other first collaborative student clients joining the first multicast group M receive the second response message about the corresponding student client ID and update their respective local mapping tables, they determine whether the number of faces of the student client corresponding to the corresponding student client ID in their respective local mapping tables is consistent with the number in the collaborative student client ID list. If they are consistent, it means that the face detection task corresponding to the third request message of the corresponding student client ID has been completed, and other first collaborative student clients other than those in the collaborative student client ID list stop counting down the random countdown timer for the corresponding student client ID. The student client needs to retrieve the local mapping table and send the analysis information to the server. The analysis information includes the student client ID, the feature vector of each face in the feature vector list, the detection box coordinates of each face in the detection box coordinate list, and the emotion category of each face in the emotion category list. After receiving the analysis information, the server retrieves the feature vector of each face from the pre-saved student information table to obtain the student name corresponding to the feature vector of each face. The student information table includes the student ID, name, and the feature vector of the face. The server overlays the student's name and emotion category onto the detection box coordinates of each face in the current video stream on the student's end, and then sends the overlaid current video stream to the teacher's end. The teacher then asks questions based on the student's name and emotion category.

6. The interactive method for AI online education as described in claim 5, characterized in that: The student client that needs to cooperate also sends the subsequent video stream of the first image frame in multicast mode through the second multicast group address. After sending the second response message about the corresponding student client ID, each first student client sends an IGMP member report message to its locally directly connected router to join the second multicast group address of the corresponding student client ID. Each first collaborative student terminal that has joined the second multicast group address receives the subsequent video stream of the first image frame through the second multicast group address. It then extracts a new frame from the subsequent video stream according to a preset first cycle and performs the following processing steps: Face detection and extraction are performed on all faces in the newly extracted frame to obtain the coordinates of the corresponding face detection bounding boxes and feature vectors. Simultaneously, the feature vectors of the faces it is responsible for, stored in its local mapping table, are compared one by one with the feature vectors of all faces in the newly extracted frame. If the comparison is unsuccessful, the test will be terminated. If the comparison is successful, each first collaborative student terminal that has been added to the second multicast group address retrieves a new frame to obtain the coordinates of the face detection box of the successfully compared face, performs sentiment analysis to obtain the sentiment category of the successfully compared face, and then sends a collaborative sentiment analysis update message to the corresponding collaborative student terminal ID. The update message includes message ID, first collaborative student terminal ID, number of faces to be compared by the student terminal, feature vector of the face currently handled by the first collaborative student terminal, coordinates of the face detection box currently handled by the first collaborative student terminal, sentiment category of the face currently handled by the first collaborative student terminal, timestamp of the face detection currently handled by the first collaborative student terminal, and a list of feature vectors of other faces besides those currently handled by the first collaborative student terminal. After receiving the update message, the collaborating student terminal needs to update its local mapping table, that is, update the face detection box coordinates, face emotion category, and face detection timestamp of the first collaborating student terminal in the face detection box coordinate list of the local mapping table. Simultaneously, the student client needs to compare the feature vector list of faces other than those currently handled by the first collaborative student client in the update message with the feature vector list of faces in the local mapping table. If the feature vectors of the faces in the former are all in the list of the latter, no processing is required. If the feature vectors of the faces in the former are not in the list of the latter, it means that a new student has been added in the newly retrieved frame. At this time, the student client needs to send a fourth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fourth request message includes a message ID, the ID of the student client that needs to collaborate, the second multicast group address, the first number of newly added faces c of the student client that needs to collaborate, and the feature vector list of the first newly added faces.

7. The interactive method for AI online education as described in claim 6, characterized in that: After receiving the fourth request message, the collaborative student terminal that has joined the first multicast group address M determines whether the preset conditions are met. The collaborative student terminal that meets the preset conditions is called the second collaborative student terminal. Each second collaborative student terminal starts a random countdown timer for a preset number of seconds locally. For the second collaborative student terminal whose countdown timer expires first, a third response message for adding collaborative sentiment analysis request is sent via multicast through the first multicast group address M. The third response message includes message ID, associated request ID, current second collaborative student terminal ID, received response status, the first number of new faces c to be added by the collaborative student terminal, and the feature vector list of the first new faces. After receiving the third response message, the other second collaborative student terminals stop the countdown timer. After receiving the third response message, the collaborative student terminals need to update their local mapping table, that is, add the current second collaborative student terminal ID c times to the collaborative student terminal ID list, update the number of faces in the collaborative student terminal to b+c, and add the feature vectors of each newly added face to the feature vector list of faces. If the second collaborative student terminal has not joined the second multicast group address, after sending the third response message, it sends an IGMP member report message to its directly connected router to join the second multicast group address of the corresponding student terminal that needs to collaborate. Then, it receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the coordinates of the face detection box corresponding to the feature vector in the feature vector list of the first newly added face, and performs sentiment analysis to obtain the sentiment category of the corresponding face. At the same time, the current second collaborative student terminal updates its local mapping table. Specifically, the current second collaborative student terminal adds the ID of the current second collaborative student terminal c times to the list of collaborative student terminal IDs corresponding to the required collaborative student terminal ID in the local mapping table, adds the feature vectors of each newly added face to the feature vector list of faces, adds the emotion category of each newly added face to the emotion category list of faces, and adds the timestamp of each newly added face when it was detected to the face detection timestamp list. Then, the current second collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal.

8. The interactive method for AI online education as described in claim 7, characterized in that: The student client needs to retrieve the local mapping table according to the preset second cycle and determine whether the difference between each face detection timestamp and the current timestamp in the face detection timestamp list is greater than the threshold. For each face detection timestamp, if it is less than the threshold, no processing is required. If it is greater than the threshold, it indicates that the current face detection timestamp has exceeded the threshold and has not been updated. It also determines whether the student client corresponding to the current face detection timestamp is offline due to a fault. If there is no offline fault, it means that the current face has left. At this time, the student terminal needs to delete the information about the current face from the mapping table. That is, delete the student terminal ID corresponding to the current face, the detection box coordinates of the current face, the feature vector of the current face, the emotion category of the current face, and the detection timestamp of the current face in the mapping table, and update the number of faces that need to be coordinated with the student terminal in the mapping table to the number of faces remaining after deletion. If there is a failure and the student terminal is offline, the collaborative student terminal needs to be replaced. In this case, the collaborative student terminal needs to send a fifth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fifth request message includes a message ID, the ID of the student terminal to be collaborated with, the second multicast group address, the second number of newly added faces of the student terminal to be collaborated with, and the list of facial feature vectors of the second newly added faces. The list of facial feature vectors of the second newly added faces is the facial feature vectors of the collaborative student terminal that is offline due to the failure.

9. The interactive method for AI online education as described in claim 8, characterized in that: After the other collaborative student terminals that have joined the first multicast group address M and are not offline due to a fault receive the fifth request message, they determine whether the preset conditions are met. Collaborative student terminals that meet the preset conditions are called third collaborative student terminals. Each third collaborative student terminal starts a random countdown timer for a preset number of seconds locally. For the third collaborative student terminal whose countdown timer expires first, a fourth response message for adding collaborative sentiment analysis request is sent via multicast through the first multicast group address M. The fourth response message includes message ID, associated request ID, current third collaborative student terminal ID, received response status, the number of second newly added faces to be added by the collaborative student terminal, and the feature vector list of the second newly added faces. After receiving the fourth response message, other third collaborative student terminals stop the countdown timer. After receiving the fourth response message, the collaborative student terminal needs to update its local mapping table, that is, replace the collaborative student terminal ID that is offline with the third collaborative student terminal ID whose countdown timer expires first in the collaborative student terminal ID list. If the current third collaborative student terminal has not joined the second multicast group address, after sending the fourth response message, it sends an IGMP member report message to its directly connected router to join the second multicast group address of the corresponding student terminal that needs to collaborate. Then, it receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the coordinates of the face detection box corresponding to the feature vector in the feature vector list of the second newly added face, and performs sentiment analysis to obtain the sentiment category of the second newly added face. At the same time, the current third collaborative student terminal updates its local mapping table. Specifically, the current third collaborative student terminal replaces the collaborative student terminal ID that is offline with the current third collaborative student terminal ID in the collaborative student terminal ID list corresponding to the collaborative student terminal ID in the local mapping table; replaces the face emotion category of the collaborative student terminal ID that is offline with the emotion category of the second newly added face in the face emotion category list; and replaces the face detection timestamp of the collaborative student terminal ID that is offline with the timestamp when the second newly added face was detected in the face detection timestamp list. Then, the current third collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal.

10. The interactive method for AI online education as described in claim 9, characterized in that: When all face detection timestamps in the face detection timestamp list of the local mapping table of the student client needing to collaborate are updated, the student client needs to collaborate to retrieve the local mapping table and send analysis information to the server. The analysis information includes the ID of the student client needing to collaborate, the feature vector of each face in the feature vector list, the detection box coordinates of each face in the face detection box coordinate list, and the emotion category of each face in the emotion category list. After receiving the analysis information, the server retrieves the feature vector of each face from the pre-saved student information table to obtain the student name corresponding to the feature vector of each face. The student information table includes the student ID, name, and the feature vector of the face. The server overlays the student's name and emotion category onto the detection box coordinates of each face in the current video stream on the student's end, and then sends the overlaid current video stream to the teacher's end. The teacher then asks questions based on the student's name and emotion category.

Citation Information

Patent Citations

  • Online education auxiliary scoring method and system

    CN112101074A

  • Education interaction live broadcast system and live broadcast method

    CN114422820A

  • Collaborative matching method, device, system and medium

    CN115204745A

  • Internet-based intelligent teaching method and system, and storage medium

    CN116994465A

  • Intelligent education management system based on multi-user cooperation

    CN119741171A