An interactive method for AI online education
By collaborating between teachers and students, and utilizing local resources on the student end for face detection and emotion analysis, the problems of excessive server load and insufficient hardware configuration in remote education are solved, thereby improving teaching efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2026-03-27
AI Technical Summary
In distance education, excessive server load can prevent face detection and emotion analysis from being completed when students' hardware configurations are insufficient, thus affecting the quality of teaching.
By collaborating between the teacher's end and the collaborating student's end, face detection and sentiment analysis are performed using local resources on the student's end, reducing the server load. A mapping table is established between the collaborating student's end and the student's end that needs to be collaborated, and collaborative sentiment analysis of the video stream is performed.
It enables face detection and emotion analysis in the video stream on the student's end, reducing server load, solving the problem of insufficient hardware configuration, and improving teaching efficiency.
Smart Images

Figure CN120852115B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of online education, and particularly relates to an AI online education interaction method. BACKGROUND
[0002] In a remote education scenario, a teacher and students are respectively located at different places, and teaching activities are carried out in a remote video conference manner through an online education platform. After starting a remote classroom, a teacher end and multiple student ends access an online classroom, and respectively unicast respective video streams (a video stream of the teacher end is teacher explanation content collected by a camera of the teacher end, and a video stream of the student end is a video picture of a student located in front of the student end collected by a camera of the student end) to a server. The server encodes and processes the video stream data and sends the video stream data to the teacher end and the student ends in a multicast manner. The teacher end receives and displays video pictures of the student ends.
[0003] For the student end, there is a case that one camera of the student end covers one to multiple students (for example, five students sit in front of one student end, and the camera of the student end covers the five students and accesses the remote education platform). Multiple student ends access the online education platform. When the teacher wants to call on a specific student (for example, the specific student looks like having understood or having confusion), the teacher wants to ask the specific student a question. In order to avoid that the teacher looks at the expressions of the students in the video pictures of the teacher end and manually judges, which is time-consuming and laborious, the teacher wants to perform expression (emotion) analysis and labeling on each student in the video pictures of the teacher end, so as to improve the calling-on and questioning efficiency.
[0004] The server performs student face detection, name matching and emotion analysis in the video pictures of the student ends, and the server uniformly performs student emotion labeling of the video pictures. In this way, because the server often bears more tasks, the server burden is increased, and then the quality of remote education is affected. Therefore, in the prior art, the server notifies each student end to be responsible for student face detection, name matching and emotion analysis in the respective video pictures, and the server uniformly performs student emotion labeling of the video pictures. In this way, for a student end with only one student, the student end can perform self-detection and analysis, but for a student end with multiple students, the student end has insufficient hardware configuration (insufficient capacity), and cannot complete the detection and analysis. SUMMARY
[0005] The application aims to solve the problems in the background art, and provides an AI online education interaction method.
[0006] To achieve the above object, the application adopts the following technical scheme:
[0007] The application provides an AI online education interaction method, which is applied to a student end to send a video stream of the student end to a teacher end through a server for playing and displaying, and in a case that the student end itself is insufficient in capability and cannot analyze a video picture in the video stream of the student end, wherein:
[0008] When a teacher asks a question to a student by playing and displaying a student video picture through the teacher end, the teacher acquires student emotion labeling by clicking a first button of the teacher end, and then the teacher end sends a first request message of the student emotion labeling to the server, and the server sends a second request message of emotion self-labeling to all student ends;
[0009] All student ends return first response messages of acceptance or rejection of the emotion self-labeling request to the server, and for a student end corresponding to the first response message of acceptance, the student end is a cooperative student end, and for a student end corresponding to the first response message of rejection, the student end is a student end needing cooperation;
[0010] When there is the first response message of rejection, the server unicasts a notification message of cooperative emotion analysis to all student ends;
[0011] After all student ends receive the notification message, all cooperative student ends perform face detection, extraction and emotion analysis on current video streams of all student ends needing cooperation, and respectively obtain face detection frame coordinates, feature vectors of faces and emotion categories in the current video streams of the student ends needing cooperation;
[0012] Each student end needing cooperation sends the face detection frame coordinates, the feature vectors of the faces and the emotion categories in the current video stream of the student end to the server, the server acquires a student name corresponding to the face according to the feature vector of the face, and labels the student name and the emotion category above the face detection frame coordinates in a video stream picture of each student end;
[0013] The server further sends the video stream labeled with the student basic information and the emotion category to the teacher end.
[0014] Preferably, after all student ends receive the second request message, each student end judges whether the capability of the student end meets a preset condition, that is, each student end judges whether the CPU core number, the memory capacity and the display memory capacity of the student end are greater than respective preset configurations, and judges whether the CPU occupancy rate, the memory occupancy rate and the display memory occupancy rate of the student end are less than respective preset values, when all the student ends meet the condition, the student end is a cooperative student end and is referred to as a first cooperative student end, otherwise, the student end is a student end needing cooperation;
[0015] The notification message comprises a message ID, a first multicast group address M for sending the first image frame of the current video stream of each student terminal to be cooperated, a multicast group address list for sending subsequent video streams of the first image frame of the current video stream of each student terminal to be cooperated, and a student terminal ID list to be cooperated, and the multicast group address list comprises a second multicast group address corresponding to each student terminal to be cooperated in the student terminal ID list;
[0016] All student terminals locally establish a mapping table related to the cooperative emotion analysis, and the mapping table comprises a record corresponding to the student terminal ID to be cooperated, which comprises the student terminal ID to be cooperated, the first multicast group address M, the second multicast group address, a cooperative student terminal ID list, the number of faces of the student terminal to be cooperated, a face detection frame coordinate list, a feature vector list of the face, and an emotion category list of the face, and a face detection timestamp corresponding to each emotion category of the face in the emotion category list of the face, and each face detection timestamp constitutes a face detection timestamp list, wherein the cooperative student terminal ID list, the face detection frame coordinate list, the feature vector list of the face, the emotion category list of the face, and the face detection timestamp list are empty when the mapping table is initially established, and the number of faces of the student terminal to be cooperated is zero when the mapping table is initially established.
[0017] Preferably, when each first cooperative student terminal performs face detection, extraction, and emotion analysis on the first image frame of the current video stream of the student terminal to be cooperated, all student terminals receive the notification message and send an IGMP member report message to the local direct-connected router to join the first multicast group address M in the notification message.
[0018] For each student terminal to be cooperated, the student terminal to be cooperated sends a third request message for cooperative emotion analysis comprising the first image frame of the current video stream through the first multicast group address M in a multicast manner, each first cooperative student terminal receiving the first image frame in the third request message starts a random countdown timer locally for a preset number of seconds, and the random countdown timers of each first cooperative student terminal locally are different in timing.
[0019] Each first cooperative student terminal performs face detection, extraction, and emotion analysis on the face in the first image frame in the order of the timing of the random countdown timer.
[0020] For the current first cooperative student terminal whose random countdown timer is timed out, the number of feature vectors in the feature vector list of the face in the record corresponding to the student terminal ID to be cooperated in the local mapping table is a, and all face detection frame coordinates in the first image frame are obtained in a preset order.
[0021] The current first collaborative student terminal extracts and performs emotion analysis on the face corresponding to the a+1th face detection frame coordinates, and obtains the feature vector and emotion category of the face corresponding to the a+1th face detection frame coordinates, i.e., the feature vector and emotion category of the a+1th face;
[0022] The current first collaborative student terminal updates the local mapping table, i.e., adds the ID of the current first collaborative student terminal to the collaborative student terminal ID list corresponding to the record of the required collaborative student terminal ID in the local mapping table, adds the feature vector of the a+1th face to the face feature vector list, and adds the timestamp when the a+1th face is detected to the face detection timestamp list;
[0023] The current first collaborative student terminal sends a second response message of the collaborative emotion analysis request in a multicast manner through the first multicast group address M, and the second response message includes a message ID, an associated request ID, an accepted response state, a current first collaborative student terminal ID, a required collaborative student terminal ID, a number of faces of the required collaborative student terminal, a number of the a+1th face currently detected, a feature vector of the a+1th face, an emotion category of the a+1th face, and a timestamp when the a+1th face is detected;
[0024] The required collaborative student terminal and other first collaborative student terminals that join the first multicast group M update their respective local mapping tables according to the second response message sent by the current first collaborative student terminal, i.e., add the ID of the current first collaborative student terminal to the collaborative student terminal ID list corresponding to the record of the required collaborative student terminal ID in the local mapping table, add the feature vector of the a+1th face to the face feature vector list, and add the timestamp when the a+1th face is detected to the face detection timestamp list. When the required collaborative student terminal updates the local mapping table, the emotion category of the a+1th face is also added to the emotion category list of the face in the mapping table.
[0025] Preferably, for the first collaborative student terminal whose random countdown timer expires first, a face detection algorithm is also used to detect all faces in the first image frame to obtain the detection frame coordinates of each face, and the number of faces is counted as b, and the actual number of required first collaborative student terminals is b;
[0026] When the first collaborative student terminal whose random countdown timer expires first updates the local mapping table, the counted number of faces b is also added to the number of faces of the required collaborative student terminal in the local mapping table;
[0027] In the second response message sent by the first collaborative student terminal whose random countdown timer expires first in a multicast manner through the first multicast group address M, the detection frame coordinates of all faces in the first image frame sorted in a preset order, i.e., a face detection frame coordinate list, are also included.
[0028] The first cooperative student terminal and other first cooperative student terminals which join the first multicast group M update the respective local mapping table according to the second response message sent by the first cooperative student terminal which arrives first, and add the number of faces b which are counted in the number of faces of the cooperative student terminal corresponding to the record of the respective local mapping table corresponding to the respective cooperative student terminal ID, and add the detection frame coordinates of all faces sorted in a preset order in the face detection frame coordinate list.
[0029] Preferably, when the cooperative student terminal and other first cooperative student terminals which join the first multicast group M receive the second response message about the respective cooperative student terminal ID and update the respective local mapping table, it is judged whether the number of faces of the cooperative student terminal corresponding to the record of the respective local mapping table corresponding to the respective cooperative student terminal ID is consistent with the number in the cooperative student terminal ID list. When it is consistent, it indicates that the face detection task corresponding to the third request message of the respective cooperative student terminal ID has been completed, and then the other first cooperative student terminals except those in the cooperative student terminal ID list stop counting the random countdown timer of the respective cooperative student terminal ID.
[0030] The cooperative student terminal retrieves the local mapping table, sends analysis information to the server, and the analysis information includes the cooperative student terminal ID, the feature vector of each face in the feature vector list of the face, the detection frame coordinates of each face in the face detection frame coordinate list, and the emotional category of each face in the emotional category list of the face.
[0031] The server receives the analysis information, retrieves each face feature vector in the pre-stored student information table, and obtains the student name corresponding to each face feature vector, wherein the student information table includes student ID, name and face feature vector.
[0032] The server superimposes the student name of the face and the emotional category of the face above the detection frame coordinates of each face in the current video stream of the cooperative student terminal according to the detection frame coordinates of the face, and sends the superimposed current video stream to the teacher terminal. The teacher terminal asks questions according to the student name and the emotional category.
[0033] Preferably, the cooperative student terminal also sends the subsequent video stream of the first image frame in a multicast manner through the second multicast group address. After each first cooperative student terminal sends the second response message about the respective cooperative student terminal ID, each first cooperative student terminal sends an IGMP member report message to the router directly connected to it locally to join the second multicast group address of the respective cooperative student terminal ID.
[0034] Each first collaborative student terminal that has joined the second multicast group address receives the subsequent video stream of the first image frame through the second multicast group address. It then extracts a new frame from the subsequent video stream according to a preset first cycle and performs the following processing steps: Face detection and extraction are performed on all faces in the newly extracted frame to obtain the coordinates of the corresponding face detection bounding boxes and feature vectors. Simultaneously, the feature vectors of the faces it is responsible for, stored in its local mapping table, are compared one by one with the feature vectors of all faces in the newly extracted frame.
[0035] If the comparison is unsuccessful, the test will be terminated.
[0036] If the comparison is successful, each first collaborative student terminal that has been added to the second multicast group address retrieves a new frame to obtain the coordinates of the face detection box of the successfully compared face, performs sentiment analysis to obtain the sentiment category of the successfully compared face, and then sends a collaborative sentiment analysis update message to the corresponding collaborative student terminal ID. The update message includes message ID, first collaborative student terminal ID, number of faces to be compared by the student terminal, feature vector of the face currently handled by the first collaborative student terminal, coordinates of the face detection box currently handled by the first collaborative student terminal, sentiment category of the face currently handled by the first collaborative student terminal, timestamp of the face detection currently handled by the first collaborative student terminal, and a list of feature vectors of other faces besides those currently handled by the first collaborative student terminal.
[0037] After receiving the update message, the collaborating student terminal needs to update its local mapping table, that is, update the face detection box coordinates, face emotion category, and face detection timestamp of the first collaborating student terminal in the face detection box coordinate list of the local mapping table.
[0038] Simultaneously, the student client needs to compare the feature vector list of faces other than those currently handled by the first collaborative student client in the update message with the feature vector list of faces in the local mapping table. If the feature vectors of the faces in the former are all in the list of the latter, no processing is required. If the feature vectors of the faces in the former are not in the list of the latter, it means that a new student has been added in the newly retrieved frame. At this time, the student client needs to send a fourth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fourth request message includes a message ID, the ID of the student client that needs to collaborate, the second multicast group address, the first number of newly added faces c of the student client that needs to collaborate, and the feature vector list of the first newly added faces.
[0039] Preferably, after receiving the fourth request message, the cooperative student terminal added with the first multicast group address M judges whether a preset condition is met, the cooperative student terminal meeting the preset condition is called a second cooperative student terminal, and each second cooperative student terminal starts a random countdown timer locally for a preset number of seconds, and the second cooperative student terminal with the countdown timer first to time sends a third response message of the added cooperative sentiment analysis request in a multicast manner through the first multicast group address M, and the third response message includes a message ID, an associated request ID, a current second cooperative student terminal ID, an accepted response state, a first added number of faces c of the cooperative student terminal, and a feature vector list of the first added faces.
[0040] After receiving the third response message, the other second cooperative student terminal stops the countdown timer, and after receiving the third response message, the cooperative student terminal updates the local mapping table, that is, adds the current second cooperative student terminal ID c times in the cooperative student terminal ID list, updates b+c in the number of faces of the cooperative student terminal, and adds the feature vector of each first added face in the feature vector list of the face.
[0041] If the current second cooperative student terminal has not joined the second multicast group address, it sends an IGMP member report message to its direct router after sending the third response message to join the second multicast group address of the corresponding cooperative student terminal, and then receives the subsequent video stream of the first image frame through the second multicast group address, takes out a frame from the subsequent video stream for face detection and extraction, obtains the face detection frame coordinates corresponding to the feature vector in the feature vector list of the first added face, and performs sentiment analysis to obtain the sentiment category of the corresponding face.
[0042] At the same time, the current second cooperative student terminal updates the local mapping table, that is, the current second cooperative student terminal adds the ID of the current second cooperative student terminal c times in the cooperative student terminal ID list corresponding to the record of the corresponding cooperative student terminal ID in the local mapping table, adds the feature vector of each first added face in the feature vector list of the face, adds the sentiment category of each first added face in the sentiment category list of the face, and adds the time stamp of detecting each first added face in the face detection time stamp list.
[0043] Then, the current second cooperative student terminal takes out a new frame from the subsequent video stream according to a preset first period, and processes it according to the processing steps of each first cooperative student terminal.
[0044] Preferably, the student terminal in cooperation needs to retrieve the local mapping table according to a preset second period, and determine whether the difference between each face detection timestamp in the face detection timestamp list and the current timestamp is greater than a threshold value. If the difference is less than the threshold value, each face detection timestamp does not need to be processed. If the difference is greater than the threshold value, it indicates that the current face detection timestamp has exceeded the threshold value without updating, and whether the student terminal in cooperation corresponding to the current face detection timestamp is offline due to a fault is determined.
[0045] If there is no offline due to a fault, it indicates that the current face has left. At this time, the student terminal in cooperation needs to delete the information about the current face from the mapping table, that is, delete the student terminal in cooperation ID corresponding to the current face, the detection frame coordinates of the current face, the feature vector of the current face, the emotional category of the current face and the current face detection timestamp in the mapping table, and update the number of faces of the student terminal in cooperation in the mapping table to the number of remaining faces after deletion.
[0046] If there is offline due to a fault, the student terminal in cooperation needs to be replaced. At this time, the student terminal in cooperation needs to send a fifth request message for adding a cooperative emotion analysis through the first multicast group address M in a multicast manner, and the fifth request message includes a message ID, a student terminal in cooperation ID, a second multicast group address, a second number of newly added faces of the student terminal in cooperation and a face feature vector list of the second newly added faces. The face feature vector list of the second newly added faces is the face feature vector responsible for by the student terminal in cooperation offline due to a fault.
[0047] Preferably, after the student terminal in cooperation receives the fifth request message, it is determined whether a preset condition is met. For the student terminal in cooperation meeting the preset condition, it is called a third student terminal in cooperation. Each third student terminal in cooperation starts a random countdown timer locally for a preset number of seconds. The third student terminal in cooperation whose countdown timer expires first sends a fourth response message for adding a cooperative emotion analysis request through the first multicast group address M in a multicast manner. The fourth response message includes a message ID, an associated request ID, a current third student terminal in cooperation ID, an accepted response state, a second number of newly added faces of the student terminal in cooperation and a feature vector list of the second newly added faces.
[0048] After the other third student terminal in cooperation receives the fourth response message, the counting of the countdown timer is stopped. After the student terminal in cooperation receives the fourth response message, the local mapping table is updated, that is, the student terminal in cooperation ID offline due to a fault is replaced by the third student terminal in cooperation ID whose countdown timer expires first in the student terminal in cooperation ID list.
[0049] If the current third collaborative student terminal has not joined the second multicast group address, the current third collaborative student terminal sends an IGMP membership report message to its direct router after sending the fourth response message to join the second multicast group address of the corresponding collaborative student terminal, and then receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the face detection frame coordinates corresponding to the feature vector in the feature vector list of the second newly added face, and performs emotion analysis to obtain the emotion category of the second newly added face;
[0050] Meanwhile, the current third collaborative student terminal updates the local mapping table, that is, the current third collaborative student terminal replaces the collaborative student terminal ID that is offline with a fault in the collaborative student terminal ID list corresponding to the record of the corresponding collaborative student terminal ID in the local mapping table with the current third collaborative student terminal ID, replaces the face emotion category responsible for the collaborative student terminal ID that is offline with a fault in the face emotion category list with the emotion category of the second newly added face, and replaces the face detection timestamp responsible for the collaborative student terminal ID that is offline with a fault in the face detection timestamp list with the timestamp when the second newly added face is detected.
[0051] Then, the current third collaborative student terminal takes a new frame from the subsequent video stream according to the preset first period and processes according to the processing steps of each first collaborative student terminal.
[0052] Preferably, when all the face detection timestamps in the face detection timestamp list of the local mapping table of the collaborative student terminal are updated once, the collaborative student terminal searches the local mapping table, sends analysis information to the server, and the analysis information includes the collaborative student terminal ID, the feature vector of each face in the feature vector list of the face, the detection frame coordinates of each face in the face detection frame coordinate list, and the emotion category of each face in the face emotion category list.
[0053] After the server receives the analysis information, the server searches the pre-stored student information table for the feature vector of each face to obtain the student name corresponding to the feature vector of each face, wherein the student information table includes the student ID, name, and feature vector of the face.
[0054] The server superimposes the student name of the face and the emotion category of the face above the detection frame coordinates of each face in the current video stream of the collaborative student terminal according to the detection frame coordinates of the face, and sends the superimposed current video stream to the teacher terminal, and the teacher terminal asks questions according to the student name and emotion category.
[0055] Compared with the prior art, the beneficial effects of the present application are:
[0056] The interactive method of the AI online education realizes face detection, extraction and emotion analysis of the video stream picture of the student end in need of cooperation, thereby reducing the server pressure, and enabling each student end to support face detection, extraction and emotion analysis of itself, solving the problem that the student end with multiple students cannot complete the problem due to insufficient hardware configuration (insufficient ability) in the prior art. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 The flowchart of the interactive method of the AI online education. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0059] In one embodiment, as shown in Figure 1 An interactive method of AI online education is provided, which applies the student end to send its own video stream to the teacher end for playing and displaying through the server, and the student end itself is insufficient in ability and cannot analyze the video picture in the student's own video stream (the present application omits the case that the student end itself is strong in ability and can perform face detection, extraction and emotion analysis on the current video stream of itself, which belongs to the prior art), wherein:
[0060] Step 1, when the teacher plays and displays the student video picture through the teacher end and asks the student a question, the teacher acquires the student emotion annotation by clicking the first button of the teacher end, and then the teacher end sends a first request message (containing request ID, request type (student emotion annotation request), teacher end ID and course information) of student emotion annotation to the server, and the server sends a second request message (containing request ID, request type (emotion self-annotation request), request server IP) of emotion self-annotation to all student ends;
[0061] Step 2, all student ends return a first response message of accepting or refusing the emotion self-annotation request to the server, and for the student end corresponding to the first response message of accepting, it is a cooperative student end; for the student end corresponding to the first response message of refusing, it is a student end in need of cooperation;
[0062] Wherein, all student terminals receive the second request message, each student terminal judges whether its own ability meets the preset condition: each student terminal judges whether its CPU core number, memory capacity and display memory capacity are greater than the respective preset configuration (such as in this embodiment, the preset configuration is that the CPU core number is 4, the memory capacity is 16 GB, and the display memory capacity is 4 GB), and judges whether its CPU occupancy rate, memory occupancy rate and display memory occupancy rate are less than the respective preset value (such as in this embodiment, the preset value is that the CPU occupancy rate is 70%, the memory capacity occupancy rate is 75%, and the display memory occupancy rate is 80%), when all the student terminals meet the condition, the student terminal is called as a collaborative student terminal, and is called as a first collaborative student terminal, and returns an accepted first response message to the server, otherwise, it is called as a student terminal needing cooperation, and returns a rejected first response message to the server.
[0063] Step 3, when there is a returned rejected first response message, the server unicast sends a notification message of collaborative sentiment analysis to all student terminals; wherein, the notification message (of collaborative sentiment analysis) includes a message ID, a first multicast group address M for sending the first image frame of the current video stream of each student terminal needing cooperation (wherein the first image frame of the current video stream is the first image frame of the video stream to be sent by the student terminal needing cooperation at the current time), a multicast group address list for sending the subsequent video stream of the first image frame of the current video stream of each student terminal needing cooperation, and a student terminal ID list needing cooperation, and the multicast group address list contains a second multicast group address corresponding to each student terminal needing cooperation in the student terminal ID list.
[0064] Step 4, after all student terminals receive the notification message, wherein all collaborative student terminals perform face detection, extraction and sentiment analysis on the current video stream of each student terminal needing cooperation, and respectively obtain the face detection frame coordinates, the feature vector of the face and the sentiment category in the current video stream of each student terminal needing cooperation (wherein, the process of face detection, extraction and sentiment analysis all belong to the prior art, such as face detection using a face detection algorithm (such as using YOLO network for face detection), obtaining the face detection frame coordinates (i.e. the coordinates of the four corners of the face detection frame) of each face in the image frame through face detection, and extracting each face to obtain the feature vector of the face, and using an expression recognition algorithm (such as a CNN model) to recognize the expression of the detected face to analyze the emotional state of the student's learning understanding (the emotion can be classified, such as: understanding, confusion, general, other)), including:
[0065] Step 4.1, when each first collaborative student end performs face detection, extraction and emotion analysis on the first image frame of the current video stream of the collaborative student end, all student ends receive a notification message, and send an IGMP member report message to the local direct connection router to join the first multicast group address M in the notification message (for the collaborative student end to join the first multicast group address M to receive the first image frame, and for the collaborative student end to join the first multicast group address M to receive subsequent information sent through the first multicast group address M. After the direct connection router of the collaborative student end and the collaborative student end receives the IGMP report, it updates the IGMP group table (including the first multicast group address M, interface information, group member state and timer, etc.), records that the router has members on the corresponding interface to join the image frame forwarding first multicast group address M. After the direct connection router updates the IGMP group table, it communicates with other routers through the multicast routing protocol to establish and maintain the multicast forwarding table item);
[0066] The mapping table is established locally by all student terminals in relation to collaborative sentiment analysis, and includes records corresponding to the student terminal IDs to be collaborated, which include the student terminal IDs to be collaborated, a first multicast group address M, a second multicast group address, a list of collaborative student terminal IDs (all collaborative student terminals corresponding to the student terminals to be collaborated are added to the list), the number of faces of the student terminal to be collaborated (i.e., the number of faces in the video stream picture of the student terminal to be collaborated), a list of face detection frame coordinates (i.e., all face detection frame coordinates of the faces in the video stream picture of the student terminal to be collaborated are added to the list, wherein each face detection frame coordinate in the list of face detection frame coordinates contains the coordinate positions of the four corners of the face detection frame, the number of face detection frame coordinates in the list of face detection frame coordinates is consistent with the number of members in the list of collaborative student terminal IDs, and the face detection frame coordinates in the list of face detection frame coordinates are sequentially and one-to-one corresponding to the collaborative student terminal IDs in the list of collaborative student terminal IDs, and the same applies below), a list of feature vectors of the faces (i.e., all feature vectors of the faces in the video stream picture of the student terminal to be collaborated are added to the list, and the number of feature vectors of the faces in the list of feature vectors of the faces is consistent with the number of members in the list of collaborative student terminal IDs, and the feature vectors of the faces in the list of feature vectors of the faces are sequentially and one-to-one corresponding to the collaborative student terminal IDs in the list of collaborative student terminal IDs), a list of emotional categories of the faces (i.e., all emotional categories of the faces in the video stream picture of the student terminal to be collaborated are added to the list, and the number of emotional categories of the faces in the list of emotional categories of the faces is consistent with the number of members in the list of collaborative student terminal IDs, and the emotional categories of the faces in the list of emotional categories of the faces are sequentially and one-to-one corresponding to the collaborative student terminal IDs in the list of collaborative student terminal IDs), and a face detection timestamp corresponding to each emotional category of the faces in the list of emotional categories of the faces, and each face detection timestamp constitutes a list of face detection timestamps (the same applies below, i.e., all the lists of face detection timestamps mentioned in the present application are the same, and no repeated description is given), wherein the list of collaborative student terminal IDs, the list of face detection frame coordinates, the list of feature vectors of the faces, the list of emotional categories of the faces, and the list of face detection timestamps are all empty when the mapping table is initially established, and the number of faces of the student terminal to be collaborated is zero when the mapping table is initially established (it should be noted that the mapping table of the student terminal to be collaborated locally has only one record about itself, and the mapping table of the collaborative student terminal has multiple records corresponding to all student terminals to be collaborated one by one).
[0067] Step 4.2, for each student terminal to be cooperated, the student terminal to be cooperated sends a third request message for cooperative sentiment analysis containing the first image frame of the current video stream in a multicast manner through the first multicast group address M (each router in the network copies and forwards the third request message to all the student terminals to be cooperated which join the first multicast group address M according to the multicast routing table and the multicast forwarding table), and each first student terminal which joins the first multicast group address M starts a random countdown timer locally (for example, a random countdown timer of 5 seconds, that is, randomly counts down within 0-5 seconds) after receiving the first image frame in the third request message, and the random countdown timers of each first student terminal are different from each other;
[0068] Each first student terminal performs face detection, extraction and sentiment analysis on the faces in the first image frame in the order of the random countdown timer (the faces are sorted according to the size of the upper left corner coordinates (X, Y) of each detection frame, that is, the face detection frame with smaller X is arranged in front, and if X is equal, the face detection frame with smaller Y is arranged in front).
[0069] For the current first student terminal whose random countdown timer is up (that is, for each first student terminal), first, the number of feature vectors a in the feature vector list of the face in the record corresponding to the ID of the student terminal to be cooperated in the local mapping table is retrieved, and the coordinates of all face detection frames sorted in a preset order in the first image frame are obtained.
[0070] The current first student terminal performs extraction and sentiment analysis on the face corresponding to the a+1th face detection frame coordinate, and obtains the feature vector and the sentiment category of the face corresponding to the a+1th face detection frame coordinate, that is, the feature vector and the sentiment category of the a+1th face.
[0071] The current first student terminal updates the local mapping table, that is, the current first student terminal adds the ID of the current first student terminal to the list of cooperative student terminal IDs in the record corresponding to the ID of the student terminal to be cooperated in the local mapping table, adds the feature vector of the a+1th face to the feature vector list of the face, and adds the timestamp when the a+1th face is detected to the face detection timestamp list.
[0072] The current first student terminal sends a second response message for cooperative sentiment analysis request in a multicast manner through the first multicast group address M, and the second response message includes message ID, associated request ID (that is, the ID of the associated third request message), accepted response status, current first student terminal ID, student terminal to be cooperated ID, number of faces of the student terminal to be cooperated, number of the a+1th face currently detected, feature vector of the a+1th face, sentiment category of the a+1th face, and timestamp when the a+1th face is detected.
[0073] The student terminal and other first cooperative student terminals that join the first multicast group M update their respective local mapping tables according to the second response message sent by the current first cooperative student terminal, that is, add the ID of the current first cooperative student terminal to the list of cooperative student terminal IDs corresponding to the record of the ID of the student terminal that needs to be cooperated in the local mapping table, add the feature vector of the a+1th face to the list of feature vectors of faces, and add the timestamp when the a+1th face is detected to the list of face detection timestamps. When the student terminal that needs to be cooperated updates the local mapping table, the a+1th face emotion category is also added to the list of face emotion categories in the mapping table. It should be noted that each first cooperative student terminal performs the above operations, but for the first cooperative student terminal whose random countdown timer expires first, there are additional operations:
[0074] Wherein, for the first cooperative student terminal whose random countdown timer expires first, a face detection algorithm is also used to detect all faces in the first image frame to obtain the detection box coordinates of each face, and the number of faces is counted as b, then the actual number of first cooperative student terminals required is b.
[0075] When the first cooperative student terminal whose random countdown timer expires first updates the local mapping table, the counted number of faces b is also added to the number of faces of the student terminal that needs to be cooperated in the local mapping table.
[0076] The second response message sent by the first cooperative student terminal whose random countdown timer expires first through the first multicast group address M in a multicast manner also includes the detection box coordinates of all faces in the first image frame sorted in a preset order, that is, a list of face detection box coordinates.
[0077] When the student terminal that needs to be cooperated and other first cooperative student terminals that join the first multicast group M update their respective local mapping tables according to the second response message sent by the first cooperative student terminal whose random countdown timer expires first, the counted number of faces b is also added to the number of faces of the student terminal that needs to be cooperated in the record corresponding to the ID of the student terminal that needs to be cooperated in the local mapping table, and the detection box coordinates of all faces sorted in a preset order are also added to the list of face detection box coordinates.
[0078] Step 4.3, when the first collaborative student terminal and other first collaborative student terminals which need to cooperate with the student terminal and join the first multicast group M receive the second response message about the corresponding student terminal ID, and update the local mapping table, the first collaborative student terminal judges whether the number of the face corresponding to the record of the corresponding student terminal ID in the local mapping table is consistent with the number in the collaborative student terminal ID list. When it is consistent, it means that the face detection task corresponding to the third request message of the corresponding student terminal ID has been completed, and other first collaborative student terminals except the collaborative student terminal ID list stop counting the random countdown timer of the corresponding student terminal ID (other first collaborative student terminals still save the mapping table locally and do not delete it);
[0079] The student terminal which needs to cooperate searches the local mapping table, sends the analysis information to the server, and the analysis information includes the student terminal ID, the feature vector of each face in the feature vector list, the detection frame coordinates of each face in the face detection frame coordinate list, and the emotional category of each face in the face emotional category list;
[0080] The server receives the analysis information, searches the pre-saved student information table for the feature vector of each face, and obtains the student name corresponding to the feature vector of each face, wherein the student information table includes the student ID, the name, and the feature vector of the face.
[0081] The server superimposes the student name and the emotional category of the face on the detection frame coordinates of each face in the current video stream of the student terminal which needs to cooperate (i.e. in each frame of the current video stream) according to the detection frame coordinates of the face, and sends the superimposed current video stream to the teacher terminal. The teacher terminal asks questions according to the student name and the emotional category.
[0082] The above is the process of detecting the face of each student terminal which needs to cooperate by the collaborative student terminal for the first image frame, and the student terminal which needs to cooperate sends the analysis information about the first image frame to the server by superimposing the video stream.
[0083] The following is for the subsequent video stream of the first image frame:
[0084] Step 4.4.1, the student terminal which needs to cooperate also sends the subsequent video stream of the first image frame in a multicast manner through the second multicast group address, and each first collaborative student terminal sends an IGMP member report message to the local direct connection router to join the second multicast group address of the corresponding student terminal ID (obtained from the local mapping table) after sending the second response message about the corresponding student terminal ID.
[0085] Step 4.4.2, after each first collaborative student terminal added with the second multicast group address receives the subsequent video stream of the first image frame through the second multicast group address, a new frame is taken out from the subsequent video stream according to a preset first period (such as 10 seconds) (since the time of detecting the first video frame is inconsistent for each first collaborative student terminal, the new frame taken out by each first collaborative student terminal is also different), and the following processing steps are performed: face detection and extraction are performed on all faces in the new frame to obtain face detection frame coordinates and feature vectors corresponding to the faces, and the face feature vectors of the students responsible for by the first collaborative student terminal are compared with the feature vectors of all faces in the new frame one by one:
[0086] Step 4.4.3, if the comparison is not successful (for example, the face (i.e. student) responsible for by the first collaborative student terminal leaves the picture, so the feature vector of the face is not detected), the detection is terminated; the next period is entered, and the step 4.4.2 is returned to continue the cycle operation of taking out a frame by the first collaborative student terminal;
[0087] Step 4.4.4, if the comparison is successful, the face detection frame coordinates of the face with a successful comparison are obtained from the new frame taken out by each first collaborative student terminal added with the second multicast group address, and emotion analysis is performed to obtain the emotion category of the face with a successful comparison, and then an update message of collaborative emotion analysis is sent to the corresponding collaborative student terminal ID, and the update message includes a message ID, a first collaborative student terminal ID, a number of faces of the collaborative student terminal, a face feature vector responsible for by the current first collaborative student terminal, a face detection frame coordinate responsible for by the current first collaborative student terminal, a face emotion category responsible for by the current first collaborative student terminal, a face detection timestamp responsible for by the current first collaborative student terminal, and a feature vector list of faces other than the face responsible for by the current first collaborative student terminal (i.e. the feature vectors of other faces identified from the image except the face responsible for by the current first collaborative student terminal);
[0088] Step 4.4.5, after the collaborative student terminal receives the update message, the local mapping table is updated, that is, the face detection frame coordinate responsible for by the current first collaborative student terminal, the face emotion category responsible for by the current first collaborative student terminal, and the face detection timestamp responsible for by the current first collaborative student terminal are updated in the face detection frame coordinate list of the local mapping table;
[0089] Step 4.4.6, at the same time, the feature vector list of the faces other than the face responsible for by the current first collaborative student terminal in the update message is compared with the feature vector list of the faces in the local mapping table, if the face feature vectors in the former list are all in the latter list (indicating that no new face appears in the picture), no processing is performed; the next preset first period is entered, and the step 4.4.2 is returned to continue the cycle operation of taking out a frame by the first collaborative student terminal;
[0090] Step 4.4.6, if the facial feature vector of the former is not in the list of the latter, it indicates that a new student is added in the frame, at this time, the fourth request message for adding new cooperative emotion analysis is sent in a multicast manner through the first multicast group address M, and the fourth request message includes message ID, student ID to be cooperated, second multicast group address, first added number of faces c of the student to be cooperated and feature vector list of the first added face.
[0091] Step 4.5: Step 4.5.1, after the cooperative student terminal added with the first multicast group address M receives the fourth request message, it is judged whether the preset condition is met (i.e. whether the self ability meets the preset condition, see step 2), the cooperative student terminal meeting the preset condition is called second cooperative student terminal, and each second cooperative student terminal starts a random countdown timer locally for a preset number of seconds, and the second cooperative student terminal whose countdown timer is first counted down sends the third response message for adding new cooperative emotion analysis request in a multicast manner through the first multicast group address M, and the third response message includes message ID, associated request ID (i.e. associated fourth request message), current second cooperative student terminal ID, accepted response state, first added number of faces c of the student to be cooperated and feature vector list of the first added face;
[0092] Step 4.5.2, after other second cooperative student terminals receive the third response message, the counting of the countdown timer is stopped, and after the student to be cooperated receives the third response message, the local mapping table is updated, i.e. the ID of the current second cooperative student terminal is added c times in the cooperative student terminal ID list, the number of faces of the student to be cooperated is updated to b+c, and the feature vector of each first added face is added in the feature vector list of the face;
[0093] Step 4.5.3, if the current second cooperative student terminal has not joined the second multicast group address, it sends an IGMP member report message to its direct router after sending the third response message to join the second multicast group address of the corresponding student to be cooperated, and then receives the subsequent video stream of the first image frame through the second multicast group address, takes out a frame from the subsequent video stream for face detection and extraction, obtains the face detection frame coordinates corresponding to the feature vector in the feature vector list of the first added face, and performs emotion analysis to obtain the emotion category of the corresponding face;
[0094] Step 4.5.4, at the same time, the current second cooperative student terminal updates the local mapping table, i.e. the ID of the current second cooperative student terminal is added c times in the cooperative student terminal ID list corresponding to the record of the ID of the corresponding student to be cooperated in the local mapping table of the current second cooperative student terminal, each first added face feature vector is added in the feature vector list of the face, each first added face emotion category is added in the emotion category list of the face, and the time stamp when each first added face is detected is added in the face detection time stamp list.
[0095] Step 4.5.6: Then, the current second collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal (that is, the second collaborative student terminal operates in the manner of steps 4.4.2-4.5 of the first collaborative student terminal).
[0096] Step 4.6: Step 4.6.1, the collaborative student terminal needs to retrieve the local mapping table according to the preset second cycle (e.g., every 10 seconds) and determine whether the difference between each face detection timestamp and the current timestamp in the face detection timestamp list is greater than the threshold (e.g., 20 seconds). For each face detection timestamp, if it is less than the threshold, no processing is required (then enter the next preset first cycle and return to step 4.4.2 for the first collaborative student terminal to continue to retrieve a frame for loop operation); if it is greater than the threshold, it indicates that the current face detection timestamp has exceeded the threshold and has not been updated, and it is determined whether the collaborative student terminal corresponding to the current face detection timestamp has a fault and is offline (the collaborative student terminal needs to send an ICMP Echo request (i.e., Ping test) to the collaborative student terminal. If the collaborative student terminal responds, there is no fault and offline; if there is no response, it indicates that there is a fault and offline).
[0097] Step 4.6.2: If there is no offline fault, it means that the current face has left. At this time, the student terminal needs to delete the information about the current face from the mapping table, that is, delete the student terminal ID corresponding to the current face, the detection box coordinates of the current face, the feature vector of the current face, the emotion category of the current face, and the detection timestamp of the current face in the mapping table, and update the number of faces that need to be coordinated with the student terminal in the mapping table to the number of faces remaining after deletion; then execute step (5);
[0098] Step 4.6.3: If there is a fault and the student terminal is offline, the collaborative student terminal needs to be replaced. At this time, the collaborative student terminal needs to send a fifth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fifth request message includes message ID, ID of the student terminal to be collaborated with, second multicast group address, second number of newly added faces of the student terminal to be collaborated with, and a list of facial feature vectors of the second newly added faces. The list of facial feature vectors of the second newly added faces is the facial feature vectors of the collaborative student terminal that is offline due to the fault.
[0099] Step 4.7: After the first group address M is added and the fifth request message is received by the rest of the collaborative student terminals except for the faulty offline one, it is determined whether the preset condition is met. The collaborative student terminals that meet the preset condition are referred to as third collaborative student terminals. Each third collaborative student terminal starts a random countdown timer locally for a preset number of seconds. The third collaborative student terminal whose countdown timer expires first sends a fourth response message of the new collaborative sentiment analysis request in a multicast manner through the first group address M. The fourth response message includes a message ID, an associated request ID (i.e., the associated fifth request message), a current third collaborative student terminal ID, an accepted response status, a second number of new faces of the student to be collaborated, and a feature vector list of the second new faces.
[0100] Step 4.7.2: After the fourth response message is received by the other third collaborative student terminals, the countdown timer is stopped. After the fourth response message is received by the student to be collaborated, the local mapping table is updated, i.e., the collaborative student terminal ID of the faulty offline one is replaced by the third collaborative student terminal ID whose countdown timer expires first in the collaborative student terminal ID list.
[0101] Step 4.7.3: If the current third collaborative student terminal has not joined the second group address, it sends an IGMP membership report message to its directly connected router after sending the fourth response message to join the second group address of the student to be collaborated. Then, it receives the subsequent video stream of the first image frame through the second group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the face detection box coordinates corresponding to the feature vectors in the feature vector list of the second new face, performs sentiment analysis, and obtains the sentiment category of the second new face.
[0102] Step 4.7.4: At the same time, the current third collaborative student terminal updates the local mapping table, i.e., the collaborative student terminal ID of the faulty offline one is replaced by the current third collaborative student terminal ID in the collaborative student terminal ID list corresponding to the student to be collaborated ID in the local mapping table, the face sentiment category responsible for by the collaborative student terminal ID of the faulty offline one is replaced by the sentiment category of the second new face in the face sentiment category list, and the face detection timestamp responsible for by the collaborative student terminal ID of the faulty offline one is replaced by the timestamp when the second new face is detected in the face detection timestamp list.
[0103] Step 4.7.5: Then, the current third collaborative student terminal takes a new frame from the subsequent video stream according to a preset first period and processes it according to the steps of the first collaborative student terminal (i.e., the third collaborative student terminal operates in the manner of steps 4.4.2-4.5 of the first collaborative student terminal).
[0104] Enter the next preset second cycle, return to step 4.6.1, and cycle the operation.
[0105] Step 5, when all the face detection timestamps in the face detection timestamp list need to be coordinated with the student end local mapping table are updated, the student end needs to be coordinated with the local mapping table, and the analysis information is sent to the server, and the analysis information includes the student end ID, the feature vector of each face in the feature vector list of the face, the detection frame coordinates of each face in the face detection frame coordinate list, and the emotional category of each face in the face emotional category list;
[0106] After the server receives the analysis information, the feature vector of each face is searched in the pre-saved student information table to obtain the student name corresponding to the feature vector of each face, wherein the student information table includes student ID, name and feature vector of the face;
[0107] The server superimposes the student name of the face and the emotional category of the face above the detection frame coordinates of each face in the current video stream of the student end to be coordinated according to the detection frame coordinates of the face, and sends the superimposed current video stream to the teacher end, and the teacher end asks questions according to the student name and the emotional category;
[0108] Enter the next preset first cycle, return to step 4.4.2, and continue to take out a frame from the first student end to be coordinated to cycle the operation.
[0109] In another embodiment, the present application also includes an AI online education interactive device, comprising a processor and a memory storing a plurality of computer instructions, which are executed by the processor to implement the steps of the method of steps 1-5. For specific limitations of the AI online education interactive device, please refer to the limitations of the AI online education interactive method described above, which will not be repeated here.
[0110] The AI online education interactive method can realize face detection, extraction and emotion analysis of the video stream picture of the student end to be coordinated through the student end to be coordinated with strong ability, thereby reducing the server pressure, supporting the face detection, extraction and emotion analysis of each student end, and solving the problem that the student end with insufficient hardware configuration (insufficient ability) cannot be completed in the prior art.
[0111] It should be understood that, although Figure 1 The steps in the flowchart of the application are displayed in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, Figure 1At least one of the steps in the above-mentioned methods can comprise a plurality of sub-steps or stages, which sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the order of the sub-steps or stages is not necessarily sequential, but can be performed in rotation or alternation with other steps or sub-steps or stages of other steps.
[0112] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the patent of the present application should be subject to the appended claims.
Claims
1. An interactive method for AI-powered online education, characterized by: The method is applied when students send their video streams to teachers via a server for playback, but the students themselves lack the capability to analyze the video frames within their own streams. When a teacher asks questions to students by playing and displaying student videos on the teacher's end, the teacher obtains student sentiment annotations by clicking the first button on the teacher's end. Then, the teacher's end sends a first request message for student sentiment annotation to the server, and the server then sends a second request message for sentiment self-annotation to all student ends. All student terminals return a first response message to the server indicating whether they accept or reject the sentiment self-labeling request. The student terminal corresponding to the first response message that returns acceptance is designated as a collaborating student terminal, and the student terminal corresponding to the first response message that returns rejection is designated as a collaborating student terminal. If a first response message returns a rejection, the server unicasts a notification message for collaborative sentiment analysis to all student terminals. After all student terminals receive the notification message, all collaborating student terminals perform face detection, extraction, and sentiment analysis on the current video stream of each collaborating student terminal, respectively obtaining the coordinates of the face detection box, the feature vector of the face, and the sentiment category in the current video stream of each collaborating student terminal; Each student terminal that needs to collaborate sends the coordinates of the face detection box, the feature vector of the face, and the emotion category in its current video stream to the server. The server obtains the student's name corresponding to the face based on the face feature vector and marks the corresponding student's name and emotion category above the coordinates of the face detection box in the video stream of each student terminal. The server then sends the video stream, labeled with the student's basic information and emotional category, to the teacher's end.
2. The interactive method for AI online education as described in claim 1, characterized in that: After all student terminals receive the second request message, each student terminal determines whether its own capabilities meet the preset conditions: each student terminal determines whether its CPU core count, memory capacity, and video memory capacity are greater than its own preset configuration, and whether its CPU utilization rate, memory utilization rate, and video memory utilization rate are less than its own preset value. If all conditions are met, the student terminal is designated as the collaborating student terminal and is called the first collaborating student terminal; otherwise, it is designated as the collaborating student terminal. The notification message includes a message ID, a first multicast group address M for sending the first image frame of the current video stream of each student terminal that needs to cooperate, a list of multicast group addresses for sending the subsequent video streams after the first image frame of the current video stream of each student terminal that needs to cooperate, and a list of student terminal IDs that need to cooperate. The list of multicast group addresses contains a second multicast group address that corresponds one-to-one with each student terminal in the list of student terminal IDs that needs to cooperate. All student terminals establish a mapping table for collaborative sentiment analysis locally. The mapping table includes records corresponding to the student terminal IDs that need to collaborate. These records include the student terminal IDs that need to collaborate, the first multicast group address M, the second multicast group address, a list of student terminal IDs that need to collaborate, a list of face detection box coordinates, a list of face feature vectors, and a list of face sentiment categories. It also includes a face detection timestamp that corresponds one-to-one with each face sentiment category in the face sentiment category list. These face detection timestamps form a face detection timestamp list. Initially, the list of student terminal IDs, the list of face detection box coordinates, the list of face feature vectors, the list of face sentiment categories, and the list of face detection timestamps are all empty, and the number of faces that need to collaborate is zero initially.
3. The interactive method for AI online education as described in claim 2, characterized in that: When each of the first collaborative student terminals performs face detection, extraction, and emotion analysis on the first image frame of the current video stream of the student terminal to be collaborated with, after receiving the notification message, all student terminals send an IGMP member report message to the router directly connected to their local network to join the first multicast group address M in the notification message; For each student terminal that needs to collaborate, the student terminal that needs to collaborate sends a third request message containing the first image frame of the current video stream via the first multicast group address M in a multicast manner. After receiving the first image frame in the third request message, each first collaborative student terminal that has joined the first multicast group address M starts a random countdown timer with a preset number of seconds on its local machine, and the timing of the random countdown timer is different between each first collaborative student terminal. Each student terminal performs face detection, extraction, and sentiment analysis on the faces in the first image frame according to the order in which the random countdown timer expires; For the first collaborating student terminal when the random countdown timer expires, first retrieve the feature vectors in the feature vector list of the face in the record corresponding to the student terminal ID to be collaborated in the local mapping table. The number of feature vectors in the feature vector list of the face is 'a'. Then, obtain the coordinates of all face detection boxes in the first image frame after being sorted in a preset order. The first collaborative student terminal extracts and performs sentiment analysis on the face corresponding to the coordinates of the (a+1)th face detection box, and obtains the feature vector and sentiment category of the face corresponding to the coordinates of the (a+1)th face detection box, that is, the feature vector and sentiment category of the (a+1)th face. The first collaborative student terminal updates its local mapping table. Specifically, the first collaborative student terminal adds its ID to the collaborative student terminal ID list corresponding to the collaborative student terminal ID in the local mapping table, adds the feature vector of the (a+1)th face to the face feature vector list, and adds the timestamp of detecting the (a+1)th face to the face detection timestamp list. The first collaborative student terminal sends a second response message for the collaborative sentiment analysis request via multicast through the first multicast group address M. The second response message includes a message ID, an associated request ID, the received response status, the current first collaborative student terminal ID, the student terminal ID that needs to collaborate, the number of faces that need to collaborate, the number of the currently detected (a+1)th face, the feature vector of the (a+1)th face, the sentiment category of the (a+1)th face, and the timestamp when the (a+1)th face was detected. The cooperating student terminal and other first cooperating student terminals that have joined the first multicast group M update their respective local mapping tables according to the second response message sent by the current first cooperating student terminal. That is, they add the current first cooperating student terminal ID to the list of cooperating student terminal IDs of the record corresponding to the cooperating student terminal ID in their respective local mapping tables, add the feature vector of the (a+1)th face to the list of face feature vectors, and add the timestamp of detecting the (a+1)th face to the list of face detection timestamps. When the cooperating student terminal updates its local mapping table, it also adds the emotion category of the (a+1)th face to the list of face emotion categories in the mapping table.
4. The interactive method for AI online education as described in claim 3, characterized in that: For the first collaborative student terminal whose random countdown timer expires, a face detection algorithm is used to detect all faces in the first image frame, obtain the detection box coordinates of each face, and count the number of faces as b. Then the actual number of first collaborative student terminals required is b. When the first collaborating student updates its local mapping table, it also adds the count of faces b to the number of faces that need to be collaborated with the student in its local mapping table. The second response message sent by the first collaborating student terminal at the first time via the first multicast group address M in a multicast manner also includes the coordinates of the detection boxes of all faces in the first image frame after being sorted in a preset order, i.e., the list of face detection box coordinates. When the student client that needs to collaborate and other first collaborative student clients that have joined the first multicast group M update their local mapping tables according to the second response message sent by the first first collaborative student client that arrives, they also add the number of faces b to the number of faces of the student client that needs to collaborate in the record corresponding to the ID of the student client that needs to collaborate in their local mapping tables, and add the detection box coordinates of all faces sorted in a preset order to the list of face detection box coordinates.
5. The interactive method for AI online education as described in claim 4, characterized in that: When the student client requiring collaboration and other first collaborative student clients joining the first multicast group M receive the second response message about the corresponding student client ID and update their respective local mapping tables, they determine whether the number of faces of the student client corresponding to the corresponding student client ID in their respective local mapping tables is consistent with the number in the collaborative student client ID list. If they are consistent, it means that the face detection task corresponding to the third request message of the corresponding student client ID has been completed, and other first collaborative student clients other than those in the collaborative student client ID list stop counting down the random countdown timer for the corresponding student client ID. The student client needs to retrieve the local mapping table and send the analysis information to the server. The analysis information includes the student client ID, the feature vector of each face in the feature vector list, the detection box coordinates of each face in the detection box coordinate list, and the emotion category of each face in the emotion category list. After receiving the analysis information, the server retrieves the feature vector of each face from the pre-saved student information table to obtain the student name corresponding to the feature vector of each face. The student information table includes the student ID, name, and the feature vector of the face. Based on the coordinates of the face detection bounding boxes, the server overlays the student's name and emotion category onto the detection bounding boxes of each face in the current video stream of the student's computer. The server then sends the overlaid video stream to the teacher's computer, where the teacher asks questions based on the student's name and emotion category.
6. The interactive method for AI online education as described in claim 5, characterized in that: The student client that needs to cooperate also sends the subsequent video stream of the first image frame in multicast mode through the second multicast group address. After sending the second response message about the corresponding student client ID, each first student client sends an IGMP member report message to its locally directly connected router to join the second multicast group address of the corresponding student client ID. Each first collaborative student terminal that has joined the second multicast group address receives the subsequent video stream of the first image frame through the second multicast group address. It then extracts a new frame from the subsequent video stream according to a preset first cycle and performs the following processing steps: Face detection and extraction are performed on all faces in the newly extracted frame to obtain the coordinates of the corresponding face detection bounding boxes and feature vectors. Simultaneously, the feature vectors of the faces it is responsible for, stored in its local mapping table, are compared one by one with the feature vectors of all faces in the newly extracted frame. If the comparison is unsuccessful, the test will be terminated. If the comparison is successful, each first collaborative student terminal that has been added to the second multicast group address retrieves a new frame to obtain the coordinates of the face detection box of the successfully compared face, performs sentiment analysis to obtain the sentiment category of the successfully compared face, and then sends a collaborative sentiment analysis update message to the corresponding collaborative student terminal ID. The update message includes message ID, first collaborative student terminal ID, number of faces to be compared by the student terminal, feature vector of the face currently handled by the first collaborative student terminal, coordinates of the face detection box currently handled by the first collaborative student terminal, sentiment category of the face currently handled by the first collaborative student terminal, timestamp of the face detection currently handled by the first collaborative student terminal, and a list of feature vectors of other faces besides those currently handled by the first collaborative student terminal. After receiving the update message, the collaborating student terminal needs to update its local mapping table, that is, update the face detection box coordinates, face emotion category, and face detection timestamp of the first collaborating student terminal in the face detection box coordinate list of the local mapping table. Simultaneously, the student client needs to compare the feature vector list of faces other than those currently handled by the first collaborative student client in the update message with the feature vector list of faces in the local mapping table. If the feature vectors of the faces in the former are all in the list of the latter, no processing is required. If the feature vectors of the faces in the former are not in the list of the latter, it means that a new student has been added in the newly retrieved frame. At this time, the student client needs to send a fourth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fourth request message includes a message ID, the ID of the student client that needs to collaborate, the second multicast group address, the first number of newly added faces c of the student client that needs to collaborate, and the feature vector list of the first newly added faces.
7. The interactive method for AI online education as described in claim 6, characterized in that: After receiving the fourth request message, the collaborative student terminal that has joined the first multicast group address M determines whether the preset conditions are met. The collaborative student terminal that meets the preset conditions is called the second collaborative student terminal. Each second collaborative student terminal starts a random countdown timer for a preset number of seconds locally. For the second collaborative student terminal whose countdown timer expires first, a third response message for adding collaborative sentiment analysis request is sent via multicast through the first multicast group address M. The third response message includes message ID, associated request ID, current second collaborative student terminal ID, received response status, the first number of new faces c to be added by the collaborative student terminal, and the feature vector list of the first new faces. After receiving the third response message, the other second collaborative student terminals stop the countdown timer. After receiving the third response message, the collaborative student terminals need to update their local mapping table, that is, add the current second collaborative student terminal ID c times to the collaborative student terminal ID list, update the number of faces in the collaborative student terminal to b+c, and add the feature vectors of each newly added face to the feature vector list of faces. If the second collaborative student terminal has not joined the second multicast group address, after sending the third response message, it sends an IGMP member report message to its directly connected router to join the second multicast group address of the corresponding student terminal that needs to collaborate. Then, it receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the coordinates of the face detection box corresponding to the feature vector in the feature vector list of the first newly added face, and performs sentiment analysis to obtain the sentiment category of the corresponding face. At the same time, the current second collaborative student terminal updates its local mapping table. Specifically, the current second collaborative student terminal adds the ID of the current second collaborative student terminal c times to the list of collaborative student terminal IDs corresponding to the required collaborative student terminal ID in the local mapping table, adds the feature vectors of each newly added face to the feature vector list of faces, adds the emotion category of each newly added face to the emotion category list of faces, and adds the timestamp of each newly added face when it was detected to the face detection timestamp list. Then, the current second collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal.
8. The interactive method for AI online education as described in claim 7, characterized in that: The student client needs to retrieve the local mapping table according to the preset second cycle and determine whether the difference between each face detection timestamp and the current timestamp in the face detection timestamp list is greater than the threshold. For each face detection timestamp, if it is less than the threshold, no processing is required. If it is greater than the threshold, it indicates that the current face detection timestamp has exceeded the threshold and has not been updated. It also determines whether the student client corresponding to the current face detection timestamp is offline due to a fault. If there is no offline fault, it means that the current face has left. At this time, the student terminal needs to delete the information about the current face from the mapping table. That is, delete the student terminal ID corresponding to the current face, the detection box coordinates of the current face, the feature vector of the current face, the emotion category of the current face, and the detection timestamp of the current face in the mapping table, and update the number of faces that need to be coordinated with the student terminal in the mapping table to the number of faces remaining after deletion. If there is a failure and the student terminal is offline, the collaborative student terminal needs to be replaced. In this case, the collaborative student terminal needs to send a fifth request message for new collaborative sentiment analysis via multicast through the first multicast group address M. The fifth request message includes a message ID, the ID of the student terminal to be collaborated with, the second multicast group address, the second number of newly added faces of the student terminal to be collaborated with, and the list of facial feature vectors of the second newly added faces. The list of facial feature vectors of the second newly added faces is the facial feature vectors of the collaborative student terminal that is offline due to the failure.
9. The interactive method for AI online education as described in claim 8, characterized in that: After the other collaborative student terminals that have joined the first multicast group address M and are not offline due to a fault receive the fifth request message, they determine whether the preset conditions are met. Collaborative student terminals that meet the preset conditions are called third collaborative student terminals. Each third collaborative student terminal starts a random countdown timer for a preset number of seconds locally. For the third collaborative student terminal whose countdown timer expires first, a fourth response message for adding collaborative sentiment analysis request is sent via multicast through the first multicast group address M. The fourth response message includes message ID, associated request ID, current third collaborative student terminal ID, received response status, the number of second newly added faces to be added by the collaborative student terminal, and the feature vector list of the second newly added faces. After receiving the fourth response message, other third collaborative student terminals stop the countdown timer. After receiving the fourth response message, the collaborative student terminal needs to update its local mapping table, that is, replace the collaborative student terminal ID that is offline with the third collaborative student terminal ID whose countdown timer expires first in the collaborative student terminal ID list. If the current third collaborative student terminal has not joined the second multicast group address, after sending the fourth response message, it sends an IGMP member report message to its directly connected router to join the second multicast group address of the corresponding student terminal that needs to collaborate. Then, it receives the subsequent video stream of the first image frame through the second multicast group address, extracts a frame from the subsequent video stream for face detection and extraction, obtains the coordinates of the face detection box corresponding to the feature vector in the feature vector list of the second newly added face, and performs sentiment analysis to obtain the sentiment category of the second newly added face. At the same time, the current third collaborative student terminal updates its local mapping table. Specifically, the current third collaborative student terminal replaces the collaborative student terminal ID that is offline with the current third collaborative student terminal ID in the collaborative student terminal ID list corresponding to the collaborative student terminal ID in the local mapping table; replaces the face emotion category of the collaborative student terminal ID that is offline with the emotion category of the second newly added face in the face emotion category list; and replaces the face detection timestamp of the collaborative student terminal ID that is offline with the timestamp when the second newly added face was detected in the face detection timestamp list. Then, the current third collaborative student terminal extracts a new frame from the subsequent video stream according to the preset first cycle, and processes it according to the processing steps of each first collaborative student terminal.
10. The interactive method for AI online education as described in claim 9, characterized in that: When all face detection timestamps in the face detection timestamp list of the local mapping table of the student client needing to collaborate are updated, the student client needs to collaborate to retrieve the local mapping table and send analysis information to the server. The analysis information includes the ID of the student client needing to collaborate, the feature vector of each face in the feature vector list, the detection box coordinates of each face in the face detection box coordinate list, and the emotion category of each face in the emotion category list. After receiving the analysis information, the server retrieves the feature vector of each face from the pre-saved student information table to obtain the student name corresponding to the feature vector of each face. The student information table includes the student ID, name, and the feature vector of the face. Based on the coordinates of the face detection bounding boxes, the server overlays the student's name and emotion category onto the detection bounding boxes of each face in the current video stream of the student's computer. The server then sends the overlaid video stream to the teacher's computer, where the teacher asks questions based on the student's name and emotion category.
Citation Information
Patent Citations
Emotion attention method for AI online education
CN120183024A
Content interaction method for AI online education
CN120183260A