Online education interaction method and system based on big data

By capturing and analyzing students' learning images in real time, combining convolutional neural networks and dynamic factor computing, personalized learning content is recommended, which solves the problem that the existing online education system cannot perceive students' learning status and adjusts learning content, and improves learning effect and interactivity.

CN120162486APending Publication Date: 2025-06-17杨倩倩
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234198.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing online education system is difficult to perceive students' learning status comprehensively and accurately, resulting in the inability to dynamically adjust the learning content and lack of targetedness and personalization.

Method used

The online education platform's camera captures students' learning images in real time, uses convolutional neural network to extract key features, such as text, symbols and student actions, recognize learning status, and combines real-time learning data to calculate dynamic factors and fusion factors, thereby recommending personalized learning content.

Benefits of technology

It realizes accurate perception of students' learning status and recommendation of personalized learning content, which significantly improves the interactivity and learning effect of online education.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162486A_ABST
    Figure CN120162486A_ABST
Patent Text Reader

Abstract

The invention provides an online education interaction method and system based on big data, and relates to the technical field of online education, and the method comprises the steps: recognizing the learning states of students according to key features, and endowing each learning state with a corresponding numerical value; acquiring real-time learning data of the student; according to the real-time learning data, dynamic factors are calculated, and the dynamic factors reflect real-time learning performance and changes of the students; calculating a fusion factor according to the learning state and the dynamic factor; comparing the fusion factor with the content in a pre-numeralized education resource library to obtain the final learning content meeting the current learning state and demand of the student; and recommending the learning content to the student in real time so that the student can learn according to the learning content and complete interaction. The learning state of the student can be comprehensively and accurately perceived, and personalized learning content recommendation can be provided according to the real-time demand of the student.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of online education, and particularly to an online education interaction method and system based on big data. Background Art

[0002] With the rapid development of network technology and multimedia technology, online education has become an important part of the modern education system. Online education platforms provide students with flexible and convenient learning channels, enabling students to access high-quality educational resources anytime and anywhere. However, some traditional online education platforms lack the ability to perceive students' learning status in real time and provide personalized recommendations, so the learning effect of students may not be fully guaranteed.

[0003] Specifically, in the existing online education system, although the learning progress of students can be evaluated through simple data such as the login time and video viewing duration of students, these data may not comprehensively reflect the real learning status of students, such as concentration and learning activity. Therefore, the existing online education system may not be able to dynamically adjust the recommended learning content according to the real-time learning status of students, making the learning process of students lack pertinence and personalization. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an online education interaction method and system based on big data, which can not only comprehensively and accurately perceive the learning status of students, but also provide personalized learning content recommendations according to the real-time needs of students, thereby significantly improving the interactivity and learning effect of online education.

[0005] To solve the above technical problem, the technical solution of the present invention is as follows:

[0006] In a first aspect, an online education interaction method based on big data, the method includes:

[0007] Real-time capture of students' learning images through the camera of the online education platform and preprocess them to obtain preprocessed images;

[0008] Use a convolutional neural network to extract key features in the preprocessed images, and the key features include text, symbols, and students' actions;

[0009] According to the key features, identify the learning status of students and assign corresponding numerical values to each learning status;

[0010] Obtain the real-time learning data of students;

[0011] According to the real-time learning data, calculate a dynamic factor, and the dynamic factor reflects the real-time learning performance and changes of students;

[0012] Calculate the fusion factor according to the learning status and dynamic factor;

[0013] Compare the fusion factor with the content in the pre-numericalized educational resource library to obtain the final learning content that meets the current learning status and needs of the students;

[0014] Recommend the learning content to the students in real time so that the students can learn according to the learning content and complete the interaction.

[0015] Furthermore, use a convolutional neural network to extract key features from the preprocessed image. The key features include text, symbols, and student actions, including:

[0016] Perform multi-scale feature extraction on the preprocessed image to obtain the extraction result;

[0017] According to the extraction result, generate text candidate boxes through a convolutional neural network, and perform non-maximum suppression on the candidate boxes to obtain the final text detection result. The final text detection result is a set of rectangular box coordinates, indicating the position of the text area in the image;

[0018] For each detected text area, perform text analysis through a convolutional recurrent neural network to obtain the text content in the preprocessed image;

[0019] According to the text content in the preprocessed image, detect and locate the symbols in the preprocessed image through a target detection convolutional neural network model to extract symbol type, position, and size features;

[0020] Detect the body key points of the student in the video through the human pose estimation model of the convolutional neural network;

[0021] The human pose estimation model outputs the coordinate positions of each key point. The coordinate positions mark the specific positions of each part of the student's body. According to the coordinate positions of the key points, the overall body pose of the student is constructed;

[0022] By analyzing the spatial position of each key point in a single frame, obtain the instantaneous pose of the student, and by analyzing the position changes of the key points in consecutive frames, identify the actions of the student.

[0023] Furthermore, according to the key features, identify the learning status of the student and assign corresponding numerical values to each learning status, including:

[0024] Analyze the student's concentration by analyzing the student's head position, eye fixation points, and body inclination;

[0025] Identify the learning activity of the student by analyzing the interaction mode between the student and the educational platform;

[0026] Based on the students' concentration and learning activity level, use a classifier to identify the learning state of each student. The classifier classifies each student into a predefined learning state according to the characteristics of the students' concentration and learning activity level, and assigns a corresponding numerical value.

[0027] Furthermore, calculate a dynamic factor based on real-time learning data, including:

[0028] Calculate the dynamic factor DF through DF = w1×(β×F 1c +(1 - β)×F 1p +α×VI)+w2×F2+w3×(β×F 3c +(1 - β)×F 3p ), where w1, w2, and w3 are weights, F 1c and F 1p are the concentration values of the current period and the previous period respectively, F 3c and F 3p are the interaction frequency values of the current period and the previous period respectively, β is the smoothing factor of EMA, 0 < β < 1, α is the coefficient to adjust the influence of concentration volatility, and F2 is the learning progress rate of the current period.

[0029] Furthermore, calculate a fusion factor based on the learning state and the dynamic factor, including:

[0030] Calculate the fusion factor CF through CF = λ×[θ×(α1×I + β1×TC)+(1 - θ)×(γ×PE - δ×NE)]+(1 - λ)×DF, where CF represents the fusion factor, λ represents the adjustment parameter, θ represents the adjustment parameter, α1 represents the weight coefficient, β1 represents the weight coefficient, I represents the number of interactions, referring to the interaction frequency of the student in the classroom, TC represents the number of tasks completed, referring to the number of learning tasks completed by the student, γ represents the weight coefficient, δ represents the weight coefficient, PE represents the number of positive emotion expressions, referring to the frequency of the student expressing positive emotions, and NE represents the number of negative emotion expressions, referring to the frequency of the student expressing negative emotions.

[0031] Furthermore, compare the fusion factor with the content in the pre-numericalized educational resource library to obtain the final learning content that meets the current learning state and needs of the student, including:

[0032] Determine the characteristic dimensions of the fusion factor CF, and the characteristic dimensions include the interest area and the emotional state;

[0033] Map the interest area value and the emotional state value of the fusion factor CF to the corresponding labels of the educational resources to obtain the mapping result;

[0034] According to the mapping result, calculate the difference between the fusion factor CF of the student and each educational resource label to obtain the matching degree calculation result;

[0035] Based on the matching degree calculation result, determine the final learning content that conforms to the current learning status and needs of the student.

[0036] Furthermore, the key points include the head, shoulders, wrists, and knees.

[0037] In a second aspect, an online education interaction system based on big data includes:

[0038] An acquisition module, configured to capture the learning image of the student in real time through the camera of the online education platform and perform preprocessing to obtain a preprocessed image; extract key features in the preprocessed image by using a convolutional neural network, where the key features include text, symbols, and student actions; identify the learning status of the student according to the key features, and assign corresponding numerical values to each learning status; obtain the real-time learning data of the student;

[0039] A processing module, configured to calculate a dynamic factor according to the real-time learning data, where the dynamic factor reflects the real-time learning performance and changes of the student; calculate a fusion factor according to the learning status and the dynamic factor; compare the fusion factor with the content in the pre-numericalized educational resource library to obtain the final learning content that conforms to the current learning status and needs of the student; recommend the learning content to the student in real time so that the student can learn according to the learning content and complete the interaction.

[0040] In a third aspect, a computing device includes:

[0041] One or more processors;

[0042] A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the one or more processors to implement the above method.

[0043] In a fourth aspect, a computer-readable storage medium stores a program, and when the program is executed by a processor, the above method is implemented.

[0044] The above solution of the present invention has at least the following beneficial effects:

[0045] The above solution of the present invention captures the learning image of the student in real time through a camera, and extracts key features including text, symbols, and student actions by using a convolutional neural network, so as to accurately identify the learning status of the student. This method can perceive important information such as the student's concentration and learning activity in real time.

[0046] Calculate dynamic factors based on real-time learning data to reflect the real-time learning performance and changes of students, fuse the learning status with the dynamic factors to generate a fusion factor, and use it to obtain learning content that meets the current learning status and needs of students, which can ensure that the recommended learning content always matches the actual learning status of students and improve the pertinence and efficiency of learning.

[0047] By comparing the fusion factor with the content in the pre-numericalized educational resource library, the present invention can provide personalized learning content recommendations for students. Such recommendations not only consider the static learning characteristics of students (such as learning styles, interests, etc.), but also fully consider the real-time learning status and needs of students.

[0048] By recommending learning content to students in real time and enabling students to learn according to the recommended content, the present invention significantly improves the interactivity and learning effect of online education. Students can obtain learning resources that meet their learning status and needs in a timely manner, thereby participating in the learning process more actively and improving learning efficiency and grades. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a schematic flowchart of an online education interaction method based on big data provided by an embodiment of the present invention.

[0050] Figure 2 is a schematic diagram of an online education interaction system based on big data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0052] As Figure 1 shown, an embodiment of the present invention proposes an online education interaction method based on big data, and the method includes:

[0053] Step 11, capturing the learning image of the student in real time through the camera of the online education platform and performing preprocessing to obtain a preprocessed image;

[0054] Step 12, using a convolutional neural network to extract key features in the preprocessed image, and the key features include characters, symbols, and student actions;

[0055] Step 13, identifying the learning status of the student according to the key features and assigning corresponding numerical values to each learning status;

[0056] Step 14, obtain the real-time learning data of the student;

[0057] Step 15, calculate a dynamic factor based on the real-time learning data, where the dynamic factor reflects the real-time learning performance and changes of the student;

[0058] Step 16, calculate a fusion factor based on the learning status and the dynamic factor;

[0059] Step 17, compare the fusion factor with the content in the pre-numericalized educational resource library to obtain the final learning content that meets the current learning status and needs of the student;

[0060] Step 18, recommend the learning content to the student in real time so that the student can learn according to the learning content and complete the interaction.

[0061] In the embodiment of the present invention, the learning image of the student is captured in real time through a camera, and a convolutional neural network is used to extract key features, including characters, symbols, and student actions, etc., so as to accurately identify the learning status of the student. This method can perceive important information such as the concentration and learning activity of the student in real time. Calculating a dynamic factor based on the real-time learning data to reflect the real-time learning performance and changes of the student, and fusing the learning status with the dynamic factor to generate a fusion factor, which is used to obtain the learning content that meets the current learning status and needs of the student from the educational resource library, can ensure that the recommended learning content always matches the actual learning status of the student, improving the pertinence and efficiency of learning. By comparing the fusion factor with the content in the pre-numericalized educational resource library, the present invention can provide personalized learning content recommendations for students. Such recommendations not only consider the static learning characteristics of students (such as learning styles, interests, etc.), but also fully consider the real-time learning status and needs of students. By recommending the learning content to the student in real time and enabling the student to learn according to the recommended content, the present invention significantly improves the interactivity and learning effect of online education. Students can obtain learning resources that meet their learning status and needs in a timely manner, thereby participating in the learning process more actively and improving learning efficiency and grades.

[0062] In another preferred embodiment of the present invention, step 11 above may include: The online education platform captures a video stream in real time through a camera connected to the student's computer or device. This video stream provides continuous images showing the student's learning environment, facial expressions, body language, etc. The real-time video stream may contain various noises, such as light changes, cluttered backgrounds, etc. The preprocessing step will first remove these noises to improve the image quality, convert the captured video stream into a format suitable for subsequent processing, such as resizing the image, cropping unnecessary background parts, etc., normalize the image to ensure that the brightness and contrast of different images are within the same range, and convert the image from the RGB color space to the grayscale or other color spaces to simplify the processing flow or improve the effect of feature extraction. For example, assume that an online education platform is conducting a remote math course. At the beginning of the course, the platform captures the student's images in real time through the student's laptop camera. These images show the student sitting at the desk facing the computer screen. The system continuously receives the video stream from the camera, which captures the student's activities at a speed of 30 frames per second. A digital filter is applied to remove the light noise caused by indoor light changes, and each frame of the image is cropped to only include the upper body and face of the student, and the image size is adjusted to a unified resolution (for example, 640x480 pixels). This preprocessed image will be used to recognize the student's facial expressions, estimate the human body posture, etc., as the basis for judging the student's learning state (such as concentration, emotional state), so that the online education platform can provide a more personalized learning experience and feedback.

[0063] In another preferred embodiment of the present invention, step 11 above may further include:

[0064] Create a grayscale histogram for the image. This histogram shows the number of pixels at each gray level (usually from 0 to 255) in the image, specifically including: traversing each pixel in the image, counting the frequency of each gray level occurrence, and recording these frequencies in an array. The index of this array corresponds to the gray level;

[0065] According to the grayscale histogram, calculate the cumulative distribution function, specifically including: creating an array of the same size as the grayscale histogram to store the values of the cumulative distribution function (CDF), accumulating the values in the histogram. Specifically, for each gray level in the grayscale histogram, its cumulative distribution function (CDF) value is the sum of the frequencies of this gray level and all lower gray levels in the histogram; normalizing the cumulative distribution function (CDF) by dividing each value in the cumulative distribution function (CDF) by the maximum value of the cumulative distribution function (CDF). Through this process, the new gray level of each pixel will be mapped based on its corresponding value in the cumulative distribution function (CDF) of its original gray level, so as to achieve the purpose of enhancing the image contrast.

[0066] In a preferred embodiment of the present invention, step 12 may include:

[0067] Step 121, perform multi-scale feature extraction on the preprocessed image to obtain an extraction result;

[0068] Step 122, according to the extraction result, generate text candidate boxes through a convolutional neural network, perform non-maximum suppression on the candidate boxes to obtain the final text detection result. The final text detection result is a set of rectangular box coordinates, representing the position of the text area in the image. For example, assume there is an image containing several paragraphs of text. The following is the specific processing flow: The image is input into a pre-trained CNN model. According to the detected features, the system generates multiple rectangular candidate boxes surrounding the detected text. If there are overlapping candidate boxes, NMS will select the box with the highest confidence and remove the other boxes. The remaining candidate boxes are determined as the areas containing text, and the four corner coordinates of each box are output, determining the exact position of the text in the image;

[0069] Step 123, for each detected text area, perform text analysis through a convolutional recurrent neural network to obtain the text content in the preprocessed image;

[0070] Step 124, according to the text content in the preprocessed image, detect and locate the symbols in the preprocessed image through a target detection convolutional neural network model to extract symbol type, position, and size features;

[0071] Step 125, detect the body key points of the student in the video through a human pose estimation model of a convolutional neural network. The key points include the head, shoulders, wrists, and knees;

[0072] Step 126, the human pose estimation model outputs the coordinate positions of each key point. The coordinate positions mark the specific positions of each part of the student's body. According to the coordinate positions of the key points, the overall body pose of the student is constructed;

[0073] Step 127, analyze the spatial position of each key point in a single frame to obtain the instantaneous pose of the student, and identify the actions of the student through the position changes of the key points in consecutive frames.

[0074] In the embodiments of the present invention, by performing multi-scale feature extraction on the preprocessed image, it is possible to capture text and symbol information of different sizes and positions in the image, improving the accuracy and robustness of recognition. A convolutional neural network is used to generate text candidate boxes, and the final text detection result is obtained through non-maximum suppression, which can accurately locate the text area in the image. Further, through a convolutional recurrent neural network to perform text analysis on the detected text area, the text content in the preprocessed image can be accurately recognized. By using a target detection convolutional neural network model to detect and locate symbols in the preprocessed image, the type, position, and size features of the symbols can be extracted. By using a human pose estimation model of a convolutional neural network to detect the body key points of students in the video and construct the overall body pose of the students, the actions and pose changes of the students can be accurately captured. By analyzing the spatial position of each key point in a single frame and the position changes in consecutive frames, the actions of the students can be recognized. These action information is closely related to the learning state. For example, the head pose and hand movements of the students can reflect their learning states such as concentration and understanding. Therefore, the learning state of the students can be more accurately judged through action recognition.

[0075] In another preferred embodiment of the present invention, the above step 121 may include:

[0076] According to the application scenario and the target, analyze the selection range of the standard deviation of the Gaussian kernel. If the target is to capture fine features, such as facial expression recognition, start from a smaller standard deviation of the Gaussian kernel; if the target is to capture a larger range of features, such as pose recognition, start from a larger standard deviation of the Gaussian kernel.

[0077] Define the range of the standard deviation of the Gaussian kernel, starting from a smaller standard deviation of the Gaussian kernel (such as 1.0), and gradually increasing. Each time it increases, the standard deviation of the Gaussian kernel may double (such as 1, 2, 4, 8,...); for larger images, use a larger standard deviation of the Gaussian kernel to obtain a similar blurring effect.

[0078] For each standard deviation of the Gaussian kernel, blur the preprocessed image through a Gaussian filter to create a scale space. Specifically, for each standard deviation of the Gaussian kernel, calculate the corresponding Gaussian kernel, convolve the preprocessed image with the Gaussian kernel to generate a blurred image. A series of blurred images constitute the scale space. Among them, each blurring level represents a different scale. In this way, the appearance of the image at different sizes can be simulated, from large scale (more blurred) to small scale (more clear).

[0079] Extract key features at each level of the scale space to capture image features from details to the whole. Specifically, it includes:

[0080] Automatically identify key features using a predetermined algorithm at each level of the scale space, including: calculating the gradient intensity and direction of each pixel point in the image through a gradient operator, where the gradient intensity reflects the speed of brightness change, i.e., the possible position of the edge; and the gradient direction provides the direction information of the edge.

[0081] Use a double-threshold method to determine true edge points. Points below the low threshold are considered non-edge points, points above the high threshold are considered strong edge points, and points between the two are determined whether they are edge points through an edge connection algorithm.

[0082] Identify corner points by analyzing the degree of brightness change when the image window moves in various directions. If the movement within a window causes significant brightness change in any direction, the point where the window is located is considered a corner point.

[0083] For each point in the image, calculate an autocorrelation matrix through the gradient of the image in the neighborhood of this point. This matrix reflects the change of the brightness structure of the image near this point; determine whether a point is a corner point by calculating a response function R. The value of the response function R is based on the eigenvalues of the autocorrelation matrix, which measures the uniformity of the gradient change of this point in various directions. A high value of the response function R indicates that this point is a corner point.

[0084] Apply a threshold to the response function R to determine the final corner points. When the response function R exceeds a certain set threshold, the corresponding points are marked as corner points. Among them, the specific calculation formula of the response function R is:

[0085]

[0086] Among them, I x and I y These two parameters respectively represent the gradients of the image in the x and y directions. The gradient is a vector that points to the direction of the fastest brightness change in the image, and its magnitude represents the rate of brightness change. In corner point detection, the gradient is used to capture structural information such as edges and corner points in the image. and These two summation terms respectively represent the sum of the squares of the gradient components of all pixels in the x and y directions in the neighborhood around the image point, which measure the intensity of the brightness change of the image in this direction. ∑I x I y This summation term represents the sum of the products of the gradient components of all pixels in the x and y directions in the neighborhood around the image point, which reflects the correlation of the brightness change of the image in different directions. k is a constant, and its value range is usually between 0.04 and 0.06.

[0087] For each pixel point in the image, a feature vector is constructed. The feature vector integrates the edge information and corner information of the pixel point. For each pixel point, the feature vector can contain the following elements:

[0088] Edge intensity, a value extracted from the edge map, representing the intensity of the point as an edge;

[0089] Corner intensity, a value extracted from the corner map, representing the intensity of the point as a corner;

[0090] Map the feature vectors of all points into the same feature space to form a unified feature. This feature space can accommodate all possible values of edge intensity and corner intensity.

[0091] For the posture and actions of the student, use a human pose estimation algorithm to automatically identify the key points of the human body in the image and track the position changes in consecutive frames, analyze the actions of the student, and thus infer their learning behavior and state. Among them, analyze the actions of the key points to identify the behavior patterns of the student, including learning activities (such as reading, writing, typing) and non-learning activities (such as distraction, yawning), and infer the learning state of the student, such as concentration, distraction, fatigue, etc., according to the behavior patterns and the duration, frequency, etc. of the actions.

[0092] In another preferred embodiment of the present invention, in the above steps 123 and 124, it may include:

[0093] According to the obtained text region coordinates, locate and crop each text region from the original image. These cropped image segments respectively contain the text in different parts of the image;

[0094] First, the convolutional part of the convolutional recurrent neural network (CRNN) extracts features from the image segments. These features capture the important visual information of the text, such as the shape and layout of characters; then, the recurrent neural network (RNN) part processes the feature sequence extracted by the convolutional layer to model the sequential relationship between characters in the text. The convolutional recurrent neural network (CRNN) can process texts of different lengths and recognize each character in the sequence; finally, the output signal of the convolutional recurrent neural network (CRNN) is converted into a text string through a decoding layer (such as CTC decoding). This layer maps the output of the network to a series of characters to obtain the final text content.

[0095] For each text region, the convolutional recurrent neural network (CRNN) model outputs a text string, representing the text content in the preprocessed image. These texts can be directly used for further text analysis or information extraction tasks;

[0096] Integrate the recognized symbol types and their position and size features together to form a comprehensive feature set of the image text content.

[0097] In the embodiment of the present invention, by cropping each text region from the original image according to the text region coordinates, the accurate extraction of the text content is ensured, and the interference of background noise or other non-text elements is avoided, thereby improving the accuracy of subsequent text recognition. The convolutional part of the Convolutional Recurrent Neural Network (CRNN) can effectively extract the key visual features of the text from the image fragments, such as the shape, size, spacing of the characters, etc. The RNN part in the Convolutional Recurrent Neural Network (CRNN) can process variable-length sequences and model the dependencies between characters. This flexibility enables the model to process text lines of different lengths, whether they are words, phrases, or complete sentences, and can be effectively recognized. The Convolutional Recurrent Neural Network (CRNN) model combines multiple steps such as feature extraction, sequence modeling, and transcription (through CTC decoding) to achieve end-to-end recognition from the original image to the text string. This integrated design simplifies the processing flow and improves the recognition efficiency.

[0098] For example, when specifically applied, assume there is an image containing a handwritten mathematical equation. First, use a convolutional neural network (such as a CNN-based text detection model) to detect the text region containing the mathematical equation in the image, crop the detected text region, and perform preprocessing, such as resizing and converting to a grayscale image, to adapt to subsequent processing. Then, input the preprocessed image into the Convolutional Recurrent Neural Network (CRNN) for text recognition; according to the text content recognized by the CRNN, use a target detection convolutional neural network model (such as Faster R-CNN) to analyze the preprocessed image again, aiming to locate the specific positions of each symbol in the text. For example, the model can recognize and locate "5", "x", "2", "+", "3", "x", "-", "7", "=", "; for each detected symbol, the model extracts its type (number, letter, operator, etc.), position (coordinates in the image), and size (the size of the symbol). For example, it may be determined that "x" is located in the central part of the image and has a size of 8×8 pixels. Therefore, through step 123 and step 124, not only the text content in the image is successfully recognized, but also the specific information of each symbol in the text is analyzed in detail.

[0099] In another preferred embodiment of the present invention, in steps 125 to 127:

[0100] Human Pose Estimation Model Using Convolutional Neural Network (CNN): Such models, like OpenPose or PoseNet, can automatically identify the key points of the human body in video frames. These key points include, but are not limited to, the head, shoulders, wrists, and knees. The model analyzes the pixel data in the video frame to identify the positions of various parts of the human body, and each identified position corresponds to a key point.

[0101] The human pose estimation model outputs a coordinate position for each detected key point. This coordinate marks the specific position of each part of the student's body in the video frame. By connecting the coordinates of these key points, an overall body pose model of the student can be constructed. This model depicts the pose of the student in the video frame, including standing, sitting, raising hands, etc.

[0102] By analyzing the spatial positions of the key points in a single frame, the instantaneous pose of the student can be obtained, such as sitting upright, reclining, etc. By comparing the position changes of the key points in consecutive frames, the actions of the student can be identified, such as raising hands, nodding, leaving the seat, etc.

[0103] For example, assume in an online education math course, the teacher wants to understand the students' reactions and understanding of new teaching content. Through the human pose estimation model, the system detects the key point positions of each student in the classroom video, including the head, shoulders, wrists, and knees. The system constructs a pose model for each student, clearly showing their sitting postures, hand positions, etc. The system observes that a student raises their hand frequently when explaining a new concept, which may indicate that they have questions or are curious about the material. At the same time, by analyzing the sitting postures and head postures of other students, the system determines that most students are maintaining focus.

[0104] In a preferred embodiment of the present invention, the above step 13 may include:

[0105] Step 131, analyze the student's focus by analyzing the student's head position, eye gaze point, and body inclination;

[0106] Step 132, identify the student's learning activity by analyzing the interaction mode between the student and the educational platform;

[0107] Step 133, according to the student's focus and learning activity, use a classifier to identify the learning state of each student. The classifier classifies each student into a predefined learning state based on the student's focus and learning activity characteristics and assigns corresponding numerical values.

[0108] In the embodiment of the present invention, step 131 can comprehensively and accurately evaluate the concentration of students by analyzing multiple aspects such as the head position, eye fixation points, and body inclination of students. This multi-dimensional analysis method can capture subtle changes such as students' attention dispersion, fatigue, or distraction, providing timely feedback to teachers or educational platforms, which helps to adjust teaching strategies or conduct personalized interventions. Step 132 can deeply identify the learning activity of students by analyzing the interaction patterns between students and the educational platform, such as click frequency, dwell time, interaction type, etc. This interaction analysis can reveal the degree of interest, participation willingness, and learning motivation of students towards learning content. Step 133 uses a classifier to accurately classify each student into a predefined learning state according to the characteristics of students' concentration and learning activity. This classification method can objectively and quantitatively evaluate the learning state of students, avoiding errors and biases caused by subjective judgments. At the same time, corresponding numerical values are assigned to each learning state, making the comparison and analysis of learning states more intuitive and convenient. By identifying the learning state of each student, the educational platform can provide personalized learning support and feedback to students. For example, for students with low concentration, more attractive learning resources can be recommended or real-time reminders and incentives can be provided; for students with high learning activity, more challenges and expansion opportunities can be given to meet their learning needs.

[0109] In another preferred embodiment of the present invention, in step 131, analyzing the head position and eye fixation points of students through human pose estimation technology or eye tracking technology can help determine whether students are looking at teaching content or other irrelevant things; at the same time, analyzing the body inclination of students to judge whether students maintain a good sitting posture, which is usually associated with high concentration.

[0110] Step 132: Analyze the learning activity of students by monitoring their interactions with the educational platform, such as click-through rate, response speed, participation in online tests, etc. High-frequency reasonable interactions usually indicate high learning activity. Step 133: Combine the data analyzed in Step 131 and Step 132 and use a machine learning classifier (such as support vector machine, decision tree, neural network, etc.) to evaluate the learning state of students. The classifier is trained to identify different learning states, such as "highly focused", "distracted", "needs rest", etc., and classify students based on their focus and learning activity characteristics. For example, in an online math class, students participate in learning through a tablet. The system analyzes the head position and eye gaze points of students through their cameras. For one student, his head position and eye gaze points are mostly focused on the teaching video, but he is occasionally distracted and looks out of the window. The system also notices that when Xiaoming is very focused, his body hardly tilts and he maintains a good sitting posture. Through the interaction records of the platform, it is found that Xiaoming has a high interaction frequency in class. He actively participates in online exercises and submits his homework in a timely manner. Considering the student's focus and learning activity comprehensively, the classifier classifies him as being in the "highly focused" state. Although he is occasionally distracted, his high participation and mostly focused behavior show his high interest and positive attitude towards the learning content. In this way, teachers can obtain real-time feedback on the learning state of students and then provide personalized guidance and support, such as providing additional challenging exercises for this student or sending a reminder when observing his distraction to help him refocus on learning.

[0111] In a preferred embodiment of the present invention, the above step 15 may include:

[0112] Step 151: Calculate the dynamic factor DF through DF = w1×(β×F 1c +(1-β)×F 1p +α×VI)+w2×F2+w3×(β×F 3c +(1-β)×F 3p ), where w1, w2, and w3 are weights, F 1c and F 1p are the focus values of the current period and the previous period respectively, F 3c and F 3p are the interaction frequency values of the current period and the previous period respectively, β is the smoothing factor of EMA, 0 < β < 1, α is the coefficient for adjusting the influence of focus volatility, and F2 is the learning progress rate of the current period.

[0113] In the embodiments of the present invention, by calculating the dynamic factor DF, the educational platform can provide personalized learning feedback for each student, identify areas that require additional support or challenge. The change in the dynamic factor DF can be used as a basis for adjusting the learning path, helping teachers or systems to timely adjust teaching strategies and better meet the learning needs of students.

[0114] For example, suppose there is a student taking an online math course. The platform calculates the dynamic factor by analyzing their learning behavior and performance: within a recent learning cycle, the student's focus F 1c has increased compared to the previous cycle F 1p , the interaction frequency F 3c remains unchanged, and the learning progress rate F2 has increased significantly; by setting appropriate weights w1, w2, w3 and parameters β and α, the dynamic factor DF of this student is calculated, showing a significant improvement in their learning performance. Based on the analysis of the dynamic factor DF, the educational platform recommends some more difficult math problems to this student to match their improved learning ability and progress. In this way, the dynamic factor helps the educational platform to achieve real-time monitoring of the student's learning status and personalized learning path adjustment.

[0115] In a preferred embodiment of the present invention, the above step 16 may include:

[0116] Step 161, calculate the fusion factor CF by CF = λ×[θ×(α1×I + β1×TC)+(1 - θ)×(γ×PE - δ×NE)]+(1 - λ)×DF, where CF represents the fusion factor, λ represents the adjustment parameter, θ represents the adjustment parameter, α1 represents the weight coefficient, β1 represents the weight coefficient, I represents the number of interactions, referring to the interaction frequency of the student in the classroom, TC represents the number of tasks completed, referring to the number of learning tasks completed by the student, γ represents the weight coefficient, δ represents the weight coefficient, PE represents the number of positive emotion expressions, referring to the frequency of the student expressing positive emotions, and NE represents the number of negative emotion expressions, referring to the frequency of the student expressing negative emotions.

[0117] In the embodiments of the present invention, the fusion factor CF is used to comprehensively evaluate the learning status of students, considering multiple dimensions such as interaction, task completion, emotion expression and real-time performance; based on the analysis of the fusion factor CF, the educational platform or teacher can provide more personalized learning support for students, such as targeted encouragement, task adjustment or emotion intervention. By accurately monitoring and timely responding to the learning status of students, it helps to improve learning efficiency and grades, and at the same time enhance the learning experience of students. For example, suppose in an online language course, the learning data of student A is as follows:

[0118] Student A had a high number of interactions with the course content in the most recent period, completed all assigned learning tasks, and frequently expressed positive feedback in the forum and assignment comments, occasionally expressing concerns about the difficulty level. Based on this information and the corresponding weight coefficients, the integration factor for Student A was calculated. According to the integration factor, the educational platform recommended some additional tutoring materials to help him overcome learning difficulties and automatically sent encouraging messages through the system to boost his confidence and motivation.

[0119] In a preferred embodiment of the present invention, step 17 may include:

[0120] Step 171, determining the characteristic dimensions of the integration factor CF, where the characteristic dimensions include the area of interest and emotional state;

[0121] Step 172, mapping the numerical value of the area of interest and the emotional state value of the integration factor CF to the corresponding labels of the educational resources to obtain a mapping result;

[0122] Step 173, calculating the difference between the integration factor CF of the student and each educational resource label according to the mapping result to obtain a matching degree calculation result;

[0123] Step 174, determining the final learning content that conforms to the student's current learning state and needs according to the matching degree calculation result.

[0124] In the embodiment of the present invention, by matching the characteristics with the educational resource labels, it is possible to provide students with learning content that better suits their personal preferences and learning states, thereby enhancing the learning experience and learning effect. By calculating the difference between the integration factor CF of the student and each educational resource label, step 17 can accurately recommend learning content that matches the student's current learning state and needs. This accurate recommendation avoids students being troubled by irrelevant or overly difficult content, improving learning efficiency. Since the interests and emotional states of students may change over time and in different situations, step 17 can flexibly adapt to these changes by continuously updating the value of the integration factor CF and adjust the recommended learning content in real time to ensure that students can always access the learning resources most suitable for their current state. When students find that the recommended learning content matches their interests and emotional states, they are more likely to actively engage in learning and maintain a high level of participation. This positive learning attitude helps to boost students' learning motivation and academic performance, forming a virtuous cycle. By effectively matching the personalized characteristics of students with educational resources, the educational platform can more reasonably allocate and utilize the content in its educational resource library, which not only helps to improve the utilization rate of educational resources but also promotes the continuous update and optimization of educational resources to meet the changing needs of students.

[0125] In another preferred embodiment of the present invention, in step 171, historical learning data of students is collected, including their learning time, grades, interaction frequency, etc. in different subjects; characteristics of students' interest areas are extracted from the collected data. For example, the degree of interest is evaluated by their learning time and grades in different subjects. At the same time, characteristics of emotional states are extracted through students' learning behaviors, feedback or emotion recognition technologies. The extracted characteristics are quantified. For example, the interest areas are divided into three levels: high, medium, and low, and the emotional states are divided into three levels: positive, neutral, and negative, and corresponding numerical values are assigned to each level.

[0126] In step 172, educational resources are labeled, including subject labels, emotional labels, etc. These labels should be able to reflect the content and characteristics of the resources, and map the numerical values of the interest areas and emotional state values in the student's fusion factor CF to the labels of the educational resources. For example, high interest is mapped to the corresponding subject in the subject label, and positive emotion is mapped to the positive emotion in the emotional label.

[0127] In step 173, for each educational resource, by calculating the correlation coefficient between feature vectors, the difference between the student's fusion factor CF and the resource label is determined, and the calculated difference is quantified to obtain a specific numerical value or score, which is used to represent the matching degree between the student and the resource.

[0128] In step 174, according to the calculation result of the matching degree, the educational resources are sorted, and the resources with the highest matching degree are selected as the recommended learning content, and the selected learning content is recommended to the students to meet their current learning status and needs.

[0129] For example, assume that student A spends a lot of time on the history subject and achieves excellent grades, while performing mediocrely in the science subject. Through feature extraction and quantification, it can be determined that among the interest area features of student A, the history subject has high interest and the science subject has low interest. In terms of emotional state, assume that through emotion recognition technology, it is detected that student A is currently in a positive emotional state. Assume that there are two resources in the educational resource library: one is a video lecture on historical battles (tags: history, medium difficulty, motivational), and the other is an operation guide for scientific experiments (tags: science, high difficulty, exploratory). According to the feature mapping results of student A, her high-interest history matches the history tag of the first resource, and the positive emotion matches the motivational tag. By calculating the correlation coefficient between student A and the two educational resources, after calculation, the similarity between student A and the video lecture on historical battles is relatively high (because both interest and emotion match), while the similarity with the operation guide for scientific experiments is relatively low (because the interest does not match). Based on the previous calculation results, the system will recommend the video lecture on historical battles as the preferred resource to student A because it best matches her interests and emotional state. The operation guide for scientific experiments may be recommended as a secondary recommendation or not recommended at all because it does not match the interest area of student A.

[0130] In another preferred embodiment of the present invention, the above step 173 may include:

[0131] By calculating the correlation coefficient between feature vectors, where represents the correlation coefficient between the feature vector of the fusion factor CF of the student and the feature vector of the educational resource label, and respectively represent the means of the feature vectors of the student's fusion factor CF and the educational resource label. The value range of the correlation coefficient is [-1, 1]. When it means complete positive correlation. When it means complete negative correlation. When it means no correlation (but it does not mean that the two are independent). The calculated correlation coefficient is used to measure the strength and direction of the linear relationship between the student's fusion factor CF and the resource label. If the correlation coefficient is close to 1 or -1, it indicates a strong linear relationship between the two; if the correlation coefficient is close to 0, it indicates a weak linear relationship between the two; cf i and r i respectively represent the values of the feature vectors of the student's fusion factor CF and the educational resource label in the i-th dimension. That is to say, if there is a set of data of the student's fusion factor and the labels of educational resources, and these data are all multi-dimensional, then cf i and r irespectively represent the specific values of these two groups of data in the i-th dimension. n represents the number of samples, that is, the number of data points. i is an index variable used to traverse all data points.

[0132] In another preferred embodiment of the present invention, step 174 may include:

[0133] Step 1741, decompose the list of educational resources to be sorted into two smaller sub-lists until the size of the sub-lists is 1 (that is, each sub-list contains only one resource); recursively sort the sub-lists and merge the sorted sub-lists into a large ordered list until merged into one complete list; when merging two sorted sub-lists, compare the first resources of the two sub-lists, select the resource with a higher matching degree and put it into the new ordered list, and remove it from the original sub-list. Repeat this process until both sub-lists are empty. The specific implementation process includes: Suppose there is a list of educational resources, and each resource has a matching degree score associated with it. The merge sort can be performed in the following way: Decompose the list of educational resources into sub-lists of individual resources. If a sub-list has only one resource, it is considered sorted; otherwise, further decompose the sub-list into smaller sub-lists and recursively sort them; when both sub-lists are sorted, merge them into an ordered list. When merging, compare the first resources of the two sub-lists and select the one with a higher matching degree and put it into the new ordered list; continue merging until all sub-lists are merged into one complete ordered list.

[0134] Step 1742, check the sorted resource list, view the list of educational resources after merge sort, and ensure that they are arranged in descending order of matching degree; select the resource with the highest matching degree from the sorted list, specifically including: It will be the first resource in the list because it has the highest matching degree score.

[0135] Step 1743, customize a recommendation message according to the selected learning resources, and send the recommendation message and the learning resource link to students through appropriate channels, including the notification system of the learning platform, email, mobile application push notifications, etc.; include a feedback mechanism in the recommendation message to allow students to provide their views on the recommended content, monitor the interaction of students with the recommended content, such as click-through rate, viewing duration, completion rate, etc., to evaluate the effectiveness of the recommendation, and adjust the recommendation strategy as needed.

[0136] In the embodiments of the present invention, by calculating the correlation coefficient between the eigenvector of the student's fusion factor CF and the eigenvector of the educational resource label, the present invention can accurately measure the matching degree between the student's current learning state and the educational resources. This correlation-based matching method is more accurate than traditional rule-based recommendations and can better meet the personalized learning needs of students. The merge sort algorithm is used to sort the educational resources to ensure that the resources are arranged in descending order of the matching degree. This sorting method is not only efficient but also ensures that resources with high matching degrees are preferentially recommended to students, thereby improving the learning efficiency and satisfaction of students. By customizing the recommended messages and providing a feedback mechanism, the present invention encourages students to interact with the recommended content. This interactivity can not only improve the learning participation of students but also provide valuable feedback data for future recommendations. The present invention monitors the interaction of students with the recommended content (such as click-through rate, viewing duration, completion rate, etc.) and evaluates the effectiveness of the recommendation in real time. This real-time monitoring and evaluation mechanism can help educators promptly discover problems and adjust the recommendation strategy, thereby ensuring that the recommendation system always remains in the best state.

[0137] As Figure 2 shown, an embodiment of the present invention further provides an online education interaction system 20 based on big data, including:

[0138] An acquisition module 21, configured to capture the learning image of the student in real time through the camera of the online education platform and perform preprocessing to obtain a preprocessed image; extract key features in the preprocessed image by using a convolutional neural network, where the key features include text, symbols, and student actions; identify the learning state of the student according to the key features and assign corresponding numerical values to each learning state; acquire the real-time learning data of the student;

[0139] A processing module 22, configured to calculate a dynamic factor according to the real-time learning data, where the dynamic factor reflects the real-time learning performance and changes of the student; calculate a fusion factor according to the learning state and the dynamic factor; compare the fusion factor with the content in the pre-numericalized educational resource library to obtain the final learning content that meets the current learning state and needs of the student; recommend the learning content to the student in real time so that the student can learn according to the learning content and complete the interaction.

[0140] Optionally, extracting key features in the preprocessed image by using a convolutional neural network, where the key features include text, symbols, and student actions, includes:

[0141] Performing multi-scale feature extraction on the preprocessed image to obtain an extraction result;

[0142] According to the extraction results, text candidate boxes are generated through a convolutional neural network, and non-maximum suppression is performed on the candidate boxes to obtain the final text detection results. The final text detection results are a set of rectangular box coordinates, indicating the positions of text regions in the image;

[0143] For each detected text region, text analysis is performed through a convolutional recurrent neural network to obtain the text content in the preprocessed image;

[0144] According to the text content in the preprocessed image, symbols in the preprocessed image are detected and located through a target detection convolutional neural network model to extract symbol type, position, and size features;

[0145] Detect the body key points of students in the video through the human pose estimation model of the convolutional neural network;

[0146] The human pose estimation model outputs the coordinate positions of each key point. The coordinate positions mark the specific positions of each part of the student's body. According to the coordinate positions of the key points, the overall body pose of the student is constructed;

[0147] By analyzing the spatial positions of each key point in a single frame, the instantaneous pose of the student is obtained, and by analyzing the position changes of the key points in consecutive frames, the actions of the student are identified.

[0148] Optionally, according to the key features, the learning states of students are identified, and corresponding numerical values are assigned to each learning state, including:

[0149] By analyzing the head position, eye fixation points, and body inclination of the student, the student's concentration is analyzed;

[0150] By analyzing the interaction mode between the student and the educational platform, the learning activity of the student is identified;

[0151] According to the student's concentration and learning activity, a classifier is used to identify the learning state of each student. The classifier classifies each student into a predefined learning state according to the concentration and learning activity characteristics of the student and assigns corresponding numerical values.

[0152] Optionally, according to the real-time learning data, dynamic factors are calculated, including:

[0153] The dynamic factor DF is calculated through DF = w1×(β×F 1c +(1 - β)×F 1p +α×VI)+w2×F2+w3×(β×F 3c +(1 - β)×F 3p ) where w1, w2, and w3 are weights, F 1c and F 1p are the concentration values of the current period and the previous period respectively, F3c and F 3p are the interaction frequency values of the current cycle and the previous cycle respectively, β is the smoothing factor of EMA, 0 < β < 1, α is the coefficient for adjusting the influence of the focus volatility, and F2 is the learning progress rate of the current cycle.

[0154] Optionally, calculate a fusion factor according to the learning state and dynamic factors, including:

[0155] Calculate the fusion factor CF according to the learning state and dynamic factors through CF = λ×[θ×(α1×I + β1×TC) + (1 - θ)×(γ×PE - δ×NE)] + (1 - λ)×DF, where CF represents the fusion factor, λ represents the adjustment parameter, θ represents the adjustment parameter, α1 represents the weight coefficient, β1 represents the weight coefficient, I represents the number of interactions, referring to the interaction frequency of the student in the classroom, TC represents the number of tasks completed, referring to the number of learning tasks completed by the student, γ represents the weight coefficient, δ represents the weight coefficient, PE represents the number of positive emotion expressions, referring to the frequency of the student expressing positive emotions, and NE represents the number of negative emotion expressions, referring to the frequency of the student expressing negative emotions.

[0156] Optionally, compare the fusion factor with the content in the pre - numericalized educational resource library to obtain the final learning content that meets the student's current learning state and needs, including:

[0157] Determine the characteristic dimensions of the fusion factor CF, and the characteristic dimensions include the interest field and the emotional state;

[0158] Map the interest field value and the emotional state value of the fusion factor CF to the corresponding labels of the educational resources to obtain a mapping result;

[0159] According to the mapping result, calculate the difference between the student's fusion factor CF and each educational resource label to obtain a matching degree calculation result;

[0160] According to the matching degree calculation result, determine the final learning content that meets the student's current learning state and needs.

[0161] Optionally, the key points include the head, shoulders, wrists, and knees.

[0162] It should be noted that this device corresponds to the above - mentioned method, and all implementation manners in the above - mentioned method embodiments are applicable to this embodiment and can also achieve the same technical effects.

[0163] An embodiment of the present invention also provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method as described above. All implementation manners in the above - mentioned method embodiments are applicable to this embodiment and can also achieve the same technical effects.

[0164] An embodiment of the present invention also provides a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to execute the method described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0165] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0166] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0167] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.

[0168] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0169] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0170] When the above-described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0171] In addition, it should be noted that in the devices and methods of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other. For those of ordinary skill in the art, it can be understood that all or any steps or components of the methods and devices of the present invention can be implemented in any computing device (including processors, storage media, etc.) or in a network of computing devices in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.

[0172] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a well-known general-purpose device. Therefore, the object of the present invention can also be achieved only by providing a program product containing program codes for implementing the method or device. That is to say, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be noted that in the devices and methods of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other.

[0173] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An online education interaction method based on big data, characterized in that: The method comprises: The learning images of students are captured in real time through the camera of the online education platform and pre-processed to obtain pre-processed images; Use convolutional neural networks to extract key features from preprocessed images, including text, symbols, and student actions; According to the key characteristics, identify the students' learning status and assign corresponding values ​​to each learning status; Obtain students’ real-time learning data; Calculate dynamic factors based on real-time learning data, which reflect students' real-time learning performance and changes; Calculate the fusion factor based on the learning state and dynamic factor; Compare the fusion factors with the content in the pre-digitized educational resource library to obtain the final learning content that meets the students' current learning status and needs; Recommend learning content to students in real time so that students can learn based on the learning content and complete the interaction.

2. The online education interaction method based on big data according to claim 1 is characterized in that: The convolutional neural network is used to extract key features from the preprocessed images. The key features include text, symbols, and student actions, including: Perform multi-scale feature extraction on the preprocessed image to obtain extraction results; According to the extraction results, a convolutional neural network is used to generate text candidate boxes, and non-maximum suppression is performed on the candidate boxes to obtain the final text detection result. The final text detection result is a set of rectangular box coordinates, which indicates the location of the text area in the image. For each detected text region, a convolutional recurrent neural network is used to perform text analysis to obtain the text content in the preprocessed image; According to the text content in the preprocessed image, the symbols in the preprocessed image are detected and located through the object detection convolutional neural network model to extract the symbol type, position and size features; Detect the key points of the students’ bodies in the video through the human posture estimation model based on convolutional neural networks; The human posture estimation model outputs the coordinate position of each key point. The coordinate position marks the specific position of each part of the student's body. Based on the coordinate position of the key point, the student's overall body posture is constructed; By analyzing the spatial position of each key point in a single frame, the student's instantaneous posture can be obtained, and the student's action can be identified through the position changes of the key points in consecutive frames.

3. The online education interaction method based on big data according to claim 2 is characterized in that: Based on key characteristics, identify students' learning status and assign corresponding values ​​to each learning status, including: Analyze students' concentration by analyzing their head position, eye gaze, and body inclination; Identify students’ learning activity by analyzing their interaction patterns with educational platforms; According to the students' concentration and learning activity, a classifier is used to identify each student's learning status. The classifier classifies each student into a predefined learning status and assigns a corresponding numerical value based on the students' concentration and learning activity characteristics.

4. The online education interaction method based on big data according to claim 3 is characterized in that: Based on real-time learning data, dynamic factors are calculated, including: By DF = w1 × (β × F 1c +(1-β)×F 1p +α×VI)+w2×F2+w3×(β×F 3c +(1-β)×F 3p ) Calculate the dynamic factor DF, where w1, w2, w3 are weights, F 1c and F 1p are the concentration values ​​of the current cycle and the previous cycle, respectively, and F 3c and F 3p are the interaction frequency values ​​of the current cycle and the previous cycle respectively, β is the smoothing factor of EMA, 0<β<1, α is the coefficient for adjusting the impact of concentration volatility, and F2 is the learning progress rate of the current cycle.

5. The online education interaction method based on big data according to claim 4 is characterized in that: Based on the learning state and dynamic factors, the fusion factor is calculated, including: The fusion factor CF is calculated by CF = λ×[θ×(α1×I+β1×TC)+(1-θ)×(γ×PE-δ×NE)]+(1-λ)×DF, where CF represents the fusion factor, λ represents the adjustment parameter, θ represents the adjustment parameter, α1 represents the weight coefficient, β1 represents the weight coefficient, I represents the number of interactions, which refers to the frequency of students' interactions in class, TC represents the number of completed tasks, which refers to the number of learning tasks completed by students, γ represents the weight coefficient, δ represents the weight coefficient, PE represents the number of positive emotion expressions, which refers to the frequency of students expressing positive emotions, and NE represents the number of negative emotion expressions, which refers to the frequency of students expressing negative emotions.

6. The online education interaction method based on big data according to claim 5 is characterized in that: Compare the fusion factors with the content in the pre-digitized educational resource library to obtain the final learning content that meets the students' current learning status and needs, including: Determine the characteristic dimensions of the fusion factor CF, which include interest areas and emotional states; Mapping the interest field value and the emotional state value of the fusion factor CF with the corresponding labels of the educational resources to obtain a mapping result; According to the mapping results, the difference between the student's fusion factor CF and each educational resource label is calculated to obtain the matching calculation result; Based on the matching calculation results, determine the final learning content that meets the students' current learning status and needs.

7. The online education interaction method based on big data according to claim 6 is characterized in that: The key points include the head, shoulders, wrists and knees.

8. An online education interactive system based on big data, characterized in that: include: An acquisition module is used to capture the student's learning image in real time through a camera of the online education platform and perform preprocessing to obtain a preprocessed image; Use convolutional neural networks to extract key features from preprocessed images, including text, symbols, and student actions; identify students' learning status based on key features, and assign corresponding values ​​to each learning status; obtain students' real-time learning data; The processing module is used to calculate the dynamic factor based on the real-time learning data, and the dynamic factor reflects the real-time learning performance and changes of the students; and calculate the fusion factor based on the learning status and the dynamic factor; The fusion factor is compared with the content in the pre-digitized educational resource library to obtain the final learning content that meets the students' current learning status and needs; the learning content is recommended to students in real time so that students can learn according to the learning content and complete the interaction.

9. A computing device, characterized in that include: one or more processors; A storage device, used for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.