Method for analyzing attention of learning object in educational metaverse, terminal and device
Patent Information
- Application Number
- CN202211436017.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-11-16
Smart Images

Figure CN115933930B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of teaching application of meta universe, and in particular to a learning object attention analysis method, terminal and device in an educational meta universe. BACKGROUND
[0002] With the continuous advancement of education digital transformation, emerging technologies such as artificial intelligence, big data, virtual reality, and learning analysis are gradually deepening in teaching applications, in order to improve the application level of teaching resources and analyze the attention of learners in online teaching processes. In recent years, meta universe has received widespread attention. As a vertical application of meta universe in the education field, educational meta universe has emerged in a series of applications in the fields of contextualized teaching, personalized learning, gamified learning, and teacher training. Automatically analyzing the learning effects of learners in such application scenarios can more intelligently perceive learner attention and more accurately identify learner emotions.
[0003] However, many existing learning effect analyses in the education field still use questionnaire surveys and manual interviews to evaluate learner learning outcomes, willingness to use, participation, and technology acceptance. Compared to data-driven analysis methods, the accuracy is not high enough, and the evidence is weak, making it difficult to serve as a convincing analysis proof. Therefore, introducing computer vision technology to analyze learner attention in educational meta universe can achieve intelligent perception and accurate analysis of learner attention, thereby improving the teaching design, teaching interaction, and teaching mode of educational meta universe and providing important support for the application of educational meta universe in various teaching scenarios.
[0004] Current learner attention analysis systems in educational meta universe still have many problems: (1) high cost of hardware devices. Using professional wearable devices and eye trackers to analyze learner attention, although the analysis results are accurate, the hardware devices are expensive, and the deployment and learning costs are high, making it difficult to be widely applied in teaching scenarios where learners are located in different places; (2) not closely integrated with teaching scenarios. Existing attention analysis systems often focus on learner eye movement frequency and gaze point position, making it difficult to provide accurate guidance for optimizing and improving educational meta universe teaching scenarios; (3) single analysis index. Most existing attention analysis systems use a single eye movement index to estimate learner attention, resulting in single attention analysis evidence and inability to generate more comprehensive attention analysis results. SUMMARY
[0005] The technical problem to be solved by the present application is to provide a learning object attention analysis method, terminal and device in an educational meta universe, which can comprehensively and accurately analyze the attention of learning objects in an educational meta universe.
[0006] To solve the above technical problems, one technical solution adopted by the present application is:
[0007] An attention analysis method of a learning object in an educational metaverse, comprising the steps of:
[0008] S1, generating a set of visible objects corresponding to the learning object according to a virtual scene of an educational metaverse space, and determining geometric-attribute information of each visible object in the set of visible objects;
[0009] S2, obtaining a learning video of the learning object, and determining a line-of-sight direction and a micro-expression of the learning object according to the learning video;
[0010] S3, determining a target visible object gazed at by the learning object according to the geometric-attribute information of each visible object and the line-of-sight direction, and associating the target visible object with the micro-expression;
[0011] S4, analyzing the attention of the learning object according to the target visible object and the micro-expression associated therewith.
[0012] To solve the above technical problems, another technical solution adopted by the present application is:
[0013] An attention analysis terminal of a learning object in an educational metaverse, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements each step of the above-mentioned attention analysis method of a learning object in an educational metaverse when executing the computer program.
[0014] To solve the above technical problems, another technical solution adopted by the present application is:
[0015] An attention analysis device of a learning object in an educational metaverse, comprising:
[0016] A virtual teaching resource organization module for generating a set of visible objects corresponding to the learning object according to a virtual scene of an educational metaverse space, and determining geometric-attribute information of each visible object in the set of visible objects;
[0017] A line-of-sight direction and micro-expression determination module for obtaining a learning video of the learning object, and determining a line-of-sight direction and a micro-expression of the learning object according to the learning video;
[0018] A gaze model association module for determining a target visible object gazed at by the learning object according to the geometric-attribute information of each visible object and the line-of-sight direction, and associating the target visible object with the micro-expression;
[0019] An attention analysis module for analyzing the attention of the learning object according to the target visible object and the micro-expression associated therewith.
[0020] The beneficial effects of the present application are that: a visual object set corresponding to a learning object is generated according to a virtual scene of an education meta-universe space; a learning video of the learning object is obtained, and a line-of-sight direction and a micro-expression of the learning object are determined according to the learning video; a target visual object that the learning object gazes at is determined according to the line-of-sight direction, and the target visual object is associated with the micro-expression; attention of the learning object is analyzed according to the target visual object and the micro-expression associated therewith; without the need for expensive hardware equipment, the target visual object that the learning object gazes at is determined through the line-of-sight direction, not only the gaze point position of the learning object is focused on, but also the virtual object of the teaching scene corresponding to the gaze point position is focused on, the close integration of the learning object and the teaching scene is realized, the relevance between the virtual and real spaces, the plane and the space, the attention of the learning object and the teaching model in the education meta-universe is improved, and the target visual object that the line-of-sight direction gazes at is associated with the corresponding micro-expression, the attention of the learning object is analyzed according to the target visual object and the micro-expression associated therewith, and the obtained attention analysis result is more comprehensive and accurate. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 A step flow chart of a learning object attention analysis method in an education meta-universe according to an embodiment of the present application;
[0022] Figure 2 A structural schematic diagram of a learning object attention analysis terminal in an education meta-universe according to an embodiment of the present application;
[0023] Figure 3 A structural schematic diagram of a learning object attention analysis device in an education meta-universe according to an embodiment of the present application;
[0024] Figure 4 A schematic diagram of mapping a triangular patch of a model surface in a virtual scene to an octree according to an embodiment of the present application;
[0025] Figure 5 A schematic diagram of displaying an image of a meta-universe scene on a terminal screen according to an embodiment of the present application;
[0026] Figure 6 A schematic diagram of collecting a learning video of a learning object through a camera of a display terminal according to an embodiment of the present application;
[0027] Figure 7 A schematic diagram of three posture deflection angle parameters in head posture positioning according to an embodiment of the present application;
[0028] Figure 8 A schematic diagram of terminal screen interest region division according to an embodiment of the present application;
[0029] Figure 9A schematic diagram of a terminal screen superimposed with a learning object attention heat map according to an embodiment of the present application;
[0030] Figure 10 A structure schematic diagram of a learning object attention analysis device in an educational meta-universe according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] To explain the technical content, achieved purposes and effects of the present application in detail, the following will be described in combination with embodiments and the accompanying drawings.
[0032] The above-mentioned learning object attention analysis method, terminal and device in an educational meta-universe according to the present application can be applied to the analysis of the attention of learning objects in an educational meta-universe scene, which will be described in detail through specific embodiments as follows:
[0033] In an optional embodiment, referring to Figure 1 A learning object attention analysis method in an educational meta-universe, comprising the steps of:
[0034] S1, generating a set of visible objects corresponding to a learning object according to a virtual scene of an educational meta-universe space, and determining the geometric-attribute information of each visible object in the set of visible objects;
[0035] Specifically, S11, subdivide the virtual scene of the educational meta-universe space to generate a set of model objects;
[0036] Among them, the virtual scene of the meta-universe space can be subdivided using an octree to organize and generate the set of model objects;
[0037] S12, determine the visible range of the view frustum corresponding to the learning object according to the viewpoint position, direction and target point position of the learning object, determine the intersection of the visible range and the set of model objects, and generate the set of visible objects corresponding to the learning object according to the intersection;
[0038] Among them, the set of visible objects corresponding to the learning object can be obtained by traversing the set of model objects contained in or intersecting with the view frustum, and rendering the image of the meta-universe scene in the graphics rendering pipeline;
[0039] S13, determine the visible object to which the gaze position on the terminal screen belongs according to the corresponding relationship between the gaze position and the coordinates of the meta-universe space, establish the corresponding relationship between the gaze position and the visible object to which it belongs, and determine the geometric-attribute information of each visible object in the set of visible objects according to the corresponding relationship, i.e. according to the gaze position of the learning object gaze point, realize the geometric-attribute query of the visible object in the meta-universe;
[0040] Step S1 is mainly to organize virtual teaching resources, in an optional embodiment, including the following steps:
[0041] (1) Generation of model object set: adopt triangle patches to represent the boundary surface of the geometric model of teaching objects (such as teaching aids, teaching components, experimental instruments, etc.) in the scene, use octree to subdivide the virtual scene of the meta-universe space, and according to the position of the barycentric coordinates, as shown in Figure 4 , map the triangle patches of the model surface in the scene to the octree, organize and generate the model object set of each scene;
[0042] (2) Space-to-screen mapping: determine the visible range of the view frustum according to the viewpoint position, direction and target point position of the learning object, traverse the model object set containing or intersecting with the view frustum, perform model, view, projection and viewport transformation in the graphics rendering pipeline, render the image of the meta-universe scene, and generate the viewable object set corresponding to the learning object, as shown in Figure 5 displayed on the terminal screen.
[0043] (3) Object picking: obtain the position of the learning object's line of sight on the terminal screen, calculate the spatial coordinates of the point in the meta-universe according to the display depth and template parameter settings, determine the triangle patch, geometric model and virtual object (virtual object is viewable object) to which the point belongs by using octree traversal method, and realize the query of virtual object geometry-attribute under the support of object-relationship database, that is, according to the gaze position of the learning object on the terminal screen, convert the gaze position into the spatial coordinates of the corresponding meta-universe, determine the triangle patch to which the spatial coordinates belong by traversing the octree, and the object-relationship database stores the corresponding relationship between the geometric model and attribute data of the virtual object, therefore, after determining the triangle patch to which the spatial coordinates belong, the virtual object to which the spatial coordinates belong can be determined by querying the geometric model of the object-relationship database, that is, the virtual object to which the gaze position of the learning object on the terminal screen belongs is determined, so that the corresponding relationship between the gaze position and the viewable object to which it belongs is established, and the corresponding relationship can be used to query the attribute information (such as object name, category, purpose, etc.) of each viewable object in the viewable object set, therefore, according to the gaze position of the gaze point of the learning object, the geometry-attribute query of the viewable object in the meta-universe can be realized, that is, according to the gaze position of the learning object and the geometry-attribute information of each viewable object in the viewable object set, the viewable object to which the gaze position belongs can be determined.
[0044] The algorithm steps for determining the spatial coordinates of the meta-universe corresponding to the gaze position are as follows:
[0045] (3-1): Convert the coordinates of the learning object's line of sight falling on the terminal screen into projection window coordinates proj x , projy and proj z Formula 1 is a conversion calculation formula:
[0046]
[0047] wherein sw and sh are the width and height of the terminal screen respectively, and (x, y) is the coordinate of the learning object's line of sight falling on the terminal screen;
[0048] (3-2): using the projection transformation matrix shown in formula 2 to convert the projection window coordinates into spatial coordinates, formula 3 is the calculation of each variable value in the projection transformation matrix from the view volume parameters, and formula 4 is the spatial point coordinates x ’ , y ’ and z ’ converted and calculated:
[0049]
[0050]
[0051]
[0052] wherein w and h are the width and height of the near clipping plane respectively, and Z n and Z f are the distances of the near clipping plane and the far clipping plane in the view volume respectively, Q is the ratio of the distance from the viewpoint to the far clipping plane to the distance from the near clipping plane to the far clipping plane in the view volume, and fov is the field of view angle.
[0053] S2, acquiring a learning video of the learning object, determining the line of sight direction and micro-expression of the learning object according to the learning video;
[0054] This step is mainly to determine the line of sight direction of the learning object and identify the micro-expression of the learning object by means of the learning video;
[0055] For determining the line of sight direction of the learning object: the learning video of the learning object is acquired through interactive data acquisition, specifically, the learning video of the learning object is collected by using a camera and is displayed synchronously in the video frame at the top left corner of the terminal; the head posture deflection angle is extracted by using a backbone network and three branch networks, and the Euler rotation angle is calculated; the four Purkinje spots coordinates are obtained by using a labeling connectivity analysis algorithm, a polynomial equation is constructed, and the line of sight direction of the learner is estimated;
[0056] In an optional embodiment, the determining the line of sight direction of the learning object according to the learning video comprises:
[0057] locating the head region of the learning object according to the learning video;
[0058] extract a head pose feature based on the head region, calculate a head Euler rotation angle according to the head pose feature;
[0059] determine an eye feature point distribution from the head region according to the head Euler rotation angle;
[0060] use a pupil center extraction algorithm to obtain a pupil center image and a corresponding coordinate in the learning video according to the eye feature point distribution, and estimate a line of sight direction of the learning object according to the pupil center image and the corresponding coordinate.
[0061] Specifically, the method comprises the following steps:
[0062] (1) Learning video data acquisition: as shown in Figure 6 , a high-definition video of the learning object participating in learning is collected using a built-in camera or an external camera in the display terminal, and the video image is displayed in the upper left corner region of the display terminal screen, if the face of the learning object is not detected, the user is reminded to turn on the camera or prompted to display the face image in the upper left corner video frame, Figure 6 , wherein 601 represents the gaze direction of the eye of the learning object to the terminal screen, and 602 represents the main axis of the camera collected video;
[0063] (2) Head pose positioning: the video image of the learning object is detected in real time using a target detection and tracking algorithm, the head region of the learning object in the image is positioned, the head pose feature is extracted using a backbone network, and three branch networks are used to obtain X, Y and Z axes from the head pose feature, respectively, as shown in Figure 7 , three kinds of attitude deflection angle parameters, the head Euler rotation angle is calculated, in the figure, 701 represents a head pose action sequence rotating around the X axis, 702 represents a head pose action sequence rotating around the Y axis, and 703 represents a head pose action sequence rotating around the Z axis;
[0064] The head pose positioning algorithm comprises the following steps:
[0065] (2-1): obtaining a high-definition video sequence of the learning object participating in learning collected by the camera;
[0066] (2-2): using a convolutional neural network, an RFB module and a down-sampling layer, stacking in turn, constructing a learner head detection network model, the input is a video sequence frame, and the output is the head region of the learning object obtained;
[0067] (2-3): using a convolutional neural network layer, a maximum pooling layer, a rectified linear unit, a Dropout layer and a full connection layer to construct a backbone network, using the backbone network to extract the head region feature of the learning object, and the output is the head pose feature of the learning object;
[0068] (2-4): Using three branch networks to estimate three kinds of posture deflection angles along X, Y and Z axes respectively from the head posture features of the learning object, and constructing three kinds of posture rotation matrices according to the deflection angles as shown in formulas 5, 6 and 7:
[0069]
[0070]
[0071]
[0072] wherein θ x , θ y and θ z are the deflection angles of the three kinds of postures respectively;
[0073] (2-5): According to the rotation along X, Y and Z axes in turn, formula 8 is the calculation formula of Euler rotation matrix:
[0074]
[0075] Let the Euler rotation matrix be formula 9:
[0076]
[0077] (2-6): The calculation of the head Euler rotation angle of the learning object is shown in formulas 10, 11 and 12:
[0078] θ x = atan2(-r yz , r zz ) (formula 10)
[0079]
[0080] θ z = atan2(-r xy , r xx ) (formula 12)
[0081] wherein atan2 is the calculation azimuth angle function;
[0082] (3) Eye line of sight tracking: using a face recognition algorithm based on geometric features to obtain the eye feature points of the learning object, then according to the distribution of the eye feature points, using a pupil center extraction algorithm to obtain the pupil center image and coordinates in the video image, using a labeled connectivity analysis algorithm to obtain the coordinates of the four Purkinje spots, using a queue to construct a coordinate series, according to the relative relationship between the current pupil center coordinates and the coordinates of the previous frame in the queue, constructing a cubic nonlinear equation to estimate the line of sight direction of the learning object;
[0083] For micro-expression recognition of the learning object, mainly according to the changes of the face eyes, corners of the mouth and eyebrows, a label category is added to the micro-expression picture set of the learning object; a convolutional neural network is used to extract the face features of the input image, and after multi-step processing, the micro-expression of the learner is located on the output image; the category of the micro-expression is recognized, a timestamp and course information are added, and the cloud server is uploaded;
[0084] In another optional implementation, the determining the micro-expression of the learning object according to the learning video comprises:
[0085] Collecting micro-expression pictures of each learning object in an online teaching process, and adding a corresponding micro-expression label category to each micro-expression picture to generate a micro-expression image annotation set;
[0086] Intercepting images of the learning video at a preset frequency, using a convolutional neural network to extract face feature key points of the learning object in the images, and locating a target micro-expression of the learning object according to the face feature key points;
[0087] Matching the target micro-expression in the micro-expression image annotation set to determine the category corresponding to the target micro-expression;
[0088] Specifically, the micro-expression recognition comprises the following steps:
[0089] (1) Micro-expression image annotation: collecting a micro-expression picture set of numerous learning objects in an online teaching process, adding five micro-expression label categories of coldness, disgust, depression, surprise and happiness to the face image according to the change amplitudes of the face eyes, corners of the mouth and eyebrows of the learning object in the image, and generating a micro-expression image annotation set;
[0090] (2) Micro-expression positioning: intercepting video sequence images at a frequency of one frame per second, using a convolutional neural network to extract face feature key points of the learning object in the images, and then inputting the face feature key points as input parameters to sequentially pass through a full connection layer and a normalization exponential function for non-linear transformation and compression of the features, and locating the micro-expression of the learning object on the output image;
[0091] (3) Micro-expression recognition: using a feature matching algorithm to realize the matching of the micro-expression image of the learning object and the annotation category, recognizing the micro-expression of the learning object, determining the micro-expression category of the current learning object, adding a timestamp and course information according to the coding category of the five micro-expression categories of coldness, disgust, depression, surprise or happiness to which the micro-expression belongs, and uploading to the cloud server;
[0092] Wherein, the specific steps of the feature matching algorithm are as follows:
[0093] (3-1): using the extreme point detection algorithm shown in formula 13 to extract the extreme points of the micro-expression of the learning object in the image:
[0094] L(x, y, s) = G(x, y, s) * I(x, y) (Equation 13)
[0095] where (x, y) is the pixel coordinate, s is the scale space factor, I(x, y) and G(x, y, s) are the original image and Gaussian function respectively, and * is the convolution operator;
[0096] (3-2): Construct the Gaussian difference pyramid shown in Equation 14 to obtain stable learning object micro-expression key points:
[0097] D(x, y, s) = L(x, y, k s) - L(x, y, s) (Equation 14)
[0098] where k is the ratio of adjacent scale factors;
[0099] (3-3): Use the scale space Taylor series expansion of the Gaussian difference pyramid function as shown in Equation 15 to interpolate the micro-expression key points and remove key points with low contrast:
[0100]
[0101] where X = (x, y, s) T ;
[0102] (3-4): Use the gradient direction features of the key point neighborhood pixels to specify the direction parameters for each key point, and Equations 16 and 17 calculate the key point gradient value and the corresponding direction:
[0103]
[0104]
[0105] (3-5): Select a square pixel region of 16x16 grid around the key point as the center, divide it into 4 sub-regions according to the 4x4 rule, calculate the gradient accumulation value of each sub-region in 8 directions, and convert each micro-expression key point of the learning object into a 4x4x8 = 128-dimensional feature descriptor;
[0106] (3-6): Use the Euclidean distance calculation formula shown in Equation 18 to obtain the distance between the micro-expression image of the learning object and the labeled category image:
[0107]
[0108] where M and N are the descriptor vectors of the images;
[0109] Select the labeled category corresponding to the minimum distance d(x, y) as the micro-expression category of the learning object;
[0110] S3, determining a target visual object that the learning object gazes at according to the geometric-attribute information of each visual object and the gaze direction, and associating the target visual object with the micro-expression;
[0111] In an optional embodiment, determining the target visual object that the learning object gazes at according to the geometric-attribute information of each visual object and the gaze direction comprises:
[0112] determining a target gaze area of the learning object according to the gaze direction, and determining the target visual object that the learning object gazes at according to the target gaze area and the geometric-attribute information of each visual object;
[0113] In another optional embodiment, determining the target visual object that the learning object gazes at according to the geometric-attribute information of each visual object and the gaze direction comprises:
[0114] dividing a terminal screen into a plurality of interest areas, and determining a corresponding target interest area according to the gaze direction;
[0115] determining three-dimensional coordinates of each corner point of the target interest area in a meta-universe space, and generating a minimum circumscribed cuboid;
[0116] extracting an associated visual object intersecting or contained in the minimum circumscribed cuboid according to the geometric-attribute information of each visual object, and generating an associated visual object set;
[0117] eliminating an associated visual object that is occluded in the associated visual object set, and generating a target visual object;
[0118] In an optional embodiment, after determining the corresponding target interest area according to the gaze direction, the method further comprises the step of:
[0119] counting a frequency of the gaze of the learning object falling in each interest area and a frequency of a corresponding micro-expression in a preset time period;
[0120] The counting the frequency of the gaze of the learning object falling in each interest area in the preset time period comprises:
[0121] calculating a maximum angle of deviation of three head posture of the learning object, i.e., nodding, shaking head and tilting head, from the terminal screen according to a length and a width of the terminal screen and eye center coordinates of the learning object;
[0122] judging whether an angle of single head deflection of the learning object falls in a range of the maximum angle and a duration is greater than a preset second, if yes, counting a frequency of an interest area where the gaze of the learning object falls, and if not, not counting;
[0123] In actual application scenarios, the display terminal screen is divided into several rectangular interest regions, which can be numbered in sequence from left to right and from top to bottom; the gaze point coordinates of the learning object are calculated according to the gaze direction and the gaze point, and the position of the learning object in the interest region is located; the sampling interval is set, and the frequency of the learner's gaze falling in each interest region in the time period is counted, thereby realizing the positioning of the learning object's attention;
[0124] Specifically, the steps include:
[0125] (1) Interest region division: obtain the length and width values of the learning object terminal screen, set a length threshold, divide the entire screen into several rectangular interest regions, as shown in FIG. 1, number the interest regions in sequence from left to right and from top to bottom, and superimpose a transparent layer on the screen display image to represent the frequency of the learner's attention to each interest region after starting the attention analysis function; Figure 8
[0126] (2) Gaze positioning: based on the internal and external orientation elements of the camera, a local coordinate system of the head of the learning object is constructed, and the position and direction of the display terminal in the coordinate system are determined; according to the gaze direction and the gaze point of the learning object, a polynomial nonlinear regression model is used to calculate the gaze point coordinates of the learner, and the interest region where the learner is located is estimated;
[0127] The polynomial nonlinear regression model shown in formula 19 is used to calculate the gaze point coordinates of the learner:
[0128]
[0129] where a0, a1, …, a 11 and b0, b1, …, b 11 are unknown numbers of the model, (x, y) is the eye center coordinates of the learner, (x e , y e , z e ) is the three orientations of the Euler transformation of the learner's head;
[0130] (3) Attention acquisition: set the sampling interval, count the frequency of the learning object's gaze falling in each interest region in the time period, use the start and end time, interest region number and gaze frequency as attributes, respectively generate the corresponding attribute values of each sampling interval, record the above results in the JSON file in time sequence, and upload to the cloud server;
[0131] Wherein, the steps of counting the frequency of the learning object's gaze falling in each interest region:
[0132] (3-1): According to the length and width of the terminal screen and the eye center coordinates of the learning object, formulas 20, 21 and 22 are used to calculate the maximum angle of the three head posture deviations from the terminal screen, i.e., nodding, shaking head and shaking head:
[0133]
[0134]
[0135]
[0136] where h and d are the length and width of the terminal screen, respectively, and x, y and z are the eye center coordinates of the learner;
[0137] (3-2): If the single head deflection angle range of the learning object is (-a, a), (-b, b) and (-g, g), and the duration exceeds the preset seconds, such as 2 seconds, the learning object's line of sight is recorded once in the interest area, otherwise it is not counted.
[0138] After determining the line of sight direction and micro-expression of the learning object, the relevant data can be sent to the cloud for unified data management, specifically:
[0139] Cloud data management: the cloud receives and analyzes attention and micro-expression recognition data, classifies them according to course and timestamp information; uses data mapping to count the number of attention times in each interest area and associate expression recognition information; according to the client request, the cloud compresses, encodes and transmits the course learning object's attention times and associated micro-expression information to the client, including the following steps:
[0140] (1) Data classification: when the cloud server listens to the attention or micro-expression recognition data received from the learning object terminal, it uses the JSON file specification to parse the file attributes and values, classifies them according to course and timestamp information, and inserts the time interval, interest area number, attention times or micro-expression category attributes and their values into the database table;
[0141] (2) Data statistics: use database query command to extract all learning object uploaded attention and micro-expression information in a certain period of time from the cloud database, use data mapping to count the number of attention times in each interest area, obtain the attention times of all participants in the course, and associate the expression recognition information;
[0142] (3) Data distribution: according to the learning object client's request for attention, the cloud server real-time statistics of the current time interval of each participant's attention focus and associated micro-expression information, using LoRa as the data transmission protocol of the attention analysis system, the statistical results are compressed, encoded and sent from the cloud server to the request terminal.
[0143] In determining the gaze direction of the learning object, the gaze model associated with the gaze direction, i.e., the visible model, can be determined. The inverse projection transformation method can be used to obtain the three-dimensional coordinates of each corner point of the interest region in the meta-universe space to generate the minimum circumscribed cuboid. The barycentric judgment method is used to extract the associated target teaching object set intersecting or contained therein. According to the object picking method in step S1(3), the occluded object model is removed to obtain the associated gaze model. Specifically:
[0144] (1) Interest region plane to space mapping: According to the number of the display terminal where the interest region is located, the horizontal and vertical coordinate values of the four corners of the rectangle are calculated. The inverse projection transformation method is used to obtain the spatial three-dimensional coordinates of each corner point screen coordinate in the meta-universe, and the maximum and minimum XYZ values are obtained by traversal to determine the minimum circumscribed cuboid.
[0145] (2) Determination of interest region associated teaching objects: According to the circumscribed cuboid range corresponding to the interest region, traverse the virtual object surface model organized by octree, and use the barycentric judgment method to judge whether the surface triangle patch contains or intersects with the circumscribed cuboid. The object model containing or intersecting is obtained to generate the associated target teaching object set.
[0146] (3) Gaze model association: According to the object picking method in step S1(3), query the model object associated with all points in the learner's gaze area, add a flag to the result object in the target teaching object set, and remove the associated objects that are occluded and do not have the flag. The remaining model is the object model that the learner gazes at, i.e., the target visible object.
[0147] S4, analyzing the attention of the learning object according to the target visible object and the micro-expression associated therewith;
[0148] Specifically:
[0149] According to the micro-expression associated with the target visible object and the frequency of the micro-expression in the preset time period, the concentration effect of the learning object is evaluated;
[0150] According to the frequency of the learning object's gaze falling on each interest region in the preset time period and the target visible object gazed at by the learning object, the frequency of the target visible object gazed at by the learning object in the preset time period is determined.
[0151] According to the frequency of the target visible object gazed at by the learning object in the preset time period, a heat map is generated, and the transparency of the heat map layer is set according to the frequency order, and the heat map layer is superimposed on the corresponding target visible object in the terminal screen.
[0152] In a specific application scenario, in the attention analysis process, the learning object is given a weight coefficient of three types of emotions: positive, neutral and negative, the interest calculation formula is used to estimate the concentration effect, the column chart is used to represent the emotions of the learning object in the learning process, and the emotion time sequence diagram of the learning object is drawn; according to the attention statistical result of the learning object, the attention heat map is generated, and the transparency of the heat map layer is added by hierarchical setting:
[0153] (1) Learning concentration effect evaluation: According to the Ekman emotion classification standard, happiness and surprise are divided into positive and neutral emotions, and indifference, aversion and depression are divided into negative emotions. Weight coefficients are assigned to each type of emotion, the time of each type of emotion during the course is classified and counted, and the learner interest calculation formula is used to estimate the concentration effect;
[0154] The learning concentration effect evaluation step includes:
[0155] (1-1) According to the micro-expression of the learning object, the weight coefficients of happiness, surprise, indifference, aversion and depression are 1.0, 0.25, -0.25, -0.75 and -0.5 respectively;
[0156] (1-2) The learning interest of the learner is calculated by formula 23:
[0157]
[0158] Where 1.0 and 0.25 are the weight coefficients of happiness and surprise, t 高兴 and t 惊讶 are the total time of the learner's happiness and surprise, and T is the total time of the learner participating in the course learning;
[0159] (1-3) The calculation of the concentration effect of the learner is shown in formula 24:
[0160]
[0161] Where t1, t2 and t3 are the total time of the positive, neutral and negative emotions of the learning object, -0.25, -0.75 and -0.5 are the weight coefficients of indifference, aversion and depression, t3, t5 and t6 are the total time of the learning object's indifference, aversion and depression, and t7 is the total time of the learning object's gaze positioning recorded as invalid;
[0162] (2) Learning object emotion time sequence diagram drawing: download the micro-expression category code of a single learning object or all participating learning objects in the course time period from the cloud server, use the course time sequence and the frequency of each period expression category as the horizontal and vertical axes of the Cartesian coordinate system, and use the column chart to represent the emotions of the learning object in the learning process;
[0163] (3) Pay attention to the generation of heat map: statistics of each time period learning object gaze point fall in each interest area frequency, according to the learning object attention frequency generation attention heat map, and according to the frequency from high to low value layer setting heat map layer transparency, the layer is superimposed to the screen image of the attention object model of the meta universe display terminal, such as Figure 9 As shown.
[0164] In another optional implementation, please refer to Figure 2 An attention analysis terminal for learning objects in an educational meta universe, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements each step of the attention analysis method for learning objects in an educational meta universe described in the above embodiments when executing the computer program.
[0165] In another optional implementation, please refer to Figure 3 An attention analysis device for learning objects in an educational meta universe, corresponding to the attention analysis method for learning objects in an educational meta universe described above, comprising:
[0166] A virtual teaching resource organization module for generating a set of visual objects corresponding to learning objects according to the virtual scene of the educational meta universe space, and determining the geometric-attribute information of each visual object in the set of visual objects;
[0167] A line of sight direction and micro-expression determination module for obtaining a learning video of the learning object, and determining the line of sight direction and micro-expression of the learning object according to the learning video;
[0168] A gaze model association module for determining the target visual object of the learning object's gaze according to the geometric-attribute information of each visual object and the line of sight direction, and associating the target visual object with the micro-expression;
[0169] An attention analysis module for analyzing the attention of the learning object according to the target visual object and its associated micro-expression;
[0170] In another optional implementation, as Figure 10 shown, the attention analysis device for learning objects in an educational meta universe can be further subdivided, comprising a virtual teaching resource organization module, an interactive data acquisition module, a micro-expression recognition module, an attention positioning module, a cloud data management module, a gaze model association module, and an attention analysis module;
[0171] The virtual teaching resource organization module is used to implement step S1 of the attention analysis method for learning objects in an educational meta universe.
[0172] The interaction data collection module is configured to implement the step of determining the line of sight of the learning object according to the learning video in the learning object attention analysis method in the educational meta-universe.
[0173] The micro-expression recognition module is configured to implement the step of recognizing the micro-expression of the learning object in the learning object attention analysis method in the educational meta-universe.
[0174] The attention positioning module is configured to implement the step of positioning the attention of the learning object in the learning object attention analysis method in the educational meta-universe.
[0175] The cloud data management module is configured to implement the step of realizing cloud data management in the learning object attention analysis method in the educational meta-universe.
[0176] The gaze model association module is configured to implement the step of obtaining an associated gaze model in the learning object attention analysis method in the educational meta-universe.
[0177] The attention analysis module is configured to implement step S4 in the learning object attention analysis method in the educational meta-universe.
[0178] In summary, the application provides a method for analyzing the attention of learning objects in an educational meta-universe, a terminal and a device. The virtual scene of the meta-universe space is divided using an octree, and a set of visible objects corresponding to the learning objects is generated according to the virtual scene of the educational meta-universe space. The learning video of the learning object is collected, the head posture deflection angle is extracted using a neural network, the Poulchin spot coordinates are obtained using an analysis algorithm, the line of sight direction of the learner is estimated, the picture set of the online learning object is collected, the micro-expression category is labeled according to the facial feature changes, the facial features are extracted using a neural network, the micro-expression is located, and the micro-expression category is identified. According to the line of sight direction, the target visible object that the learning object gazes at is determined, the three-dimensional coordinates of each corner point of the interest area in the meta-universe space are obtained using an inverse projection transformation method, a minimum circumscribed cuboid is generated, the associated target teaching object set intersecting or contained in the center of gravity is extracted using a center of gravity judgment method, the associated object model that is blocked is removed, the target visible object is obtained, and the target visible object is associated with the micro-expression. According to the target visible object and the micro-expression associated therewith, the attention of the learning object is analyzed, the column chart is used to represent the emotions of each category in the learning process, the emotion time sequence diagram of the learning object is drawn, the attention heat map is generated according to the attention statistical result, and the transparency of the heat map layer is added by layering. Without the need for high-priced hardware equipment, the target visible object that the learning object gazes at is determined according to the line of sight direction, not only the gaze point position of the learning object is focused on, but also the virtual object of the teaching scene corresponding to the gaze point position is focused on, the close integration of the learning object and the teaching scene is realized, the correlation between the virtual and real spaces, the plane and the space, the attention of the learning object and the teaching model in the educational meta-universe is improved, the target visible object that the line of sight direction gazes at is associated with the corresponding micro-expression, the attention of the learning object is analyzed according to the target visible object and the micro-expression associated therewith, and the obtained attention analysis result is more comprehensive and accurate. With the rapid application and integration of the meta-universe in the education field, the learner attention analysis technology based on computer vision and data driving has the non-invasive advantage and has a broad application prospect in the education field.
[0179] The above description is only an embodiment of the application, and does not limit the patent scope of the application. Any equivalent transformation or direct or indirect application in the related technical field based on the content of the specification and drawings is also included in the patent protection scope of the application.
Claims
1. A method of analyzing attention of a learning object in an educational metaverse, characterized by, The method comprises the steps of: S1, generating a visual object set corresponding to a learning object according to a virtual scene of an education meta-universe space, and determining geometric-attribute information of each visual object in the visual object set; S2, obtaining a learning video of the learning object, and determining a line-of-sight direction and a micro-expression of the learning object according to the learning video; S3, determining a target visual object gazed at by the learning object according to the geometric-attribute information of each visual object and the line-of-sight direction, and associating the target visual object with the micro-expression; S4, analyzing attention of the learning object according to the target visual object and the micro-expression associated therewith; The step S1 comprises: subdividing a virtual scene of an education meta-universe space to generate a model object set; determining a visual range of a viewing frustum corresponding to a learning object according to a viewpoint position, a direction and a target point position of the learning object, determining an intersection of the visual range and the model object set, and generating a visual object set corresponding to the learning object according to the intersection; determining a visual object to which a gaze position on a terminal screen belongs according to a corresponding relationship between the gaze position and a coordinate in a meta-universe space of a line-of-sight of the learning object, establishing a corresponding relationship between the gaze position and the visual object to which the gaze position belongs, and determining geometric-attribute information of each visual object in the visual object set according to the corresponding relationship; The step S3 comprises: determining a target gaze area of the learning object according to the line-of-sight direction, and determining a target visual object gazed at by the learning object according to the target gaze area and the geometric-attribute information of each visual object; dividing a terminal screen into a plurality of interest areas, and determining a corresponding target interest area according to the line-of-sight direction; determining three-dimensional coordinates of each corner point of the target interest area in a meta-universe space, and generating a minimum circumscribed cuboid; extracting an associated visual object intersecting or contained in the minimum circumscribed cuboid according to the geometric-attribute information of each visual object, and generating an associated visual object set; removing an associated visual object that is occluded in the associated visual object set, and generating a target visual object.
2. The method of claim 1, wherein the method is characterized by: The step of determining a visual object to which a gaze position on a terminal screen belongs according to a corresponding relationship between the gaze position and a coordinate in a meta-universe space of a line-of-sight of a learning object comprises: According to formula 1, the gaze position coordinates of the line of sight of the learning object on the terminal screen are converted into projection window coordinates proj x , proj y , and proj z : wherein sw and sh respectively represent a width and a height of the terminal screen, and (x, y) represents a gaze position coordinate of the line-of-sight of the learning object on the terminal screen; The projection transformation matrix according to formula 2 converts the projection window coordinates into the corresponding coordinates x, y and z of the meta-universe space through formula 3 and formula 4: ’ ’ ’ where w and h represent the width and height of the near clipping plane in the view frustum, respectively n and Z f represent the distance of the near clipping plane and the distance of the far clipping plane in the view frustum, respectively, Q represents the ratio of the distance from the gaze position to the far clipping plane to the distance of the near clipping plane to the far clipping plane in the view frustum, and fov represents the field of view angle; determining a coordinate x ’ , y ’ and z ’ of the meta-universe space according to the determined coordinate of the meta-universe space determining a visual object to which the coordinates of the meta-universe space belong.
3. The method of claim 1 or 2, wherein the method is characterized by: The step of determining a line-of-sight direction of a learning object according to a learning video comprises: locating a head region of the learning object according to the learning video; extracting a head posture feature based on the head region, and calculating a head Euler rotation angle according to the head posture feature; determining an eye feature point distribution from the head region according to the head Euler rotation angle; According to the eye feature point distribution, a pupil center extraction algorithm is used to obtain pupil center images and corresponding coordinates in the learning video, and a line of sight direction of the learning object is estimated according to the pupil center images and the corresponding coordinates.
4. The method of claim 1 to 2, wherein, The micro-expression of the learning object is determined according to the learning video, which includes: Micro-expression image annotation sets are generated by collecting micro-expression pictures of each learning object in an online teaching process and adding corresponding micro-expression label categories to each micro-expression picture. Images of the learning video are intercepted at a preset frequency, and a convolutional neural network is used to extract facial feature key points of the learning object in the images, and a target micro-expression of the learning object is located according to the facial feature key points. The target micro-expression is matched in the micro-expression image annotation set to determine the corresponding category of the target micro-expression. 5.The method of claim 1, wherein After determining the corresponding target interest region according to the line of sight direction, the following steps are further included: The frequency of the line of sight of the learning object falling on each interest region and the frequency of the corresponding micro-expression in a preset time period are counted. The step S4 includes: The concentration effect of the learning object is evaluated according to the micro-expression and the frequency of the micro-expression associated with the target visual object in the preset time period; The frequency of the target visual object gazed at by the learning object in the preset time period is determined according to the frequency of the line of sight of the learning object falling on each interest region in the preset time period and the target visual object gazed at by the learning object; A heat map is generated according to the frequency of the target visual object gazed at by the learning object in the preset time period, and the transparency of the heat map layer is set in a hierarchical order according to the frequency, and the heat map layer is superimposed on the corresponding target visual object in the terminal screen.
6. The method of claim 5, wherein the method further comprises: determining a degree of attention of the learning object based on the attention score. The frequency of the line of sight of the learning object falling on each interest region in a preset time period includes: The maximum angle of deviation of the three head postures of the learning object, i.e., nodding, shaking head and shaking head, from the terminal screen is calculated according to the length and width of the terminal screen and the eye center coordinates of the learning object; It is judged whether the angle of single head deflection of the learning object falls within the above maximum angle range and the duration is greater than a preset number of seconds, if yes, the frequency of the interest region where the line of sight of the learning object falls is counted, otherwise, it is not counted.
7. An attention analysis terminal for learning objects in an educational metaverse, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize each step in the attention analysis method of the learning object in the educational meta-universe as claimed in any one of claims 1 to 6. 8.An attention analysis device for a learning object in an educational meta universe, characterized by, It includes: A virtual teaching resource organization module is used to generate a set of visual objects corresponding to the learning object according to the virtual scene of the educational meta-universe space, and to determine the geometric-attribute information of each visual object in the set of visual objects; A line of sight direction and micro-expression determination module is used to obtain a learning video of the learning object, and to determine the line of sight direction and the micro-expression of the learning object according to the learning video; A gaze model association module is used to determine the target visual object gazed at by the learning object according to the geometric-attribute information of each visual object and the line of sight direction, and to associate the target visual object with the micro-expression. an attention analysis module configured to analyze attention of the learning object to the target visual object and its associated micro-expression; the virtual teaching resource organization module is further configured to: subdivide a virtual scene of the education meta-universe space to generate a model object set; determine a visual range of a corresponding view volume of the learning object according to a viewpoint position, a direction and a target point position of the learning object, determine an intersection of the visual range and the model object set, and generate a corresponding visual object set of the learning object according to the intersection; determine a visual object to which a gaze position of the learning object belongs according to a corresponding relationship between the gaze position on a terminal screen and a coordinate in the meta-universe space, establish a corresponding relationship between the gaze position and the visual object to which the gaze position belongs, determine geometric-attribute information of each visual object in the visual object set according to the corresponding relationship; the determination of the target visual object gazed at by the learning object according to the geometric-attribute information of each visual object and the gaze direction includes: determine a target gaze area of the learning object according to the gaze direction, and determine the target visual object gazed at by the learning object according to the target gaze area and the geometric-attribute information of each visual object; divide the terminal screen into a plurality of interest areas, and determine a corresponding target interest area according to the gaze direction; determine three-dimensional coordinates of each corner point of the target interest area in the meta-universe space, and generate a minimum circumscribed cuboid; extract an associated visual object intersecting or contained in the minimum circumscribed cuboid according to the geometric-attribute information of each visual object, and generate an associated visual object set; remove an associated visual object that is occluded in the associated visual object set, and generate a target visual object.
Citation Information
Patent Citations
Teaching system in education of universe and working method thereof
CN114092290A
Online learning concentration degree monitoring method and system based on machine vision and medium
CN115205764A