State detection method, device, equipment, medium and product

By combining edge detection and deep learning models, the system can accurately monitor students' online classroom status in real time, solving the problem of reliance on students' subjective factors in existing technologies and improving the accuracy of status detection and teaching effectiveness.

CN120876801APending Publication Date: 2025-10-31CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510984819.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing technologies, judging the status of students in online classes relies on subjective factors of students, which cannot accurately obtain the true classroom status, resulting in poor teaching effectiveness.

Method used

Student images are acquired through image acquisition devices, preprocessed using edge detection algorithms, facial feature points and distances are extracted, and a deep learning model is used to determine the student's attention state, thus achieving non-contact, passive state recognition.

Benefits of technology

It enables real-time and accurate monitoring of students' status, reduces reliance on students' active cooperation, improves the accuracy and stability of status detection, and allows for timely adjustments to teaching methods to enhance teaching effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876801A_ABST
    Figure CN120876801A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a state detection method and device, equipment, a medium and a product. The method comprises the following steps: firstly, when a target user is in a course state, acquiring a first user image through image acquisition equipment; then, preprocessing the first user image through an edge detection algorithm to obtain a second user image; then, acquiring a plurality of target feature points from the second user image according to the preset feature points, and acquiring a feature distance between each target feature point and the image acquisition equipment; further, inputting the feature distance into a network model based on deep learning, and determining an attention value output by the network model based on deep learning; and finally, determining the user state of the target user according to the attention value. Through the method, the real-time accurate monitoring of the student state is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of educational informatization technology, and in particular to a state detection method, device, equipment, medium and product. Background Technology

[0002] 5G smart technology is a series of communication applications developed based on the high bandwidth and high speed of 5G internet technology, combined with practical applications across various industries. This has positive implications for education. Currently, 5G smart classrooms and online education are developing rapidly, with more and more students choosing online learning as a way to acquire knowledge and education. However, a common problem in online classrooms is whether students are truly paying attention and participating in learning. Therefore, quickly assessing students' classroom engagement can help teachers understand students' attention levels in real time, allowing for timely adjustments to teaching methods and effectively improving teaching outcomes.

[0003] In existing technologies, students' status in online classes typically relies on self-reporting or manual operation. For example, students manually report their status (such as "focused," "fatigued," "away," etc.) to the system through an interface; or, students demonstrate their "active status" through speaking, chatting, Q&A interaction, etc.

[0004] However, existing methods for judging students' classroom status are affected by students' subjective factors and cannot accurately obtain students' classroom status. Summary of the Invention

[0005] This application provides a status detection method, apparatus, device, medium, and product to solve the problem in the prior art that it is impossible to accurately obtain the classroom status of students.

[0006] In a first aspect, embodiments of this application provide a state detection method, including:

[0007] When the target user is in a course state, the first user image is acquired through an image acquisition device;

[0008] The first user image is preprocessed using an edge detection algorithm to obtain the second user image;

[0009] Multiple target feature points are obtained from the second user image based on preset feature points, and the feature distance between each target feature point and the image acquisition device is obtained.

[0010] The feature distance is input into a deep learning-based network model to determine the attention value output by the deep learning-based network model. The deep learning-based network model is pre-trained based on the sample feature distance and the corresponding sample attention value.

[0011] The user state of the target user is determined based on the attention value.

[0012] In one possible implementation, the step of preprocessing the first user image using an edge detection algorithm to determine the second user image includes:

[0013] The user grayscale image is obtained by performing harmonic average grayscale processing on the RGB values ​​of the pixels in the first user image.

[0014] Based on the user's grayscale image, image convolution processing is performed using a first-direction convolution kernel and a second-direction convolution kernel to obtain a horizontal edge image and a vertical edge image;

[0015] The horizontal edge image and the vertical edge image are superimposed by absolute value processing to obtain the user edge image;

[0016] The second spatial derivative of the user edge image is calculated to obtain the enhancement value of the user edge image;

[0017] The user edge image is added to the enhancement value of the user edge image to obtain the second user image.

[0018] In one possible implementation, before acquiring multiple target feature points from the second user image based on preset feature points, and acquiring the feature distance between each target feature point and the image acquisition device, the method further includes:

[0019] The second user image and each user image in the preset user image library are subjected to grayscale normalization processing to obtain the first grayscale probability density distribution histogram corresponding to the second user image and the second grayscale probability density distribution histogram corresponding to each user image in the preset user image library. Each user image in the preset user image library is an image that has been preprocessed by an edge detection algorithm.

[0020] The first gray-level probability density distribution histogram is compared with the second gray-level probability density distribution histogram corresponding to each user image to obtain multiple similarity values.

[0021] The visual status of the target user is determined based on the multiple similarity values.

[0022] In one possible implementation, determining the visual status of the target user based on multiple similarity values ​​includes:

[0023] The multiple similarity values ​​are sorted from high to low, and it is determined whether the highest similarity value is greater than a preset value.

[0024] If so, the target user is determined to be in a visible state, and multiple target feature points are obtained from the second user image based on preset feature points;

[0025] If not, the target user is determined to be in an invisible state, and adjustment information is output. The adjustment information is used to prompt the target user to adjust their pose to a visible state.

[0026] In one possible implementation, the process of determining whether a user is in a course session includes...

[0027] Voice data is acquired through a voice acquisition device, and the voice data includes multiple feature words;

[0028] Determine whether at least one feature word in the speech data belongs to a preset feature word library;

[0029] If so, it is determined that the user is not in a course session.

[0030] If not, it is determined that the user is in a course status.

[0031] In one possible implementation, determining the user state of the target user based on the attention value includes:

[0032] Determine whether the attention value is greater than a preset attention value;

[0033] If so, confirm that the target user is in a normal state;

[0034] If not, determine that the target user's state is abnormal and output a reminder message, which is used to remind the target user to improve their learning attention.

[0035] Secondly, embodiments of this application provide a state detection device, including:

[0036] The first acquisition module is used to acquire the image of the first user through an image acquisition device when the target user is in a course state;

[0037] The processing module is used to preprocess the first user image using an edge detection algorithm to obtain the second user image;

[0038] The second acquisition module is used to acquire multiple target feature points from the second user image based on preset feature points, and to acquire the feature distance between each target feature point and the image acquisition device;

[0039] The first determining module is used to input the feature distance into the deep learning-based network model and determine the attention value output by the deep learning-based network model. The deep learning-based network model is pre-trained based on the sample feature distance and the corresponding sample attention value.

[0040] The second determining module is used to determine the user state of the target user based on the attention value.

[0041] In one possible implementation, the processing module is specifically used for:

[0042] The user grayscale image is obtained by performing harmonic average grayscale processing on the RGB values ​​of the pixels in the first user image.

[0043] Based on the user's grayscale image, image convolution processing is performed using a first-direction convolution kernel and a second-direction convolution kernel to obtain a horizontal edge image and a vertical edge image;

[0044] The horizontal edge image and the vertical edge image are superimposed by absolute value processing to obtain the user edge image;

[0045] The second spatial derivative of the user edge image is calculated to obtain the enhancement value of the user edge image;

[0046] The user edge image is added to the enhancement value of the user edge image to obtain the second user image.

[0047] In one possible implementation, before acquiring multiple target feature points from the second user image based on preset feature points and acquiring the feature distance between each target feature point and the image acquisition device, the processing module is further configured to:

[0048] The second user image and each user image in the preset user image library are subjected to grayscale normalization processing to obtain the first grayscale probability density distribution histogram corresponding to the second user image and the second grayscale probability density distribution histogram corresponding to each user image in the preset user image library. Each user image in the preset user image library is an image that has been preprocessed by an edge detection algorithm.

[0049] The first gray-level probability density distribution histogram is compared with the second gray-level probability density distribution histogram corresponding to each user image to obtain multiple similarity values.

[0050] The visual status of the target user is determined based on the multiple similarity values.

[0051] In one possible implementation, the processing module is specifically used for:

[0052] The multiple similarity values ​​are sorted from high to low, and it is determined whether the highest similarity value is greater than a preset value.

[0053] If so, the target user is determined to be in a visible state, and multiple target feature points are obtained from the second user image based on preset feature points;

[0054] If not, the target user is determined to be in an invisible state, and adjustment information is output. The adjustment information is used to prompt the target user to adjust their pose to a visible state.

[0055] In one possible implementation, the state detection device further includes a judgment module, used for:

[0056] Voice data is acquired through a voice acquisition device, and the voice data includes multiple feature words;

[0057] Determine whether at least one feature word in the speech data belongs to a preset feature word library;

[0058] If so, it is determined that the user is not in a course session.

[0059] If not, it is determined that the user is in a course status.

[0060] In one possible implementation, the second determining module is specifically used for:

[0061] Determine whether the attention value is greater than a preset attention value;

[0062] If so, confirm that the target user is in a normal state;

[0063] If not, determine that the target user's state is abnormal and output a reminder message, which is used to remind the target user to improve their learning attention.

[0064] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0065] The memory stores computer-executed instructions;

[0066] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0067] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0068] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.

[0069] This application provides a state detection method, apparatus, device, medium, and product. The method includes: First, when the target user is in a course state, acquiring a first user image through an image acquisition device to avoid acquiring images when the user is not in a course state, thus saving computational resources; then, preprocessing the first user image using an edge detection algorithm to obtain a second user image, thereby transforming the complex original image into an edge image with clear structure and less background interference; next, acquiring multiple target feature points from the second user image based on preset feature points, and obtaining the feature distance between each target feature point and the image acquisition device; further, inputting the feature distance into a deep learning-based network model to determine the attention value output by the deep learning-based network model; finally, determining the user state of the target user based on the attention value. Thus, this method achieves real-time and accurate monitoring of student state without requiring manual user intervention, reducing dependence and improving the accuracy of student state detection. Attached Figure Description

[0070] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0071] Figure 1 Flowchart of the state detection method provided in the embodiments of this application Figure 1 ;

[0072] Figure 2 Flowchart of the state detection method provided in the embodiments of this application Figure 2 ;

[0073] Figure 3 This is a schematic diagram of the structure of the state detection device provided in the embodiments of this application;

[0074] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0075] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0076] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0077] 5G smart technology is a series of communication applications developed based on the high bandwidth and high speed of 5G internet technology, combined with practical applications across various industries. This has positive implications for education. Currently, 5G smart classrooms and online education are developing rapidly, with more and more students choosing online learning as a way to acquire knowledge and education. However, a common problem in online classrooms is whether students are truly paying attention and participating in learning. Therefore, quickly assessing students' classroom engagement can help teachers understand students' attention levels in real time, allowing for timely adjustments to teaching methods and effectively improving teaching outcomes.

[0078] In existing technologies, students' status in online classes typically relies on self-reporting or manual operation. For example, students manually report their status (such as "focused," "fatigued," "away," etc.) to the system through an interface; or, students demonstrate their "active status" through speaking, chatting, Q&A interaction, etc.

[0079] However, existing methods for judging students' classroom status are affected by students' subjective factors. Some students may deliberately choose the wrong status to hide their true situation (such as pretending to be focused), making it impossible to accurately obtain students' classroom status.

[0080] Based on this, this application proposes a user state detection method. Considering that in existing technologies, student state detection in online learning scenarios mainly relies on students' active operations, which cannot obtain accurate and true state information, if state recognition can be shifted from "relying on students' active input" to "automatic perception by the device," that is, achieving non-contact, passive recognition of student state through the integration of computer vision and sensing technology, it is possible to effectively determine whether the student is currently focused, thereby avoiding reliance on students' active cooperation and improving the accuracy and stability of detection. Specifically, by using an edge detection algorithm to process the camera image (i.e., the first user image), the human portrait contour (i.e., the second user image) is extracted to improve the accuracy of portrait recognition. Then, multiple feature points in the portrait contour and the feature distances between these feature points and the image acquisition device are extracted. Finally, a trained model is used to predict the feature distances to determine the current learning state of the student.

[0081] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0082] Figure 1 Flowchart of the state detection method provided in the embodiments of this application Figure 1 ;like Figure 1 As shown, the method includes:

[0083] S101. When the target user is in a course state, acquire the first user image through the image acquisition device.

[0084] In one possible approach, before acquiring the first user image, it is also necessary to determine whether the user is in a course state. Specifically:

[0085] First, voice data is acquired through a voice acquisition device; then, it is determined whether at least one feature word in the voice data belongs to a preset feature word library; if it does, it is determined that the user is not in a course state; if it does not, it is determined that the user is in a course state.

[0086] The voice data includes multiple feature words; the feature words in the preset feature word library are feature words that indicate non-course status, such as "taking notes", "break between classes", "opening the textbook", etc.

[0087] It should be understood that voice data of the user's current environment is collected in real time through voice acquisition devices (such as microphones and voice recognition modules). Then, the voice recognition engine can be used to convert it into text, extract the feature words contained therein, and then match these feature words with a preset feature word library. If the feature words identified in the voice data exist in the word library, it is considered that the user is not currently in a course state; if the above feature words do not appear, it means that the user is listening to a course and can be judged to be in a course state.

[0088] Understandably, by determining whether the current user is in a course state, it is possible to avoid frequently calling computationally intensive image processing tasks (such as edge detection, face recognition, etc.) and reduce the burden on network resources.

[0089] S102. The first user image is preprocessed using an edge detection algorithm to obtain the second user image.

[0090] In one possible approach, firstly, a harmonic average grayscale processing is performed on the RGB values ​​of pixels in the first user image to obtain a user grayscale image; then, based on the user grayscale image, image convolution processing is performed using a first-direction convolution kernel and a second-direction convolution kernel respectively to obtain a horizontal edge image and a vertical edge image; next, the horizontal edge image and the vertical edge image are superimposed by absolute value processing to obtain a user edge image; further, the second-order spatial derivative of the user edge image is calculated to obtain an enhancement value for the user edge image; finally, the user edge image and the enhancement value of the user edge image are added together to obtain a second user image.

[0091] It should be noted that this application embodiment uses image judgment instead of video judgment, mainly to reduce bandwidth consumption and avoid excessive bandwidth consumption after using this system, which would cause lag in online classes. Furthermore, the faster the transmitted image is transmitted per unit time, the stronger the real-time performance of the user status judgment.

[0092] It should be understood that due to environmental factors, the original RGB image is greatly affected by the intensity of ambient light and changes in the color of the light source. Conventional grayscale processing (such as weighted averaging) is prone to distortion in areas of strong light or shadow, while harmonic average grayscale processing has less impact on highlight values ​​and can better highlight edge contour information. Subsequently, using convolutional kernels in different directions can detect the contour information in the horizontal and vertical directions of the image respectively, effectively extracting edge positions such as faces, corners of eyes, contours, and shoulders. Furthermore, by superimposing the absolute values ​​of the two directions, it can more comprehensively reflect all edges of the image, resulting in a more stable and comprehensive edge image (i.e., structural features in all directions). In addition, the second derivative can capture the parts of accelerated change, that is, the locations where the image intensity undergoes "drastic changes," further enhancing edge details, making the contours clearer, better distinguishing real edges from blurred areas, and facilitating subsequent feature point localization.

[0093] Specifically, the formula for harmonic average grayscale processing is:

[0094]

[0095] Where B, G, and R are the RGB values ​​of pixels in the first user image, respectively; G(x, y) is the harmonic average grayscale value.

[0096] The formula for the convolution kernel in the first direction (i.e., the horizontal direction) is:

[0097]

[0098] The formula for the convolution kernel in the second direction (i.e., the vertical direction) is:

[0099]

[0100] in, The kernel is the convolution kernel in the first direction; To harmonize the average grayscale value; This is the kernel for the second direction of convolution.

[0101] Furthermore, user edge images:

[0102] The formula for the second spatial derivative is:

[0103]

[0104] Second user image:

[0105] in, For user edge images; For the second user image; is the second spatial derivative; k is the coefficient for extracting details from the edge contour. Adjusting its size can yield grayscale images with different edge details.

[0106] Understandably, preprocessing the original user image using edge detection algorithms can effectively enhance the clarity of key structural contours in the image, reduce the impact of background interference and lighting changes on image recognition, thereby improving the accuracy and robustness of subsequent feature point extraction and attention state judgment.

[0107] S103. Obtain multiple target feature points from the second user image based on preset feature points, and obtain the feature distance between each target feature point and the image acquisition device.

[0108] Among them, the preset feature points are pre-defined facial feature points of the user, such as features like the forehead, eyes, chin, and nose.

[0109] It is understandable that by extracting key structural points on a user's face (such as the corners of the eyes, the tip of the nose, and the corners of the mouth) and obtaining their spatial geometric distance from the image acquisition device, it is possible to accurately perceive the user's facial posture and behavioral state, enhance the accuracy, stability, and real-time performance of attention judgment, and provide input data for subsequent deep learning models, thereby significantly improving the reliability and practicality of user state detection.

[0110] S104. Input the feature distance into the deep learning-based network model to determine the attention value output by the deep learning-based network model.

[0111] Among them, the deep learning-based network model is pre-trained based on the sample feature distance and the corresponding sample attention value.

[0112] During the training phase, the model learns by fitting a large amount of sample data (each set of samples consists of four distance features and their corresponding real attention labels) to identify the attention level corresponding to different spatial postures. For example, when the forehead distance is extremely small while the chin distance is extremely large and the distance between the two eyes is inconsistent, the model identifies it as a state of extremely low attention (such as 20%), thereby realizing the judgment of the student's attention state.

[0113] It should be understood that this application constructs a deep learning-based neural network model, using the feature distances between four key facial feature points extracted from the user image (such as forehead a, left eye b, right eye c, and chin d) and the screen as input features, corresponding to distances 1 to 4 respectively. These geometric features reflect the user's facial pose and attention state. Furthermore, the feature distances are input into the deep learning-based network model, which (e.g., multilayer perceptron, convolutional neural network (CNN), Transformer, etc.) captures the non-linear relationship between the feature distances and the student's attention state, thereby outputting attention values ​​that more closely resemble the student's true state and improving the accuracy of user state detection.

[0114] It should also be noted that the network structure uses CNN to process each distance input: the first to fourth convolutional layers process individual features from distance 1 to distance 4 respectively. Each layer uses a different number and size of convolutional kernels for feature extraction, and the ReLU activation function is used to enhance the non-linear expressive power. Pooling operation further reduces dimensionality and extracts key features. Finally, the data output from multiple convolutional layers are pooled into a fully connected layer for deep fusion analysis. Dropout is used to prevent overfitting. Finally, the output layer of the Sigmoid activation function generates an attention concentration value between 0 and 1.

[0115] S105. Determine the user status of the target user based on the attention value.

[0116] In one possible approach, it is determined whether the attention value is greater than a preset attention value; if so, the target user's state is determined to be normal; if not, the target user's state is determined to be abnormal, and a reminder message is output.

[0117] The reminder message is used to encourage the target user to improve their learning focus.

[0118] Understandably, by determining whether the attention value output by the deep learning model exceeds a preset attention threshold, the system can quickly assess and respond to the target user's current learning state. Specifically, the attention value output by the model is compared with a preset threshold (e.g., 60%). If the value is greater than or equal to the threshold, it indicates that the user is currently in a relatively focused state, and the system determines that the user's state is "normal." Conversely, if the attention value is lower than the threshold, the system determines that the user's state is "abnormal," meaning that the user may be in a non-learning state such as being distracted, fatigued, inattentive, or having improper posture, and immediately generates a reminder message.

[0119] It should be noted that reminders can be presented in various ways, such as on-screen text prompts like "Please concentrate," audio announcements like "You are daydreaming," or interface highlighting and flashing, to guide users to adjust their state and regain focus in a timely manner.

[0120] This application provides a state detection method, which includes: First, when the target user is in a course state, acquiring a first user image through an image acquisition device to avoid acquiring images when the user is not in a course state, thus saving computing resources; then, preprocessing the first user image using an edge detection algorithm to obtain a second user image, thereby transforming the complex original image into an edge image with clear structure and less background interference; then, acquiring multiple target feature points from the second user image based on preset feature points, and acquiring the feature distance between each target feature point and the image acquisition device; further, inputting the feature distance into a deep learning-based network model to determine the attention value output by the deep learning-based network model; finally, determining the user state of the target user based on the attention value. Thus, this method achieves real-time and accurate monitoring of student state without requiring manual user cooperation, reducing dependence and improving the accuracy of student state detection.

[0121] Figure 2 Flowchart of the state detection method provided in the embodiments of this application Figure 2 ,like Figure 2 As shown, in this embodiment... Figure 2 Based on the embodiments, the process of determining the user's availability status using a second user image is described in detail. The method includes:

[0122] S201. Perform grayscale normalization processing on the second user image and each user image in the preset user image library respectively to obtain the first grayscale probability density distribution histogram corresponding to the second user image and the second grayscale probability density distribution histogram corresponding to each user image in the preset user image library.

[0123] Each user image in the preset user image library is an image that has been preprocessed by an edge detection algorithm.

[0124] Specifically, the formula for grayscale normalization is:

[0125]

[0126] Where 'a' represents the image grayscale level, where 0 is black and 1 is white, limiting the image grayscale pixel values ​​to the range [0,1]. Represents discrete grayscale values. Indicates the appearance In this grayscale value representation, 'b' represents the total number of pixels, and 'c' represents the total number of grayscale values, which is generally less than or equal to 256. This represents the probability density function.

[0127] The above grayscale normalization formula can be used to obtain different grayscale values. Indicated The probability density histogram of occurrence, namely the first gray-level probability density distribution histogram corresponding to the second user image, and the second gray-level probability density distribution histogram corresponding to each user image in the preset user image library.

[0128] Understandably, transforming user images with clear structures and prominent outlines into a unified and comparable statistical grayscale distribution enhances the stability of similarity judgments between different images.

[0129] S202. Compare the similarity between the first gray-level probability density distribution histogram and the second gray-level probability density distribution histogram corresponding to each user image to obtain multiple similarity values.

[0130] It should be noted that similarity comparison methods can use statistical measures such as Euclidean distance, cosine similarity, and chi-square distance. The specific similarity comparison method is not limited in this embodiment of the application; for example, the calculation formula for Euclidean distance is:

[0131]

[0132] The corresponding similarity value calculation formula is:

[0133] in, These represent the probability density values ​​of the second user image and the image in the image library at grayscale value i, respectively.

[0134] S203. Determine the visual status of the target user based on multiple similarity values.

[0135] In one possible approach, firstly, multiple similarity values ​​are sorted from high to low, and it is determined whether the highest similarity value is greater than a preset value. If so, the target user is determined to be in a visible state, and multiple target feature points are obtained from the second user image based on preset feature points. If not, the target user is determined to be in an invisible state, and adjustment information is output.

[0136] The adjustment information is used to prompt the target user to adjust their pose to a visible state.

[0137] It should be understood that the calculated similarity values ​​are sorted from high to low, and it is determined whether the maximum similarity value is higher than a set threshold (e.g., 80%). If it is higher than the threshold, it means that the current image structure is very close to some standard images in the image library, and it is determined to be in a visible state. The user's feature points are then extracted for attention analysis. If it is lower than the threshold, it means that the image cannot be matched with any standard image (e.g., looking down, side profile, occlusion, etc.), and it is determined to be in a non-visible state. Feature point extraction is no longer performed, and adjustment reminder information, such as "Please sit up straight" or "Please face the screen directly", is output to guide the user to adjust their posture.

[0138] Understandably, by judging the user's visual state, low-quality images caused by factors such as shooting angle deviation, occlusion, blurriness, or looking down can be filtered out first, avoiding erroneous information from interfering with subsequent facial feature extraction and attention analysis. Secondly, by judging whether the image meets the preset visual conditions, the user can be promptly prompted to adjust their sitting posture or face the screen, further improving the user's attention. In addition, it reduces the system's ineffective use of computing resources, enabling image processing, feature matching, and deep learning algorithms to run under high-quality input, improving the overall response speed and robustness.

[0139] Figure 3 This is a schematic diagram of the structure of the state detection device provided in the embodiments of this application; as shown below. Figure 3 As shown, the device includes:

[0140] The first acquisition module 301 is used to acquire the image of the first user through an image acquisition device when the target user is in a course state.

[0141] Processing module 302 is used to preprocess the first user image using an edge detection algorithm to obtain the second user image;

[0142] The second acquisition module 303 is used to acquire multiple target feature points from the second user image based on preset feature points, and to acquire the feature distance between each target feature point and the image acquisition device;

[0143] The first determining module 304 is used to input the feature distance into the deep learning-based network model and determine the attention value output by the deep learning-based network model. The deep learning-based network model is pre-trained based on the sample feature distance and the corresponding sample attention value.

[0144] The second determining module 305 is used to determine the user status of the target user based on the attention value.

[0145] In one possible implementation, the processing module 302 is specifically used for:

[0146] The user grayscale image is obtained by performing harmonic average grayscale processing on the RGB values ​​of the pixels in the first user image.

[0147] Based on the user's grayscale image, image convolution processing is performed using a first-direction convolution kernel and a second-direction convolution kernel respectively to obtain horizontal edge images and vertical edge images;

[0148] The horizontal edge image and the vertical edge image are superimposed by absolute value processing to obtain the user edge image;

[0149] The second spatial derivative of the user edge image is calculated to obtain the enhancement value of the user edge image;

[0150] The second user image is obtained by adding the user edge image to the enhanced value of the user edge image.

[0151] In one possible implementation, before acquiring multiple target feature points from the second user image based on preset feature points and acquiring the feature distance between each target feature point and the image acquisition device, the processing module 303 is further configured to:

[0152] The second user image and each user image in the preset user image library are subjected to grayscale normalization processing to obtain the first grayscale probability density distribution histogram corresponding to the second user image and the second grayscale probability density distribution histogram corresponding to each user image in the preset user image library. Each user image in the preset user image library is an image that has been preprocessed by the edge detection algorithm.

[0153] The first gray-level probability density distribution histogram is compared with the second gray-level probability density distribution histogram corresponding to each user image to obtain multiple similarity values;

[0154] The visual status of the target user is determined based on multiple similarity values.

[0155] In one possible implementation, the processing module 302 is specifically used for:

[0156] Sort multiple similarity values ​​from high to low and determine whether the highest similarity value is greater than a preset value;

[0157] If so, the target user is determined to be in a visible state, and multiple target feature points are obtained from the second user image based on preset feature points;

[0158] If not, the target user is determined to be in a non-visual state, and adjustment information is output to prompt the target user to adjust their pose to a visible state.

[0159] In one possible implementation, the state detection device further includes a judgment module, used for:

[0160] Voice data is acquired through a voice acquisition device, and the voice data includes multiple feature words;

[0161] Determine whether at least one feature word in the speech data belongs to a preset feature word library;

[0162] If so, it confirms that the user is not in a course session.

[0163] If not, it confirms that the user is in a course session.

[0164] In one possible implementation, the second determining module 305 is specifically used for:

[0165] Determine if the attention value is greater than the preset attention value;

[0166] If so, confirm that the target user is in a normal state;

[0167] If not, determine that the target user's state is abnormal and output a reminder message to remind the target user to improve their learning attention.

[0168] The state detection device provided in this application embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0169] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 40 provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the device 40 further includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.

[0170] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.

[0171] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0172] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0173] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0174] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0175] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0176] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0177] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0178] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0179] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0180] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0181] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0182] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0183] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0184] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A state detection method, characterized in that, include: When the target user is in a course state, the first user image is acquired through the image acquisition device; The first user image is preprocessed using an edge detection algorithm to obtain the second user image; Multiple target feature points are obtained from the second user image based on preset feature points, and the feature distance between each target feature point and the image acquisition device is obtained. The feature distance is input into a deep learning-based network model to determine the attention value output by the deep learning-based network model. The deep learning-based network model is pre-trained based on the sample feature distance and the corresponding sample attention value. The user state of the target user is determined based on the attention value.

2. The method according to claim 1, characterized in that, The step of preprocessing the first user image using an edge detection algorithm to determine the second user image includes: The user grayscale image is obtained by performing harmonic average grayscale processing on the RGB values ​​of the pixels in the first user image. Based on the user's grayscale image, image convolution processing is performed using a first-direction convolution kernel and a second-direction convolution kernel to obtain a horizontal edge image and a vertical edge image; The horizontal edge image and the vertical edge image are superimposed by absolute value processing to obtain the user edge image; The second spatial derivative of the user edge image is calculated to obtain the enhancement value of the user edge image; The user edge image is added to the enhancement value of the user edge image to obtain the second user image.

3. The method according to claim 1 or 2, characterized in that, Before obtaining multiple target feature points from the second user image based on preset feature points, and obtaining the feature distance between each target feature point and the image acquisition device, the method further includes: The second user image and each user image in the preset user image library are subjected to grayscale normalization processing to obtain the first grayscale probability density distribution histogram corresponding to the second user image and the second grayscale probability density distribution histogram corresponding to each user image in the preset user image library. Each user image in the preset user image library is an image that has been preprocessed by an edge detection algorithm. The first gray-level probability density distribution histogram is compared with the second gray-level probability density distribution histogram corresponding to each user image to obtain multiple similarity values. The visual status of the target user is determined based on the multiple similarity values.

4. The method according to claim 3, characterized in that, Determining the visual status of the target user based on multiple similarity values ​​includes: The multiple similarity values ​​are sorted from high to low, and it is determined whether the highest similarity value is greater than a preset value. If so, the target user is determined to be in a visible state, and multiple target feature points are obtained from the second user image based on preset feature points; If not, the target user is determined to be in an invisible state, and adjustment information is output. The adjustment information is used to prompt the target user to adjust their pose to a visible state.

5. The method according to claim 1, characterized in that, The process of determining whether a user is in a course session includes Voice data is acquired through a voice acquisition device, and the voice data includes multiple feature words; Determine whether at least one feature word in the speech data belongs to a preset feature word library; If so, it is determined that the user is not in a course session. If not, it is determined that the user is in a course status.

6. The method according to any one of claims 1, characterized in that, Determining the user state of the target user based on the attention value includes: Determine whether the attention value is greater than a preset attention value; If so, confirm that the target user is in a normal state; If not, determine that the target user's state is abnormal and output a reminder message, which is used to remind the target user to improve their learning attention.

7. A state detection device, characterized in that, include: The first acquisition module is used to acquire the image of the first user through an image acquisition device when the target user is in a course state; The processing module is used to preprocess the first user image using an edge detection algorithm to obtain the second user image; The second acquisition module is used to acquire multiple target feature points from the second user image based on preset feature points, and to acquire the feature distance between each target feature point and the image acquisition device; The first determining module is used to input the feature distance into the deep learning-based network model and determine the attention value output by the deep learning-based network model. The deep learning-based network model is pre-trained based on the sample feature distance and the corresponding sample attention value. The second determining module is used to determine the user state of the target user based on the attention value.

8. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it is used to implement the method as described in any one of claims 1 to 6.