Image recognition method and system based on deep learning

By identifying the teaching content and students' facial angles in the teaching area, combining the seat coordinate array and learning probability, and calculating the learning attention coefficient, the problem of inaccurate classroom learning status assessment is solved, and a more accurate classroom attention assessment is achieved.

CN120340095APending Publication Date: 2025-07-18SHANDONG JIANZHU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510459299.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art classroom learning attention monitoring, the image recognition-based method has low accuracy, especially when students' facial expressions cannot be collected, the learning status cannot be accurately evaluated.

Method used

Content recognition is performed by collecting teaching images from the teaching area, cropping student images with seat coordinate arrays, identifying facial angles and calculating learning coefficients, cross-verification is performed based on the probability of attention and recording probability to obtain learning attention coefficients.

Benefits of technology

It realizes more accurate assessment of students' learning status under different teaching links, improves the accuracy and applicability of classroom learning status monitoring, and breaks through the limitations of traditional single visual feature judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340095A_ABST
    Figure CN120340095A_ABST
Patent Text Reader

Abstract

The invention relates to an image recognition method and system based on deep learning, and relates to the technical field of image processing, and the method comprises the steps: collecting a teaching image of a teaching region in a classroom, carrying out the teaching content recognition, and obtaining an attention probability and a recording probability; collecting classroom images of a student area, performing student image cutting according to the seat coordinate array to obtain a student image array, performing face angle recognition to obtain a face angle array, performing screening to obtain a recording angle array and an attention angle array, and performing calculation to obtain a recording learning coefficient; verifying the attention angle array and a standard face angle array corresponding to the seat coordinate array to obtain an attention learning coefficient; and performing learning attention verification calculation according to the record learning coefficient, the attention learning coefficient, the attention probability and the record probability to obtain a learning attention coefficient as an image recognition result. According to the method, the technical problem that classroom attention analysis is inaccurate by adopting image recognition in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image recognition method and system based on deep learning. Background Art

[0002] In the prior art for classroom learning attention monitoring, some use deep learning technology for image recognition, mainly analyzing students' facial expressions to determine their degree of concentration, and the analysis accuracy is relatively low.

[0003] Moreover, in the classroom, students may be in various states of understanding learning content and taking notes, resulting in the inability to collect students' facial expression images for some time, and thus the attention monitoring analysis cannot be carried out. Therefore, there is a technical problem in the prior art that the use of image recognition for classroom attention analysis is inaccurate. Summary of the Invention

[0004] The present invention aims at the technical problem of inaccurate classroom attention analysis using image recognition in the prior art, and provides an image recognition method and system based on deep learning to solve it.

[0005] The technical solution of the present invention to solve the above technical problems is as follows: In the first aspect, the present invention provides an image recognition method based on deep learning, including: in a classroom, collecting teaching images of a teaching area, performing teaching content recognition, and obtaining an attention probability and a recording probability; Collecting classroom images of a student area, cropping student images according to a seat coordinate array to obtain a student image array, performing facial angle recognition to obtain a facial angle array, screening to obtain a recording angle array and an attention angle array, and calculating to obtain a recording learning coefficient; Verifying the attention angle array with a standard facial angle array corresponding to the seat coordinate array to obtain an attention learning coefficient; According to the recording learning coefficient and the attention learning coefficient, performing learning attention verification calculation with the attention probability and the recording probability to obtain a learning attention coefficient as the image recognition result.

[0006] In the second aspect, the present invention provides an image recognition system based on deep learning, including: a teaching image recognition module for collecting teaching images of a teaching area in a classroom, performing teaching content recognition, and obtaining an attention probability and a recording probability; A recording coefficient analysis module for collecting classroom images of a student area, cropping student images according to a seat coordinate array to obtain a student image array, performing facial angle recognition to obtain a facial angle array, screening to obtain a recording angle array and an attention angle array, and calculating to obtain a recording learning coefficient; An attention coefficient analysis module, configured to verify the standard facial angle array corresponding to the attention angle array and the seat coordinate array, and obtain an attention learning coefficient; An image recognition result acquisition module, configured to perform learning attention verification calculation on the recorded learning coefficient and the attention learning coefficient, and the attention probability and the recorded probability, to obtain a learning attention coefficient as the image recognition result.

[0007] The beneficial effects of the present invention are as follows: By collecting teaching images in the teaching area and performing teaching content recognition, the present solution can obtain an attention probability and a recorded probability, and thus, in combination with the complexity of the actual teaching content, ensure a reasonable assessment of the learning status of students in different teaching links, and avoid misjudgment caused by relying solely on facial expressions. By collecting classroom images in the student area and performing student image cropping in combination with the seat coordinate array, the present solution effectively addresses the challenge of multi-person monitoring in the classroom environment, ensures the accurate extraction of each student's image, and improves the accuracy of face recognition. On this basis, by using face angle recognition technology, a facial angle array of students is extracted, and a recorded angle array and an attention angle array are further screened, and a recorded learning coefficient is calculated, so as to more accurately distinguish whether a student is listening attentively or lowering the head to take notes, and solve the problem that existing methods cannot handle multiple learning states of students. At the same time, by verifying the attention angle array with the standard facial angle array and calculating the attention learning coefficient, the present solution ensures that the facial orientation of the student conforms to the spatial relationship between his seat position and the podium, further improving the rationality of attention judgment. Finally, through learning attention verification calculation, in combination with the recorded learning coefficient and the attention learning coefficient, cross-verification is performed with the attention probability and the recorded probability, and a learning attention coefficient is calculated as the final image recognition result. The present invention breaks through the limitation of traditional judgment based on a single visual feature, realizes a more accurate and scientific classroom attention assessment, and improves the accuracy and applicability of classroom learning status monitoring. Description of the Drawings

[0008] Figure 1 It is a schematic flowchart of an image recognition method based on deep learning provided by the present invention; Figure 2 It is a schematic structural diagram of an image recognition system based on deep learning provided by the present invention.

[0009] In the drawings, the components represented by the reference numerals are described as follows: A teaching image recognition module 11, a recording coefficient analysis module 12, an attention coefficient analysis module 13, and an image recognition result acquisition module 14. Detailed Embodiments

[0010] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0011] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0012] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or having more advantages than other embodiments. In order for any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope that conforms to the principles and features disclosed in the present invention.

[0013] Embodiment 1, as Figure 1 shown, the embodiment of the present invention provides an image recognition method based on deep learning, which specifically includes the following steps: S10: In the classroom, collect teaching images of the teaching area, perform teaching content recognition, and obtain the attention probability and recording probability.

[0014] In the embodiments of the present application, in the classroom, teaching images of the teaching area are collected and teaching content recognition is performed. Among them, during the teaching process in the classroom, some teaching content requires students to pay attention and understand, such as the teaching content of example questions, and some teaching content requires students to record, such as basic concepts, etc.

[0015] Perform teaching content recognition to obtain the probability that the teaching content needs to be paid attention to and understood or recorded, and establish a data basis for the subsequent analysis of students' attention.

[0016] The step S10 in the method provided by the embodiments of the present application includes: In the classroom, collect teaching images of the teaching area; Pre-trained teaching content recognizer; Input the teaching image into the teaching content recognizer to obtain the attention probability and the recording probability.

[0017] In the embodiment of the present application, in the classroom environment, first, a high-definition imaging device is used to collect real-time images of the teaching area. The teaching area is, for example, the blackboard area or the projection screen area on the podium. For example, images are collected at a resolution of 1080P or higher, and automatic exposure and white balance adjustment are adopted to ensure clear teaching images can be obtained under different lighting conditions.

[0018] Furthermore, pre-train a teaching content recognizer for recognizing the teaching content in the teaching image to identify the probability that the teaching content in the teaching image needs to be focused on, understood, or recorded.

[0019] In the embodiment of the present application, the pre-trained teaching content recognizer includes: According to the teaching content within a historical time, collect a set of sample teaching images, and collect the proportions of students' understanding and learning and recording when the teaching content appears in different sample teaching images, and label them as a set of sample attention probabilities and a set of sample recording probabilities; Use the set of sample teaching images, the set of sample attention probabilities, and the set of sample recording probabilities as the supervised training data for teaching content recognition. Based on the convolutional neural network in deep learning, train the teaching content recognizer and conduct tests, and stop training after meeting the requirements.

[0020] In the training stage of the teaching content recognizer in the embodiment of the present application, first, it is necessary to construct a set of sample teaching images, that is, collect a large amount of teaching image data based on the classroom teaching content within a historical time. These images can be from different courses and cover various teaching methods, such as teacher's blackboard writing and explanation, projection of PPT presentation, experimental demonstration, etc., to ensure that the recognizer can be generalized to different classroom environments. The collected teaching images are used as sample teaching images to obtain a set of sample teaching images.

[0021] Furthermore, collect the proportions of students' attention, understanding, and learning (i.e., watching the teaching area) and taking notes (i.e., lowering the head to take notes) when the teaching content appears in the above different sample teaching images. Specifically, by collecting the number of students paying attention to the teaching area and the number of students taking notes, calculate the ratios to the total number of students respectively as the proportional coefficients, and then use them as the sample attention probability and the sample recording probability to obtain a set of sample attention probabilities and a set of sample recording probabilities. For example, if the proportion of students for attention, understanding, and learning is 55% and the proportion of students taking notes is 45%, then the sample attention probability and the sample recording probability are 55% and 45%.

[0022] Further, the sample teaching image set, the sample attention probability set, and the sample recording probability set are used as the supervised training data for teaching content recognition, and a teaching content recognizer is trained based on the convolutional neural network in deep learning.

[0023] Exemplarily, based on the convolutional neural network, a teaching content recognizer is constructed, which includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The supervised training data is divided, for example, divided into a training set and a test set according to a ratio of 8:2. There are 5 layers each for the convolutional layer and the pooling layer. The size of the convolutional kernel in the convolutional layer is 3×3. Different dimensional features are extracted through multiple layers of convolution and pooling and input into the fully connected layer, and the attention probability and the recording probability are output. During the training process, the sample teaching image is input, the output attention probability and recording probability are obtained, the differences from the corresponding sample attention probability and sample recording probability are calculated, and then the loss is calculated based on the mean square error loss function. The Adam optimizer is used to optimize the network parameters, and iterative training is performed until the loss is less than the requirement, for example, less than 0.05. The test set is used for testing. If the loss still meets the requirement, the training is completed; otherwise, iterative training continues until the requirement is met.

[0024] After the test meets the requirement, the training is stopped, and the construction and training of the teaching content recognizer are completed.

[0025] The currently collected teaching image is input into the teaching content recognizer for recognition, and the recognized and predicted attention probability and recording probability are obtained. That is, the proportion of students who theoretically pay attention to and take notes on the teaching content in the current teaching image.

[0026] In the embodiment of the present application, by recognizing the attention probability and recording probability of the teaching content, the limitation of traditional classroom monitoring relying only on students' facial expressions is broken through, and more scientific learning attention monitoring can be combined with the characteristics of the teaching content itself to ensure obtaining the proportion of students' different learning states under different teaching contents, improving the accuracy and applicability of classroom learning monitoring.

[0027] S20: Collect classroom images of the student area, crop the student images according to the seat coordinate array to obtain a student image array, perform facial angle recognition to obtain a facial angle array, and screen to obtain a recording angle array and an attention angle array, and calculate to obtain a recording learning coefficient.

[0028] In the embodiment of the present application, classroom images of the student area are collected, and then the student images of each student are cropped according to the seat coordinate array.

[0029] Then, facial angle recognition is performed to obtain the facial angles of multiple students, forming a facial angle array. Then, the recording angles for taking notes are screened out to obtain a recording angle array, and the attention angles for paying attention to and understanding teaching content are obtained to obtain an attention angle array. Among them, different learning states have different facial angles. For example, when paying attention to learning, the vertical angle in the head angle is horizontal or upward, while when taking notes, the vertical angle in the facial angle is downward. Further, calculate the proportion of the number of the recording angle array for recording learning as the recording learning coefficient.

[0030] Step S20 in the method provided by the embodiment of the present application includes: Obtain the coordinates of multiple seats in the classroom within the classroom image, and construct a seat coordinate array; Collect the classroom image of the student area, and crop the classroom image according to the seat coordinate array to obtain a student image array; According to the student learning monitoring data within the historical time, collect a set of sample student images, and label the facial angles of the students in each sample student image to obtain a set of sample facial angles. Each sample facial angle includes a vertical facial angle and a horizontal facial angle; Use the set of sample student images and the set of sample facial angles as the supervised training data for facial angle recognition, and train a facial angle recognizer based on the convolutional neural network in deep learning; Input the multiple student images within the student image array into the facial angle recognizer, and output to obtain a facial angle array.

[0031] In the embodiment of the present application, first, obtain the coordinates of multiple seats in the classroom within the classroom image. Among them, the classroom image of the student area in the classroom is collected according to a fixed position and a fixed angle. Multiple seats in the classroom have corresponding pixel coordinates within the classroom image. For example, in a classroom image of 1920×1080, there are corresponding coordinates, forming the coordinates of multiple seats in the classroom image, that is, a seat coordinate array. For example, if there are 5 rows and 6 seats in each row in the classroom, then 30 seat coordinates are formed, such as {(x1,y1),(x2,y2),...,(x30,y30)}. Among them, one seat coordinate may include multiple pixel coordinates, and multiple pixel points form all the images of the student image on one seat. For example, select all the pixel points of each seat and the students on the seat to form all the seat coordinates of one seat.

[0032] Optionally, a space coordinate system in the classroom can also be constructed. For example, taking the corner of the classroom as the origin of the coordinate system, construct seat coordinates according to the positions of multiple seats to form a seat coordinate array, and then collect the classroom image, and perform student image cropping according to the mapping relationship between the seat coordinates and the pixel coordinates of the image formed by the seats in the classroom image to obtain a student image array.

[0033] Further, collect the classroom images of the student area, and crop the classroom images according to the seat coordinate array. For example, crop the images at each seat coordinate to obtain the images of each seat and the students on the seats, forming a student image array, that is, {I1, I2,..., I30}. Optionally, a rectangular frame can also be used to frame and crop the student images centered on each seat coordinate. Further, identify the student images in the student image array to identify the facial angles of the students in the student images, where a convolutional neural network is used to train the facial angle recognizer for facial angle recognition.

[0034] Specifically, according to the student learning monitoring data within the historical time, collect the student learning images in different classroom environments. For example, when the students bow their heads to take notes (horizontal angle is 0 degrees, vertical angle is -45 degrees, that is, tilted downward by 45 degrees) and when they focus on the teaching content on the blackboard (both the horizontal angle and the vertical angle in the facial angle are 0 degrees), collect the classroom images and crop the student images as sample student images to form a sample student image set.

[0035] Further, label the facial angles of the students in each sample student image, specifically label the vertical angle and the horizontal angle of the student's face. Among them, the horizontal plane and the vertical plane directly in front of the student's face are the horizontal angle of 0 degrees and the vertical angle of 0 degrees, and the vertical angle and the horizontal angle are marked according to the angles between the student's facial angle and the horizontal plane and the vertical plane. Among them, when the facial angle is tilted upward, the angle with the horizontal plane is a positive value, when it is tilted downward, the angle with the horizontal plane is a negative value, when the facial angle is tilted to the left, the angle with the vertical plane is a positive value, and when the facial angle is tilted to the right, the angle with the vertical plane is a negative value.

[0036] Exemplarily, when the student looks straight ahead, both the vertical facial angle and the horizontal facial angle are 0 degrees. When the student usually turns his head 30 degrees to the left, the vertical facial angle is 0 degrees and the horizontal facial angle is 30 degrees. When the student usually turns his head 45 degrees to the right, the vertical facial angle is 0 degrees and the horizontal facial angle is -45 degrees. When the student's head is directly in front and he lowers or raises his head by 30 degrees, the vertical facial angle is -30 degrees or 30 degrees and the horizontal facial angle is 0 degrees.

[0037] In this way, label the facial angles in each student image to obtain the facial angle set, and each sample facial angle includes the vertical facial angle and the horizontal facial angle.

[0038] Use the sample student image set and the sample facial angle set as the supervised training data for facial angle recognition, and train the facial angle recognizer based on the convolutional neural network in deep learning.

[0039] Exemplarily, based on a convolutional neural network, a facial angle recognizer is constructed, which includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The sample student image set and the sample facial angle set are divided, for example, divided into a training set and a test set according to a ratio of 8:2. There are 3 layers each for the convolutional layer and the pooling layer. The size of the convolutional kernel in the convolutional layer is 3×3. Features within student images of different dimensions are extracted through multiple layers of convolution and pooling and input into the fully connected layer, and the vertical facial angle and the horizontal facial angle within the facial angle are output. During the training process, the sample student images are input, the vertical facial angle and the horizontal facial angle within the output facial angle are obtained, the amplitude of the angle difference from the corresponding vertical facial angle and horizontal facial angle within the sample facial angle is calculated, and then the loss is calculated based on the mean square error loss function. The Adam optimizer is used to optimize the network parameters, and iterative training is performed until the loss is less than the requirement, for example, less than 0.05. The test set is used for testing. If the loss still meets the requirement, the training is completed; otherwise, iterative training continues until the requirement is met.

[0040] Based on the trained facial angle recognizer, multiple student images within the student image array are input into the facial angle recognizer for facial angle recognition, and multiple facial angles obtained as output are combined with the seat coordinates within the corresponding seat coordinate array to form a facial angle array, where the facial angle array includes a vertical facial angle array and a horizontal facial angle array.

[0041] In the embodiment of the present application, through deep learning, student images are recognized, and then the facial angles of the students are recognized, which can further analyze whether the students are in a state of concentrating on learning or in a state of taking notes, and then further learning attention analysis is performed by combining the attention probability and the recording probability, improving the accuracy and reliability of learning attention monitoring.

[0042] Step S20 in the method provided in the embodiment of the present application further includes: Obtaining a recording vertical angle range and a focusing vertical angle range; Obtaining the vertical facial angle array within the facial angle array; According to the recording vertical angle range and the focusing vertical angle range, the vertical facial angle array is screened, and the facial angles corresponding to the vertical facial angles falling within the recording vertical angle range or the focusing vertical angle range are respectively used as recording angles or focusing angles to obtain a recording angle array and a focusing angle array; Calculating the proportion of the recording angle array in the facial angle array to obtain a recording learning coefficient.

[0043] In the embodiment of the present application, the facial angles of the students are further classified to classify and obtain the number of students in the note-taking state, and then the recording learning coefficient is calculated.

[0044] Among them, first, obtain the distribution intervals of the vertical facial angles when the student is in the note-taking state or the learning attention state, that is, the note-taking vertical angle interval and the attention vertical angle interval.

[0045] Exemplarily, when the student is in the note-taking state, the head generally tilts downward, and the vertical facial angle is generally negative, for example, from -60 degrees to -15 degrees. As the note-taking vertical angle interval, it can be constructed by actually collecting the maximum and minimum values of the facial vertical angle distribution when the student takes notes.

[0046] When the student is in the state of paying attention to the blackboard or projection content during learning, the head is generally close to horizontal or tilted upward, for example, from -15 degrees to 45 degrees, as the attention vertical angle interval. The attention vertical angle interval can be constructed by actually collecting the maximum and minimum values of the facial vertical angle distribution when the student pays attention to learning. There is no intersection between the note-taking vertical angle interval and the attention vertical angle interval.

[0047] Furthermore, extract the vertical facial angle array within the facial angle array identified in the foregoing content.

[0048] According to the note-taking vertical angle interval and the attention vertical angle interval, classify and screen the vertical facial angle array. The facial angles corresponding to the vertical facial angles falling within the note-taking vertical angle interval are used as note-taking angles, and the facial angles corresponding to the vertical facial angles falling within the attention vertical angle interval are used as attention angles. In this way, the classification and screening of the facial angle array are completed, and the note-taking angle array and the attention angle array are obtained.

[0049] Exemplarily, if a vertical facial angle is 0 degrees, and it falls within the attention vertical angle interval, then the facial angle corresponding to this vertical facial angle is recorded as an attention angle.

[0050] Furthermore, calculate the ratio of the number of note-taking angles in the note-taking angle array to the number of all facial angles in the facial angle array, as the note-taking learning coefficient, that is, the proportion of the currently identified students in the note-taking state. Exemplarily, if the number of note-taking angles in the note-taking angle array is 12, and the number of all facial angles in the facial angle array is 30, then the note-taking learning coefficient is 12 / 30 = 0.4.

[0051] In the embodiments of the present application, by using the distinguishing features of the vertical facial angles of students in the note-taking or learning attention state, the classification of the recording angle and the attention angle is carried out, that is, the classification of students in the note-taking or learning attention state is carried out, and the recording angle array and the attention angle array are obtained. The proportion of students currently in the note-taking state can be accurately calculated, and subsequent learning attention analysis is carried out in combination with the recording probability, improving the analysis accuracy and efficiency. The learning attention can also be monitored when the students are looking down and their facial expressions cannot be analyzed.

[0052] S30: Verify the attention angle array with the standard facial angle array corresponding to the seat coordinate array to obtain the attention learning coefficient.

[0053] In the embodiments of the present application, the multiple attention angles in the attention angle array indicate that the students are not in the note-taking state and may be in the learning attention state. However, when the students observe other directions, their attention angles may also fall within the attention vertical angle range. Therefore, it is necessary to further analyze the proportion of students in the learning attention state.

[0054] When students are in different seats, there are different angles between them and the teaching area. If the students are in the learning attention state, their facial angles should point to the teaching area. If the facial angles of the students do not coincide with the angles of the teaching area, it is possible that the students are in a non-learning attention state such as being distracted.

[0055] Based on this, the standard facial angle array is obtained based on the seat coordinate array, that is, the standard facial angle when the students are in the attention state. The attention angles in the attention angle array are verified and calculated to determine whether they are consistent with the standard facial angles, and then the proportion of students in the learning attention state, that is, the attention learning coefficient, is obtained to accurately monitor the learning attention.

[0056] Step S30 in the method provided by the embodiments of the present application includes: Obtain the teaching coordinates of the teaching area, and calculate the theoretical horizontal facial angles of multiple seat coordinates according to the seat coordinate array to obtain the theoretical horizontal facial angle array; Obtain the attention horizontal facial angle array in the attention angle array; Calculate the angle similarity between the attention horizontal facial angle array and the corresponding theoretical horizontal facial angles in the theoretical horizontal facial angle array, as shown in the following formula: ; Wherein, is the angle similarity, is the attention horizontal facial angle, is the theoretical horizontal facial angle corresponding to the attention horizontal facial angle; Filter the number of the attention-level facial angles whose angular similarity is greater than the similarity threshold, calculate the ratio of the number to the number of the facial angle array, and obtain the attention learning coefficient.

[0057] In the embodiments of the present application, in order to accurately evaluate the attention learning state of students, it is necessary to calculate whether the horizontal facial angle is consistent with the direction of the teaching area (such as the podium and the blackboard), and determine the attention degree through angular similarity calculation, and finally obtain the attention learning coefficient.

[0058] First, obtain the teaching coordinates of the teaching area, that is, the coordinates of the teaching area in the classroom, and then calculate the horizontal angles of the teaching coordinates relative to the coordinates of multiple seats in the seat coordinate array to form the theoretical horizontal facial angles of multiple seat coordinates, and obtain the theoretical horizontal facial angle array.

[0059] Exemplarily, if the teaching coordinates are (XT, YT) and a certain seat coordinate is (XT, YZ), the coordinate values of the X coordinate axes of the teaching coordinates and the seat coordinates are the same, and the teaching coordinates and the seat coordinates are on a parallel line parallel to the length direction of the classroom, then the theoretical horizontal facial angle of the teaching coordinates relative to the seat coordinates is 0 degrees. Among them, based on trigonometric functions, the horizontal angle of the teaching coordinates relative to the seat coordinates can also be calculated based on the teaching coordinates and the seat coordinates to obtain the theoretical horizontal facial angle array. For example, if the teaching coordinates are (0, 5) and a certain seat coordinate is (−2, 1), then the theoretical horizontal facial angle is -26 degrees.

[0060] Further, extract the horizontal facial angles of each attention angle in the attention angle array to obtain the attention-level facial angle array.

[0061] Further, calculate the angular similarity between each attention-level facial angle in the attention-level facial angle array and the theoretical horizontal facial angle under the corresponding seat coordinates, as shown in the following formula: ; Among them, is the angular similarity, is the attention-level facial angle, is the theoretical horizontal facial angle of the seat coordinates corresponding to the attention-level facial angle.

[0062] For example, if the attention-level facial angle recognized from the student image with seat coordinates (−2, 1) is -25 degrees, then the angular similarity between it and the theoretical horizontal facial angle of -26 degrees is 96%.

[0063] In this way, the angular similarity of all the attention-level face angles within the attention-level face angle array is calculated. Further, the number of attention-level face angles with an angular similarity greater than the similarity threshold is screened, and the ratio of this number to the number of all the attention-level face angles within the attention-level face angle array is calculated as the attention learning coefficient.

[0064] Among them, if the angular similarity between the attention-level face angle and the theoretical horizontal face angle is greater, the greater the probability that the student is in the attention learning state. The similarity threshold is, for example, 80%, and it can be obtained by calculating the average angular similarity between the attention-level face angles of the students in the attention learning state and the theoretical horizontal face angle. If the angular similarity is greater than this similarity threshold, it is highly probable that the student is in the attention learning state rather than in a state of being distracted and looking in other directions.

[0065] Further, calculate the ratio of the number of attention-level face angles with an angular similarity greater than the similarity threshold to the number of all the face angles within the face angle array, that is, the ratio to the total number of students, to obtain the proportion of the number of students in the attention learning state to the total number of students as the attention learning coefficient. For example, if the number of attention-level face angles with an angular similarity greater than the similarity threshold is 15 and the total number of students is 30, then the attention learning coefficient is 15 / 30 = 50%.

[0066] The embodiment of the present application combines seat coordinates, face horizontal angle detection, and similarity calculation, and improves the accuracy of classroom attention monitoring by quantitatively analyzing the attention direction of students. It can be applied to an intelligent teaching analysis system to provide teaching feedback for teachers and optimize classroom interaction.

[0067] S40: According to the recorded learning coefficient and the attention learning coefficient, perform learning attention verification calculation with the attention probability and the recorded probability to obtain the learning attention coefficient as the image recognition result.

[0068] In the embodiment of the present application, according to the recorded learning coefficient and the attention learning coefficient obtained by face angle processing in the foregoing content, that is, the proportion of the number of students in the note-taking state and the attention learning state, perform learning attention verification calculation with the attention probability and the recorded probability of the predicted teaching content to obtain the learning attention coefficient as the image recognition result.

[0069] Step S40 in the method provided by the embodiment of the present application includes: Calculate the similarity between the recorded learning coefficient and the recorded probability to obtain the recorded attention coefficient; Calculate the similarity between the attention learning coefficient and the attention probability to obtain the attention attention coefficient; According to the recorded attention coefficient and the attention attention coefficient, calculate to obtain the learning attention coefficient as the image recognition result.

[0070] In the embodiments of the present application, if the recorded learning coefficient is closer to the recorded probability, and if the attention coefficient is closer to the attention probability, it indicates that the proportion of the number of students in the note-taking state and the attention learning state in the current monitoring is more consistent with the recorded probability and the attention probability predicted by the teaching content. Then, the more concentrated the attention of the current students is during learning, and the larger the learning attention coefficient is.

[0071] Among them, to calculate the similarity between the recorded learning coefficient and the recorded probability, specifically, calculate the difference between the recorded learning coefficient and the recorded probability, then calculate the ratio of the absolute value of the difference to the recorded probability as the deviation amplitude between the two, and use 1 minus this deviation amplitude as the recorded attention coefficient. Exemplarily, if the recorded learning coefficient is 40% and the recorded probability is 45%, then the recorded attention coefficient is 1 - 5% / 45% = 89%.

[0072] Calculate the similarity between the attention learning coefficient and the attention probability to obtain the attention attention coefficient. Specifically, calculate the difference between the attention learning coefficient and the attention probability, then calculate the ratio of the absolute value of the difference to the attention probability, use 1 minus this ratio as the deviation amplitude between the two, and use 1 minus this deviation amplitude as the attention attention coefficient. Exemplarily, if the attention learning coefficient is 50% and the attention probability is 55%, then the attention attention coefficient is 1 - 5% / 55% = 91%.

[0073] Furthermore, according to the recorded attention coefficient and the attention attention coefficient, calculate the comprehensive learning attention coefficient as the image recognition result, that is, the result of currently monitoring the learning attention of students by image recognition. Exemplarily, calculate the mean value of the recorded attention coefficient and the attention attention coefficient, and combine the degree of compliance of students' different learning states in the two dimensions to obtain the learning attention coefficient, which reflects the degree of concentration of the current students' attention during learning. For example, the learning attention coefficient is (89% + 91%) / 2 = 90%.

[0074] The learning attention coefficient can be used as a reference for the classroom to adjust teaching strategies or control teaching quality. For example, if the learning attention coefficient is greater than 85%, it indicates that the classroom learning state is good and most students have a high degree of concentration. On the contrary, the classroom learning state may be poor and supervision is required.

[0075] In the embodiments of the present application, by combining the facial angle recognized by deep learning, the recording behavior analysis, and the teaching content model, and through quantitative analysis of students' attention and recording behavior, the final learning attention coefficient is calculated, so as to provide an objective classroom concentration assessment for teachers, which helps to optimize teaching strategies and improve classroom effects.

[0076] The image recognition method based on deep learning provided by the embodiments of the present invention has at least the following technical effects: In the embodiments of the present invention, by collecting teaching images in the teaching area and performing teaching content recognition, this solution can obtain the attention probability and recording probability, so as to combine the complexity of the actual teaching content, ensure the reasonable evaluation of students' learning status under different teaching links, and avoid misjudgment caused by relying solely on facial expressions. This solution collects classroom images in the student area, performs student image cropping in combination with the seat coordinate array, effectively copes with the challenge of multi-person monitoring in the classroom environment, ensures the accurate extraction of each student image, and improves the accuracy of face recognition. On this basis, by using the face angle recognition technology, the face angle array of students is extracted, and the recording angle array and attention angle array are further screened to calculate the recording learning coefficient, so as to more accurately distinguish whether the student is listening attentively or lowering the head to take notes, and solves the problem that the existing methods cannot handle various learning states of students. At the same time, this solution verifies the attention angle array with the standard face angle array corresponding to the seat coordinate array, calculates the attention learning coefficient, ensures that the face orientation of the student conforms to the spatial relationship between his seat position and the podium, and further improves the rationality of attention judgment. Finally, through the learning attention verification calculation, combining the recording learning coefficient and the attention learning coefficient, cross-verifying with the attention probability and the recording probability, and calculating the learning attention coefficient as the final image recognition result. The present invention breaks through the limitation of traditional judgment based on a single visual feature, realizes a more accurate and scientific classroom attention assessment, and improves the accuracy and applicability of classroom learning status monitoring.

[0077] Embodiment 2, as Figure 2 shown, based on the same inventive concept as the image recognition method based on deep learning provided in Embodiment 1, the embodiments of the present invention also provide an image recognition system based on deep learning, including: A teaching image recognition module 11, configured to collect teaching images in the teaching area in the classroom, perform teaching content recognition, and obtain the attention probability and the recording probability; A recording coefficient analysis module 12, configured to collect classroom images in the student area, perform student image cropping according to the seat coordinate array to obtain a student image array, perform face angle recognition to obtain a face angle array, screen to obtain a recording angle array and an attention angle array, and calculate to obtain a recording learning coefficient; An attention coefficient analysis module 13, configured to verify the attention angle array with the standard face angle array corresponding to the seat coordinate array to obtain an attention learning coefficient; An image recognition result obtaining module 14, configured to perform learning attention verification calculation according to the recording learning coefficient and the attention learning coefficient, and cross-verify with the attention probability and the recording probability to obtain a learning attention coefficient as the image recognition result.

[0078] Further, the deep learning-based image recognition system is also used to: in the classroom, collect teaching images of the teaching area; Pre-train the teaching content recognizer; Input the teaching images into the teaching content recognizer to obtain the attention probability and recording probability.

[0079] Further, the deep learning-based image recognition system is also used to: according to the teaching content within a historical time, collect a set of sample teaching images, and collect the proportions of students' understanding and learning and recording when the teaching content appears in different sample teaching images, and label them as a set of sample attention probabilities and a set of sample recording probabilities; Use the set of sample teaching images, the set of sample attention probabilities, and the set of sample recording probabilities as the supervised training data for teaching content recognition, and based on the convolutional neural network in deep learning, train the teaching content recognizer and conduct tests, and stop training after meeting the requirements.

[0080] Further, the deep learning-based image recognition system is also used to: obtain the coordinates of multiple seats in the classroom image during class and construct a seat coordinate array; Collect the classroom image of the student area, and crop the classroom image according to the seat coordinate array to obtain an array of student images; According to the student learning monitoring data within a historical time, collect a set of sample student images, and label the facial angles of students in each sample student image to obtain a set of sample facial angles, and each sample facial angle includes a vertical facial angle and a horizontal facial angle; Use the set of sample student images and the set of sample facial angles as the supervised training data for facial angle recognition, and based on the convolutional neural network in deep learning, train the facial angle recognizer; Input the multiple student images in the array of student images into the facial angle recognizer to output an array of facial angles.

[0081] Further, the deep learning-based image recognition system is also used to: obtain the recording vertical angle interval and the attention vertical angle interval; Obtain the vertical facial angle array in the array of facial angles; According to the recording vertical angle interval and the attention vertical angle interval, screen the vertical facial angle array, and respectively use the facial angles corresponding to the vertical facial angles falling within the recording vertical angle interval or the attention vertical angle interval as the recording angles or attention angles to obtain an array of recording angles and an array of attention angles; Calculate the proportion of the array of recording angles in the array of facial angles to obtain the recording learning coefficient.

[0082] Further, the deep learning-based image recognition system is also used to: obtain the teaching coordinates of the teaching area, calculate the theoretical horizontal facial angles of multiple seat coordinates according to the seat coordinate array, and obtain the theoretical horizontal facial angle array; Obtain the attention horizontal facial angle array within the attention angle array; Calculate the angle similarity between the attention horizontal facial angle array and the corresponding theoretical horizontal facial angle in the theoretical horizontal facial angle array, as shown in the following formula: ; Where, is the angle similarity, is the attention horizontal facial angle, is the theoretical horizontal facial angle corresponding to the attention horizontal facial angle; Screen the number of attention horizontal facial angles with an angle similarity greater than the similarity threshold, calculate the ratio to the number of the facial angle array, and obtain the attention learning coefficient.

[0083] Further, the deep learning-based image recognition system is also used to: calculate the similarity between the recorded learning coefficient and the recorded probability to obtain the recorded attention coefficient; Calculate the similarity between the attention learning coefficient and the attention probability to obtain the attention attention coefficient; Calculate the learning attention coefficient according to the recorded attention coefficient and the attention attention coefficient as the image recognition result.

[0084] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept.

[0085] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. An image recognition method based on deep learning, characterized in that, The method includes: In the classroom, collect teaching images of the teaching area, perform teaching content recognition, and obtain the attention probability and recording probability; Collect classroom images of the student area, crop the student images according to the seat coordinate array to obtain a student image array, perform facial angle recognition to obtain a facial angle array, screen to obtain a recording angle array and an attention angle array, and calculate to obtain a recording learning coefficient; Verify the attention angle array with the standard facial angle array corresponding to the seat coordinate array to obtain an attention learning coefficient; According to the recording learning coefficient and the attention learning coefficient, perform learning attention verification calculation with the attention probability and the recording probability to obtain a learning attention coefficient as the image recognition result.

2. The image recognition method based on deep learning according to claim 1, wherein, In the classroom, collect teaching images of the teaching area, perform teaching content recognition, and obtain the attention probability and the recording probability, including: In the classroom, collect teaching images of the teaching area; Pre-train a teaching content recognizer; Input the teaching images into the teaching content recognizer to recognize and obtain the attention probability and the recording probability.

3. The image recognition method based on deep learning according to claim 2, wherein Pre-training the teaching content recognizer includes: According to the teaching content within the historical time, collect a set of sample teaching images, and collect the proportion of students' understanding and learning and the proportion of recording when the teaching content appears in different sample teaching images, and label them as a set of sample attention probabilities and a set of sample recording probabilities; Use the set of sample teaching images, the set of sample attention probabilities, and the set of sample recording probabilities as the supervised training data for teaching content recognition, and based on the convolutional neural network in deep learning, train the teaching content recognizer and perform tests, and stop training after meeting the requirements.

4. The image recognition method based on deep learning according to claim 1, wherein, Collect classroom images of the student area, crop the student images according to the seat coordinate array to obtain a student image array, and perform facial angle recognition to obtain a facial angle array, including: Obtain the coordinates of multiple seats in the classroom image during class, and construct a seat coordinate array; Collect classroom images of the student area, and crop the classroom image according to the seat coordinate array to obtain a student image array; According to the student learning monitoring data within the historical time, collect a set of sample student images, and label the facial angles of the students in each sample student image to obtain a set of sample facial angles, and each sample facial angle includes a vertical facial angle and a horizontal facial angle; Use the set of sample student images and the set of sample facial angles as the supervised training data for facial angle recognition, and based on the convolutional neural network in deep learning, train the facial angle recognizer; Input multiple student images in the student image array into the facial angle recognizer, and output to obtain a facial angle array.

5. The image recognition method based on deep learning according to claim 1, characterized in that, And screen to obtain a recording angle array and an attention angle array, and calculate to obtain a recording learning coefficient, including: Obtain the recording vertical angle interval and the attention vertical angle interval; Obtain the vertical facial angle array in the facial angle array; According to the recorded vertical angle range and the concerned vertical angle range, screen the vertical facial angle array, and respectively take the facial angles corresponding to the vertical facial angles falling within the recorded vertical angle range or the concerned vertical angle range as the recorded angles or the concerned angles, so as to obtain a recorded angle array and a concerned angle array; Calculate the proportion of the recorded angle array in the facial angle array to obtain a recorded learning coefficient.

6. The image recognition method based on deep learning according to claim 1, wherein Verify the concerned angle array with the standard facial angle array corresponding to the seat coordinate array to obtain a concerned learning coefficient, including: Obtain the teaching coordinates of the teaching area, and calculate the theoretical horizontal facial angles of multiple seat coordinates according to the seat coordinate array to obtain a theoretical horizontal facial angle array; Obtain the concerned horizontal facial angle array within the concerned angle array; Calculate the angular similarity between the concerned horizontal facial angle array and the corresponding theoretical horizontal facial angles in the theoretical horizontal facial angle array, as shown in the following formula: ; Among them, is the angular similarity, is the attention level facial angle, is the theoretical horizontal facial angle corresponding to the attention level facial angle; Screen the number of concerned horizontal facial angles with angular similarity greater than the similarity threshold, and calculate the ratio to the number of the facial angle array to obtain a concerned learning coefficient.

7. The image recognition method based on deep learning according to claim 1, characterized in that, According to the recorded learning coefficient and the concerned learning coefficient, perform a learning attention verification calculation with the concerned probability and the recorded probability to obtain a learning attention coefficient as the image recognition result, including: Calculate the similarity between the recorded learning coefficient and the recorded probability to obtain a recorded attention coefficient; Calculate the similarity between the concerned learning coefficient and the concerned probability to obtain a concerned attention coefficient; Calculate and obtain a learning attention coefficient according to the recorded attention coefficient and the concerned attention coefficient as the image recognition result.

8. An image recognition system based on deep learning, characterized in that, Steps for implementing the deep learning-based image recognition method according to any one of claims 1 to 7, including: A teaching image recognition module, configured to collect teaching images of a teaching area in a classroom, perform teaching content recognition, and obtain a concerned probability and a recorded probability; A recorded coefficient analysis module, configured to collect classroom images of a student area, crop student images according to a seat coordinate array to obtain a student image array, perform facial angle recognition to obtain a facial angle array, screen to obtain a recorded angle array and a concerned angle array, and calculate to obtain a recorded learning coefficient; A concerned coefficient analysis module, configured to verify the concerned angle array with the standard facial angle array corresponding to the seat coordinate array to obtain a concerned learning coefficient; An image recognition result acquisition module, configured to perform a learning attention verification calculation with the concerned probability and the recorded probability according to the recorded learning coefficient and the concerned learning coefficient to obtain a learning attention coefficient as the image recognition result.