Attention monitoring method and device, storage medium and electronic equipment
Through iterative optimization of the distributed stream processing framework and image processing model, the problems of low efficiency and poor accuracy in student attention monitoring in existing technologies have been solved, and efficient and accurate attention monitoring and timely alarm processing have been achieved.
Patent Information
- Application Number
- CN202510628467.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-23
AI Technical Summary
When monitoring the attention of a large number of students, existing technologies have problems such as long image or video data processing time, low recognition efficiency, poor accuracy, and teachers are unable to handle warning information in a timely manner.
A distributed stream processing framework is used to collect student images collected by IoT devices, and multiple computing nodes are used to iteratively optimize the image processing model. Combined with Kafka cluster storage and a custom sending mechanism, recognition efficiency and accuracy are improved.
It achieves efficient and accurate recognition of a large number of student images and timely warning information processing, improves the efficiency and accuracy of attention monitoring, and reduces teachers' response time.
Smart Images

Figure CN120689915A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an attention monitoring method, device, storage medium and electronic device. Background Art
[0002] Monitoring students' attention refers to evaluating and improving students' concentration through various methods and technical means, thereby optimizing students' learning outcomes.
[0003] At present, monitoring of students' attention can be carried out based on facial recognition of students through cameras in the classroom. The facial recognition results can be used to monitor students' attention and then determine whether students are concentrating on learning.
[0004] However, when using this attention monitoring method, if the attention of a large number of students needs to be monitored, a large amount of image or video data will need to be processed for face recognition, which will cause the recognition process to take a long time, affecting the efficiency of student attention monitoring. In the case of a large amount of data that needs to be processed, it will also lead to missed recognition and misidentification, which will lead to low accuracy of the monitoring results. In addition, when monitoring a large number of students, multiple alarm messages will be generated, which will make it impossible for teachers to check each alarm message in time, and thus make it impossible for teachers to remind students in time. Summary of the Invention
[0005] In view of this, the present application provides an attention monitoring method, device, storage medium and electronic device, the main purpose of which is to improve the current existing technology. If it is necessary to monitor the attention of a large number of students, a large amount of image or video data will need to be processed for face recognition, resulting in the recognition process taking a long time, affecting the efficiency of student attention monitoring. In the case of a large amount of data that needs to be processed, it will also lead to missed recognition and misidentification, which will lead to low accuracy of the monitoring results. In addition, when monitoring a large number of students, multiple alarm messages will be generated, resulting in the teacher being unable to check each alarm message in time, which in turn leads to the teacher being unable to remind students in time. Technical problem.
[0006] In a first aspect, the present application provides an attention monitoring method, comprising:
[0007] Collect student images collected in real time by multiple IoT devices through a distributed stream processing framework to obtain a set of student images to be processed, and store the set of student images to be processed in a Kafka cluster;
[0008] Utilizing multiple computing nodes in the distributed stream processing framework to respectively read the set of student images to be processed, so that each computing node reads a portion of the images included in the set of student images to be processed;
[0009] In each computing node, an image processing model is used to perform image processing on a portion of the read image to identify a target image where the student's attention is abnormal, wherein the image processing model is obtained by iteratively optimizing a pre-trained initial image processing model using a predetermined gradient optimizer;
[0010] The target image set identified by the plurality of computing nodes is determined, and the target image set is sent to the teacher terminal device according to a target sending mode corresponding to the number of images in the target image set.
[0011] In a second aspect, the present application provides an attention monitoring device, comprising:
[0012] A collection module is configured to collect student images collected in real time by multiple IoT devices through a distributed stream processing framework, obtain a set of student images to be processed, and store the set of student images to be processed in a Kafka cluster;
[0013] a reading module configured to respectively read the set of student images to be processed using a plurality of computing nodes in the distributed stream processing framework, so that each computing node reads a portion of images included in the set of student images to be processed;
[0014] A recognition module is configured to perform image processing on the read portion of the image in each computing node using an image processing model to identify a target image in which the student's attention is abnormal, wherein the image processing model is obtained by iteratively optimizing a pre-trained initial image processing model using a predetermined gradient optimizer;
[0015] The sending module is configured to determine the target image set identified by the multiple computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set.
[0016] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the attention monitoring method described in the first aspect.
[0017] In a fourth aspect, the present application provides an electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the attention monitoring method described in the first aspect when executing the computer program.
[0018] By means of the above technical solution, the present application provides an attention monitoring method, device, storage medium and electronic device. Compared with the current existing technology, the present application collects student images collected in real time by multiple Internet of Things devices through a distributed stream processing framework to obtain a set of student images to be processed, and stores the set of student images to be processed in a kafka cluster, so that the present application can collect a large number of student images at the same time for unified storage and analysis; by utilizing multiple computing nodes in the distributed stream processing framework to read the set of student images to be processed respectively, so that each computing node reads part of the images contained in the set of student images to be processed, so that the present application can avoid the situation of missed recognition and misrecognition due to the large amount of data, thereby improving the recognition accuracy; and use the image processing model in each computing node to perform image processing on the read part of the image to identify the student's attention. The target image has abnormal force, and the image processing model is obtained by iteratively optimizing the pre-trained initial image processing model through a predetermined gradient optimizer. A large number of student images can be processed in parallel, and the optimized image processing model is used to recognize the image at each computing node, which can improve the efficiency of image recognition. The use of the iteratively optimized image processing model for recognition can also improve the accuracy of image recognition. The present application also needs to determine the target image set recognized by multiple computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set. That is, the present application can determine the most suitable sending method according to the size of the image number to improve the real-time sending, so that the teacher can determine the attention monitoring result more quickly. The present application can save the teacher's time while also improving the efficiency and accuracy of attention monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A flow chart of an attention monitoring method provided in an embodiment of the present application is shown;
[0022] Figure 2 A schematic diagram showing an example provided by an embodiment of the present application is shown;
[0023] Figure 3A flow chart of an attention monitoring method provided in an embodiment of the present application is shown;
[0024] Figure 4 A schematic diagram of an example process provided by an embodiment of the present application is shown;
[0025] Figure 5 A schematic structural diagram of an attention monitoring device provided in an embodiment of the present application is shown;
[0026] Figure 6 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0027] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other.
[0028] In order to improve the existing technology, if a large number of students need to be monitored for attention, a large amount of image or video data will need to be processed for face recognition, which will result in a long time for the recognition process, affecting the efficiency of student attention monitoring. In the case of a large amount of data that needs to be processed, it will also lead to missed recognition and misidentification, which will lead to low accuracy of the monitoring results. In addition, when a large number of students are monitored, multiple alarm messages will be generated, which will make it impossible for teachers to view each alarm message in time, and thus lead to technical problems such as teachers being unable to remind students in time. This embodiment provides an attention monitoring method, such as Figure 1 As shown, the method includes:
[0029] Step 101: Collect student images collected in real time by multiple IoT devices through a distributed stream processing framework to obtain a set of student images to be processed, and store the set of student images to be processed in a Kafka cluster.
[0030] In the embodiments of the present application, the distributed stream processing framework may be Apache Flink. Specifically, Apache Flink is an open-source distributed stream processing framework designed for high-performance, high-throughput, and low-latency real-time data processing. It supports event-driven applications and batch data processing tasks, and provides powerful state management, fault tolerance, and complex event processing capabilities.
[0031] In some examples, an Apache Kafka cluster is a distributed stream processing platform that allows you to build real-time data pipelines and streaming applications. A Kafka cluster consists of multiple Kafka brokers, each of which is an independent Kafka server instance. These brokers work together to provide high availability and scalability.
[0032] In this example, Apache Flink and Apache Kafka are two crucial components in the big data ecosystem. They are often used together to build real-time data processing pipelines. Kafka, as a high-performance distributed message queue system, efficiently collects, stores, and transmits large data streams, while Flink excels at performing complex real-time processing on these data streams. Combining the two enables efficient data ingestion, processing, and analysis.
[0033] It should be noted that the advantages of combining Flink and Kafka can include: 1. High throughput and low latency: Kafka supports high-throughput data input and output, while Flink can process this data with low latency. 2. Scalability: Both support horizontal expansion, and nodes can be added as needed to improve performance. 3. Fault tolerance: Flink provides a powerful fault tolerance mechanism, including checkpointing and savepoints, to ensure that the state can be restored in the event of a failure; Kafka itself also provides data replication capabilities to ensure that data is not lost. 4. Flexibility: Through Flink's various APIs (such as DataStream API, SQL / Table API, etc.), complex transformations and analyses of data from Kafka can be easily performed.
[0034] For example, the multiple IoT devices that collect student images in real time in the embodiments of the present application may include but are not limited to: IP cameras, gateways, sensors, storage devices, mobile devices, security and privacy protection devices, edge computing units, etc.
[0035] As an optional method, if this application needs to monitor the attention of students in multiple classes in a school, the students' images can be collected in real time based on the cameras of each class, and the collected student images can be uploaded to the distributed stream processing framework for aggregation to obtain a set of student images to be processed, and then the student image set can be stored in the Kafka cluster.
[0036] Step 102: Utilize multiple computing nodes in the distributed stream processing framework to read the set of student images to be processed respectively, so that each computing node reads part of the images included in the set of student images to be processed.
[0037] In the embodiments of this application, Figure 2As shown, a distributed stream processing framework can include multiple compute nodes, each of which can read a portion of the student image collection to be processed. Specifically, in a distributed stream processing framework such as Apache Flink or Apache Spark Streaming, a compute node (also called a worker node or execution node) is the entity responsible for the actual data processing task.
[0038] In some examples, if the set of student images to be processed includes image 1, image 2...image 100, the distributed stream processing framework includes computing node A, computing node B, computing node C, computing node D and computing node E, wherein computing node A can read a part of the images in the set of student images to be processed, such as image 1, image 2...image 20; computing node B can read a part of the images in the set of student images to be processed, such as image 21, image 22...image 40; computing node C can read a part of the images in the set of student images to be processed, such as image 41, image 42...image 60; computing node D can read a part of the images in the set of student images to be processed, such as image 61, image 62...image 80; computing node A can read a part of the images in the set of student images to be processed, such as image 81, image 82...image 100, and so on, which are not listed one by one here.
[0039] Step 103: Use the image processing model in each computing node to perform image processing on the read partial image, and identify the target image where the student's attention is abnormal.
[0040] The image processing model is obtained by iteratively optimizing a pre-trained initial image processing model through a predetermined gradient optimizer.
[0041] In this embodiment, each computing node includes an image processing model, and the image processing model is obtained by iteratively optimizing a pre-trained initial image processing model using a predetermined gradient optimizer.
[0042] In some examples, the pre-trained initial image processing model can be a faster region-based convolutional neural network (Faster Region-based Convolutional Neural Network, Faster R-CNN) is a deep learning model for target detection, mainly used for target detection in images, that is, identifying the location and category of objects in the picture. Specifically, Faster R-CNN introduces the Region Proposal Network (RPN), which greatly improves the speed and efficiency of candidate region generation, combines the powerful feature extraction capability of convolutional neural networks, and achieves high-precision target detection, achieving a good balance between speed and accuracy, and is suitable for real-time applications.
[0043] In some examples, the predetermined gradient optimizer can be an AdaDelta optimizer. AdaDelta is an adaptive learning rate optimization algorithm that can address the problem of rapid learning rate drop in AdaGrad when training deep neural networks. AdaDelta replaces the global historical gradient accumulation by maintaining the squared gradient accumulation within a window, allowing the learning rate to adjust based on the recent gradient magnitude. This approach reduces sensitivity to the initial learning rate setting and eliminates the need for manual learning rate setting.
[0044] It should be noted that the iterative optimization in the embodiment of the present application can save the results obtained after each image processing, optimize the model used based on the saved results and the teacher's marking based on the results, and obtain an optimized image processing model for the next image recognition.
[0045] As an optional method, the embodiment of the present application can focus on collecting, cropping, and processing students' body movements (such as hands, head, face) and micro-expressions (subtle emotional changes such as joy, happiness, anger, sadness, etc.) according to the characteristics and features of students in different grades. According to the data tags of Internet devices (such as school and class information), the focus is on the active behaviors such as whispering, looking around, and distraction in lower grades, and other behaviors. The focus is on monitoring the behaviors of higher grades such as dozing off, depressed expressions, and lowering the head for a long time during class, and conducting subsequent accurate analysis of students' attention.
[0046] Step 104: Determine the target image set identified by the plurality of computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set.
[0047] In an embodiment of the present application, the target image set is a set of target images calculated by multiple computing nodes. The target image can specifically be an image that requires attention reminder identified by an image processing model in the computing node. The target image set is then sent to the teacher's terminal device to remind the teacher to remind the students based on the target image set.
[0048] Compared with the current existing technology, this embodiment collects student images collected in real time by multiple Internet of Things devices through a distributed stream processing framework to obtain a set of student images to be processed, and stores the set of student images to be processed in a Kafka cluster, so that this embodiment can collect a large number of student images at the same time for unified storage and analysis; by utilizing multiple computing nodes in the distributed stream processing framework to read the set of student images to be processed respectively, so that each computing node reads part of the images contained in the set of student images to be processed, so that the present application can avoid the situation of missed recognition and misrecognition due to the large amount of data, thereby improving the recognition accuracy; and use the image processing model in each computing node to process the read part of the image, and identify the target image with abnormal student attention. The image processing model is through The pre-trained initial image processing model is iteratively optimized by a predetermined gradient optimizer, which can perform parallel processing on a large number of student images. The optimized image processing model is used to recognize images at each computing node, which can improve the efficiency of image recognition. The use of the iteratively optimized image processing model for recognition can also improve the accuracy of image recognition. This embodiment also needs to determine the target image set recognized by multiple computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set. That is, this embodiment can determine the most suitable sending method according to the size of the image number to improve the real-time sending, so that the teacher can determine the attention monitoring results more quickly. This embodiment can save the teacher's time while also improving the efficiency and accuracy of attention monitoring.
[0049] As a refinement and extension of the above embodiment, the iterative optimization process of the image processing model can adopt but is not limited to the following methods: Figure 3 As shown, the method includes:
[0050] Step 201: For any optimization process, determine the target training parameter combination required for the current optimization process from the target training parameters required for image processing training of the image processing model.
[0051] In an embodiment of the present application, the target training parameters may include but are not limited to: the number of GPUs used (num_workers) = G0 (such as 1), the number of images in each batch (ims_per_batch) = ims0 (such as 2), the learning rate (BASE_LR) LR0 = 0.0025, the maximum number of iterations: (max_iter) = mi0 (such as 6000), the number of ROIs required for training each image = R0 (such as 128), the number of monitoring categories required C0 (such as 4 types of monitoring targets), etc.
[0052] In some examples, the target training parameter combination required for the current optimization process can be a combination of the above-mentioned target training parameters. For example, it can be a combination of the number of GPUs used (num_workers) = G0 (such as 1), the number of images in each batch (ims_per_batch) = ims0 (such as 2), and the learning rate (BASE_LR) LR0 = 0.0025. It can also be a combination of learning rate (BASE_LR) LR0 = 0.0025, the maximum number of iterations: (max_iter) = mi0 (such as 6000), the number of ROIs required for training each image = R0 (such as 128), the number of required monitoring categories C0 (such as 4 types of monitoring targets), and so on.
[0053] Step 202: Determine the cumulative square gradient of the current optimization process in the predetermined gradient optimizer based on the target training parameter combination, and update the target training parameter based on the cumulative square gradient to obtain an updated target training parameter.
[0054] Optionally, step 202 may specifically include: determining the gradient value of the predetermined gradient optimizer at the current moment, and a weighted moving average corresponding to the parameter update difference of each parameter update; determining the cumulative square gradient based on the gradient value at the current moment; and updating the target training parameter based on the cumulative square gradient and the weighted moving average to obtain the updated target training parameter.
[0055] Exemplarily, based on the example in step 201, let G0, ims0, LR0, mi0, r0, c0 be the gradient optimizer sample x i , let the parameter combination required for the i-th training be y i , linear regression is used to describe the corresponding relationship, as shown in the following formula 1:
[0056] h θ (x i )=θ0+θ1x i (Formula 1)
[0057] In formula 1, θ (x i) represents the parameter combination optimized for the (i)th time, θ0 represents the bias term, θ1 represents the weight coefficient, and further, the corresponding model loss function is:
[0058]
[0059] Among them, the gradient update parameter of the i-th time is Δθ i,t , g i,t is the gradient at time t, as shown in the following formula 3:
[0060]
[0061] Furthermore, the cumulative squared gradient can be calculated using Formula 4, which is as follows:
[0062]
[0063] Furthermore, the parameter update can be calculated using Formula 5, which is specifically as follows:
[0064]
[0065] Furthermore, the exponentially weighted moving average of the square of each parameter update difference, that is, the weighted moving average corresponding to the parameter update difference in the embodiment of the present application, can be calculated by Formula 6. Formula 6 is specifically shown as follows:
[0066]
[0067] It should be noted that AdaDelta does not need to set a global learning rate α, avoiding manual adjustment of the learning rate to control the training process, making the training smoother and improving the training efficiency; ε is a very small number (usually 1 / e 8 )Avoid the denominator being 0.
[0068] Step 203: Obtain an optimized image processing model based on the updated target training parameters.
[0069] It should be noted that the embodiment of the present application uses the control variable method to adjust some samples and some variables, and continuously iterates through AdaDelta gradient parameter adjustment to accelerate the convergence speed. Some samples are shown below. By adjusting the optimized gradient through AdaDelta, rapid convergence from parameter configuration to encoding time is achieved. The training effect is shown in the following trend chart. Through training effect and learning rate adjustment, it is found that by using Flink's distributed computing cluster, by reasonably adjusting the number of GPU / CPU cores and memory, data reading and writing are accessed in memory as much as possible, reducing the disk IO time of data, improving the utilization efficiency of CPU and GPU, reducing busy waiting time, and thus improving the cluster's greater throughput efficiency, reducing feature encoding, and algorithm training time; based on this method, compared with traditional algorithm model encoding, the concurrent processing efficiency of the algorithm encoding model can be greatly improved, and the multi-processor concurrency advantage that a single process cannot utilize can be fully utilized. Based on a distributed cluster method, more computing and storage resources can be managed, and large-scale features and massive data that cannot be processed by a single machine or a single process can be processed. It can support more encoding and training methods, has stronger horizontal expansion capabilities, flexible scalability, and can cope with the processing of different data samples, and is sufficient to meet the timeliness requirements of feature encoding, algorithm training, etc. Therefore, through the adaptive adjustment of this method, machine resources can be utilized more flexibly, the processing efficiency of the machine can be improved, the algorithm training time can be shortened, and the encoding quality and efficiency can be improved more efficiently.
[0070] Optionally, after executing step 103, the method of this embodiment further includes: determining the number of images to be processed in the student image set to be processed; and determining the target sending number corresponding to the target image set identified by the image processing model of the student image set to be processed based on the Poisson distribution probability value of the number of images to be processed.
[0071] In some examples, the Poisson distribution is a discrete probability distribution in statistics that describes the probability of an event occurring multiple times within a fixed time or space. It is particularly well-suited for situations where the probability of an event occurring is relatively small but the event is likely to occur multiple times, and the events are independent (i.e., the occurrence of one event does not affect the occurrence of another).
[0072] In an embodiment of the present application, a custom connector and a custom sending mechanism can be used based on Flink. For the identified image samples that require intervention, relevant information can be sent to the teacher based on the classroom and teacher information of the reading device. The efficiency of each sending is judged based on the currently collected data, ensuring both real-time sending and maximum utilization of machine resources. The number of data records to be sent is N, which is the target sending number in the embodiment of the present application. Within one second, the probability of N occurring follows a Poisson distribution, where λ is the average number of records sent per second. The optimal number of records sent is obtained based on the probability, as shown in the following formula 7:
[0073]
[0074] Optionally, when executing step 104 of "sending the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set", the following methods may be used but are not limited to: comparing the number of images in the target image set with the target sending number; when it is determined that the number of images in the target image set is less than or equal to the target sending number, sending the target image set to the teacher terminal device based on the sending interface of the distributed stream processing framework; when it is determined that the number of images in the target image set is greater than the target sending number, sending the target image set to the teacher terminal device through the Kafka queue of the Kafka cluster.
[0075] For example, based on the above example, if during the off-peak period, the number of student records to be intervened is no more than N, the sending interface is used to send them directly; if during the off-peak period, the number of students to be intervened is greater than N, the data to be sent is put into the kafka queue and sent once.
[0076] Optionally, after executing "determining the target image set identified by multiple computing nodes" in step 104, the method of this embodiment also includes: determining the size of the target image set, and comparing the size of the target image set with a predetermined storage preset; when it is determined that the size of the target image set is less than or equal to the predetermined storage preset, storing the target image set in the memory corresponding to the task manager in the distributed stream processing framework; when it is determined that the size of the target image set is greater than the predetermined storage preset, storing the target image set in a predetermined embedded storage system in the distributed stream processing framework.
[0077] Exemplarily, the embodiment of the present application uses Flink to read image data samples, utilizes cluster computing resources to process the images, uses Faster-CNN for recognition and classification, and performs pooling and convolution optimization. Specifically, a Flink custom function is used, and Faster-CNN is introduced to pool and convolve the data to recognize the images. In order to improve processing efficiency and stability, a threshold M0 is set for the size of the results recognized by Flink and Faster-CNN, which is the predetermined storage threshold in the embodiment of the present application. If the size of the target image set is less than or equal to the threshold M0, it is stored in the memory where the Flink task manager (JobManager) is located (that is, the memory corresponding to the task manager in the distributed stream processing framework in the embodiment of the present application). If the size of the target image set is greater than the threshold M0, it is stored in RocksDB (that is, the predetermined embedded storage system in the embodiment of the present application).
[0078] For this example, RocksDB is an embeddable persistent key-value store suitable for situations requiring high-performance data access. It is built on a Log-Structured Merge Tree (LSM Tree) architecture, providing efficient write performance and excellent read performance.
[0079] Optionally, after executing step 104, the method of this embodiment further includes: collecting feedback information of the teacher reminding the target students corresponding to the target image set to pay attention; based on the student image set to be processed, the target image set and the feedback information, iteratively optimizing the image processing model to obtain an optimized image processing model.
[0080] Among them, the optimized image processing model is used to perform image processing on the collected student images next time.
[0081] It should be noted that the embodiment of the present application can collect teacher feedback data after each recognition, and continuously monitor the prediction accuracy and real-time performance, and continuously optimize according to teacher feedback. The above-mentioned Flin+Kafka+Faster CNN parallel method is based on the secondary development and transformation of Flink, and the AdaDelta optimizer is used to obtain the optimal algorithm model super-parameter combination strategy. Through the Flink distributed and high-performance cluster architecture, the correct storage medium is selected to improve training efficiency and early warning real-time performance. Using Flink real-time processing technology, monitoring data can be read and processed in real time to ensure the real-time performance of algorithm model recognition. At the same time, relying on the custom-developed Connector, the data to be intervened will be sent to the relevant teachers in real time, and the autonomous learning of the neural network is relied on to greatly improve the efficiency of teacher intervention and improve the quality of intervention.
[0082] Compared with the current existing technology, this embodiment collects student images collected in real time by multiple Internet of Things devices through a distributed stream processing framework to obtain a set of student images to be processed, and stores the set of student images to be processed in a Kafka cluster, so that this embodiment can collect a large number of student images at the same time for unified storage and analysis; by utilizing multiple computing nodes in the distributed stream processing framework to read the set of student images to be processed respectively, so that each computing node reads part of the images contained in the set of student images to be processed, so that the present application can avoid the situation of missed recognition and misrecognition due to the large amount of data, thereby improving the recognition accuracy; and use the image processing model in each computing node to process the read part of the image, and identify the target image with abnormal student attention. The image processing model is through The pre-trained initial image processing model is iteratively optimized by a predetermined gradient optimizer, which can perform parallel processing on a large number of student images. The optimized image processing model is used to recognize images at each computing node, which can improve the efficiency of image recognition. The use of the iteratively optimized image processing model for recognition can also improve the accuracy of image recognition. This embodiment also needs to determine the target image set recognized by multiple computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set. That is, this embodiment can determine the most suitable sending method according to the size of the image number to improve the real-time sending, so that the teacher can determine the attention monitoring results more quickly. This embodiment can save the teacher's time while also improving the efficiency and accuracy of attention monitoring.
[0083] In order to illustrate the specific implementation process of this embodiment, the following specific application examples are given: Figure 4 As shown, but not limited to:
[0084] Step 1: Initialize the Flink computing environment and algorithm training environment, read the dataset, cleanse and process the data rows, and focus on collecting, cropping, and processing images of student body movements (such as hands, head, and face) and micro-expressions (subtle emotional changes such as joy, happiness, anger, and sadness) based on the characteristics and features of students of different grades. This allows for subsequent, accurate analysis of student attention. Required resources include the CPU, memory, and parallelism configurations required by Flink for coding, as well as the GPU and network configurations required for Faster-CNN.
[0085] Step 2: Set Faster-CNN based on the optimization algorithm. During the iterative training process, set the sample x required for gradient optimization. i(Number of images, learning journey, maximum number of iterations, number of ROIs required for single-image training, and the category book that needs to be monitored). Adjust the algorithm parameter configuration combination by calculating the cumulative square gradient and parameter update.
[0086] Step 3: Develop a custom component using Flink custom functions, embed the Faster-CNN algorithm, and read image data. During training, use the AdaDelta optimization algorithm to set the training strategy. Iterate the training process, gradually adjusting the algorithm based on the degree of convergence to determine whether it has reached convergence. If convergence is achieved, export the model for subsequent real-time data prediction.
[0087] Step 4: If the training reaches convergence, the algorithm model is exported and then deployed to the Flink cluster. Then, subsequent real-time predictions are performed. During the prediction process, the number and correctness of predictions are continuously collected for subsequent iterations and optimizations.
[0088] Step 5: Use the exported model with Flink to perform real-time predictions, process and predict the read data in real time. Send the real-time data to be sent in real time. At the same time, based on the resource allocation, use the Poisson distribution to calculate the number of data records to be sent, so as to ensure the real-time transmission of the prediction and monitoring data and maximize the utilization of resources.
[0089] In actual applications, the first to fifth steps mentioned above can be executed in sequence according to the specific coding situation to complete the work of monitoring and alerting students' attention.
[0090] Compared with the current existing technology, this embodiment collects student images collected in real time by multiple Internet of Things devices through a distributed stream processing framework to obtain a set of student images to be processed, and stores the set of student images to be processed in a Kafka cluster, so that this embodiment can collect a large number of student images at the same time for unified storage and analysis; by utilizing multiple computing nodes in the distributed stream processing framework to read the set of student images to be processed respectively, so that each computing node reads part of the images contained in the set of student images to be processed, so that the present application can avoid the situation of missed recognition and misrecognition due to the large amount of data, thereby improving the recognition accuracy; and use the image processing model in each computing node to process the read part of the image, and identify the target image with abnormal student attention. The image processing model is through The pre-trained initial image processing model is iteratively optimized by a predetermined gradient optimizer, which can perform parallel processing on a large number of student images. The optimized image processing model is used to recognize images at each computing node, which can improve the efficiency of image recognition. The use of the iteratively optimized image processing model for recognition can also improve the accuracy of image recognition. This embodiment also needs to determine the target image set recognized by multiple computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set. That is, this embodiment can determine the most suitable sending method according to the size of the image number to improve the real-time sending, so that the teacher can determine the attention monitoring results more quickly. This embodiment can save the teacher's time while also improving the efficiency and accuracy of attention monitoring.
[0091] Further, as Figure 1 and Figure 3 The embodiment provides a method for monitoring attention. Figure 5 As shown, the device includes: a collecting module 31 , a reading module 32 , an identifying module 33 , and a sending module 34 .
[0092] The collection module 31 is configured to collect student images collected in real time by multiple IoT devices through a distributed stream processing framework, obtain a set of student images to be processed, and store the set of student images to be processed in a Kafka cluster;
[0093] A reading module 32 is configured to use multiple computing nodes in the distributed stream processing framework to read the set of student images to be processed respectively, so that each computing node reads a portion of the images included in the set of student images to be processed;
[0094] The recognition module 33 is configured to perform image processing on the read portion of the image in each computing node using an image processing model to identify a target image in which the student's attention is abnormal, wherein the image processing model is obtained by iteratively optimizing a pre-trained initial image processing model using a predetermined gradient optimizer;
[0095] The sending module 34 is configured to determine the target image set identified by the multiple computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set.
[0096] In some examples of this embodiment, the iterative optimization process of the image processing model is specifically configured to determine, for any optimization process, the target training parameter combination required for the current optimization process from the target training parameters required for image processing training of the image processing model; determine the cumulative square gradient of the current optimization process in the predetermined gradient optimizer based on the target training parameter combination, and update the target training parameters based on the cumulative square gradient to obtain updated target training parameters; and obtain the optimized image processing model based on the updated target training parameters.
[0097] In some examples of this embodiment, the iterative optimization process of the image processing model is further configured to determine the gradient value of the predetermined gradient optimizer at the current moment, and the weighted moving average corresponding to the parameter update difference for each parameter update; determine the cumulative square gradient based on the gradient value at the current moment; and update the target training parameters based on the cumulative square gradient and the weighted moving average to obtain the updated target training parameters.
[0098] In some examples of this embodiment, the sending module 34 is also configured to determine the number of images to be processed in the student image set to be processed; based on the Poisson distribution probability value of the number of images to be processed, determine the target sending number corresponding to the target image set identified by the image processing model of the student image set to be processed.
[0099] In some examples of this embodiment, the sending module 34 is specifically configured to compare the number of images in the target image set with the target sending number; when it is determined that the number of images in the target image set is less than or equal to the target sending number, the target image set is sent to the teacher terminal device based on the sending interface of the distributed stream processing framework; when it is determined that the number of images in the target image set is greater than the target sending number, the target image set is sent to the teacher terminal device through the Kafka queue of the Kafka cluster.
[0100] In some examples of this embodiment, the sending module 34 is also configured to determine the size of the target image set and compare the size of the target image set with a predetermined storage preset; when it is determined that the size of the target image set is less than or equal to the predetermined storage preset, the target image set is stored in the memory corresponding to the task manager in the distributed stream processing framework; when it is determined that the size of the target image set is greater than the predetermined storage preset, the target image set is stored in a predetermined embedded storage system in the distributed stream processing framework.
[0101] In some examples of this embodiment, the sending module 34 is also configured to collect feedback information from the teacher on reminding the target students corresponding to the target image set to pay attention; based on the student image set to be processed, the target image set and the feedback information, the image processing model is iteratively optimized to obtain an optimized image processing model, and the optimized image processing model is used for image processing of the collected student images next time.
[0102] It should be noted that for other corresponding descriptions of the functional units involved in the attention monitoring device provided in this embodiment, please refer to Figure 1 and Figure 3 The corresponding description in will not be repeated here.
[0103] Based on the above Figure 1 and Figure 3 The method shown in FIG. 1 is a method for performing the above-mentioned steps. Accordingly, this embodiment further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program can realize the above-mentioned steps. Figure 1 and Figure 3 The method shown.
[0104] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of the present application.
[0105] like Figure 6 FIG. 1 is a schematic diagram of the hardware structure of an electronic device of the present invention, comprising:
[0106] at least one processor 401; and,
[0107] A memory 402 in communication with at least one of the processors 401; wherein,
[0108] The memory 402 stores instructions that can be executed by at least one of the processors. The instructions are executed by at least one of the processors to enable the at least one of the processors to perform the attention monitoring method as described above.
[0109] Figure 6 A processor 401 is taken as an example.
[0110] The electronic device may further include an input device 403 and a display device 404 .
[0111] The processor 401, the memory 402, the input device 403 and the display device 404 may be connected via a bus or other means. Figure 6 The bus connection is taken as an example.
[0112] The memory 402 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the attention monitoring method in the embodiment of the present application, for example, Figure 1 and Figure 3 The processor 401 executes the non-volatile software programs, instructions and modules stored in the memory 402 to perform various functional applications and data processing, that is, to implement the attention monitoring method in the above embodiment.
[0113] The memory 402 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the attention monitoring method, etc. In addition, the memory 402 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 402 may optionally include a memory remotely located relative to the processor 401, and these remote memories may be connected to the device executing the attention monitoring method via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0114] The input device 403 can receive user clicks and generate signal inputs related to user settings and function control of the attention monitoring method. The display device 404 can include a display device such as a display screen.
[0115] The one or more modules are stored in the memory 402 and, when executed by the one or more processors 401 , perform the attention monitoring method in any of the above method embodiments.
[0116] Optionally, the physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a Wi-Fi module, and the like. The user interface may include a display, an input unit such as a keyboard, and the like. The optional user interface may also include a USB interface, a card reader interface, and the like. The network interface may optionally include a standard wired interface, a wireless interface (such as a Wi-Fi interface), and the like.
[0117] Those skilled in the art will understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of certain components, or different component arrangements.
[0118] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the physical device, supporting the execution of information processing programs and other software and / or programs. The network communication module is used to enable communication between components within the storage medium, as well as with other hardware and software within the physical information processing device.
[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware platforms, or by hardware. By applying the solution of this embodiment, compared with the current existing technology, this embodiment collects student images collected in real time by multiple IoT devices through a distributed stream processing framework to obtain a set of student images to be processed, and stores the set of student images to be processed in a kafka cluster, so that this embodiment can collect a large number of student images at the same time for unified storage and analysis; by utilizing multiple computing nodes in the distributed stream processing framework to read the set of student images to be processed respectively, so that each computing node reads part of the images contained in the set of student images to be processed, and uses an image processing model in each computing node to perform image processing on the read part of the images, and identifies the target image with abnormal student attention, and the image processing model is a pre-trained image processed by a predetermined gradient optimizer. The initial image processing model obtained by iterative optimization can process a large number of student images in parallel. The optimized image processing model is used to recognize images at each computing node, which can improve the efficiency of image recognition. Using the iteratively optimized image processing model for recognition can also improve the accuracy of image recognition. This embodiment also needs to determine the target image set recognized by multiple computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set. That is, this embodiment can determine the most suitable sending method according to the size of the image number to improve the real-time sending, so that the teacher can determine the attention monitoring results more quickly. This embodiment can save teacher time while also improving the efficiency and accuracy of attention monitoring.
[0120] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0121] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand and implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments described herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.
Claims
1. A method for monitoring attention, characterized in that: include: Collect student images collected in real time by multiple IoT devices through a distributed stream processing framework to obtain a set of student images to be processed, and store the set of student images to be processed in a Kafka cluster; Utilizing multiple computing nodes in the distributed stream processing framework to respectively read the set of student images to be processed, so that each computing node reads a portion of the images included in the set of student images to be processed; In each computing node, an image processing model is used to perform image processing on a portion of the read image to identify a target image where the student's attention is abnormal, wherein the image processing model is obtained by iteratively optimizing a pre-trained initial image processing model using a predetermined gradient optimizer; The target image set identified by the plurality of computing nodes is determined, and the target image set is sent to the teacher terminal device according to a target sending mode corresponding to the number of images in the target image set.
2. The method according to claim 1, characterized in that The iterative optimization process of the image processing model includes: For any optimization process, determining a target training parameter combination required for the current optimization process from the target training parameters required for image processing training of the image processing model; Determining a cumulative square gradient of a current optimization process in the predetermined gradient optimizer based on the target training parameter combination, and updating the target training parameter based on the cumulative square gradient to obtain an updated target training parameter; An optimized image processing model is obtained based on the updated target training parameters.
3. The method according to claim 2, characterized in that The step of determining the cumulative square gradient of the current optimization process in the predetermined gradient optimizer based on the target training parameter combination, and updating the target training parameter based on the cumulative square gradient to obtain the updated target training parameter, includes: Determine the gradient value of the predetermined gradient optimizer at the current moment, and the weighted moving average corresponding to the parameter update difference of each parameter update; Determining the cumulative square gradient based on the gradient value at the current moment; Based on the accumulated square gradient and the weighted moving average, the target training parameter is updated to obtain an updated target training parameter.
4. The method according to claim 1, wherein After performing image processing on the read portion of the image using a pre-trained image processing model in each computing node and identifying a target image where the student's attention is abnormal, the method further includes: Determining the number of images to be processed in the set of student images to be processed; Based on the Poisson distribution probability value of the number of images to be processed, the target sending number corresponding to the target image set identified by the image processing model through the student image set to be processed is determined.
5. The method according to claim 4, characterized in that According to the target sending method corresponding to the number of images in the target image set, the target image set is sent to the teacher terminal device. Comparing the number of images in the target image set with the target sending number; When it is determined that the number of images in the target image set is less than or equal to the target sending number, sending the target image set to the teacher terminal device based on the sending interface of the distributed stream processing framework; When it is determined that the number of images in the target image set is greater than the target sending number, the target image set is sent to the teacher terminal device through the Kafka queue of the Kafka cluster.
6. The method according to claim 1, wherein After determining the target image set identified by the plurality of computing nodes, the method further includes: determining a size of the target image set, and comparing the size of the target image set with a predetermined stored preset; If it is determined that the size of the target image set is less than or equal to the predetermined storage preset, storing the target image set in a memory corresponding to a task manager in the distributed stream processing framework; In a case where it is determined that the size of the target image set is larger than the predetermined storage preset, the target image set is stored in a predetermined embedded storage system in the distributed stream processing framework.
7. The method according to claim 1, characterized in that After determining the target image set identified by the plurality of computing nodes and sending the target image set to the teacher terminal device according to the target sending mode corresponding to the number of images in the target image set, the method further includes: Collecting feedback information on the teacher's attention reminder to the target students corresponding to the target image set; Based on the set of student images to be processed, the target image set and the feedback information, the image processing model is iteratively optimized to obtain an optimized image processing model, which is used for image processing of the collected student images next time.
8. An attention monitoring device, characterized in that: include: A collection module is configured to collect student images collected in real time by multiple IoT devices through a distributed stream processing framework, obtain a set of student images to be processed, and store the set of student images to be processed in a Kafka cluster; a reading module configured to respectively read the set of student images to be processed using a plurality of computing nodes in the distributed stream processing framework, so that each computing node reads a portion of images included in the set of student images to be processed; A recognition module is configured to perform image processing on the read portion of the image in each computing node using an image processing model to identify a target image in which the student's attention is abnormal, wherein the image processing model is obtained by iteratively optimizing a pre-trained initial image processing model using a predetermined gradient optimizer; The sending module is configured to determine the target image set identified by the multiple computing nodes, and send the target image set to the teacher terminal device according to the target sending method corresponding to the number of images in the target image set.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.