Business administration teaching model identification and feedback method and system based on machine vision
By using machine vision-based methods to capture students' facial expressions and micro-expressions in real time, the problem of quantifying learning status in business administration teaching has been solved, enabling precise adaptation of teaching content and personalized guidance, and improving the pertinence of teaching management.
Patent Information
- Application Number
- CN202511914763.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-12-18
AI Technical Summary
Existing technologies cannot effectively quantify students' learning status in business administration education, especially in complex business case analysis and cross-departmental collaborative simulation scenarios. They are unable to capture subtle emotional characteristics and behavioral data, resulting in a lack of targeted adjustments to teaching content and an inability to meet the needs of refined management.
Using a machine vision-based approach, students' facial expressions and micro-expressions are captured in real time through a dual-camera image acquisition device. Combined with Laplacian edge enhancement and a convolutional neural network model, emotional features are extracted and linked to a digital interactive system to generate real-time feedback and personalized teaching content adaptation suggestions.
It enables precise quantification and real-time feedback of students' learning status, improves the adaptability of teaching content and personalized guidance, and optimizes the allocation and management decisions of teaching resources.
Smart Images

Figure CN121330748A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of computer vision and adaptive education technology, specifically involving a method and system for identifying and providing feedback on the quantification of student learning status, teaching decision support, and optimization of teaching resources in a business administration teaching scenario based on machine vision. Background Technology
[0002] In modern business administration education, the impact of classroom interaction (such as business case discussions and corporate decision-making simulations) and student learning status on teaching effectiveness is receiving increasing attention, especially in management aspects such as teaching quality assessment, course content iteration, and resource allocation for business administration courses. Students' emotional responses (such as confusion about business cases and engagement in decision-making simulations) and participation are core indicators for measuring the effectiveness of teaching management. However, two major problems exist in current business administration teaching management: First, it is difficult to quantify students' actual learning status in business administration-specific teaching scenarios (such as complex business case analysis and cross-departmental collaboration simulation). The traditional method of relying on teachers' subjective observation cannot provide objective data for management decisions such as whether to increase class hours for a certain type of business case or whether to optimize the decision-making simulation module. Secondly, existing technologies cannot deeply integrate student emotional data with the management objectives of business administration teaching (such as evaluating the effectiveness of cultivating students' business thinking and analyzing the satisfaction of course modules), resulting in a lack of targeted adjustments to teaching content and resource allocation, making it difficult to meet the refined needs of business administration teaching that is "guided by management decisions".
[0003] Existing technology: A smart classroom teaching system based on the Internet of Things (IoT), as described in CN202410817148.8, monitors students' physiological data such as heart rate and integrates multiple terminals to achieve teaching management and adjustment. However, this technology cannot capture subtle emotional features such as facial micro-expressions, making it difficult to meet the real-time and accurate emotional response requirements in classroom discussions. Especially in dynamic interactive scenarios, it lacks support for the fusion and analysis of emotional and behavioral data, resulting in insufficient targeted adaptation of teaching content.
[0004] The above solutions cannot capture subtle emotional features such as facial micro-expressions, making it difficult to accurately judge students' level of understanding and engagement in business administration courses (such as discussions of controversial points in business cases and risk decision analysis). They also cannot provide data for business administration-specific teaching and management needs such as adjusting the difficulty of business cases and optimizing decision simulation processes. The design lacks connection with the teaching and management goals of business administration, and only stays at the basic logic of monitoring physiological data and adjusting teaching content. It does not involve the management loop of student learning status, course module evaluation, and teaching resource allocation. It cannot provide business administration teaching managers with decision support on which course modules need to be strengthened and which student groups need targeted tutoring, which does not meet the core requirements of management decision support.
[0005] Therefore, there is an urgent need for a technical solution that can balance the accuracy of emotional data and real-time feedback in a dynamic classroom environment, quantify students' learning status through technical means, and directly serve the decision-making of business administration teaching management, filling the gap in existing technology in terms of data support for business administration teaching management. Summary of the Invention
[0006] This invention discloses a machine vision-based method and system for recognizing and responding to teaching models in business administration, which can effectively solve the technical problem that existing teaching systems lack the ability to perceive and respond to unstructured and weak-signal student states.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A machine vision-based method for recognizing and providing feedback on business administration teaching models includes the following steps: Step 1: Parallel image sequence acquisition of students' faces is carried out using a dual-camera image acquisition device. The wide-angle camera continuously captures the overall expression trend, while the telephoto camera specifically captures the micro-expression changes in key areas such as the eyes and mouth. The two video streams are synchronized through timestamp alignment to obtain raw image data containing emotional feature extraction and determine the initial data granularity layering. Step 2: Based on the original image data, sharpening processing is performed on the facial areas captured in local detail. An edge enhancement method based on the Laplacian operator is adopted, and the sharpening parameter intensity factor is set to 1.2-1.8 to generate enhanced and refined image data. The image quality score is calculated by the peak signal-to-noise ratio (PSNR) algorithm, and a threshold of 30dB is set to determine whether the enhanced and refined image data meets the accuracy requirements for emotional feature extraction. Step 3: If the enhanced and refined image data reaches a preset threshold, then subtle emotional response features are obtained through emotional feature extraction, and the subtle emotional response features are associated with the emotional response mapping in the digital interaction system to obtain a preliminary emotional response distribution. Step 4: Based on the preliminary distribution of emotional responses, combined with the student participation records and interaction frequency analysis in the digital interaction system, generate real-time feedback transmission data to determine the emotional participation level in the current discussion session; Step 5: Through the real-time feedback transmission data, based on the emotional fluctuations obtained from the overall trend analysis, integrate the discussion data integration function to generate teaching content adaptation suggestions under multi-level content recognition, and determine whether the teaching content adaptation suggestions meet the personalized guidance needs. Step 6: If the teaching content adaptation suggestion does not meet the preset emotional engagement standard, the image acquisition parameters are adjusted cyclically, and new image sequence data is re-acquired through digital zoom and physical position adjustment. Combined with the digital interaction system to update the student participation record, the optimized emotional response distribution is obtained. Step 7: Based on the optimized emotional response distribution, integrate the results of the multi-level content recognition and the real-time feedback transmission to generate a personalized feedback report that combines students' micro-expressions with discussion and interaction in education and teaching, and determine the final emotional data support.
[0008] In one embodiment of the present invention, step 1 specifically includes: The student's face is continuously acquired using an image acquisition device. A focus switching mechanism is used to acquire multi-level image data in different modes to obtain a preliminary set of raw images. Based on the original image set, local detail segmentation is performed on the student's facial region to separate key regions containing emotional features and determine the distribution range of emotion-related regions. By initially extracting the emotional features of key region images and matching them using preset feature templates, the initial classification results of the emotional features are obtained. If the confidence level of the initial classification result is lower than the preset threshold, the key area image is re-acquired after a second focus adjustment to obtain a clearer local detail image. Based on the re-acquired local detail images and combined with the overall dynamic image data, in-depth analysis of emotional features is performed. A convolutional neural network model is used to refine the features and determine the emotional category.
[0009] By analyzing the results of emotion category judgments and combining them with data granularity hierarchy, the classification results are mapped to the original image set to obtain the final emotion feature distribution data. Based on the emotion feature distribution data, a time series analysis is performed on the overall dynamic changes of students' faces to determine the trends and patterns of emotion changes.
[0010] In one embodiment of the present invention, step 2 specifically includes: By performing layered processing on the original image data, a preliminary sharpening enhancement operation is performed on the local fine facial areas to obtain the first enhanced image set; Based on the enhanced first image set, a preset sharpening tool is used to further optimize the details of the facial area, generating a clearer second image set; If the detail clarity of the second image set does not meet the accuracy standard for emotion feature extraction, then targeted data augmentation processing is performed on the facial region to obtain an optimized third image set. By applying a convolutional neural network model to facial region data from a third image set, preliminary classification of emotional features is performed to determine the category distribution of emotional features. Based on the category distribution of emotional features, for the data in the classification results where the confidence level is lower than the preset threshold standard, local image detail comparison processing is performed to obtain more accurate feature classification data; If the accuracy of the feature classification data still does not reach the preset accuracy standard, the classification data will be filtered a second time to obtain the final refined sentiment feature data. By dynamically comparing refined emotional feature data over time, we can determine the emotional change patterns of facial regions at different points in time and identify the continuity characteristics of emotional changes.
[0011] In one embodiment of the present invention, step 3 specifically includes: If the enhanced and refined image data reaches the preset threshold standard, the refined image data will be processed by a pre-established emotional feature extraction tool to obtain subtle emotional response data. Based on the acquired subtle emotional response data, a pre-set mapping table is used to compare the response data with the emotional response categories in the digital interaction system to obtain preliminary emotional response distribution results. Based on the preliminary distribution results of emotional responses, data comparison tools were used to group the emotional response categories in the distribution results to determine the main distribution range of each emotional response category. If the distribution range of a certain type of emotional response overlaps with other categories, cluster analysis tools can be used to further subdivide the data in the overlapping areas to obtain clearer category classification results. Based on the subdivision results, a data filtering mechanism is used to prioritize and sort the emotional response data of each category, resulting in a sorted emotional response data set. For the sorted sentiment response dataset, high-priority sentiment response data are highlighted using data integration tools to determine the final core sentiment response data; Based on the final core data of emotional response, a record-keeping mechanism is used to update and compare the core data with the mapping table in the digital interaction system to obtain the latest emotional response mapping results.
[0012] In one embodiment of the present invention, step 4 specifically includes: By acquiring emotional response data and distribution data through digital interactive platforms, analyzing the statistical records of student participation, and obtaining preliminary emotional response patterns; Based on the preliminary emotional response patterns and the data on interaction frequency, statistical tools are used to calculate the activity index in the discussion session to determine the current participation status in the discussion. If the current level of participation in the discussion is below a preset threshold, a real-time feedback mechanism is triggered to generate transmission information and determine the areas where emotional participation is insufficient. By analyzing the content of the transmitted information, we can identify areas where emotional engagement is insufficient, obtain specific interaction records from the discussion sessions, and pinpoint the weaknesses in student participation. Based on the weaknesses in student participation and the distribution data of emotional responses, the fluctuations in interaction frequency were analyzed to obtain the changing trend of emotional participation. By analyzing the changing trends in emotional engagement, dynamic adjustment strategies can be generated for the current discussion segment to determine the direction for improving subsequent interactions. By acquiring the output of the dynamic adjustment strategy and combining it with the transmission information from real-time feedback, we can analyze the room for improvement in the discussion process and determine the final optimization plan for emotional participation.
[0013] In one embodiment of the present invention, step 5 specifically includes: By collecting data through a real-time feedback mechanism, emotional fluctuation data is obtained from user interactions. A pre-established emotion recognition model is used for preliminary processing to obtain the classification results of emotional states. Based on the classification results of emotional states and the analysis results of overall trends, the long-term change characteristics reflected in data transmission are extracted to determine the emotional tendencies of users in different time periods. If the sentiment trend shows significant fluctuations, then we will conduct in-depth analysis of the integrated discussion data to obtain details of user feedback in specific scenarios and determine potential sentiment drivers. By using multi-layer recognition technology, emotional driving factors are associated and matched with teaching content to generate a preliminary content adaptation plan and obtain teaching adjustment directions that are consistent with the user's emotional state. If the initial content adaptation plan deviates from the goal of personalized guidance, the plan will be optimized a second time based on the logical rules of demand judgment to determine the final adaptation result. Based on the final adaptation results, the output module generated by the linkage suggestion presents the adjustment direction of the teaching content in a structured form, and obtains a personalized guidance plan for users. Through the continuous processing of the above steps, multiple attributes such as real-time feedback, emotional fluctuations, and overall trends are integrated into the business process to achieve the adaptation of teaching content and the realization of personalized guidance.
[0014] In one embodiment of the present invention, step 6 specifically includes: By linking teaching content with adaptation suggestions, initial emotional engagement data is obtained, and a pre-established emotional assessment model is used to determine whether emotional engagement meets the preset engagement standards. If emotional engagement does not meet the preset engagement standard, the image acquisition parameters are adjusted, and the image sequence data is reacquired through digital zoom and camera physical position adjustment to obtain clearer sequence data content. Based on the newly acquired image sequence data, and combined with the correlation between the sequence data and digital interaction, the data is synchronized through the digital interaction system to update students' participation records and determine whether the participation records reflect emotional changes. If the participation records reflect emotional changes, the distribution features of emotional responses are extracted by associating the interaction system with the participation records. The support vector machine algorithm is then used to classify the response distribution to obtain the classified emotional response results. Based on the classified emotional response results, and combined with the correlation between emotional participation and response distribution, the dynamic trend of emotional participation is analyzed to determine the optimization direction of emotional participation; By optimizing the direction of emotional engagement, combining the connection between teaching content and digital interaction, adjusting the specific parameters of the adaptation suggestions, obtaining the updated teaching content adaptation scheme, and determining whether it meets the emotional engagement standards; If the updated adaptation scheme still does not meet the emotional participation standard, the acquisition method will be optimized by adjusting the image acquisition parameters in a loop to obtain more accurate sequence data and determine the final emotional response distribution.
[0015] In one embodiment of the present invention, step 7 specifically includes: By acquiring raw image data of students' micro-expressions from classroom video streams, and using a pre-established image processing module for preliminary cleaning, processed facial feature images are obtained. Based on the processed facial feature images, the support vector machine algorithm is used to classify students' micro-expressions and determine the corresponding emotional response categories. Based on the emotional response categories obtained from the classification, audio segments of discussions and interactions are obtained from real-time feedback transmissions, and the text content is extracted using a speech-to-text tool to obtain interactive text records; If the interactive text record contains preset emotional keywords, then a matching analysis is performed based on the emotional response category to determine the students' emotional tendencies in the discussion interaction; By continuously monitoring emotional tendencies, we can obtain the trend of emotional response changes over a period of time and generate a dynamic distribution chart of emotional data. Based on the dynamic distribution chart and the specific requirements of the teaching and learning scenario, personalized feedback content is generated for each student to determine the final emotional data support basis. By integrating personalized feedback content with sentiment data support, a structured feedback report document is automatically generated for subsequent analysis and reference.
[0016] Furthermore, this invention also discloses a machine vision-based business administration teaching model recognition and feedback system, used to execute the aforementioned machine vision-based business administration teaching model recognition and feedback method, characterized in that it includes: The image acquisition device is used to acquire parallel image sequences of students' faces, and simultaneously captures the overall expression trend and local micro-expression changes through parallel wide-angle and telephoto cameras to obtain raw image data containing emotional feature extraction. The image processing module is communicatively connected to the image acquisition device and is used to perform sharpening processing on the facial areas captured in local detail based on the original image data, generate enhanced and refined image data, and determine whether the enhanced and refined image data meets the accuracy requirements for emotional feature extraction. The sentiment analysis module, which is communicatively connected to the image processing module, is used to extract subtle emotional response features by means of emotional feature extraction when the enhanced and refined image data reaches a preset threshold, and associate the subtle emotional response features with the emotional response mapping in the digital interaction system to obtain a preliminary emotional response distribution. The real-time feedback module is connected to the sentiment analysis module and is used to transmit data through real-time feedback. Based on the sentiment fluctuations obtained from the overall trend analysis, it integrates the discussion data integration function to generate teaching content adaptation suggestions under multi-level content recognition and determines whether the teaching content adaptation suggestions meet the personalized guidance needs. The closed-loop control module is communicatively connected to the real-time feedback module and the image acquisition device. When the teaching content adaptation suggestion does not meet the preset emotional participation standard, it cyclically adjusts the image acquisition parameters, controls the image acquisition device to re-acquire new image sequence data, and updates the student participation record in conjunction with the digital interactive system to obtain the optimized emotional response distribution. The report generation module, which is communicatively connected to the closed-loop control module and the sentiment analysis module, is used to generate a personalized feedback report that combines students' micro-expressions with discussion and interaction in education and teaching, based on the optimized sentiment response distribution and integrating the results of the multi-level content recognition and the real-time feedback transmission.
[0017] Furthermore, the image acquisition device includes a parallel wide-angle camera and a telephoto camera, an image synchronization unit, and a quality control unit. The quality control unit is configured to dynamically adjust the acquisition parameters based on an image quality assessment algorithm, or to control the dual-camera collaborative working mode in response to emotion recognition confidence commands.
[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention aims to address the challenge of capturing and optimizing student emotional engagement in real-time during digital interactive teaching scenarios, leading to inaccurate content adaptation and insufficient personalized guidance. The invention utilizes an image acquisition device with a focus-switching mechanism to adjust the acquisition mode in real-time, acquiring raw image data containing emotional features. This data is then sharpened to generate enhanced images. If the accuracy requirements are met, subtle emotional response features are extracted and correlated with the emotional response mapping mechanism to generate a preliminary emotional response distribution. Combined with student participation records and interaction frequency analysis, real-time feedback data is generated, and multi-level content recognition is integrated to produce teaching content adaptation suggestions. If the standards are not met, the focus is cyclically adjusted for re-acquisition to optimize the emotional distribution. Finally, the results are integrated to generate a personalized feedback report. This invention improves the adaptability of teaching content and the support for emotional data, achieving efficient optimization of educational interaction. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 This is a flowchart of the image acquisition and preliminary processing of the present invention.
[0021] Figure 2 This is a flowchart of the image enhancement and quality judgment process of the present invention.
[0022] Figure 3 Flowchart for adapting and optimizing the teaching content of this invention. Detailed Implementation
[0023] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0024] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings. Example
[0025] The technical solution of this embodiment is effective under the following boundary conditions: classroom light intensity of 100-1000 lux, distance between student and camera of 1-5 meters, angle between face and camera of no more than ±30 degrees, and single continuous teaching time of no more than 90 minutes.
[0026] See Figures 1-3 This embodiment discloses a machine vision-based method for recognizing and providing feedback on business administration teaching models, specifically including the following steps: Step 1: Parallel image sequence acquisition of students' faces is carried out using a dual-camera image acquisition device. The wide-angle camera continuously captures the overall expression trend, while the telephoto camera specifically captures the micro-expression changes in key areas such as the eyes and mouth. The two video streams are synchronized through timestamp alignment to obtain raw image data containing emotional feature extraction and determine the initial data granularity layering. Step 2: Based on the original image data, sharpening processing is performed on the facial areas captured in local detail. An edge enhancement method based on the Laplacian operator is adopted, and the sharpening parameter intensity factor is set to 1.2-1.8 to generate enhanced and refined image data. The image quality score is calculated by the peak signal-to-noise ratio (PSNR) algorithm, and a threshold of 30dB is set to determine whether the enhanced and refined image data meets the accuracy requirements for emotional feature extraction. Step 3: If the enhanced and refined image data reaches a preset threshold, then subtle emotional response features are obtained through emotional feature extraction, and the subtle emotional response features are associated with the emotional response mapping in the digital interaction system to obtain a preliminary emotional response distribution. Step 4: Based on the preliminary distribution of emotional responses, combined with the student participation records and interaction frequency analysis in the digital interaction system, generate real-time feedback transmission data to determine the emotional participation level in the current discussion session; Step 5: Through the real-time feedback transmission data, based on the emotional fluctuations obtained from the overall trend analysis, integrate the discussion data integration function to generate teaching content adaptation suggestions under multi-level content recognition, and determine whether the teaching content adaptation suggestions meet the personalized guidance needs. Step 6: If the teaching content adaptation suggestion does not meet the preset emotional engagement standard, the image acquisition parameters are adjusted cyclically, and new image sequence data is re-acquired through digital zoom and physical position adjustment. Combined with the digital interaction system to update the student participation record, the optimized emotional response distribution is obtained. Step 7: Based on the optimized emotional response distribution, integrate the results of the multi-level content recognition and the real-time feedback transmission to generate a personalized feedback report that combines students' micro-expressions with discussion and interaction in education and teaching, and determine the final emotional data support.
[0027] Furthermore, step 1 specifically includes: The student's face is continuously acquired using an image acquisition device. A focus switching mechanism is used to acquire multi-level image data in different modes to obtain a preliminary set of raw images. Based on the original image set, local detail segmentation is performed on the student's facial region to separate key regions containing emotional features and determine the distribution range of emotion-related regions. By initially extracting the emotional features of key region images and matching them using preset feature templates, the initial classification results of the emotional features are obtained. If the confidence level of the initial classification result is lower than the preset threshold, the key area image is re-acquired after a second focus adjustment to obtain a clearer local detail image. Based on the re-acquired local detail images and combined with the overall dynamic image data, in-depth analysis of emotional features is performed. A convolutional neural network model is used to refine the features and determine the emotional category.
[0028] By analyzing the results of emotion category judgments and combining them with data granularity hierarchy, the classification results are mapped to the original image set to obtain the final emotion feature distribution data. Based on the emotion feature distribution data, a time series analysis is performed on the overall dynamic changes of students' faces to determine the trends and patterns of emotion changes.
[0029] In specific implementation, for the technical implementation of student facial image sequence acquisition and emotional feature extraction, a dual-camera parallel acquisition system is first used. The total processing capacity of the dual-camera system on the embedded device is 10 frames per second. The wide-angle camera is allocated 6 frames for overall trend analysis, and the telephoto camera is allocated 4 frames for micro-expression detail capture. The allocation of computing resources is optimized through frame scheduling algorithm, and the dual video streams are time-stamp aligned through hardware synchronization signal to ensure the comprehensiveness and real-time nature of data acquisition.
[0030] The acquired image sequences were processed using a deep learning model. A resolution of 640×480 pixels was used for face detection, while a resolution of 1280×720 pixels was used for the ROI region in the micro-expression analysis stage. The lightweight convolutional neural network MobileNetV3 was used to replace ResNet-50 to balance accuracy and speed. 68 key points on the face were located, and feature parameters such as the upward angle of eyebrows (range 0-25 degrees) and the upward angle of the corners of the mouth (range -10 to +10 degrees) were calculated to generate emotion vector data. The recognition accuracy of six basic emotions was 70-75% in a controlled laboratory environment.
[0031] The initial data granularity was determined, and the emotional feature data was divided into three granularity levels by time window aggregation: real-time level (one data point every 2 seconds, for immediate feedback), conversation level (aggregated once every 300 seconds, for single course analysis), and long-term level (aggregated once every 3600 seconds, for learning behavior assessment). The sliding window mean (window size 5-10 samples) and coefficient of variation of the emotional intensity of each level were calculated to support subsequent multi-dimensional analysis.
[0032] The above steps form a closed-loop logic through automated system processing, and are also linked with classroom interaction data (such as students answering twice per minute) to build a correlation model between emotion and behavior, ensuring the comprehensiveness and scientific nature of the analysis results.
[0033] Furthermore, step 2 specifically includes: By performing layered processing on the original image data, a preliminary sharpening enhancement operation is performed on the local fine facial areas to obtain the first enhanced image set; Based on the enhanced first image set, a preset sharpening tool is used to further optimize the details of the facial area, generating a clearer second image set; If the detail clarity of the second image set does not meet the accuracy standard for emotion feature extraction, then targeted data augmentation processing is performed on the facial region to obtain an optimized third image set. By applying a convolutional neural network model to facial region data from a third image set, preliminary classification of emotional features is performed to determine the category distribution of emotional features. Based on the category distribution of emotional features, for the data in the classification results where the confidence level is lower than the preset threshold standard, local image detail comparison processing is performed to obtain more accurate feature classification data; If the accuracy of the feature classification data still does not reach the preset accuracy standard, the classification data will be filtered a second time to obtain the final refined sentiment feature data. By dynamically comparing refined emotional feature data over time, we can determine the emotional change patterns of facial regions at different points in time and identify the continuity characteristics of emotional changes.
[0034] In practice, for the subtle facial regions obtained from the original image data, the key regions are first sharpened using an image enhancement algorithm. Specifically, an edge enhancement method based on the Laplacian operator is adopted. When processing the key regions, a resolution of 1280×720 pixels is used, and the sharpening parameter intensity factor is dynamically adjusted according to the image quality (range 1.2-1.8) to highlight the subtle texture changes in areas such as the eyes and corners of the mouth. The processed image structure similarity index (SSIM) is improved to above 0.8, thus providing a reliable visual basis for subsequent analysis.
[0035] The enhanced and refined image data is automatically evaluated for quality. The peak signal-to-noise ratio (PSNR) algorithm is used to calculate the image quality score. The threshold is set at 30dB. If the score is lower than this value, the adaptive histogram equalization (CLAHE) algorithm is used for further optimization. The contrast gain is limited to 2.0 to ensure that the image details are not distorted. The processed PSNR value needs to be recalculated and meet the threshold requirement.
[0036] To determine whether the enhanced image data meets the accuracy requirements for emotion feature extraction, the system introduces an evaluation mechanism based on the structural similarity index (SSIM), with a target value of 0.85 or higher. Simultaneously, it combines image entropy analysis, requiring an entropy value greater than 6.5 to ensure sufficient information. If the target is not met, a secondary sharpening process is automatically triggered, adjusting the sharpening factor to 1.8 and re-evaluating until the conditions are met.
[0037] The evaluation results are correlated with classroom environment data (such as light intensity of 500 lux per square meter). If insufficient lighting leads to a decrease in image quality, the exposure parameters of the image acquisition device are automatically adjusted to +0.5 EV to ensure data stability. This forms a closed-loop processing logic from sharpening to evaluation to optimization.
[0038] In one embodiment of the present invention, step 3 specifically includes: If the enhanced and refined image data reaches the preset threshold standard, the refined image data will be processed by a pre-established emotional feature extraction tool to obtain subtle emotional response data. Based on the acquired subtle emotional response data, a pre-set mapping table is used to compare the response data with the emotional response categories in the digital interaction system to obtain preliminary emotional response distribution results. Based on the preliminary distribution results of emotional responses, data comparison tools were used to group the emotional response categories in the distribution results to determine the main distribution range of each emotional response category. If the distribution range of a certain type of emotional response overlaps with other categories, cluster analysis tools can be used to further subdivide the data in the overlapping areas to obtain clearer category classification results. Based on the subdivision results, a data filtering mechanism is used to prioritize and sort the emotional response data of each category, resulting in a sorted emotional response data set. For the sorted sentiment response dataset, high-priority sentiment response data are highlighted using data integration tools to determine the final core sentiment response data; Based on the final core data of emotional response, a record-keeping mechanism is used to update and compare the core data with the mapping table in the digital interaction system to obtain the latest emotional response mapping results.
[0039] In practice, once the enhanced and refined image data reaches the preset quality threshold (PSNR > 30dB), the system automatically initiates the emotion feature extraction process. Using the lightweight convolutional neural network MobileNetV3 architecture, feature vectors are extracted from key facial regions. The input image size is set to 224x224 pixels to meet network input requirements, and the extracted feature dimension is 128 to reduce computational complexity. The system focuses on analyzing subtle changes that conform to facial anatomy constraints, such as eyebrow tilt angle (range -12 degrees to 12 degrees) and mouth corner curvature (curvature value 0.3 to 0.7), to generate an emotion response feature set.
[0040] The extracted feature set is compared with the emotion response mapping database in the digital interaction system. The method for constructing the emotion response mapping relationship includes: collecting annotated facial micro-expression datasets, determining six basic emotion categories (joy, surprise, disgust, anger, fear, and sadness) through expert annotation, learning the emotion category boundaries in the feature space using a metric learning algorithm, establishing a mapping function from feature vectors to emotion categories, optimizing the mapping threshold through cross-validation, calculating the matching degree between feature vectors and database templates using a cosine similarity algorithm, setting the similarity threshold to 0.6, and automatically recording "neutral" emotion if the matching degree is lower than the threshold. The classification reliability is improved through a multi-frame voting mechanism (at least 2 out of 3 consecutive frames are consistent). The classification accuracy is expected to be 60-70% in a real classroom environment.
[0041] A preliminary emotional response distribution map is generated, which outputs the probability of each emotion as a percentage, such as joy at 40% and surprise at 30%. The results are then correlated with the user's historical emotional data to calculate the emotional volatility (with the standard deviation controlled within 0.3). If the volatility is abnormal, the system log recording module is triggered to automatically save the current feature data and environmental variables (such as timestamps and device IDs) for subsequent traceability, forming a complete automated logical chain from feature extraction to distribution mapping to anomaly monitoring.
[0042] Furthermore, step 4 specifically includes: By acquiring emotional response data and distribution data through digital interactive platforms, analyzing the statistical records of student participation, and obtaining preliminary emotional response patterns; Based on the preliminary emotional response patterns and the data on interaction frequency, statistical tools are used to calculate the activity index in the discussion session to determine the current participation status in the discussion. If the current level of participation in the discussion is below a preset threshold, a real-time feedback mechanism is triggered to generate transmission information and determine the areas where emotional participation is insufficient. By analyzing the content of the transmitted information, we can identify areas where emotional engagement is insufficient, obtain specific interaction records from the discussion sessions, and pinpoint the weaknesses in student participation. Based on the weaknesses in student participation and the distribution data of emotional responses, the fluctuations in interaction frequency were analyzed to obtain the changing trend of emotional participation. By analyzing the changing trends in emotional engagement, dynamic adjustment strategies can be generated for the current discussion segment to determine the direction for improving subsequent interactions. By acquiring the output of the dynamic adjustment strategy and combining it with the transmission information from real-time feedback, we can analyze the room for improvement in the discussion process and determine the final optimization plan for emotional participation.
[0043] In practical implementation, when constructing real-time feedback transmission data to determine the emotional participation level in the current discussion session, the system collects data on the distribution of students' emotional responses through a digital interaction system. The system records the distribution of students' emotional feedback during the discussion session and calculates an emotional index using time weighting, with higher weights for more recent time periods (current minute 0.6, previous minute 0.3, earlier time period 0.1). The emotional index calculation formula is: Σ(emotional category weight × time decay coefficient × confidence level). The emotional participation threshold is set within an adjustable range of 0.3-0.7. Combining student participation records and interaction frequency analysis, a total of 200 student speeches were extracted in the past 10 minutes, averaging 2 speeches per student. High-frequency participants (speaking more than 5 times) accounted for 10%, and the interaction activity level was calculated as 200 / 100 / 10 = 0.2, reflecting an uneven distribution of interaction.
[0044] The emotional index and interaction activity are analyzed comprehensively using the following formula: Emotional engagement = Emotional Index * 0.6 + Interaction Activity * 0.4, resulting in an emotional engagement score of 0.45 * 0.6 + 0.2 * 0.4 = 0.35, indicating that the emotional engagement in the current discussion session is moderately low.
[0045] The system automatically generates real-time feedback data, including an emotion index of 0.45, an interaction activity level of 0.2, and an emotion participation level of 0.35, and transmits this data to the teaching analysis module via a data interface.
[0046] To form a logical closed loop, the course content difficulty coefficient (assumed to be 0.7) was also linked. By comparing with historical data (average emotional engagement of 0.5), it was inferred that the current low engagement might be related to the difficulty of the content, which automatically triggered the generation of adjustment suggestion data, such as reducing the difficulty of the explanation or increasing interactive elements. In the end, a complete feedback chain was formed to ensure that data analysis and business improvement were seamlessly connected.
[0047] Furthermore, step 5 specifically includes: By collecting data through a real-time feedback mechanism, emotional fluctuation data is obtained from user interactions. A pre-established emotion recognition model is used for preliminary processing to obtain the classification results of emotional states. Based on the classification results of emotional states and the analysis results of overall trends, the long-term change characteristics reflected in data transmission are extracted to determine the emotional tendencies of users in different time periods. If the sentiment trend shows significant fluctuations, then we will conduct in-depth analysis of the integrated discussion data to obtain details of user feedback in specific scenarios and determine potential sentiment drivers. By using multi-layer recognition technology, emotional driving factors are associated and matched with teaching content to generate a preliminary content adaptation plan and obtain teaching adjustment directions that are consistent with the user's emotional state. If the initial content adaptation plan deviates from the goal of personalized guidance, the plan will be optimized a second time based on the logical rules of demand judgment to determine the final adaptation result. Based on the final adaptation results, the output module generated by the linkage suggestion presents the adjustment direction of the teaching content in a structured form, and obtains a personalized guidance plan for users. Through the continuous processing of the above steps, multiple attributes such as real-time feedback, emotional fluctuations, and overall trends are integrated into the business process to achieve the adaptation of teaching content and the realization of personalized guidance.
[0048] In practice, by transmitting data in real time, the system can use sensors and online learning platforms to collect students' emotional data during the learning process. For example, by using voice tone analysis and facial expression recognition technology, the system can detect students' emotional fluctuations in class. Suppose that during a 10-minute class, the system records that a student's emotion changes from positive (score 0.8) to negative (score 0.3). The data transmission frequency is once per second, totaling 600 data points. Using a time series analysis algorithm to calculate the emotional fluctuation trend, the system finds that the slope of the emotion decline is -0.005 / second, indicating that the emotion is gradually decreasing.
[0049] Next, based on the overall trend analysis of emotional fluctuations, clustering algorithms (such as K-means) were used to classify students' emotional fluctuations into three categories: stable, slight fluctuations, and severe fluctuations. Assuming that the student was classified as a severe fluctuation student, combined with historical data analysis, the fluctuations were found to be related to the difficulty of the course. When the difficulty index rose from 3.5 to 4.2, the emotional fluctuations decreased significantly.
[0050] Then, by integrating the discussion data, keywords such as "confused" and "does not understand" are extracted from the classroom discussion text using natural language processing (NLP) technology. The frequency is 5 times per minute. Combined with sentiment data, it is determined that students' understanding of the current content is only 40%.
[0051] Subsequently, teaching content adaptation suggestions are generated under multi-level content recognition. Based on the above analysis, the content database is called to match simplified teaching resources with a difficulty index of 3.0, and personalized practice questions are generated. The number of questions is reduced by 30%, from 10 questions to 7 questions, to ensure that the learning pressure is reduced.
[0052] Finally, the system determines whether the teaching content matching suggestions meet the personalized guidance needs. It evaluates the suggestions based on preset rules (such as a difficulty matching degree greater than 80%) and user history feedback (a satisfaction score greater than 0.7). The current suggestion matching degree is calculated to be 85%, and the predicted satisfaction score is 0.75, which meets the requirements. If it does not meet the requirements, the system will automatically adjust the resource difficulty to 3.2 and recalculate until the conditions are met.
[0053] The entire process forms a closed-loop logic, with emotion data and discussion data complementing each other. Content adaptation and demand judgment are automatically iterated through algorithms to ensure optimal teaching results.
[0054] Furthermore, step 6 specifically includes: By linking teaching content with adaptation suggestions, initial emotional engagement data is obtained, and a pre-established emotional assessment model is used to determine whether emotional engagement meets the preset engagement standards. If emotional engagement does not meet the preset engagement standard, the image acquisition parameters are adjusted, and the image sequence data is reacquired through digital zoom and camera physical position adjustment to obtain clearer sequence data content. Based on the newly acquired image sequence data, and combined with the correlation between the sequence data and digital interaction, the data is synchronized through the digital interaction system to update students' participation records and determine whether the participation records reflect emotional changes. If the participation records reflect emotional changes, the distribution features of emotional responses are extracted by associating the interaction system with the participation records. The support vector machine algorithm is then used to classify the response distribution to obtain the classified emotional response results. Based on the classified emotional response results, and combined with the correlation between emotional participation and response distribution, the dynamic trend of emotional participation is analyzed to determine the optimization direction of emotional participation; By optimizing the direction of emotional engagement, combining the connection between teaching content and digital interaction, adjusting the specific parameters of the adaptation suggestions, obtaining the updated teaching content adaptation scheme, and determining whether it meets the emotional engagement standards; If the updated adaptation scheme still does not meet the emotional participation standard, the acquisition method will be optimized by adjusting the image acquisition parameters in a loop to obtain more accurate sequence data and determine the final emotional response distribution.
[0055] In practice, when the teaching content adaptation suggestion fails to meet the preset emotional engagement standard, the image quality optimization process is automatically initiated. First, the region of interest is magnified by 1.5-2 times using digital zoom. When the image quality score (based on the BRISQUE no-reference image quality assessment algorithm, with a score range of 0-100, where a higher score indicates lower quality) is higher than 40, a command to adjust the physical position of the camera is triggered, with an adjustment step size of 10-20 cm. Image sequence data is then re-acquired at a frequency of 5-10 frames per second for 20-40 seconds. The above data is used to extract features from a pre-trained convolutional neural network model to identify the student's emotional state, such as focus and confusion, and output the emotional classification probability, for example, focus is 0.65, confusion is 0.25, and boredom is 0.10.
[0056] Data exchange is conducted with the digital interaction system through a defined standard RESTful API interface. The specific implementation of the digital interaction system interface includes: defining a unified student identifier management mechanism, establishing a real-time data stream processing pipeline, setting a data synchronization time window (default 5 seconds) and a retry mechanism (maximum 3 times) to ensure the temporal consistency of emotional data and interaction records. The data format adopts JSON Schema and includes fields such as timestamp, anonymous student ID, emotional feature vector, and confidence level to ensure interoperability between systems. Student participation records are updated, and the collected emotional data is compared and analyzed with historical interaction data. The average emotional participation score of the current class is calculated using a weighted average algorithm. Assuming the historical score is 0.70 and the current score is 0.68, with weighting coefficients of 0.6 and 0.4, the final score is updated to 0.692. If it is still lower than the preset standard of 0.75, a new round of data collection parameter adjustment is triggered, increasing the magnification through digital zoom and repeating the above process.
[0057] Meanwhile, based on the emotional response distribution optimization algorithm, the K-means clustering method is used to divide students' emotional responses into three categories. The emotional intensity value of the center point of each category is calculated. For example, the center of the focused category is 0.72, and the center of the confused category is 0.28. An optimized emotional response distribution map is generated. Combined with the teaching content database, the teaching speed is automatically adjusted. For example, the speaking speed is reduced from 120 words per minute to 100 words per minute to improve students' comprehension. This forms a closed-loop optimization logic to ensure that the emotional participation gradually approaches the standard value.
[0058] Furthermore, step 7 specifically includes: By acquiring raw image data of students' micro-expressions from classroom video streams, and using a pre-established image processing module for preliminary cleaning, processed facial feature images are obtained. Based on the processed facial feature images, the support vector machine algorithm is used to classify students' micro-expressions and determine the corresponding emotional response categories. Based on the emotional response categories obtained from the classification, audio segments of discussions and interactions are obtained from real-time feedback transmissions, and the text content is extracted using a speech-to-text tool to obtain interactive text records; If the interactive text record contains preset emotional keywords, then a matching analysis is performed based on the emotional response category to determine the students' emotional tendencies in the discussion interaction; By continuously monitoring emotional tendencies, we can obtain the trend of emotional response changes over a period of time and generate a dynamic distribution chart of emotional data. Based on the dynamic distribution chart and the specific requirements of the teaching and learning scenario, personalized feedback content is generated for each student to determine the final emotional data support basis. By integrating personalized feedback content with sentiment data support, a structured feedback report document is automatically generated for subsequent analysis and reference.
[0059] In practical implementation, in education and teaching, for the generation of personalized feedback reports combining students' micro-expressions and discussion interactions, the system first uses an optimized emotional response distribution and integrates visual emotion recognition and voice / text emotion analysis using an ensemble learning approach. Assuming that the system collects data from multiple modalities in a lesson, a comprehensive emotional assessment is generated through time alignment and confidence weighting. When the visual emotion weight is greater than 0.7, visual analysis takes precedence; otherwise, voice / text analysis takes precedence. The weights for focus are 0.5, confusion 0.3, and boredom 0.2. The final emotional score is (focus × 0.6 + confusion × 0.3 + boredom × 0.1) × 100 = (0.6 × 0.6 + 0.3 × 0.3 + 0.1 × 0.1) × 100 = 46 points, which is lower than the preset threshold of 50 points, indicating that the teaching content may need adjustment.
[0060] By integrating multi-level content recognition results, the voice data of classroom discussions and interactions were analyzed using natural language processing (NLP) technology to extract keyword frequencies. For example, "question" appeared 20 times and "understanding" appeared 10 times. The discussion participation index was calculated as (20+10) / total number of speeches 50=0.6. Combined with micro-expression emotion scores, the comprehensive interaction quality score was obtained as (46*0.4+0.6*60)=54.4 points, reflecting that the students' participation was acceptable but the emotional feedback was low.
[0061] Through a real-time feedback transmission mechanism, the above data is updated to the cloud analysis platform every 5 minutes. Using time series analysis algorithms, the trend of emotional fluctuations is predicted. Assuming that the emotional scores of the past three 5-minute data segments are 40, 42, and 41, the predicted score for the next segment is 40.5, prompting teachers to pay attention to the persistently low mood.
[0062] Finally, a personalized feedback report is generated. The system automatically generates suggestions based on the emotional data (overall score of 54.4, predicted value of 40.5), such as adding interactive elements or adjusting the difficulty of the explanation. The report specifically points out that the emotional low point occurs at the 15th minute of the course, and suggests inserting a Q&A session at that time based on the discussion data to improve participation.
[0063] Through the above steps, a complete logical chain is formed from data collection to analysis and then to recommendations, ensuring the accuracy and real-time nature of the feedback. The system also includes a periodic calibration mechanism, which calculates the system's recognition accuracy and consistency index by collecting standard test samples with known emotional states. When the accuracy drops by more than 15%, the model retraining process is triggered to ensure the long-term reliability of the system.
[0064] This embodiment establishes a dynamic balance between local micro-expression capture (high-resolution detail extraction) and macro-behavioral trend analysis (overall dynamic perception) through a focal length switching mechanism, breaking through the limitations of traditional single-focal-length vision systems. In local mode, optical magnification enables facial muscle displacement detection accuracy to reach 0.1mm, capturing micro-expressions lasting only 250ms. In overall mode, through a mapping model of head posture and attention state, the accuracy of emotion recognition is improved from 66.3% of traditional methods to 89.7%.
[0065] Based on the principle of multimodal data fusion, visual emotion data and behavioral data from digital interactive systems are spatiotemporally aligned and weighted to construct a complete teaching participation evaluation system. This enables the system to trigger a teaching content replacement mechanism within 800ms when it detects a confused expression lasting 3 seconds. Its rapid response capability stems from the direct coupling between the emotion recognition module and the content scheduling module.
[0066] This embodiment introduces a closed-loop control principle, using emotional engagement as a feedback signal to dynamically adjust image acquisition parameters, forming an optimized cycle of "perception-evaluation-adjustment-re-perception." When emotional engagement falls below a threshold of 0.75, the system automatically adjusts the focal length parameters. After a maximum of three iterations of optimization, the recognition confidence level is increased from the initial 0.68 to 0.83, demonstrating the system's self-learning and adaptive capabilities. Finally, based on the principles of feature quantization and time series analysis, continuous dynamic emotional data is transformed into a structured evaluation report. Through dimensionality reduction processing using 7-dimensional feature vectors and 512-dimensional feature tensors, the report generation delay is controlled within 2 seconds while ensuring information integrity, achieving efficient transformation from raw data to actionable insights.
[0067] Example 2: This embodiment discloses a machine vision-based business administration teaching model recognition and feedback system, used to execute the aforementioned machine vision-based business administration teaching model recognition and feedback method, characterized by comprising: The image acquisition device is used to acquire parallel image sequences of students' faces, and simultaneously captures the overall expression trend and local micro-expression changes through parallel wide-angle and telephoto cameras to obtain raw image data containing emotional feature extraction. The image processing module is communicatively connected to the image acquisition device and is used to perform sharpening processing on the facial areas captured in local detail based on the original image data, generate enhanced and refined image data, and determine whether the enhanced and refined image data meets the accuracy requirements for emotional feature extraction. The sentiment analysis module, which is communicatively connected to the image processing module, is used to extract subtle emotional response features by means of emotional feature extraction when the enhanced and refined image data reaches a preset threshold, and associate the subtle emotional response features with the emotional response mapping in the digital interaction system to obtain a preliminary emotional response distribution. The real-time feedback module is connected to the sentiment analysis module and is used to transmit data through real-time feedback. Based on the sentiment fluctuations obtained from the overall trend analysis, it integrates the discussion data integration function to generate teaching content adaptation suggestions under multi-level content recognition and determines whether the teaching content adaptation suggestions meet the personalized guidance needs. The closed-loop control module is communicatively connected to the real-time feedback module and the image acquisition device. When the teaching content adaptation suggestion does not meet the preset emotional participation standard, it cyclically adjusts the image acquisition parameters, controls the image acquisition device to re-acquire new image sequence data, and updates the student participation record in conjunction with the digital interactive system to obtain the optimized emotional response distribution. The report generation module, which is communicatively connected to the closed-loop control module and the sentiment analysis module, is used to generate a personalized feedback report that combines students' micro-expressions with discussion and interaction in education and teaching, based on the optimized sentiment response distribution and integrating the results of the multi-level content recognition and the real-time feedback transmission.
[0068] Furthermore, the image acquisition device includes a parallel wide-angle camera and a telephoto camera, an image synchronization unit, and a quality control unit. The quality control unit is configured to dynamically adjust the acquisition parameters based on an image quality assessment algorithm, or to control the dual-camera collaborative working mode in response to emotion recognition confidence commands.
[0069] The system also includes a failure handling mechanism. When the average confidence score of emotion recognition is lower than a preset threshold (e.g., 0.5) within multiple consecutive time windows (usually 3-5 windows), it automatically switches to an alternative emotion assessment mode based on voice and text analysis. Natural language processing technology is used to analyze the emotional tendency of the discussion content to ensure that the system can still provide basic emotion participation assessment when visual analysis is limited.
[0070] The system also includes a data security and privacy protection module, which performs real-time anonymization processing on all collected facial images, deletes the original image data immediately after feature extraction, and retains only the desensitized emotional feature vectors, in compliance with personal information protection regulations.
[0071] The system also includes an adaptive degradation mechanism that automatically switches to a single-modal emotion assessment mode based on voice and text analysis when the ambient light level is detected to be below 100 lux or the student is more than 5 meters away from the camera, ensuring that the system can still provide basic emotion engagement assessment services under adverse conditions.
[0072] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0073] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A machine vision-based method for recognizing and providing feedback on teaching models in business administration, characterized in that, Specifically, the steps include the following: Step 1: Parallel image sequence acquisition of students' faces is carried out using a dual-camera image acquisition device. The wide-angle camera continuously captures the overall expression trend, while the telephoto camera specifically captures the micro-expression changes in key areas such as the eyes and mouth. The two video streams are synchronized through timestamp alignment to obtain raw image data containing emotional feature extraction and determine the initial data granularity layering. Step 2: Based on the original image data, sharpening processing is performed on the facial areas captured in local detail. An edge enhancement method based on the Laplacian operator is adopted, and the sharpening parameter intensity factor is set to 1.2-1.8 to generate enhanced and refined image data. The image quality score is calculated by the peak signal-to-noise ratio (PSNR) algorithm, and a threshold of 30dB is set to determine whether the enhanced and refined image data meets the accuracy requirements for emotional feature extraction. Step 3: If the enhanced and refined image data reaches a preset threshold, then subtle emotional response features are obtained through emotional feature extraction, and the subtle emotional response features are associated with the emotional response mapping in the digital interaction system to obtain a preliminary emotional response distribution. Step 4: Based on the preliminary distribution of emotional responses, combined with the student participation records and interaction frequency analysis in the digital interaction system, generate real-time feedback transmission data to determine the emotional participation level in the current discussion session; Step 5: Through the real-time feedback transmission data, based on the emotional fluctuations obtained from the overall trend analysis, integrate the discussion data integration function to generate teaching content adaptation suggestions under multi-level content recognition, and determine whether the teaching content adaptation suggestions meet the personalized guidance needs. Step 6: If the teaching content adaptation suggestion does not meet the preset emotional engagement standard, the image acquisition parameters are adjusted cyclically, and new image sequence data is re-acquired through digital zoom and physical position adjustment. Combined with the digital interaction system to update the student participation record, the optimized emotional response distribution is obtained. Step 7: Based on the optimized emotional response distribution, integrate the results of the multi-level content recognition and the real-time feedback transmission to generate a personalized feedback report that combines students' micro-expressions with discussion and interaction in education and teaching, and determine the final emotional data support.
2. The machine vision-based business administration teaching model recognition and feedback method according to claim 1, characterized in that, Step 1 specifically includes: The student's face is continuously acquired using an image acquisition device. A focus switching mechanism is used to acquire multi-level image data in different modes to obtain a preliminary set of raw images. Based on the original image set, local detail segmentation is performed on the student's facial region to separate key regions containing emotional features and determine the distribution range of emotion-related regions. By initially extracting the emotional features of key region images and matching them using preset feature templates, the initial classification results of the emotional features are obtained. If the confidence level of the initial classification result is lower than the preset threshold, the key area image is re-acquired after a second focus adjustment to obtain a clearer local detail image. Based on the re-acquired local detail images and combined with the overall dynamic image data, the emotional features are analyzed in depth. A convolutional neural network model is used to refine the features and determine the emotional category. By analyzing the results of emotion category judgments and combining them with data granularity hierarchy, the classification results are mapped to the original image set to obtain the final emotion feature distribution data. Based on the emotion feature distribution data, a time series analysis is performed on the overall dynamic changes of students' faces to determine the trends and patterns of emotion changes.
3. The machine vision-based business administration teaching model recognition and feedback method according to claim 1, characterized in that, Step 2 specifically includes: By performing layered processing on the original image data, a preliminary sharpening enhancement operation is performed on the local fine facial areas to obtain the first enhanced image set; Based on the enhanced first image set, a preset sharpening tool is used to further optimize the details of the facial area, generating a clearer second image set; If the detail clarity of the second image set does not meet the accuracy standard for emotion feature extraction, then targeted data augmentation processing is performed on the facial region to obtain an optimized third image set. By applying a convolutional neural network model to facial region data from a third image set, preliminary classification of emotional features is performed to determine the category distribution of emotional features. Based on the category distribution of emotional features, for the data in the classification results where the confidence level is lower than the preset threshold standard, local image detail comparison processing is performed to obtain more accurate feature classification data; If the accuracy of the feature classification data still does not reach the preset accuracy standard, the classification data will be filtered a second time to obtain the final refined sentiment feature data. By dynamically comparing refined emotional feature data over time, we can determine the emotional change patterns of facial regions at different points in time and identify the continuity characteristics of emotional changes.
4. The machine vision-based business administration teaching model recognition and feedback method according to claim 1, characterized in that, Step 3 specifically includes: If the enhanced and refined image data reaches the preset threshold standard, the refined image data will be processed by a pre-established emotional feature extraction tool to obtain subtle emotional response data. Based on the acquired subtle emotional response data, a pre-set mapping table is used to compare the response data with the emotional response categories in the digital interaction system to obtain preliminary emotional response distribution results. Based on the preliminary distribution results of emotional responses, data comparison tools were used to group the emotional response categories in the distribution results to determine the main distribution range of each emotional response category. If the distribution range of a certain type of emotional response overlaps with other categories, cluster analysis tools can be used to further subdivide the data in the overlapping areas to obtain clearer category classification results. Based on the subdivision results, a data filtering mechanism is used to prioritize and sort the emotional response data of each category, resulting in a sorted emotional response data set. For the sorted sentiment response dataset, high-priority sentiment response data are highlighted using data integration tools to determine the final core sentiment response data; Based on the final core data of emotional response, a record-keeping mechanism is used to update and compare the core data with the mapping table in the digital interaction system to obtain the latest emotional response mapping results.
5. The machine vision-based business administration teaching model recognition and feedback method according to claim 1, characterized in that, Step 4 specifically includes: By acquiring emotional response data and distribution data through digital interactive platforms, analyzing the statistical records of student participation, and obtaining preliminary emotional response patterns; Based on the preliminary emotional response patterns and the data on interaction frequency, statistical tools are used to calculate the activity index in the discussion session to determine the current participation status in the discussion. If the current level of participation in the discussion is below a preset threshold, a real-time feedback mechanism is triggered to generate transmission information and determine the areas where emotional participation is insufficient. By analyzing the content of the transmitted information, we can identify areas where emotional engagement is insufficient, obtain specific interaction records from the discussion sessions, and pinpoint the weaknesses in student participation. Based on the weaknesses in student participation and the distribution data of emotional responses, the fluctuations in interaction frequency were analyzed to obtain the changing trend of emotional participation. By analyzing the changing trends in emotional engagement, dynamic adjustment strategies can be generated for the current discussion segment to determine the direction for improving subsequent interactions. By acquiring the output of the dynamic adjustment strategy and combining it with the transmission information from real-time feedback, we can analyze the room for improvement in the discussion process and determine the final optimization plan for emotional participation.
6. The machine vision-based business administration teaching model recognition and feedback method according to claim 1, characterized in that, Step 5 specifically includes: By collecting data through a real-time feedback mechanism, emotional fluctuation data is obtained from user interactions. A pre-established emotion recognition model is used for preliminary processing to obtain the classification results of emotional states. Based on the classification results of emotional states and the analysis results of overall trends, the long-term change characteristics reflected in data transmission are extracted to determine the emotional tendencies of users in different time periods. If the sentiment trend shows significant fluctuations, then we will conduct in-depth analysis of the integrated discussion data to obtain details of user feedback in specific scenarios and determine potential sentiment drivers. By using multi-layer recognition technology, emotional driving factors are associated and matched with teaching content to generate a preliminary content adaptation plan and obtain teaching adjustment directions that are consistent with the user's emotional state. If the initial content adaptation plan deviates from the goal of personalized guidance, the plan will be optimized a second time based on the logical rules of demand judgment to determine the final adaptation result. Based on the final adaptation results, the output module generated by the linkage suggestion presents the adjustment direction of the teaching content in a structured form, and obtains a personalized guidance plan for users. Through the continuous processing of the above steps, real-time feedback, emotional fluctuations, and overall trends are integrated into the business process to achieve the adaptation of teaching content and the realization of personalized guidance.
7. The machine vision-based business administration teaching model recognition and feedback method according to claim 1, characterized in that, Step 6 specifically includes: By linking teaching content with adaptation suggestions, initial emotional engagement data is obtained, and a pre-established emotional assessment model is used to determine whether emotional engagement meets the preset engagement standards. If emotional engagement does not meet the preset engagement standard, the image acquisition parameters are adjusted, and the image sequence data is reacquired through digital zoom and camera physical position adjustment to obtain clearer sequence data content. Based on the newly acquired image sequence data, and combined with the correlation between the sequence data and digital interaction, the data is synchronized through the digital interaction system to update students' participation records and determine whether the participation records reflect emotional changes. If the participation records reflect emotional changes, the distribution features of emotional responses are extracted by associating the interaction system with the participation records. The support vector machine algorithm is then used to classify the response distribution to obtain the classified emotional response results. Based on the classified emotional response results, and combined with the correlation between emotional participation and response distribution, the dynamic trend of emotional participation is analyzed to determine the optimization direction of emotional participation; By optimizing the direction of emotional engagement, combining the connection between teaching content and digital interaction, adjusting the specific parameters of the adaptation suggestions, obtaining the updated teaching content adaptation scheme, and determining whether it meets the emotional engagement standards; If the updated adaptation scheme still does not meet the emotional participation standard, the acquisition method will be optimized by adjusting the image acquisition parameters in a loop to obtain more accurate sequence data and determine the final emotional response distribution.
8. The machine vision-based method for recognizing and providing feedback on business administration teaching models according to claim 1, characterized in that, Step 7 specifically includes: By acquiring raw image data of students' micro-expressions from classroom video streams, and using a pre-established image processing module for preliminary cleaning, processed facial feature images are obtained. Based on the processed facial feature images, the support vector machine algorithm is used to classify students' micro-expressions and determine the corresponding emotional response categories. Based on the emotional response categories obtained from the classification, audio segments of discussions and interactions are obtained from real-time feedback transmissions, and the text content is extracted using a speech-to-text tool to obtain interactive text records; If the interactive text record contains preset emotional keywords, then a matching analysis is performed based on the emotional response category to determine the students' emotional tendencies in the discussion interaction; By continuously monitoring emotional tendencies, we can obtain the trend of emotional response changes over a period of time and generate a dynamic distribution chart of emotional data. Based on the dynamic distribution chart and the specific requirements of the teaching and learning scenario, personalized feedback content is generated for each student to determine the final emotional data support basis. By integrating personalized feedback content with sentiment data support, a structured feedback report document is automatically generated for subsequent analysis and reference.
9. A machine vision-based business administration teaching model recognition and feedback system, used to execute the machine vision-based business administration teaching model recognition and feedback method according to any one of claims 1-8, characterized in that, include: The image acquisition device is used to acquire parallel image sequences of students' faces, and simultaneously captures the overall expression trend and local micro-expression changes through parallel wide-angle and telephoto cameras to obtain raw image data containing emotional feature extraction. The image processing module is communicatively connected to the image acquisition device and is used to perform sharpening processing on the facial areas captured in local detail based on the original image data, generate enhanced and refined image data, and determine whether the enhanced and refined image data meets the accuracy requirements for emotional feature extraction. The sentiment analysis module, which is communicatively connected to the image processing module, is used to extract subtle emotional response features by means of emotional feature extraction when the enhanced and refined image data reaches a preset threshold, and associate the subtle emotional response features with the emotional response mapping in the digital interaction system to obtain a preliminary emotional response distribution. The real-time feedback module is connected to the sentiment analysis module and is used to transmit data through real-time feedback. Based on the sentiment fluctuations obtained from the overall trend analysis, it integrates the discussion data integration function to generate teaching content adaptation suggestions under multi-level content recognition and determines whether the teaching content adaptation suggestions meet the personalized guidance needs. The closed-loop control module is communicatively connected to the real-time feedback module and the image acquisition device. When the teaching content adaptation suggestion does not meet the preset emotional participation standard, it cyclically adjusts the image acquisition parameters, controls the image acquisition device to re-acquire new image sequence data, and updates the student participation record in conjunction with the digital interactive system to obtain the optimized emotional response distribution. The report generation module, which is communicatively connected to the closed-loop control module and the sentiment analysis module, is used to generate a personalized feedback report that combines students' micro-expressions with discussion and interaction in education and teaching, based on the optimized sentiment response distribution and integrating the results of the multi-level content recognition and the real-time feedback transmission.
10. The machine vision-based business administration teaching model recognition and feedback system according to claim 9, characterized in that: The image acquisition device includes a parallel wide-angle camera and a telephoto camera, an image synchronization unit, and a quality control unit. The quality control unit is configured to dynamically adjust the acquisition parameters based on an image quality assessment algorithm, or to control the dual-camera collaborative working mode in response to emotion recognition confidence commands.
Citation Information
Patent Citations
Teaching quality evaluation method and system based on facial expressions and human body actions
CN114971971A
Robot personalized teaching method based on digital human
CN118537182A
Camera and alarm system for network teaching
CN118553017A
Smart classroom teaching system based on Internet of Things
CN118691434A
Course content feedback system and method based on big data
CN118898003A