An immersive aesthetic education cabin with integrated audio and visual elements driven by facial emotion recognition
Through the integrated audio-visual immersive aesthetic edifice driven by facial emotion recognition, combined with deep learning models and personalized content recommendations, the problem of real-time recognition and response of students' emotional state in the field of education is solved, and the students' emotional intervention effect and learning experience are improved.
Patent Information
- Application Number
- CN202411837511.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-12-13
AI Technical Summary
The existing technology is difficult to realize real-time identification and response to students' emotional states in the field of education, resulting in poor recommendation of personalized audio-visual content, low efficiency of traditional psychological counseling, and lack of real-time response to students' current emotional states.
The audio-visual integrated immersive aesthetic education cabin is adopted to generate facial emotion recognition-driven audio-visual integrated aesthetic education cabin. It collects facial expressions and physiological signals through the student client, combines deep learning models to identify students' emotional states, and screens personalized audio and video content from the aesthetic education database. It uses the administrator client for management and data analysis to provide personalized artistic works recommendations.
It improves the effect of positive emotional intervention, reduces the burden of psychological counseling, improves student satisfaction and learning experience, and realizes real-time response to students' emotional state and personalized content recommendations.
Smart Images

Figure CN119694172B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of emotion recognition and recommendation technology, and in particular to an audio-visual integrated immersive aesthetic education cabin driven by facial emotion recognition. Background Art
[0002] Although all types of schools have psychological counseling rooms and are actively promoting mental health counseling, this type of conversational intervention requires a large number of professional psychological counselors, and students need a certain amount of courage to enter the psychological counseling room, so the positive guidance effect on students is limited.
[0003] Traditional psychological counseling often relies on face-to-face, interpersonal interaction. Without established trust, it's difficult for students to open up, and thus it often takes a long time to achieve the desired results. With the development of artificial intelligence, particularly with the advancement of big data and large-scale models, it's become possible to provide students with a solitary, relaxing environment for personalized learning and emotional experiences. Creating immersive environments that effectively integrate audiovisual perception with art therapy, and thus categorize students' emotional states and provide timely intervention and positive guidance for negative emotions, remains a thorny issue that needs to be addressed.
[0004] First, emotions play a crucial role in the learning process. Research shows that students' emotional states directly influence their behavioral motivations, daily lives, and spiritual lives. Therefore, accurately identifying students' emotional states, adjusting audiovisual content accordingly, and guiding their emotional experiences are crucial. While current emotion recognition technology has made progress in multiple fields, its application in education faces unique challenges. For example, students' emotional expressions in educational settings can be influenced by a variety of factors, making emotion recognition more complex and difficult to achieve high accuracy. Furthermore, these technologies often require processing large data sets, and the need for real-time performance places even higher demands on data processing speed.
[0005] Secondly, the core human perception of beauty primarily focuses on the visual and auditory dimensions. Classical paintings and music embody the artist's emotional expression of beauty. Furthermore, the purpose of aesthetic education is not simply to impart knowledge and artistic skills; more importantly, it is to inspire students' creativity, enhance their artistic literacy and aesthetic abilities, and enrich their emotional expression. Therefore, aesthetic education requires a greater focus on students' individual development and emotional experiences. Educators must be able to provide customized artistic aesthetic content based on students' individual emotional states and personality preferences, allowing students to achieve their aesthetic needs in a relaxed state. Currently, personalized recommendation systems are quite mature for recommending music, movies, or shopping products. However, applying these systems to education, particularly incorporating emotion recognition technology to recommend audiovisual content that matches students' emotions and has a positive impact, remains an open area of research. Existing recommendation systems mostly rely on static user preferences or historical data, lacking the ability to respond to users' current emotional states in real time, resulting in poor matching.
[0006] Thirdly, art therapy, a theory of emotional intervention that spans the fields of art and psychology, fully aligns with human instinct and innate talent. Art education can effectively guide and stimulate inner experiences and emotional expression, offering a means of emotional and psychological expression beyond, and even beyond, verbal communication. This immersive aesthetic education experience system, constructed through integrated audio-visual art, uses the beauty of art to transform the human inner world, instantly channeling negative or negativistic emotions in a positive or active way, improving students' mental health and fostering an aesthetic experience.
[0007] Furthermore, with the development of information technology, multimedia and internet technologies have been widely used in education, providing new tools and platforms for aesthetic education. In school settings, in particular, leveraging technology to create immersive and interactive learning experiences has become a key area of educational innovation. However, effectively integrating these technologies, particularly how to leverage them to identify and respond to students' emotional states in real time, remains a technical challenge.
[0008] To sum up, there is an urgent need to develop an immersive aesthetic education experience system that spans the two major fields of art and psychology and can combine emotion recognition and personalized recommendation technologies. Summary of the Invention
[0009] The purpose of this invention is to provide an audio-visual integrated immersive aesthetic education cabin driven by facial emotion recognition, which can provide personalized aesthetic education content recommendations based on students' emotional state and behavioral characteristics, and at the same time improve students' learning experience and effects through a comfortable learning environment and intuitive interactive interface.
[0010] To achieve the above object, the present invention provides the following solutions:
[0011] An audio-visual integrated immersive aesthetic education cabin driven by facial emotion recognition, including: a student client, an administrator client, and a server;
[0012] The student client is used to collect students' facial expression pictures, physiological signals and behavioral data, identify the students' current emotional state based on the facial expression pictures, physiological signals and behavioral data, and interact with the students;
[0013] The server is configured to filter out audio and video content from an aesthetic education database based on the current emotional state;
[0014] The administrator client is used for the administrator to manage the aesthetic education cabin, the aesthetic education database and the students' emotional state data.
[0015] Optionally, the student client includes a data acquisition module, an emotion recognition module, a content playback module and an information interaction module;
[0016] The data acquisition module is used to collect the facial expression pictures, physiological signals and behavioral data of students;
[0017] The emotion recognition module is used to perform feature splicing on the facial expression image and the physiological signal to obtain a comprehensive feature vector, and identify the student's current emotional state based on the comprehensive feature vector and the behavioral data;
[0018] The content playing module is used to play the filtered audio and video content;
[0019] The information interaction module is used for students to provide feedback on the audio and video content and select preferences.
[0020] Optionally, performing feature splicing on the facial expression image and the physiological signal to obtain a comprehensive feature vector includes:
[0021] The facial expression image and the physiological signal are subjected to feature splicing using a gradient boosting decision tree model to obtain the comprehensive feature vector.
[0022] Optionally, identifying the student's current emotional state based on the comprehensive feature vector and behavioral data includes:
[0023] Inputting the comprehensive feature vector into an improved ResNet-50 model for recognition, and outputting a high-dimensional feature vector, wherein the improved ResNet-50 model is obtained based on depthwise separable convolution;
[0024] Input the high-dimensional feature vector into a first Bi-LSTM network with an attention mechanism to output a first feature, and input the first feature into a second Bi-LSTM network with an attention mechanism to obtain a sentiment feature sequence;
[0025] Inputting the behavior data into a Deep Q-Learning behavior decision network and outputting a behavior feature sequence;
[0026] The current emotional state is acquired based on the emotional feature sequence and the behavioral feature sequence.
[0027] Optionally, obtaining the current emotional state based on the emotional feature sequence and the behavioral feature sequence includes:
[0028] Inputting the emotion feature sequence and the behavior feature sequence into a long short-term memory network and outputting a feature matrix;
[0029] The feature of the last time step in the feature matrix is selected and input into the Transformer Encoder model to obtain the feature representation of the current emotional state, and the feature representation of the current emotional state is input into the fully connected layer to obtain the current emotional state.
[0030] Optionally, filtering out audio and video content from the aesthetic education database according to the current emotional state includes:
[0031] The similarity between users is calculated based on the students' rating data and the current emotional state, the recommendation algorithm is optimized based on the similarity, and the audio and video content is screened based on the optimized recommendation algorithm.
[0032] Optionally, a method for calculating the similarity between users based on the student's rating data and the current emotional state is:
[0033]
[0034] Among them, r ui and r vi Represents the ratings of user u and v on item i, e ui and e vi They represent the current emotional states of users u and v towards item i, respectively. w1 and w2 are the weights of the rating data and the current emotional state, respectively.
[0035] Optionally, the administrator client includes: a device management module, a content management module and a data analysis module;
[0036] The device management module is used to check and set the status of the aesthetic education cabin, send device inspection instructions to the student client, and receive status feedback information;
[0037] The content management module is used to add, delete, modify and check the audio and video content in the aesthetic education database;
[0038] The data analysis module is used to perform data statistics and big data analysis on students' emotional states, and to perform statistics and analysis on content interaction data.
[0039] The beneficial effects of the present invention are as follows: By combining emotion recognition and behavioral psychological correction mechanisms, the present invention can more accurately understand each student's emotional state and psychological needs, thereby providing more personalized art content recommendations. This highly personalized work can better meet students' unique needs, enhance the effectiveness of positive emotional intervention and improve student satisfaction. The present invention can intelligently recommend the most appropriate art resources based on students' emotional state and learning progress, avoiding interference from similar resources and ensuring efficient utilization of art resources. Furthermore, by collecting and analyzing student feedback and experience data, the art database and its effectiveness can be further optimized and enhanced in iterations. The management interface provided by the present invention helps the system observe students' experience states and emotional changes in real time, allowing for more precise adjustments to experience methods and content. This not only reduces the workload of psychological counseling but also improves the relevance and efficiency of independent experience. The present invention applies the latest artificial intelligence technology to the field of aesthetic education experience, bringing innovation and transformation to traditional art audio-visual methods. By effectively improving the quality and quantity of art works through technical means, the present invention also paves a path for new developments in aesthetic education.
[0040] To sum up, the present invention provides a new auxiliary means for aesthetic education experience and art therapy through highly integrated and innovative technical solutions, combined with painting, photography and music art creation. It embodies a significant level of personalization, integrates the effects of artistic experience and the efficiency of aesthetic education resource utilization, and is of great significance to promoting the development of educational technology and improving the quality of education. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 This is a flowchart of a working method of an audio-visual integrated immersive aesthetic education cabin driven by facial emotion recognition according to an embodiment of the present invention;
[0043] Figure 2 Schematic diagram of a zero-gravity interaction and audio-visual playback device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] like Figure 1 As shown, the present invention discloses an audio-visual integrated immersive aesthetic education cabin based on facial emotion recognition, comprising: a student client, an administrator client and a server;
[0047] The student client is used to collect students' facial expression images, physiological signals and behavioral data, identify the students' current emotional state based on the facial expression images, physiological signals and behavioral data, and interact with the students; the student client is integrated into the interactive equipment in the immersive integrated aesthetic education cabin.
[0048] Specifically, the student client includes a data acquisition module, an emotion recognition module, a content playback module and an information interaction module; the data acquisition module is used to collect students' facial expression pictures, physiological signals and behavioral data; the emotion recognition module is used to perform feature splicing on facial expression pictures and physiological signals, obtain a comprehensive feature vector, and identify the student's current emotional state based on the comprehensive feature vector and behavioral data; the content playback module is used to play the selected audio and video content; the information interaction module is used for students to provide feedback on the audio and video content and select preferences.
[0049] In this embodiment, a high-resolution camera (1080P, 60fps) and a depth sensor (such as Intel RealSense D455) are used to obtain students' behavior and expression data, and the key areas are located based on the YOLOv5 target detection network, and the shooting angle and range of the camera are adjusted in real time to ensure the diversity and accuracy of the collection.
[0050] Furthermore, the emotion recognition module is used to accurately identify the student’s current emotional state by analyzing the student’s facial expressions, heart rate and blood pressure physiological data, and behavioral psychological data obtained from behavioral video observations;
[0051] The emotion recognition module includes facial basic expression recognition, micro-expression recognition, and emotion-assisted recognition of electrocardiogram and blood pressure based on deep learning. It further captures facial expressions and micro-expressions through a high-definition camera, combines electrocardiogram equipment to record electrocardiogram signals, and blood pressure monitoring equipment to obtain blood pressure data, and uses a gradient boosting decision tree model to conduct a comprehensive analysis of these data. Features such as heart rate, heart rate variability, facial expression features, and blood pressure changes are extracted, and facial expression features, electrocardiogram features, and blood pressure features are spliced into a comprehensive feature vector through a feature splicing method. The gradient boosting decision tree model can handle complex feature relationships and nonlinear data, gradually construct multiple decision trees, and improve classification performance. In this embodiment, the gradient boosting decision tree model is used to perform feature splicing on facial expression images and physiological signals to obtain a comprehensive feature vector.
[0052] To convert the comprehensive feature vector into basic emotional states, a model combining a convolutional neural network (CNN) and a recurrent LSTM with an attention mechanism is employed, with reinforcement learning incorporated to optimize the decision-making process. First, a modified ResNet-50 is constructed by replacing the convolutional layers with depthwise separable convolutions for feature recognition. This outputs a 2048-dimensional high-dimensional feature vector that comprehensively captures subtle changes in behavior and expression. To enhance robustness for small-sample training, group normalization is introduced in the convolutional layers, effectively improving training stability with small batch sizes. These high-dimensional feature vectors are then fed into the first Bi-LSTM layer with a multi-head attention mechanism. The 2048-dimensional input feature vector is encoded into a 1024-dimensional contextual representation. Subsequently, a multi-head attention mechanism focuses on key features in different time segments, with each attention head having a dimension of 64 and a total of 8 heads. Ultimately, a 1024-dimensional feature representation is generated that incorporates temporal information. Next, these features are input into the second-layer Bi-LSTM with a multi-head attention mechanism to further extract the deep dynamic relationships in the time series, and finally output a 512-dimensional emotional feature sequence to describe the changing trend of the emotional state over time. This can effectively process the feature sequence and capture subtle changes in expression, generating a more accurate representation of the emotional state. The multi-head attention mechanism is added to the output of the Bi-LSTM.
[0053] The behavioral psychology correction function collects student behavioral data using machine vision technology. This data is then processed based on Deep Q-Learning's behavioral decision network to obtain a sequence of behavioral features. Specifically, the features are first reduced to 128 dimensions through a fully connected layer and then input into a two-layer fully connected Q network. The first layer consists of 64 neurons and uses Leaky ReLU as the activation function. The second layer outputs the specific action space dimensions for generating behavioral adjustment strategies. To ensure efficient operation in an edge computing environment, the Deep Q-Learning behavioral decision network undergoes weight pruning and quantization, significantly reducing computational complexity while ensuring decision accuracy.
[0054] The behavioral and emotional feature sequences are mapped to a 256-dimensional feature space, with each sample sequence consisting of 30 time steps. The feature representation of each time step is a 256-dimensional vector that contains dynamic information related to behavior and emotion. To capture the dependencies and dynamic changes between these time steps, the behavioral and emotional feature sequences are input into an LSTM. This model incorporates a reinforcement learning strategy to optimize the processing of feature sequences and outputs serialized feature representations that capture temporal dynamic relationships. When the behavioral and emotional feature sequences are processed, the LSTM outputs a complete feature matrix of shape 30×256, where each row is a 256-dimensional feature representation of a time step. Furthermore, to capture the current emotional state, the features of the last time step are selected from these 30 time steps as the primary representation of the current state, as they summarize the dynamic changes from the beginning to the end of the sequence. The features of the last time step output by the LSTM are fed into a six-layer Transformer Encoder to output the feature representation of the current emotional state. This feature representation is then fed into a fully connected layer to obtain the current emotional state. Each layer of the six-layer Transformer Encoder includes a multi-head attention module (eight heads, each with a dimension of 64) and a feedforward neural network (two hidden layers, with a dimension of 1024). This structure not only captures inter-feature dependencies but also optimizes the semantic consistency of the fused features through global modeling.
[0055] Ultimately, the entire system is continuously optimized by applying a unified decision-making framework and feedback mechanism to the data collection, emotion recognition, and behavior adjustment modules. Specifically, in an edge computing environment, the accuracy and reliability of data collection are improved by adjusting camera positions and sensor settings. The ResNet-50, a two-layer Bi-LSTM optimized for deep separable convolution, is seamlessly combined with a multi-head attention mechanism, Transformer Encoder, and a deep reinforcement learning framework to improve model performance while ensuring efficient operation in an edge computing environment. The content playback module, based on the analysis results of the emotion recognition module, selects audio and video content from the preset audio-visual integrated aesthetic education database that can positively guide or intervene in students' emotions, and presents the content to students in audio-visual formats such as MP4.
[0056] The information interaction module provides a user interface that allows students to select preferences and adjust recommended settings, and provide feedback on the audio-visual experience. It also receives device inspection instructions sent by the server and provides feedback on the device status.
[0057] Specifically, the server is used to filter audio and video content from the aesthetic education database based on the current emotional state;
[0058] Furthermore, the server side includes:
[0059] Content recommendation module: Based on the emotion recognition analysis results sent by the student client, a recommendation algorithm is used to filter out a list of audio and video content recommendations that can positively guide or intervene in students' emotions from the audio-visual integrated aesthetic education database and send it back to the corresponding student client;
[0060] The content recommendation module integrates real-time emotion recognition technology and a dynamic recommendation algorithm to achieve real-time responsiveness. Leveraging multiple data sources, including facial expressions, behavioral pattern analysis, and physiological signals, it captures and analyzes the user's current emotional state in real time. Based on this current emotional state, the recommendation algorithm dynamically adjusts content recommendations to ensure that the audio and video content provided is highly relevant to the user's immediate emotional state. This improves the relevance of recommended content, enhances user interaction and satisfaction, and provides more personalized and accurate content playback.
[0061] The audio and video content includes art resources of various types and styles, such as music, painting, photography, aerial images and music of students' hometown scenes familiar to them, etc., aiming to provide different types of content according to user preferences to enrich students' aesthetic education or emotional intervention experience.
[0062] The information interaction module allows students to set their own preferences, evaluate the relevance and satisfaction of recommended content, and adjust the recommendation algorithm by collecting students' feedback and behavioral data, processing and analyzing these data. The recommendation algorithm includes existing collaborative filtering and content-based algorithms, as well as improved deep learning and hybrid recommendation algorithms. Specifically, based on the students' ratings and emotional states, the similarity weights and content feature weights are recalculated to optimize the parameters and structure of the recommendation model. Advanced technologies such as deep neural networks are used to dynamically adjust the weights and preference layers of the algorithm. Regularly evaluate the effectiveness of the recommendation system, compare the performance of different adjustment strategies through A / B testing, and continuously collect new feedback data for iterative optimization, so as to achieve a personalized and intelligent recommendation system to meet the ever-changing needs of students and achieve continuous learning and optimization. Among them, the current emotional state is integrated into the user-item matrix and used together with the rating data to calculate the similarity between users. The user-item matrix is a two-dimensional matrix with rows representing users and columns representing items. Each element r in the matrix ui The current emotional state is obtained by analyzing the user's physiological characteristics such as heart rate, skin electrical response, facial expression, and behavioral data, and quantified as a value e ui , and integrated into the matrix. When calculating similarity, the weighted cosine similarity method is used, and the rating data and current emotional state are given different weights w1 and w2 respectively. The calculation formula for user similarity based on student ratings and current emotional state is as follows:
[0063]
[0064] Among them, r ui and r vi Represents the ratings of user u and v on item i, e ui and e vi They represent the current emotional states of users u and v towards item i, respectively. w1 and w2 are the weights of the rating data and the current emotional state, respectively.
[0065] Using advanced technologies such as deep neural networks, the algorithm dynamically adjusts its weights and biases. Specifically, a convolutional neural network (CNN) architecture is employed, consisting of multiple convolutional and pooling layers, followed by a final fully connected layer. The Adam optimizer and cross-entropy loss function are used, which facilitates rapid convergence and accurately measures the error in class predictions. The network learns from the training data, optimizing its weights and biases based on the loss function to reduce error. After deployment, model performance is monitored in real time, and fine-tuning or full algorithm updates may be performed based on performance, continuing training or adjusting the network architecture. This process involves a feedback loop system to ensure the model can adapt to new data or environmental changes, maintaining its effectiveness and accuracy through evaluation and iteration. This system design not only meets initial task requirements but also adapts to potential future changes, ensuring long-term stability and high performance.
[0066] This approach allows the content recommendation module to more accurately calculate similarities between users and recommend items favored by similar users. Regularly collecting new ratings and current sentiment, recalculating similarities, and adjusting algorithm parameters ensures the recommendation system can dynamically adapt to changes in user sentiment and preferences, enabling personalized and intelligent recommendations.
[0067] Student information recording module: interacts with the student client and stores the student's settings, preferences, audio and video tracks, etc. in the server for easy use and modification by students;
[0068] Data storage module: provides secure storage for audio-visual content and student information. When a student client sends a content request, the requested content is sent to the corresponding student client; when an administrator client sends a content management request, the corresponding operation is performed.
[0069] The server also stores encrypted audio-visual content and student-related information;
[0070] The server also includes an authority verification module, which is used to verify the identity of students and administrators; and a content distribution module, which is used to send content to students or perform administrator management operations based on requests.
[0071] Administrator client, used by administrators to manage the aesthetic education cabin, aesthetic education database and students' emotional status data.
[0072] The Administrator Client includes:
[0073] The equipment management module provides an interface that allows administrators to manage the aesthetic education cabins distributed in a certain space, view and set the status of the aesthetic education cabins, send equipment inspection instructions to student clients, and receive feedback on the equipment status of the aesthetic education cabins from student clients; the equipment management module allows administrators to manage audio-visual cabins in different locations, retrieve real-time data from the aesthetic education cabins, check the equipment status, and realize monitoring and management of the equipment.
[0074] The content management module provides an interface that allows administrators to add, delete, modify, and check video content;
[0075] The data analysis module provides an interface that allows experts with psychological education qualifications to conduct statistics and big data analysis on student information, namely student emotional state data, and conduct data statistics and analysis on content interaction in order to further promote analysis and strategy adjustment. The data on content interaction refers to student ratings and student selection preferences.
[0076] like Figure 1 As shown in the figure, the operation process of an audio-visual integrated immersive aesthetic education cabin driven by facial emotion recognition includes:
[0077] Step 1: Setup and implementation of emotion recognition module:
[0078] 1.1. Install high-resolution cameras in the aesthetic education cabin to capture students’ facial expressions and behaviors.
[0079] 1.2. Install a portable electrocardiograph on the interactive equipment in the aesthetic education cabin to record students’ blood pressure and electrocardiogram data.
[0080] 1.3. Configure emotion recognition software, which can analyze students' facial expressions, blood pressure and electrocardiogram data in real time and identify their basic emotional states, such as happiness, sadness, anger, etc.
[0081] 1.4. Introduce behavioral analysis algorithms, analyze students’ behavioral patterns through video data of student behavior captured by cameras, combined with psychological principles, to assist in the judgment of emotional state.
[0082] 1.5. Integrate the emotion recognition results and behavior analysis results to form a comprehensive assessment of the student’s emotional state, providing a basis for the content recommendation module.
[0083] Step 2: Setting up and implementing the content recommendation module:
[0084] 2.1. Establish an integrated audio-visual aesthetic education content database, which includes various types of audio-visual aesthetic education resources, and each resource is marked with the appropriate emotional state.
[0085] 2.2. Develop a content recommendation algorithm that selects the most matching content from the aesthetic education content database for recommendation based on the student emotional state provided by the emotion recognition module.
[0086] 2.3. Design an autonomous learning mechanism for the recommendation system. Based on students’ feedback on the recommended content, such as viewing time and interaction frequency, adjust the recommendation algorithm to improve the accuracy and satisfaction of future recommendations.
[0087] Step 3: Setup and implementation of user interaction module, such as Figure 2 As shown:
[0088] 3.1. Zero-gravity chairs are equipped in the aesthetic education cabin to provide students with a comfortable emotional adjustment environment. The zero-gravity chairs can adjust their posture according to the student's body shape and preferences, creating a human-computer interaction experience in the most relaxed physiological state.
[0089] 3.2. A touch screen is installed next to the zero-gravity chair, through which students can interact with the system, such as selecting their favorite content, adjusting the volume and playback speed, and providing evaluation and feedback on the content.
[0090] 3.3. Develop a user interaction interface. The interface design should be simple and intuitive to ensure that students can easily operate it even when they are in a learning state.
[0091] 3.4. Collect students’ feedback and preference information through the user interaction module and feed it back to the content recommendation module to optimize the recommendation strategy.
[0092] Step 4: Setting up and implementing the content playback module, such as Figure 2 As shown:
[0093] 4.1. Equipped with high-definition display screens and high-quality surround sound systems to ensure the audio-visual effects of aesthetic education content.
[0094] 4.2. Develop content playback control software that supports multiple media formats and can automatically play corresponding audio-visual integrated art content based on the recommendations of the content recommendation module.
[0095] 4.3. Design the playback interface, including basic control functions such as play, pause, fast forward, and fast rewind, and display relevant information of the content, such as the author, historical background, artistic characteristics, etc.
[0096] 4.4. Realize the linkage between the playback module and the user interaction module. For example, students can control the playback progress through the touch screen, and the system adjusts the playback mode according to the students' interaction.
[0097] Step 5: System Integration and Testing:
[0098] 5.1. Integrate the emotion recognition module, content recommendation module, user interaction module, and content playback module into a unified system to ensure the efficiency and stability of data transmission and processing between modules.
[0099] 5.2. Carry out system installation and debugging in the aesthetic education cabin to ensure that all equipment and software work together to achieve the expected functions.
[0100] 5.3. Conduct system testing, including functional testing, performance testing, and user experience testing, to ensure that the system is stable and reliable, user-friendly, and can meet students' learning needs.
[0101] 5.4. Optimize the system based on test feedback, adjust parameters, improve interaction design, and enhance user experience.
[0102] Through the above steps, the present invention realizes a highly integrated and personalized aesthetic education learning system, which can provide personalized aesthetic education content recommendations based on students' emotional state and behavioral characteristics, and at the same time improve students' learning experience and effects through a comfortable learning environment and intuitive interactive interface.
[0103] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. An audio-visual integrated immersive aesthetic education cabin driven by facial emotion recognition, characterized by: include: Student client, administrator client and server; The student client is used to collect students' facial expression pictures, physiological signals and behavioral data, identify the students' current emotional state based on the facial expression pictures, physiological signals and behavioral data, and interact with the students; The server is configured to filter out audio and video content from an aesthetic education database based on the current emotional state; The administrator client is used for the administrator to manage the aesthetic education cabin, the aesthetic education database and the students' emotional state data; The student client includes a data acquisition module, an emotion recognition module, a content playback module and an information interaction module; The data acquisition module is used to collect the facial expression pictures, physiological signals and behavioral data of students; The emotion recognition module is used to perform feature splicing on the facial expression image and the physiological signal to obtain a comprehensive feature vector, and identify the student's current emotional state based on the comprehensive feature vector and the behavioral data; The content playing module is used to play the filtered audio and video content; The information interaction module is used for students to provide feedback on the audio and video content and select preferences; Identifying the student's current emotional state based on the comprehensive feature vector and behavioral data includes: Inputting the comprehensive feature vector into an improved ResNet-50 model for recognition, and outputting a high-dimensional feature vector, wherein the improved ResNet-50 model is obtained based on depthwise separable convolution; Input the high-dimensional feature vector into a first Bi-LSTM network with an attention mechanism to output a first feature, and input the first feature into a second Bi-LSTM network with an attention mechanism to obtain a sentiment feature sequence; Inputting the behavior data into a Deep Q-Learning behavior decision network and outputting a behavior feature sequence; The current emotional state is acquired based on the emotional feature sequence and the behavioral feature sequence.
2. The audio-visual integrated immersive aesthetic education cabin based on facial emotion recognition according to claim 1 is characterized in that: Performing feature splicing on the facial expression image and the physiological signal to obtain a comprehensive feature vector includes: The facial expression image and the physiological signal are subjected to feature splicing using a gradient boosting decision tree model to obtain the comprehensive feature vector.
3. The audio-visual integrated immersive aesthetic education cabin based on facial emotion recognition according to claim 1 is characterized in that: Acquiring the current emotional state based on the emotional feature sequence and the behavioral feature sequence includes: Inputting the emotion feature sequence and the behavior feature sequence into a long short-term memory network and outputting a feature matrix; The feature of the last time step in the feature matrix is selected and input into the Transformer Encoder model to obtain the feature representation of the current emotional state, and the feature representation of the current emotional state is input into the fully connected layer to obtain the current emotional state.
4. The audio-visual integrated immersive aesthetic education cabin based on facial emotion recognition according to claim 2 is characterized in that: The audio and video contents selected from the aesthetic education database according to the current emotional state include: The similarity between users is calculated based on the students' rating data and the current emotional state, the recommendation algorithm is optimized based on the similarity, and the audio and video content is screened based on the optimized recommendation algorithm.
5. The audio-visual integrated immersive aesthetic education cabin based on facial emotion recognition according to claim 4 is characterized in that: The method for calculating the similarity between users based on the student's rating data and the current emotional state is: , Among them, r ui and r vi Represents the ratings of user u and v on item i, e ui and e vi They represent the current emotional states of users u and v towards item i, respectively. w1 and w2 are the weights of the rating data and the current emotional state, respectively.
6. The audio-visual integrated immersive aesthetic education cabin based on facial emotion recognition according to claim 1 is characterized in that: The administrator client includes: a device management module, a content management module and a data analysis module; The device management module is used to check and set the status of the aesthetic education cabin, send device inspection instructions to the student client, and receive status feedback information; The content management module is used to add, delete, modify and check the audio and video content in the aesthetic education database; The data analysis module is used to perform data statistics and big data analysis on students' emotional states, and to perform statistics and analysis on content interaction data.
Citation Information
Patent Citations
Advertisement information pushing method and system
CN106997549A
Online learning system based on cloud fusion multi-modal analysis
CN112907406A