Religious teaching AR virtual reality experience system

By introducing AR virtual reality technology and deep learning models in religious teaching, a 3D virtual environment and perceptual interaction system is built, the problems of environmental simulation and user interaction in the existing technology are solved, personalized emotional recognition and feedback are achieved, and the effect and fun of religious teaching are improved.

CN120103978AInactive Publication Date: 2025-06-06CHENGDU UNIV

Patent Information

Application Number
CN202510216630.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

There are problems in existing religious teaching that complex environment simulation is difficult, user-perceived interaction accuracy and effect need to be improved, and the lack of speech and text understanding modules in specific religious fields.

Method used

Design a religious teaching AR virtual reality experience system, including the construction of a 3D virtual environment module, a perceptual interaction module, an intelligent feedback module, a multimodal fusion module and a speech and text understanding module in a specific religious field, and use a deep learning model to identify emotions and understand specific religious terms.

Benefits of technology

It realizes the construction of highly realistic religious scenes, real-time personalized emotional recognition and feedback, improves the fun and efficiency of teaching, and provides a rich, three-dimensional and interactive religious teaching experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103978A_ABST
    Figure CN120103978A_ABST
Patent Text Reader

Abstract

The invention discloses a religious teaching AR virtual reality experience system. According to the system, a 3D virtual environment module is constructed, a 3D environment conforming to the characteristics of a religious place is designed, and a model is established by using 3D modeling software. And through the perception interaction module, facial expressions, voices and physiological states of the user are captured and analyzed in real time. And through the intelligent feedback module, the virtual environment and the teaching content are dynamically adjusted according to the emotional state of the user, and personalized learning experience is provided. The system further comprises a multi-modal fusion module which provides a more comprehensive and accurate result for emotion recognition. The voice and text understanding module in the specific religious field enables the system to capture the emotion and demand of the user more accurately. Based on the cloud computing technology, the system realizes multi-user online interactive learning, shares a virtual environment and teaching resources, can perform deep analysis and evaluation optimization of user learning behaviors, and greatly improves the efficiency and effect of religious teaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of teaching, and more specifically relates to an AR virtual reality experience system for religious teaching. Background Art

[0002] Religious teaching is a special way of cultural inheritance. However, traditional religious teaching methods mainly rely on one-on-one or one-to-a-few face-to-face teaching, or knowledge imparting through books, audio and video, etc. This teaching method has low information transmission efficiency, limited teaching resources, lack of fun and participation, and greatly limits the teaching effect.

[0003] With the continuous development of science and technology, the emergence of virtual reality (VR) and augmented reality (AR) technologies has provided new possibilities for religious teaching. By constructing virtual religious scenes and combining virtual information with the actual environment using the characteristics of AR, learners can be placed in an immersive learning environment, improving their interest and effectiveness in learning. In addition, the perceptual interaction system, which uses facial expressions and voice as the main input information, can capture and identify user emotions in real time during the religious learning process, providing a personalized learning experience.

[0004] However, existing AR technology still faces some problems in religious teaching applications. For example, it is difficult to simulate and reconstruct complex religious environments, the accuracy and effect of user perception interaction need to be further improved, and there is a lack of voice and text understanding modules for specific religious fields.

[0005] Therefore, it is necessary to design an AR virtual reality experience system that can effectively solve the above problems and provide a rich, three-dimensional, and interactive religious teaching experience. Summary of the invention

[0006] The present invention aims to solve some problems in existing religious teaching, such as how to construct a complex 3D religious environment model, how to embed an effective perceptual interaction system to integrate multimodal information such as user emotions into real-time teaching feedback, how to understand specific religious terms and rhetoric, and how to use cloud computing and big data technology to achieve multi-user online interactive learning, share virtual environments and teaching resources, and ultimately improve the efficiency and quality of religious teaching.

[0007] In order to achieve the above object, the present invention adopts the following technical solution: the religious teaching AR virtual reality experience system comprises: To build a 3D virtual environment module, first design a 3D environment based on the historical background, architectural features, symbolic meaning, and cultural value of the religious site, and use 3D modeling software to build a model. In the scene layout stage, the created 3D model is placed in the virtual space; Perception and interaction module: captures and analyzes the user's facial expressions, voice, and physiological state in real time, and performs emotion recognition through a multimodal fusion deep learning model; Intelligent feedback module: dynamically adjusts the virtual environment and teaching content according to the user's emotional state, helping to provide a personalized learning experience; Multimodal fusion module: Integrates and analyzes emotion recognition information from different sources, such as facial expressions, voice, and physiological feedback, to obtain more comprehensive and accurate emotion recognition results; Religious domain-specific speech and text understanding modules: Understand specific religious terms, as well as religious rhetoric and metaphors embedded in speech and text to more accurately capture user emotions and needs.

[0008] In one solution, the 3D virtual environment module performs mapping and material processing after modeling is completed, so that the 3D model shows complex visual effects under lighting, thereby enhancing the user's visual experience.

[0009] In one embodiment, the facial expression recognition and speech recognition are performed by deep learning methods, convolutional neural networks (CNN) and recurrent neural networks (RNN), and the physiological signals are processed by suitable methods.

[0010] In one embodiment, the system fuses features extracted from facial expressions, speech, and physiological signals through a multimodal fusion model, and predicts emotional states through a multi-class classifier.

[0011] In one embodiment, the speech and text understanding module for a specific religious field uses the BERT model to perform semantic understanding, effectively understand and distinguish various religious terms, and accurately capture religious rhetoric and metaphors in speech and text.

[0012] In one embodiment, the multimodal fusion deep learning model includes: (1) In the feature extraction stage, for facial expression image data, a convolutional neural network (CNN) is used to extract features; the input image is X, and after being processed by the convolution layer, pooling layer, and fully connected layer, the feature representation is: ; For speech signals, recurrent neural networks (RNNs), especially long short-term memory networks (LSTMs), are used to process time series data; speech signals are sequences , the features after processing by the LSTM layer are: ; For physiological signals, Fourier transform (FFT) is used to extract frequency domain features; the heart rate signal is , its frequency domain characteristics are: ; (2) Multimodal feature fusion: represents the eigenvector of the i-th mode, then the independent modal features are written as ; fuse them, use the serial fusion method to merge the feature vectors from different modes to capture the association between them; there are n modes, and the feature vector of each mode is , and merge them using the serial fusion method: ; in, is the weight matrix, is the bias vector. This fused feature vector Integrates information from all modalities; (3) Emotion recognition: Train a multi-class classifier to predict the emotional state; use a fully connected neural network (FCNN) to classify the fused features and predict the user's emotional state; the emotion category is C, and the output of the fully connected layer is: ; in, and are the weights and biases of the output layer; the cross entropy loss function is used to measure the difference between the predicted results and the true labels:

[0013] in, It is a real emotional label. is the probability predicted by the model.

[0014] In one embodiment, the multimodal fusion module includes: (1) Data collection and feature extraction: First, emotion-related data is collected from multiple sources, including facial expressions, voice features, and physiological feedback. A camera is used to capture the user's facial image sequence, and a convolutional neural network is used to extract facial expression features. ,in, represents the input facial image; The user's voice signal is collected through a microphone, and the voice features are extracted using a long short-term memory network (LSTM). ,in, is a speech signal sequence; Acquire the user's physiological signals through sensors and extract features using signal processing technology; ,in, Represents physiological signals; (2) Feature fusion: The multimodal fusion system adopts a variety of fusion strategies, including feature-level fusion and decision-level fusion, to integrate information from different modalities. (3) Emotion recognition: The fused features are input into the classifier for emotion recognition.

[0015] Beneficial effects of the present invention: The present invention can construct a highly realistic religious scene by introducing AR and 3D virtual environment model construction technology, allowing participants to experience the religious environment as if they were there. The perception interaction module can capture and analyze the user's facial expressions, voice and physiological state in real time, making the teaching process more humane and achieving dynamic matching between the teaching process and the user's emotional state. The intelligent feedback module dynamically adjusts the virtual environment and teaching content according to the user's emotional state, so that everyone can enjoy a personalized learning experience.

[0016] The speech and text understanding module in the specific religious field of the present invention can accurately understand and process the symbolic meaning, cultural value and special terms used by a specific religion, demonstrating a strong language processing and understanding capability. The entire system is based on cloud computing technology and can be used and interacted with by multiple users online. It can not only provide a rich, three-dimensional and interactive religious teaching experience, but also realize the sharing of religious teaching resources, thus improving the efficiency and quality of religious teaching. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a system block diagram of the present invention; Figure 2 This is a flow chart for implementing interactive perception of the present invention; Figure 3 This is a flow chart of the multimodal fusion model of the present invention. DETAILED DESCRIPTION

[0018] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Typical embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0019] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used in the present invention in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Typical embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0020] like Figure 1 As shown, an improved religious teaching AR virtual reality experience system is provided, and the following are the main components of the system: (1) Constructing 3D virtual environment module: The system will create a 3D environment rich in historical and cultural connotations according to different religious and teaching needs to provide a realistic and immersive teaching experience.

[0021] In the process of creating a 3D virtual environment, a lot of preliminary research is needed first. Religious places are spaces full of historical and cultural connotations. It is necessary to have a deep understanding of the historical background, architectural features, symbolic meanings, cultural values, etc. of each religious place. Only in this way can these details be reproduced in the virtual environment, making it both educational and ornamental. The research includes reading relevant literature, browsing pictures, interviewing experts, and even visiting and experiencing these religious places in person.

[0022] Then, based on these studies, a 3D model is created. The process of creating a model includes planning, design, modeling, texturing, lighting, and rendering. In the planning stage, a sketch of the venue is drawn to determine its basic elements such as size, shape, layout, and color tone. In the design stage, the sketch is further refined to add details such as carvings, decorations, and furniture. In the modeling stage, 3D modeling software is used to start from basic shapes based on design drawings and gradually build up complex 3D shapes. In the texturing stage, color and texture are added to the model to make it more realistic. In the lighting stage, light sources are set to make the model produce shadows and reflections to enhance the three-dimensional effect. In the rendering stage, the model is converted into a 2D image for display on the screen.

[0023] After the 3D model is completed, it will move to the construction phase of the virtual environment. This includes steps such as scene layout, animation design, and sound effect addition. In the scene layout phase, the created 3D model will be placed in the virtual space according to the arrangement of the design drawings. In the animation design phase, actions will be added to certain elements, such as floating candlelight, melodious bells, walking crowds, etc., to make the scene more vivid. In the sound effect phase, appropriate music, singing, prayers, narration, etc. will be recorded or found to increase the atmosphere of the environment.

[0024] Use 3D modeling software such as Autodesk's 3ds Max or Maya, Blender, etc. to start the creation of 3D models. After the modeling is completed, the mapping and material processing will be carried out. This step is to make the 3D model look more realistic. Various properties such as surface texture, color, reflection, transparency, etc. will be added to make the appearance of the object show complex visual effects under light.

[0025] (2) Perception and interaction module: The system uses devices such as cameras and microphones to capture and analyze the user's facial expressions, voice, and physiological state (such as heart rate, fluctuating brain waves, eye movements, etc.) in real time, and performs emotion recognition through a multimodal fusion deep learning model.

[0026] The basic idea of ​​multimodal fusion is to extract features from multiple data sources and then fuse these features at a certain level so that the model can understand the input from different perspectives and make more accurate predictions.

[0027] In this AR virtual reality experience system, the user's facial expressions and voice are captured in real time through devices such as cameras and microphones, while the user's physiological state such as heart rate, brain waves and eye movement is monitored. These data are used to extract features and then processed through a multimodal fusion deep learning model. Finally, the system judges the user's emotional state based on this.

[0028] Modal information such as facial expressions and voice are considered to be very important features in emotion recognition, while physiological signals such as heart rate, brain waves and eye tracking can provide a richer source of information for emotion recognition.

[0029] like Figure 2 As shown, the following is the implementation process: S201. Feature extraction: First, for facial expressions and speech signals, we can use deep learning methods, such as convolutional neural networks (CNN) and recurrent neural networks (RNN), to transform sound waves into useful feature sequences. Face images can be used by CNN to extract facial expression features, and speech signals can be converted into characteristic time series features through RNN; at the same time, suitable methods can also be used to process and analyze physiological signals, such as heartbeat signals and the frequency domain characteristics of brain telegraph speech.

[0030] In the feature extraction stage, deep learning technology is mainly used to extract useful features from data of different modalities. For image data of facial expressions, convolutional neural network (CNN) is used to extract features. Assuming that the input image is X, after processing by convolution layer, pooling layer and fully connected layer, the feature representation is: ; For speech signals, recurrent neural networks (RNNs), especially long short-term memory networks (LSTMs), are used to process time series data. Assume that the speech signal is a sequence , the features after processing by the LSTM layer are: ; For physiological signals (such as heart rate and brain waves), Fourier transform (FFT) can be used to extract frequency domain features. Assume that the heart rate signal is , its frequency domain characteristics are: ;

[0031] S202, multimodal feature fusion: Considering the correlation between these modal information, these features are merged through multimodal fusion technology. Specifically, assuming that represents the eigenvector of the i-th mode, then the independent modal features can be written as . To fuse them, use the serial fusion method to merge the feature vectors from different modes to capture the association between them. Assume there are n modes, and the feature vector of each mode is , they can be merged using the serial fusion method: ; in, is the weight matrix, is the bias vector. This fused feature vector Integrates information from all modalities.

[0032] S203, Emotion Recognition: Determine the emotion category through a hardware network such as a fully connected layer. Predict the emotional state by training a multi-class classifier. Use a fully connected neural network (FCNN) to classify the fused features to predict the user's emotional state. Assuming the emotion category is C, the output of the fully connected layer is: ; in, and are the weights and biases of the output layer. The cross entropy loss function is used to measure the difference between the predicted results and the true labels: ; in, It is a real emotional label. is the probability predicted by the model. Through back-propagation and optimization algorithms such as Adam or SGD, the model can be trained to minimize the loss function, thereby improving the accuracy of emotion recognition.

[0033] Through the above steps, the system can effectively extract and fuse features from multimodal data and accurately identify the user's emotional state, providing support for personalized AR teaching experience.

[0034] After training, including back-propagation and optimization algorithms such as stochastic gradient descent (SGD), Adam, etc., the model can achieve unified management and emotion recognition of multimodal information such as facial expressions, voice and physiological status, thereby providing a personalized AR teaching experience.

[0035] (3) Multimodal fusion module: The system can flexibly integrate and analyze emotion recognition information from different sources, such as facial expressions, voice, and physiological feedback, to obtain more comprehensive and accurate emotion recognition results.

[0036] like Figure 3 As shown, S301, data collection and feature extraction, the multimodal fusion system first collects emotion-related data from multiple sources, including facial expressions, voice features, and physiological feedback.

[0037] Facial expression: A camera is used to capture a sequence of user’s facial images, and a convolutional neural network (CNN) is used to extract facial expression features.

[0038] ;

[0039] in, Represents the input face image.

[0040] Voice features: The user's voice signal is collected through a microphone, and the voice features are extracted using Mel-frequency cepstral coefficients (MFCC) or long short-term memory networks (LSTM).

[0041] ; in, is a speech signal sequence.

[0042] Physiological feedback: The user’s physiological signals, such as heart rate and skin electrical response, are acquired through sensors, and features are extracted using signal processing technology.

[0043] ;

[0044] in, Represents physiological signals.

[0045] S302, feature alignment and standardization, because data of different modalities may differ in time and scale, they need to be aligned and standardized. Each feature vector is normalized so that it can be fused at the same scale: ;

[0046] in, and are the characteristic mean and standard deviation of mode i respectively.

[0047] S303, feature fusion, the multimodal fusion system adopts a variety of fusion strategies, including feature-level fusion and decision-level fusion, to integrate information from different modalities.

[0048] Feature-level fusion: The standardized feature vectors are concatenated or weighted averaged to form a comprehensive feature vector: ;

[0049] Decision-level fusion: Emotion recognition is performed on each modality separately, and then the recognition results of each modality are fused by voting or weighted averaging.

[0050] Emotion recognition, the fused features are input into the classifier for emotion recognition. Deep neural network (DNN) can be used for classification and output the probability distribution of emotion categories: ; in, and are the weights and biases of the classifier.

[0051] S304, model training and optimization, by using the cross entropy loss function and optimization algorithm (such as Adam optimizer), adjust the model parameters to improve the accuracy of emotion recognition: ; Through the above steps, the multimodal fusion system can effectively integrate and analyze emotional information from different sources, providing accurate and comprehensive emotion recognition results for the intelligent feedback system. This multimodal approach improves the system's ability to cope with complex emotional states and enhances the personalization and adaptability of the user experience.

[0052] (4) Speech and text understanding modules for specific religious fields: The system can understand specific religious terms, as well as religious rhetoric and metaphors embedded in speech and text, thereby more accurately capturing users’ emotions and needs.

[0053] First, we use the BERT model to build a good semantic understanding model for a specific religious field. This model can effectively distinguish and understand various religious terms and accurately capture religious rhetoric and metaphors in speech and text.

[0054] The BERT model is a pre-trained deep learning model that can understand the semantics of words based on context. Through pre-training, the BERT model is able to understand language patterns in a large text corpus consisting of millions of sentences. During the pre-training phase, the BERT model first "predicts" certain words in a sentence through a method called a masked language model. The model then predicts the relationship between two sentences through another task called next sentence prediction. After pre-training on these two tasks, the BERT model has learned to understand the complex structure and patterns of language, so it can provide more accurate results in subsequent tasks (such as sentiment analysis, question answering, etc.).

[0055] For speech processing, speech recognition technology can be used to convert speech into text and then process it. Speech recognition is performed using the deep learning method mentioned earlier.

[0056] For BERT, its pre-training process can be summarized as follows: For the argument prediction task, a masked language model is used. Given a sentence, 15% of the words are randomly masked. Assuming there are m masked words, the loss function is: ; For the next sentence prediction task, given two sentences, determine whether the second sentence is the context of the first sentence. Let the loss function be: ; The above two losses are added together to get the final loss function, which is then trained through backpropagation and optimization algorithms such as Adam.

[0057] (5) Intelligent feedback module: The system will adjust the virtual environment and teaching process in real time based on the captured emotional information to match the user's emotional needs. For example, when the user feels uneasy or afraid, the system can help the user calm down by changing factors such as lighting and music in the environment; when the user feels confused or puzzled, the system will provide automated answer reminders and even provide personalized teaching consultation.

[0058] The intelligent feedback system is an important part of the religious teaching AR virtual reality experience system. It analyzes the user's emotional state in real time and dynamically adjusts the virtual environment and teaching process to provide a personalized learning experience. The following is the detailed implementation process of the intelligent feedback system: Emotional state analysis: First, the system obtains the user's emotional state through the perception interaction module. Through multimodal feature fusion and emotion recognition model, the system obtains the probability distribution of the user's current emotion. , where each Represents the probability that the user is in emotional state i.

[0059] Emotional state mapping, based on the emotion recognition results, the system maps the emotional state with the highest probability to the corresponding feedback action. For example, suppose the system recognizes that the user's emotional state is "anxiety", recorded as The system predefines a set of mapping rules from emotions to feedback actions: ; Virtual environment adjustment, according to the mapping rules, the system adjusts the parameters of the virtual environment in real time. Assuming that the vector Represents various parameters of the virtual environment, such as lighting, music, colors, etc. The system adjusts these parameters based on the emotional state: ;

[0060] in, is the environmental parameter adjustment vector calculated based on the "anxiety" state. For example, Maybe lower the light intensity and increase the volume of soothing music.

[0061] In addition to adjusting the environment, the system will also adjust the teaching content and process according to the emotional state. If the user shows a "confused" state, it is recorded as , the system will provide additional teaching support: ;

[0062] In this case, the system may insert additional prompt information or provide a virtual assistant for personalized consultation. The adjustment of teaching content can be achieved by dynamically selecting different teaching modules.

[0063] Through the above steps, the intelligent feedback system can provide timely and appropriate adjustments during the user's learning process, helping users get the best learning experience in the virtual environment. This dynamic adjustment mechanism not only improves user participation and satisfaction, but also enhances learning effects.

[0064] Example: Suppose you need to create a virtual religious place environment for a religious learner. First, collect a lot of information related to the religious place from the library and the Internet, including historical background, architectural features, symbolic meanings, etc. Then, based on this information, the designer creates a similar 3D model in SketchUp software.

[0065] Use the Unity 3D game engine to create a virtual 3D environment, import the 3D model into Unity, and use the built-in lighting and material tools to render the model to make it look more realistic. Next, add animation and sound effects to the virtual environment, such as floating candlelight, melodious bells, etc.

[0066] On the other hand, the system integrates multiple information such as facial expressions, voice and physiological feedback, and performs emotion recognition through a modal deep fusion model based on convolutional neural network (CNN) and recurrent neural network (RNN). For example, the user's heart rate is 72 bpm at the beginning, and after watching a religious scene, his heart rate rises to 78 bpm, and his facial expression becomes tense, and his voice becomes tense. This information is sent to CNN and RNN for processing, and the output is the emotion that the user may be experiencing.

[0067] Then, we scan the user's facial expressions and use the face API to get the expression change value, such as: eyebrows raised by 10%, eyes widened by 15%. For voice feedback, we convert it into text through the speech recognition API, and use NLP tools to perform semantic analysis to get the sentiment analysis results of the sentence.

[0068] Based on the analysis results, the system will adjust the virtual environment to suit the user’s mood. For example, in the above scenario, if the user is nervous and anxious, the system may reduce the light intensity of the environment and play some soothing music to help the user relax.

[0069] The system then adjusts the teaching process based on the stored user learning data, such as providing more background knowledge explanations when the user is anxious, or giving more prompts when the user is confused.

[0070] Through the above methods, the religious teaching AR virtual reality experience system can provide a real, interactive and personalized learning environment, enabling users to learn effectively in an immersive environment.

[0071] The present invention is mainly used for religious teaching. It designs a 3D environment and builds a model using 3D modeling software. The present invention deeply identifies user emotions from three levels: facial expression, voice and physiological state. The present invention also includes a voice and text understanding module for specific religious fields, which can understand specific religious terms, as well as religious rhetoric and metaphors embedded in voice and text, so as to capture the user's emotions and needs.

[0072] The present invention uses a multimodal deep learning model for emotion recognition, integrating and analyzing emotion recognition information from different sources; the multimodal fusion deep learning model of the present invention includes image data of facial expressions, voice signals, and physiological signals, etc.; The present invention also focuses on the intelligent feedback module, which can dynamically adjust the virtual environment and teaching content according to the user's emotional state. In summary, the present invention integrates emotion recognition in multiple aspects, including facial expressions, voice, and physiological feedback, understands voice and text in specific religious fields, and makes intelligent feedback adjustments based on the user's emotional state, thereby providing a personalized learning experience.

[0073] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0074] It should be understood that the detailed description of the technical solutions of the present invention by means of the preferred embodiments is illustrative rather than restrictive. A person skilled in the art may modify the technical solutions described in the embodiments, or replace some of the technical features by equivalents, based on reading the specification of the present invention; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A religious teaching AR virtual reality experience system, characterized by: The system comprises: To construct the 3D virtual environment module, firstly, the 3D environment is designed according to the historical background, architectural features, symbolic meaning and cultural value of the religious site, and the model is built using 3D modeling software; in the scene arrangement stage, the created 3D model is placed in the virtual space; Perception and interaction module: captures and analyzes the user's facial expressions, voice, and physiological state in real time, and performs emotion recognition through a multimodal fusion deep learning model; Multimodal fusion module: Integrates and analyzes emotion recognition information from different sources, such as facial expressions, voice, and physiological feedback, to obtain more comprehensive and accurate emotion recognition results; Religious domain-specific speech and text understanding modules: Understand specific religious terms, as well as religious rhetoric and metaphors embedded in speech and text to more accurately capture user emotions and needs; Intelligent Feedback Module: Dynamically adjusts the virtual environment and teaching content according to the user's emotional state, helping to provide a personalized learning experience.

2. A religious teaching AR virtual reality experience system according to claim 1, characterized in that: The 3D virtual environment module performs mapping and material processing after modeling is completed, so that the 3D model shows complex visual effects under lighting, thereby enhancing the user's visual experience.

3. A religious teaching AR virtual reality experience system according to claim 1, characterized in that: The facial expression recognition and speech recognition are achieved through deep learning methods, convolutional neural networks (CNN) and recurrent neural networks (RNN), and the physiological signals are processed through suitable methods.

4. The AR virtual reality experience system for religious teaching according to claim 1, characterized in that: The system fuses features extracted from facial expressions, speech and physiological signals through a multimodal fusion model, and predicts emotional states through a multi-class classifier.

5. The AR virtual reality experience system for religious teaching according to claim 1, characterized in that: The speech and text understanding module for specific religious fields adopts the BERT model for semantic understanding, effectively understands and distinguishes various religious terms, and accurately captures religious rhetoric and metaphors in speech and text.

6. The AR virtual reality experience system for religious teaching according to claim 4, characterized in that: The multimodal fusion deep learning model includes: (1) In the feature extraction stage, for facial expression image data, a convolutional neural network (CNN) is used to extract features; the input image is X, and after being processed by the convolution layer, pooling layer, and fully connected layer, the feature representation is: ; For speech signals, recurrent neural networks (RNNs), especially long short-term memory networks (LSTMs), are used to process time series data; speech signals are sequences , the features after processing by the LSTM layer are: ; For physiological signals, Fourier transform (FFT) is used to extract frequency domain features; the heart rate signal is , its frequency domain characteristics are: ; (2) Multimodal feature fusion: represents the eigenvector of the i-th mode, then the independent modal features are written as ; fuse them, use the serial fusion method to merge the feature vectors from different modes to capture the association between them; there are n modes, and the feature vector of each mode is , and merge them using the serial fusion method: ; in, is the weight matrix, is the bias vector, this fused feature vector Integrates information from all modalities; (3) Emotion recognition: Train a multi-class classifier to predict the emotional state; use a fully connected neural network (FCNN) to classify the fused features and predict the user's emotional state; the emotion category is C, and the output of the fully connected layer is: ; in, and are the weights and biases of the output layer; the cross entropy loss function is used to measure the difference between the predicted results and the true labels: ; in, It is a real emotional label. is the probability predicted by the model.

7. The AR virtual reality experience system for religious teaching according to claim 1, characterized in that: The multimodal fusion module comprises: (1) Data collection and feature extraction: First, emotion-related data is collected from multiple sources, including facial expressions, voice features, and physiological feedback. A camera is used to capture the user's facial image sequence, and a convolutional neural network is used to extract facial expression features. ,in, represents the input facial image; The user's voice signal is collected through a microphone, and the voice features are extracted using a long short-term memory network (LSTM). ,in, is a speech signal sequence; Acquire the user's physiological signals through sensors and extract features using signal processing technology; in, Represents physiological signals; (2) Feature fusion: The multimodal fusion system adopts a variety of fusion strategies, including feature-level fusion and decision-level fusion, to integrate information from different modalities. (3) Emotion recognition: The fused features are input into the classifier for emotion recognition.

Citation Information

Patent Citations

  • Low-resource speech recognition method and system, and speech model training method

    CN114242071A

  • Automatic classification of emotional awareness

    CN116982048A

  • Multi-mode XR emotion interaction method, system and device and storage medium

    CN119251438A

  • An educational platform that provides video content in an immersive format

    KR102391074B1

  • A system and method for providing chatbot services using a manufacturing-specific small language model-based generative ai

    KR102731386B1

Cited By

  • Situation interactive foreign language training system based on virtual reality

    CN120544439A