An interactive teaching system based on virtual reality
By introducing scene perception and mode switching, multi-domain audio management and hierarchical spatial audio rendering modules into the virtual reality teaching system, the problems of communication mode switching and audio rendering in multi-user environments are solved, achieving an efficient immersive teaching experience and real-time performance.
Patent Information
- Application Number
- CN202511747589.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-11-26
AI Technical Summary
Existing virtual reality teaching systems lack scene-adaptive communication mode switching and hierarchical spatial audio rendering in multi-user environments, resulting in problems such as speech delay, audio overlap, and spatial positioning distortion, which affect the immersion and real-time performance of teaching.
Through the scene perception and mode switching module, multi-domain audio management module, hierarchical spatial audio rendering module, and edge collaboration and network optimization module, adaptive communication mode switching and high-fidelity spatial audio interaction are achieved, including semantic region segmentation, multimodal behavior recognition, dynamic sparse matrix audio management and edge collaborative network optimization.
Balancing immersion, real-time performance, and system resource efficiency in multi-user interaction, it achieves smooth communication mode switching and high-precision spatial audio rendering, enhancing the immersive experience and interactive efficiency of virtual teaching.
Smart Images

Figure CN121214739B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of virtual reality and educational information technology, and more specifically, to an interactive teaching system based on virtual reality for realizing multi-user immersive teaching, real-time interaction and adaptive resource scheduling. Background Technology
[0002] With the rapid development of Virtual Reality (VR) technology, immersive teaching has gradually become an important direction for educational informatization. Traditional online teaching systems typically transmit information based on two-dimensional interfaces, lacking spatial immersion and interactive realism, making it difficult to effectively stimulate learners' interest and motivation. In recent years, the introduction of virtual reality technology has enabled teaching scenarios to be presented in three-dimensional space, allowing teachers and students to interact immersively through virtual avatars, providing new pathways for distance education, experimental teaching, and collaborative learning.
[0003] However, existing virtual teaching systems mostly focus on building visual presentation effects, lacking sophisticated management and dynamic scheduling mechanisms for audio interaction. Since teaching activities include various scenarios such as lecturing, discussion, Q&A, and self-study, the audio propagation range, communication modes, and resource requirements differ significantly across these scenarios. If the communication topology and audio rendering strategies are not adaptively adjusted according to the scenario, problems such as speech delay, audio overlap, and spatial positioning distortion often occur, thus affecting the immersive teaching experience and communication efficiency.
[0004] Furthermore, in multi-user concurrent virtual teaching environments, the system needs to process a large amount of audio streams and behavioral data simultaneously. Current technologies often employ a single mixing strategy or global broadcasting for audio data, which fails to reflect spatial hierarchy and interactive relationships, and also results in excessive bandwidth consumption and wasted computing resources. Simultaneously, due to differences in terminal device performance and network environment fluctuations, it is difficult to balance audio rendering quality with system smoothness, affecting the real-time performance and stability of the teaching system.
[0005] In summary, how to achieve scene-adaptive communication mode switching and hierarchical spatial audio rendering based on multimodal behavior recognition in virtual reality teaching environments, so as to balance immersion, real-time performance and system performance optimization in multi-user interaction, has become an urgent technical problem to be solved. Summary of the Invention
[0006] In order to overcome a series of defects in the existing technology, the purpose of this application is to provide an interactive teaching system based on virtual reality, which includes a scene perception and mode switching module, a multi-domain audio management module, a hierarchical spatial audio rendering module, a resource scheduling and performance optimization module, and an edge collaboration and network optimization module. The implementation of the interactive teaching system includes the following steps.
[0007] Obtain initial configuration information for the virtual teaching space, including the number of users, device performance parameters, and network topology status, and perform semantic region division on the virtual teaching space.
[0008] The scene perception and mode switching module collects users' multimodal teaching behavior data in real time, combines the semantic region segmentation information to identify the current teaching scene status, and classifies the teaching scene status into teaching status, discussion status, self-study status, or Q&A status.
[0009] The target communication mode is determined based on the teaching scenario state. The target communication mode includes a broadcast mode, a group discussion mode, a neighborhood communication mode, and a private dialogue mode. A smooth switch from the current communication mode to the target communication mode is then performed.
[0010] An audio domain matrix is established using the multi-domain audio management module. The audio domain matrix is used to represent the participation status of each user in different communication domains, and a corresponding weight coefficient is assigned to each communication domain.
[0011] The source location, volume, and direction information of all audio streams in the virtual teaching space are collected. Based on the audio domain matrix and the weighting coefficients, the multi-domain audio is weighted and mixed. Audio conflicts are eliminated through a competition suppression strategy to generate personalized mixed audio output for each user.
[0012] The hierarchical spatial audio rendering module obtains the distance between each audio source and the target user, the role importance parameter, and the user attention data. Based on the distance, the role importance parameter, and the user attention data, the audio stream is divided into fine-level LOD, standard-level LOD, and economic-level LOD.
[0013] The system allocates complete HRTF spatial audio rendering resources to the fine-level LOD audio stream, simplified stereo positioning resources to the standard-level LOD audio stream, and basic volume attenuation resources to the economy-level LOD audio stream, thereby achieving low-cost and high-precision three-dimensional spatial audio positioning.
[0014] Furthermore, the semantic region division of the virtual teaching space includes the following steps.
[0015] The virtual teaching space is meshed in three dimensions, and the space is divided into several cubic units, each with a side length of 0.5 meters to 2 meters, which are used to form the basic analysis units of the virtual space.
[0016] Based on the teaching function requirements, the podium area, student seating area, experimental operation area, group discussion area and free activity area are marked on the three-dimensional grid, and each area is assigned a corresponding semantic label and priority value for subsequent spatial semantic recognition.
[0017] The system tracks the coordinates of the user's virtual avatar in the 3D grid in real time, combines the user's dwell time and movement trajectory characteristics to determine the semantic region where the user is currently located, and uses the information of this region as the spatial context for scene state recognition.
[0018] When multiple users gather in the same semantic region, the attribute parameters of that region are automatically extracted, including the maximum number of users, the recommended audio rendering mode, and the network transmission strategy.
[0019] Furthermore, the collection of the multimodal teaching behavior data includes...
[0020] The head-mounted display device continuously acquires the user's head pitch angle, yaw angle, and roll angle using built-in attitude sensors at a sampling frequency of no less than 90Hz, which is used to determine the direction of gaze and focus of attention.
[0021] The controller tracks the position and gestures of the user's hands in the virtual teaching space in real time, and recognizes teaching interaction behaviors including raising hands, pointing, writing and operating virtual objects.
[0022] The system collects the user's audio signals through a voice input device, distinguishes between speaking and silent states, and extracts features such as speaking duration, volume, and speech rate.
[0023] Record the virtual avatar's movement speed, stopping position, and relative distance to other users to characterize the user's spatial movement and interaction behavior.
[0024] The head posture data, hand movement data, voice data, and spatial behavior data are timestamped, aligned, and fused to generate a comprehensive behavior vector that reflects the level of teaching participation and interaction intent.
[0025] Furthermore, the identification of the teaching scenario status includes the following steps.
[0026] A predefined set of feature rules for teaching scenario states is used to describe the spatial, posture, voice, and group behavior characteristics of different teaching scenario states. Among them, the teaching state rules include: users in the podium area maintain a standing posture and have a voice activity level of more than 0.7, while other users are in the seat area and more than 60% of their heads are facing the user; the discussion state rules include: forming groups of three to five people according to semantic regions, with members within the group being less than 1.5 meters apart and multiple people speaking at the same time, with an average spatial distance between groups greater than 4 meters and an audio crossover degree between groups less than 15%.
[0027] Based on the aforementioned feature rules, a preliminary judgment is made on the current multimodal user behavior data to generate candidate teaching scenario states.
[0028] The multimodal user behavior vector sequence collected in the past ten seconds is input into a lightweight temporal convolutional network model. This model contains three one-dimensional convolutional layers, each with 32 convolutional kernels, which are used to extract temporal features and output the probability distribution of each scene state.
[0029] Based on the probability distribution output by the model, the scenario state with the highest probability is selected as the machine learning judgment result.
[0030] The rule-based judgment results are combined with the machine learning judgment results to generate the final candidate teaching scenario state.
[0031] A state transition is only executed when the confidence level of a newly identified state exceeds a preset threshold of 0.8 for three consecutive seconds, in order to ensure the stability and accuracy of scene recognition.
[0032] The confirmed scene state is used as the current teaching scene identification result to drive the interaction strategy, audio rendering, and resource scheduling of the virtual teaching system.
[0033] Furthermore, the smooth switching of the communication mode includes the following steps.
[0034] Determine whether the current communication mode needs to be switched based on the teaching scenario or user needs.
[0035] When switching to group discussion mode, the volume of users outside the same group is linearly reduced to 15% during a 0.5-second transition period, while the volume of users within the same group remains unchanged.
[0036] When switching to the proximity communication mode, a spatial range expansion strategy is adopted to linearly expand the effective radius of audio propagation from the initial 1 meter to 3 meters, with the expansion speed controlled at 0.8 meters per second.
[0037] The audio domain matrix is recalculated based on the new communication model, and an independent audio domain identifier and weight parameter are assigned to each discussion group or spatial range.
[0038] During the audio domain matrix update, timestamps are added to the audio data packets being transmitted to ensure that the audio frames before and after the switch are played in the correct timing, avoiding audio jumps or repetitions.
[0039] Throughout the switching process, the audio stream is kept continuous to ensure uninterrupted teaching and communication, achieving a smooth transition.
[0040] Furthermore, the audio domain matrix is stored and computed using a dynamic sparse matrix structure: the rows of the matrix represent all users in the virtual teaching space, the columns represent currently active communication domains, and the value of each matrix element ranges from 0 to 1, representing the corresponding user's participation intensity in that communication domain; when a user is fully involved in a communication domain, the element value is 1, when not involved at all, it is 0, and when partially involved, it takes an intermediate value; the matrix elements are dynamically updated based on the user's semantic region location, the current interaction object, and historical communication records, with an update frequency of five times per second; when performing audio mixing computation, the corresponding matrix row vector is extracted for each target user, and the vector is multiplied by the audio streams of all communication domains, automatically filtering out domains with zero participation intensity, and only performing weighted summation on the audio of non-zero domains; at the same time, when no user participates or audio activity occurs in a communication domain for five consecutive seconds, the domain is automatically removed from the matrix, releasing computing resources and reducing the matrix dimension, maintaining a lightweight operating state.
[0041] Furthermore, the implementation of the interactive teaching system also includes the following steps.
[0042] During the audio rendering process, the resource scheduling and performance optimization module monitors the device's frame rate and computing load in real time and dynamically allocates audio processing resources according to a preset computing budget pool.
[0043] If the frame rate of the device is lower than the preset frame rate threshold, the LOD level of some audio streams will be reduced or the number of audio streams participating in the mixing will be reduced, and an adaptive bitrate control strategy will be activated to compress the audio data.
[0044] If the device's frame rate is higher than the preset frame rate threshold and there is sufficient computational budget, then the LOD level of some audio streams will be increased or the accuracy of spatial audio rendering will be improved.
[0045] The edge collaboration and network optimization module is used to determine the spatial distance between each audio source and the target user. Audio streams with a distance less than a preset proximity threshold are marked as near-distance audio, and audio streams with a distance greater than the preset proximity threshold are marked as far-distance audio.
[0046] The near-field audio is rendered locally in real time on the client, while the far-field audio is pre-mixed on the edge server and then sent to the client.
[0047] The network topology is dynamically switched based on the current teaching scenario and user distribution density. If users are densely distributed and in group discussion mode, a P2P network topology is used to achieve direct connection within the group; if users are scattered or in a broadcast mode, a server relay network topology is used to achieve unified distribution.
[0048] Repeat the steps of teaching scenario status recognition, communication mode switching, audio mixing processing, spatial audio rendering, and dynamic resource scheduling until the teaching activities in the virtual teaching space end.
[0049] Furthermore, the network topology is dynamically switched according to the current teaching scenario status and user distribution density. If users are densely distributed and in group discussion mode, a P2P network topology is used to achieve direct connection within the group; if users are scattered or in a broadcast mode, a server relay network topology is used to achieve unified distribution, including the following steps.
[0050] Based on the three-dimensional coordinates and distribution density information of users in the current virtual teaching space, the user aggregation situation in each semantic region is calculated in real time, providing basic data for network topology switching.
[0051] Based on the current teaching scenario, determine whether the user is in a group discussion, lecturing, or broadcasting to all, and determine the corresponding network topology switching strategy.
[0052] When users are densely distributed and in group discussions, a P2P direct connection channel is established between users in the same group to enable direct transmission of audio and data within the group, thereby improving the real-time nature of interaction.
[0053] When users are scattered or in a state of broadcasting to all users, a server relay network topology is enabled. The regional access server and the main server mix the audio stream and then distribute it to each client in a unified manner to ensure global consistency.
[0054] During topology switching, the connection status between each client and the P2P node or server is updated in real time to maintain the continuity of data transmission and the synchronous playback of audio streams.
[0055] Continuously monitor network latency, bandwidth usage, and packet loss rate. If any abnormalities or excessive load occur, automatically adjust network topology parameters or select auxiliary relay nodes to ensure the stability and communication quality of virtual teaching activities.
[0056] Compared with existing technologies, this application has the following advantages: through the collaborative design of scene perception and mode switching, audio domain matrix management, hierarchical spatial audio rendering and edge collaborative network optimization, it realizes adaptive communication mode switching and high-fidelity spatial audio interaction based on multimodal behavior recognition in virtual teaching environment, thereby taking into account immersion, real-time performance and system resource efficiency in multi-user real-time interaction. Attached Figure Description
[0057] Figure 1 This is a schematic diagram illustrating the implementation process of an interactive teaching system based on virtual reality disclosed in an embodiment of this application. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some embodiments of this invention, but not all embodiments.
[0059] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] The embodiments and directional terms described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0061] like Figure 1 As shown, an implementation method for an interactive teaching system based on virtual reality includes a scene perception and mode switching module, a multi-domain audio management module, a hierarchical spatial audio rendering module, a resource scheduling and performance optimization module, and an edge collaboration and network optimization module. The implementation of the interactive teaching system includes the following steps.
[0062] Step 1: Obtain the initial configuration information of the virtual teaching space, including the number of users, device performance parameters, and network topology status, and perform semantic region division on the virtual teaching space.
[0063] The semantic region division of the virtual teaching space includes the following steps.
[0064] The virtual teaching space is meshed in three dimensions, and the space is divided into several cubic units, each with a side length of 0.5 meters to 2 meters, which are used to form the basic analysis units of the virtual space.
[0065] Based on the teaching function requirements, the podium area, student seating area, experimental operation area, group discussion area and free activity area are marked on the three-dimensional grid, and each area is assigned a corresponding semantic label and priority value for subsequent spatial semantic recognition.
[0066] The system tracks the coordinates of the user's virtual avatar in the 3D grid in real time, combines the user's dwell time and movement trajectory characteristics to determine the semantic region where the user is currently located, and uses the information of this region as the spatial context for scene state recognition.
[0067] When multiple users gather in the same semantic region, the attribute parameters of that region are automatically extracted, including the maximum number of users, the recommended audio rendering mode, and the network transmission strategy.
[0068] This embodiment achieves dynamic perception and precise modeling of the teaching environment by acquiring the initial configuration information of the virtual teaching space and performing refined 3D meshing and semantic region division. It can label functional areas such as the podium, student seats, experimental operations, group discussions, and free activities, and assign semantic labels and priorities to each area, enabling precise acquisition of spatial context information. By tracking the coordinates, dwell time, and movement trajectory of the user's virtual avatar in real time, it can accurately determine the semantic region where the user is located, providing a reliable basis for recognizing the state of the teaching scene. When multiple users are concentrated in the same area, it automatically extracts area attribute parameters, including the maximum capacity, recommended audio rendering mode, and network transmission strategy, thereby providing precise support for subsequent interactive communication, audio processing, and resource scheduling, effectively improving the adaptability, interaction efficiency, and immersive experience of the virtual teaching environment.
[0069] Step 2: Collect users' multimodal teaching behavior data in real time through the scene perception and mode switching module, identify the current teaching scene status by combining the semantic region segmentation information, and classify the teaching scene status into teaching status, discussion status, self-study status, or Q&A status.
[0070] The collection of the multimodal teaching behavior data includes...
[0071] The head-mounted display device continuously acquires the user's head pitch angle, yaw angle, and roll angle using built-in attitude sensors at a sampling frequency of no less than 90Hz, which is used to determine the direction of gaze and focus of attention.
[0072] The controller tracks the position and gestures of the user's hands in the virtual teaching space in real time, and recognizes teaching interaction behaviors including raising hands, pointing, writing and operating virtual objects.
[0073] The system collects the user's audio signals through a voice input device, distinguishes between speaking and silent states, and extracts features such as speaking duration, volume, and speech rate.
[0074] Record the virtual avatar's movement speed, stopping position, and relative distance to other users to characterize the user's spatial movement and interaction behavior.
[0075] The head posture data, hand movement data, voice data, and spatial behavior data are timestamped, aligned, and fused to generate a comprehensive behavior vector that reflects the level of teaching participation and interaction intent.
[0076] This embodiment uses a scene perception and mode switching module to collect user head posture, hand movements, voice signals, and virtual avatar spatial behavior in real time. It then timestamps and fuses this multimodal data to form a comprehensive behavior vector, enabling precise representation of user attention, interactive behavior, and engagement levels. Combined with semantic region information from the virtual teaching space, it can dynamically identify the current teaching scenario state and classify it into lecture, discussion, self-study, or Q&A modes, thus providing a reliable basis for subsequent communication mode selection, audio management, and resource scheduling.
[0077] The identification of the teaching scenario status includes the following steps.
[0078] A predefined set of feature rules for teaching scenario states is used to describe the spatial, posture, voice, and group behavior characteristics of different teaching scenario states. Among them, the teaching state rules include: users in the podium area maintain a standing posture and have a voice activity level of more than 0.7, while other users are in the seat area and more than 60% of their heads are facing the user; the discussion state rules include: forming groups of three to five people according to semantic regions, with members within the group being less than 1.5 meters apart and multiple people speaking at the same time, with an average spatial distance between groups greater than 4 meters and an audio crossover degree between groups less than 15%.
[0079] Based on the aforementioned feature rules, a preliminary judgment is made on the current multimodal user behavior data to generate candidate teaching scenario states.
[0080] The multimodal user behavior vector sequence collected in the past ten seconds is input into a lightweight temporal convolutional network model. This model contains three one-dimensional convolutional layers, each with 32 convolutional kernels, which are used to extract temporal features and output the probability distribution of each scene state.
[0081] Based on the probability distribution output by the model, the scenario state with the highest probability is selected as the machine learning judgment result.
[0082] The rule-based judgment results are combined with the machine learning judgment results to generate the final candidate teaching scenario state.
[0083] A state transition is only executed when the confidence level of a newly identified state exceeds a preset threshold of 0.8 for three consecutive seconds, in order to ensure the stability and accuracy of scene recognition.
[0084] The confirmed scene state is used as the current teaching scene identification result to drive the interaction strategy, audio rendering, and resource scheduling of the virtual teaching system.
[0085] This embodiment uses predefined spatial, gestural, vocal, and group behavior feature rules for teaching scenario states, combined with a lightweight temporal convolutional network to extract temporal features from multimodal user behavior data, achieving accurate identification of scenario states such as lecturing, discussion, self-study, and Q&A. The rule-based judgments are fused with machine learning judgments, and a three-second confidence threshold verification is used to ensure the stability and reliability of state transitions. This embodiment can reflect the interaction patterns of multiple users in a virtual teaching space in real time and accurately, providing a reliable basis for communication strategy selection, audio rendering allocation, and resource scheduling, significantly improving the interactive response accuracy and overall immersive experience of the virtual teaching environment.
[0086] Step 3: Determine the target communication mode based on the teaching scenario status. The target communication mode includes broadcast mode, group discussion mode, neighborhood communication mode and private dialogue mode, and perform a smooth switch from the current communication mode to the target communication mode.
[0087] The smooth switching of the communication mode includes the following steps.
[0088] Determine whether the current communication mode needs to be switched based on the teaching scenario or user needs.
[0089] When switching to group discussion mode, the volume of users outside the same group is linearly reduced to 15% during a 0.5-second transition period, while the volume of users within the same group remains unchanged.
[0090] When switching to the proximity communication mode, a spatial range expansion strategy is adopted to linearly expand the effective radius of audio propagation from the initial 1 meter to 3 meters, with the expansion speed controlled at 0.8 meters per second.
[0091] The audio domain matrix is recalculated based on the new communication model, and an independent audio domain identifier and weight parameter are assigned to each discussion group or spatial range.
[0092] During the audio domain matrix update, timestamps are added to the audio data packets being transmitted to ensure that the audio frames before and after the switch are played in the correct timing, avoiding audio jumps or repetitions.
[0093] Throughout the switching process, the audio stream is kept continuous to ensure uninterrupted teaching and communication, achieving a smooth transition.
[0094] This embodiment dynamically determines communication mode switching based on the teaching scenario status or user needs, and combines transition time control and volume adjustment strategies to achieve smooth switching between modes such as group discussions and proximity communication. During the switching process, the volume of non-target users is linearly adjusted to expand the audio propagation range, and independent audio domain identifiers and weight parameters are reallocated to each discussion group or spatial range. At the same time, timestamps are added to audio data packets to ensure correct frame timing, ensuring the continuity of audio streams and uninterrupted interaction. This achieves a smooth and natural transition of communication modes in the virtual teaching environment, improving the stability and immersion of multi-user interaction.
[0095] The audio domain matrix is stored and computed using a dynamic sparse matrix structure: the rows of the matrix represent all users in the virtual teaching space, the columns represent currently active communication domains, and the value of each matrix element ranges from 0 to 1, representing the corresponding user's participation intensity in that communication domain; when a user is fully involved in a communication domain, the element value is 1, when not involved at all, it is 0, and when partially involved, it takes an intermediate value; the matrix elements are dynamically updated based on the user's semantic region location, the current interaction object, and historical communication records, with an update frequency of five times per second; when performing audio mixing computation, the corresponding matrix row vector is extracted for each target user, and the vector is multiplied by the audio streams of all communication domains, automatically filtering out domains with zero participation intensity, and only performing weighted summation on the audio of non-zero domains; at the same time, when there is no user participation or audio activity in a communication domain for five consecutive seconds, the domain is automatically removed from the matrix, releasing computing resources and reducing the matrix dimension to maintain a lightweight operating state.
[0096] This embodiment employs a dynamic sparse matrix structure to store and compute the audio domain matrix, enabling precise characterization of user and communication domain participation intensity in the virtual teaching space. Matrix elements are dynamically updated five times per second based on the user's semantic region location, interaction objects, and historical communication records, reflecting the user's real-time participation status. During audio mixing computation, the row vector corresponding to the target user is extracted, and weighted summation is performed only on the audio streams of non-zero participation domains, automatically filtering out irrelevant audio and ensuring mixing accuracy and computational efficiency. Simultaneously, communication domains with no user participation or audio activity for five consecutive seconds are automatically removed, freeing up computational resources and reducing matrix dimensionality. This achieves lightweight and efficient dynamic audio management, improving the real-time performance and stability of multi-user interaction in virtual teaching.
[0097] Step 4: Use the multi-domain audio management module to establish an audio domain matrix. The audio domain matrix is used to represent the participation status of each user in different communication domains and assign a corresponding weight coefficient to each communication domain.
[0098] The allocation of the weighting coefficients includes the following steps.
[0099] First, a basic weight value is preset for each communication domain type, with the basic weight being 0.8 for the teacher lecture domain, 0.6 for the group discussion domain, 0.4 for the neighboring communication domain, and 0.2 for the background environment domain.
[0100] The base weights are dynamically adjusted based on the user's role in the audio source. The weight coefficient for the teacher role is multiplied by a gain factor of 1.5, the weight coefficient for the teaching assistant role is multiplied by a gain factor of 1.2, while the base weight value for the ordinary student role remains unchanged. This is to reflect the priority of different roles in teaching and communication.
[0101] When a user is asking or answering a question, the weight coefficient of the communication domain to which the user belongs is temporarily increased by 0.3 during the speaking period. After the speaking ends, the weight coefficient is restored to the original value after a two-second decay period to ensure that the speaking user is prominently presented in the mixed audio.
[0102] The sound type is determined based on the spectral characteristics of the audio signal. Audio containing speech components is given a higher weight, while ambient sound effects or background music are given a lower weight. The weight ratio is controlled at approximately 3:1 to ensure that speech information dominates the mixed audio.
[0103] After all weighting coefficients are adjusted, they are normalized to ensure that the total energy of the mixed audio remains within a reasonable range, avoiding volume overload or distortion due to weight stacking.
[0104] The audio from each communication domain is mixed according to the normalized weight coefficients, so that the sound of different communication domains presents a distinct auditory effect to the user's ears, thereby improving the overall auditory experience.
[0105] Step 5: Collect the source location, volume, and direction information of all audio streams in the virtual teaching space; perform weighted mixing calculation on the multi-domain audio based on the audio domain matrix and the weight coefficients; and eliminate audio conflicts through a competition suppression strategy to generate personalized mixed audio output for each user.
[0106] This embodiment achieves refined management of multi-domain audio in a virtual teaching space by setting preset basic weight values for each communication domain and dynamically adjusting them in conjunction with user roles, speaking status, and audio types. The weight differences for teachers, teaching assistants, and students reflect the priority of different roles in teaching communication. Temporarily increasing the weight of speaking users ensures prominent speech presentation, while spectral analysis distinguishes speech from ambient sound, guaranteeing the dominant position of speech information. Weight normalization avoids mixing overload or distortion, achieving overall energy balance. In audio mixing calculation, a dynamic audio domain matrix and normalized weight coefficients are combined to perform weighted summation of each audio stream and apply a competition suppression strategy to eliminate conflicts, thereby generating personalized, layered mixed audio output. This allows each user to obtain a clear, natural, and highly immersive auditory experience in the virtual teaching environment, while ensuring real-time multi-user interaction and accurate audio information transmission.
[0107] Eliminating audio conflicts through a competition suppression strategy includes the following steps.
[0108] The instantaneous loudness values of all audio streams to be mixed are monitored in real time. When it is detected that the loudness of more than three audio streams exceeds the negative 12 LUFS threshold at the same time, an audio competition conflict is determined to have occurred.
[0109] The overall priority score of each audio stream is calculated based on the role weight of the audio source, the importance of the communication domain, the spatial distance from the target user, and the duration of the current speech, and the audio streams are sorted according to priority.
[0110] The two to three highest priority audio streams are retained for normal mixing, while the remaining audio streams are gated and attenuated, and their volume is smoothly reduced to 25% of the original value within 0.3 seconds.
[0111] Before performing gating attenuation, the suppressed audio is buffered for a short period of two seconds. After the high-priority audio finishes, the buffer is checked for any important audio segments that have not been played. If they are found, they are played at 70% volume with a delay to ensure that important information is not lost.
[0112] Analyze the spectral distribution characteristics of each audio stream to determine the frequency bands where the main energy of different audio streams is concentrated, and adjust the gain of each frequency band through a multi-band dynamic compressor so that the sound of different frequency bands remains recognizable after mixing.
[0113] This embodiment manages audio conflicts in the virtual teaching space in real time through a competition suppression strategy. When multiple audio streams simultaneously exceed the loudness threshold, a comprehensive priority is calculated based on the audio source role weight, communication domain importance, spatial distance, and speaking duration. The audio stream with the highest priority is retained for mixing, while the remaining audio streams undergo smooth gating attenuation. Simultaneously, a short-term buffering mechanism delays the playback of suppressed important speech segments to ensure that critical information is not lost. Furthermore, the audio spectrum distribution is analyzed, and a multi-band dynamic compressor is used to adjust the gain of each band, ensuring that the different audio streams remain clearly distinguishable after mixing. This guarantees the continuity, information integrity, and auditory layering of audio output in a high-density multi-user interactive environment, effectively enhancing the interactive effect and immersive experience of virtual teaching.
[0114] Step 6: Obtain the distance between each audio source and the target user, the role importance parameter, and the user attention data through the hierarchical spatial audio rendering module. Based on the distance, the role importance parameter, and the user attention data, divide the audio stream into fine-level LOD, standard-level LOD, and economic-level LOD.
[0115] The classification of audio stream LOD levels includes the following steps.
[0116] Calculate the Euclidean distance between each audio source and the target user. Audio sources with a distance of less than 3 meters receive 5 points, audio sources with a distance between 3 and 8 meters receive 3 points, and audio sources with a distance greater than 8 meters receive 1 point.
[0117] The scores are based on the importance of the audio source in the teaching activity: 5 points for the teacher or the questioner who is speaking, 3 points for the active member of the group discussion, and 1 point for the silent user.
[0118] The user's attention focus is obtained by eye tracking or head orientation analysis. When the audio source is located in a 30° cone-shaped area in the center of the user's field of vision and the gaze lasts for more than 1 second, the user's attention is scored 5 points. When the audio source is located in the peripheral field of vision, the user is scored 2 points. When the audio source is located outside the field of vision, the user is scored 0 points.
[0119] The three scores are weighted and summed, with distance accounting for 40%, role importance accounting for 35%, and attention accounting for 25%, to obtain the overall priority score of the audio stream, with a maximum score of 15 points.
[0120] The audio stream is divided into different LOD levels based on the overall score. A total score of 11 or higher is classified as Fine LOD, a total score between 6 and 11 is classified as Standard LOD, and a total score less than 6 is classified as Economy LOD. A hysteresis interval is set so that the level switch is only performed when the score changes by more than 2 points and the duration reaches 1 second, in order to ensure a smooth transition in audio rendering quality and improve the efficiency of computing resource utilization.
[0121] This embodiment prioritizes and assigns Level of Detail (LOD) levels to each audio stream by comprehensively considering the spatial distance between the audio source and the target user, the importance of the teaching role, and the user's attention focus. Distance, role, and attention scores are weighted and summed to form an overall score, which determines whether the audio stream should be assigned to a Fine, Standard, or Economic LOD level. A hysteresis interval is set, and the LOD level is switched only after the score change exceeds a threshold and persists for a certain period, thereby achieving a smooth transition in audio rendering and efficient resource utilization. This ensures that key audio sources receive priority presentation in the user experience while reducing the rendering load of non-critical audio, improving the accuracy and real-time performance of spatial audio in the virtual teaching environment, optimizing computational resource allocation, and enhancing overall immersion and interactive effects.
[0122] Step 7: Allocate complete HRTF spatial audio rendering resources to the fine-level LOD audio stream, allocate simplified stereo positioning resources to the standard-level LOD audio stream, and allocate basic volume attenuation resources to the economy-level LOD audio stream to achieve low-cost and high-precision three-dimensional spatial audio positioning.
[0123] Allocating complete HRTF spatial audio rendering resources for the fine-level LOD audio stream includes the following steps.
[0124] Step S1: Preload a high-precision head-related transfer function database containing 250 azimuth measurement points. The database covers a spatial range of 360 degrees in the horizontal direction and -40 degrees to 90 degrees in the vertical direction. Each measurement point contains the impulse response filter coefficients of the left and right ears. At the same time, load the geometric structure information of the virtual teaching space and the acoustic characteristics of the wall materials. The reflection coefficient of the hard wall material is set between 0.6 and 0.85, and the reflection coefficient of the soft sound-absorbing material is set between 0.15 and 0.4.
[0125] Step S2: Real-time detection of the three-dimensional spatial coordinates of the audio source corresponding to the fine-level audio stream relative to the user's head, including azimuth, elevation, and distance parameters; and calculation of the head-related transfer function coefficients in the corresponding direction from the HRTF database based on the real-time spatial position of the audio source.
[0126] Step S3: The fine-level audio stream signal is processed by the HRTF filter through a time-domain convolution method to simulate the time difference and intensity difference of the sound waves arriving at the left and right ears, generating a binaural audio signal with orientation perception characteristics.
[0127] Step S4: Calculate the first reflection path of the sound based on the wall material and geometry of the virtual teaching space, and generate no more than five early reflection components. The attenuation coefficient of the reflected sound energy is set according to the acoustic properties of the material, and the reflection delay time is calculated based on the sound path difference. The early reflection components are superimposed on the binaural audio signal to form an intermediate audio signal with a sense of spatial immersion.
[0128] Step S5: Determine whether the audio source is in a moving state. If it is in a moving state, re-execute step S2 for each frame to calculate new interpolation coefficients and apply a new filter. If it is in a stationary state, keep the current filter parameters unchanged.
[0129] Step S6: Output the fine-level audio stream that has undergone full HRTF spatial audio rendering processing, enabling the user to accurately perceive the position and motion trajectory of the sound source in three-dimensional space.
[0130] This embodiment achieves high-precision 3D sound source localization and spatial perception in a virtual teaching space by allocating complete HRTF spatial audio rendering resources to a fine-grained LOD audio stream. A comprehensive high-precision HRTF database is pre-loaded, and combined with the acoustic characteristics of the virtual space geometry and wall materials, the real-time 3D position of the audio source relative to the user is calculated. The audio signal is processed through temporal convolution to generate binaural audio, simulating the time difference and intensity difference of sound wave arrival. Simultaneously, early reflection paths are calculated and superimposed onto the binaural signal to enhance the sense of spatial immersion. Filter parameters are updated in real-time for moving audio sources to ensure the accuracy of dynamic sound source localization. Through this embodiment, users can accurately perceive the position, direction, and trajectory of sound sources in a virtual teaching environment, achieving a highly immersive 3D spatial audio experience, while providing clear, natural, and refined audio presentation for highly interactive, multi-user teaching scenarios.
[0131] Step 8: During the audio rendering process, the resource scheduling and performance optimization module monitors the device's frame rate and computing load in real time, and dynamically allocates audio processing resources according to the preset computing budget pool.
[0132] Step 9: If the frame rate of the device is lower than the preset frame rate threshold, reduce the LOD level of some audio streams or reduce the number of audio streams participating in the mixing, and start the adaptive bitrate control strategy to compress the audio data.
[0133] The adaptive bitrate control strategy implements multi-level dynamic compression for audio data transmission: high-definition audio uses 128 kilobits per second stereo encoding, standard audio uses 64 kilobits per second mono encoding, economy audio uses 32 kilobits per second compression encoding, and the lowest level uses 16 kilobits per second ultra-compression encoding. When the bandwidth utilization exceeds 75% or the loss rate exceeds 3%, the network is determined to be in a congested state. At this time, long-distance audio streams and economy-level LOD audio streams are automatically switched to a lower bitrate. If the network condition continues to deteriorate, the bitrate level is further reduced, and forward error correction encoding is enabled, adding 20% more bandwidth to each audio data packet. Redundant verification data is implemented. For fine-level LOD and near-field audio streams, priority is given to ensuring that the bitrate does not decrease, freeing up bandwidth resources for critical audio by reducing the transmission of other audio streams. The quantization precision is dynamically adjusted according to the spectral characteristics of the audio content, maintaining high-precision encoding for mid-frequency speech segments that are sensitive to the human ear, and appropriately reducing the precision for high-frequency and low-frequency parts, maximizing compression efficiency while ensuring speech clarity. When network conditions improve and bandwidth utilization drops below 50%, the audio bitrate is gradually restored, with the restoration speed controlled at one level every five seconds to avoid network fluctuations caused by sudden bitrate jumps, achieving intelligent matching and smooth adaptation between audio quality and network conditions.
[0134] This embodiment achieves dynamic compression and transmission optimization of audio data through an adaptive bitrate control strategy. It adjusts the audio bitrate level in real time based on network bandwidth utilization and packet loss rate, prioritizing critical audio streams. For long-distance and economic-level LOD audio streams, the bitrate is automatically reduced. Simultaneously, forward error correction coding is activated during periods of continuous network congestion to increase redundancy and ensure information integrity. For fine-level LOD and short-distance audio streams, a high bitrate is maintained first, freeing up bandwidth for critical audio by reducing the transmission of other audio streams. Quantization precision is dynamically adjusted based on audio spectrum characteristics, maintaining high precision for mid-frequency speech and moderately compressing high and low frequencies, improving transmission efficiency while ensuring speech clarity. When network conditions improve, the audio bitrate is gradually restored smoothly to avoid sudden fluctuations, achieving intelligent matching and continuous adaptation between audio quality and network conditions, ensuring stability and an immersive experience for multi-user audio interaction in virtual teaching.
[0135] Step 10: If the frame rate of the device is higher than the preset frame rate threshold and the calculation budget has a margin, then increase the LOD level of some audio streams or increase the accuracy of spatial audio rendering.
[0136] Step 11: Use the edge collaboration and network optimization module to determine the spatial distance between each audio source and the target user, mark audio streams with a distance less than a preset proximity threshold as near-distance audio, and mark audio streams with a distance greater than the preset proximity threshold as far-distance audio.
[0137] The determination of the proximity threshold includes the following steps.
[0138] The basic proximity threshold is set according to the teaching scenario status information. Specifically: when in a teaching state, the basic proximity threshold is set to 2 meters; when in a discussion state, the basic proximity threshold is set to 4 meters; and when in a self-study state, the basic proximity threshold is set to 1.5 meters.
[0139] Obtain the spatial size information of the semantic region where the user is located, and adjust the basic proximity threshold according to the spatial size information: when the semantic region is a narrow region, adjust the basic proximity threshold to half the diagonal length of the region; when the semantic region is an open region, adjust the basic proximity threshold to 5 meters.
[0140] The system detects the user distribution density in a local area. When the detection results show that the density exceeds 3 people per 4 square meters, a hysteresis band with a width of 0.5 meters is set at the near-distance threshold boundary. The processing method for the audio source is only changed when the location of the audio source clearly crosses the boundary of the hysteresis band.
[0141] The audio source is classified and processed according to the determined proximity threshold: when the distance between the audio source and the user is less than the proximity threshold, near-field audio processing is performed; when the distance between the audio source and the user is greater than or equal to the proximity threshold, far-field audio processing is performed.
[0142] This embodiment dynamically determines the near-field threshold based on the teaching scenario state, semantic region size, and user distribution density to achieve refined classification and processing of audio sources. A basic near-field threshold is set for different teaching states such as lectures, discussions, and self-study, and adjusted according to the spatial size to adapt to the acoustic characteristics of narrow or open areas. A hysteresis band is set for high-density user areas to avoid frequent switching, thus achieving smooth processing when the audio source enters or leaves the near-field range. Based on the determined threshold, the audio source is divided into near-field and far-field for separate processing, allowing near-field audio to be rendered in real-time on the client side, while far-field audio is pre-mixed and transmitted via an edge server.
[0143] Step 12: The near-field audio is rendered locally in real time on the client, and the far-field audio is pre-mixed on the edge server and then sent to the client.
[0144] Step 13: Dynamically switch the network topology based on the current teaching scenario status and user distribution density. If users are densely distributed and in group discussion mode, use a P2P network topology to achieve direct connection within the group; if users are scattered or in a broadcast mode, use a server relay network topology to achieve unified distribution. This includes the following steps.
[0145] Based on the three-dimensional coordinates and distribution density information of users in the current virtual teaching space, the user aggregation situation in each semantic region is calculated in real time, providing basic data for network topology switching.
[0146] Based on the current teaching scenario, determine whether the user is in a group discussion, lecturing, or broadcasting to all, and determine the corresponding network topology switching strategy.
[0147] When users are densely distributed and in group discussions, a P2P direct connection channel is established between users in the same group to enable direct transmission of audio and data within the group, thereby improving the real-time nature of interaction.
[0148] When users are scattered or in a state of broadcasting to all users, a server relay network topology is enabled. The regional access server and the main server mix the audio stream and then distribute it to each client in a unified manner to ensure global consistency.
[0149] During topology switching, the connection status between each client and the P2P node or server is updated in real time to maintain the continuity of data transmission and the synchronous playback of audio streams.
[0150] Continuously monitor network latency, bandwidth usage, and packet loss rate. If any abnormalities or excessive load occur, automatically adjust network topology parameters or select auxiliary relay nodes to ensure the stability and communication quality of virtual teaching activities.
[0151] This embodiment achieves efficient management of audio and data transmission in the virtual teaching space by dynamically switching the network topology based on the teaching scenario and user distribution density. When users are concentrated in group discussions, a direct P2P connection channel is established within the group to achieve low-latency, high-real-time communication. When users are dispersed or in a broadcast mode, a server relay topology is adopted. The audio stream is mixed by the access server and the main server and then uniformly distributed to ensure global consistency. The system calculates the three-dimensional coordinates and cluster density of users in real time, and determines the topology switching strategy based on the teaching status. During the switching process, client connections are dynamically updated to maintain the continuity and synchronous playback of the audio stream. At the same time, by continuously monitoring network latency, bandwidth usage, and packet loss rate, and automatically adjusting topology parameters or activating auxiliary relay nodes when there are abnormalities or excessive load, the system ensures the stability, communication quality, and immersive experience of multi-user interaction in the virtual teaching environment.
[0152] Step 14: Repeat the steps of teaching scenario status recognition, communication mode switching, audio mixing processing, spatial audio rendering, and dynamic resource scheduling until the teaching activities in the virtual teaching space end.
[0153] In summary, this virtual reality-based interactive teaching system achieves efficient matching of user behavior and system resources in the virtual teaching environment through the organic combination of real-time teaching scene recognition, dynamic switching of communication modes, fine-grained multi-domain audio management, hierarchical spatial audio rendering, dynamic resource scheduling, and edge collaborative network optimization. This ensures personalized, low-latency, and high-precision transmission of audio information, significantly enhancing the immersiveness and interactive experience of virtual teaching. Simultaneously, it optimizes system performance, network efficiency, and resource consumption, adapting to differences in user distribution density and terminal device performance. It provides a highly flexible, stable, and reliable virtual reality interactive teaching solution, greatly improving the shortcomings of traditional virtual teaching systems in multi-user interaction, spatial audio management, and resource scheduling.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An implementation method of a virtual reality-based interactive teaching system, characterized by, The interactive teaching system comprises a scene perception and mode switching module, a multi-domain audio management module and a hierarchical spatial audio rendering module, and the implementation method of the interactive teaching system comprises the following steps: Obtaining initial configuration information of a virtual teaching space, including the number of users, device performance parameters and network topology state, and performing semantic region division on the virtual teaching space; Real-time collection of multi-modal teaching behavior data of users by the scene perception and mode switching module, identification of the current teaching scene state in combination with the semantic region division information, and classification of the teaching scene state into a teaching state, a discussion state, a self-study state or a question-answering state; Determination of a target communication mode according to the teaching scene state, the target communication mode comprising a full-group broadcast mode, a group discussion mode, a nearby communication mode and a private conversation mode, and performing smooth switching from a current communication mode to the target communication mode; Establishment of an audio domain matrix by the multi-domain audio management module, the audio domain matrix being used to represent the participation state of each user in different communication domains, and each communication domain being assigned a corresponding weight coefficient; Collection of source position, volume and direction information of all audio streams in the virtual teaching space, weighted mixing calculation of multi-domain audio based on the audio domain matrix and the weight coefficient, elimination of audio conflicts by a competition inhibition strategy, and generation of individualized mixed audio output for each user; Obtaining, by the hierarchical spatial audio rendering module, distance between each audio source and a target user, role importance parameters and user attention data, dividing audio streams into fine-level LOD, standard-level LOD and economic-level LOD according to the distance, the role importance parameters and the user attention data; Assigning complete HRTF spatial audio rendering resources to the fine-level LOD audio stream, assigning simplified stereo positioning resources to the standard-level LOD audio stream, and assigning basic volume attenuation resources to the economic-level LOD audio stream, so as to realize low-cost and high-precision three-dimensional spatial audio positioning.
2. The implementation method of the virtual reality-based interactive teaching system according to claim 1, characterized in that, The semantic region division of the virtual teaching space comprises the following steps: Three-dimensional gridding of the virtual teaching space, dividing the space into a plurality of cubic units, each unit having a side length of 0.5-2 meters, and serving as a basic analysis unit of the virtual space; Labeling a podium area, a student seat area, an experimental operation area, a group discussion area and a free activity area on the basis of the three-dimensional grid according to teaching function requirements, and assigning corresponding semantic labels and priority values to each area for subsequent spatial semantic recognition; Real-time tracking of the coordinate position of a user virtual avatar in the three-dimensional grid, determination of the semantic area in which the user is currently located in combination with user stay duration and movement trajectory features, and taking the area information as a spatial context basis for scene state recognition; When a plurality of users gather in the same semantic area, automatically extracting attribute parameters of the area, including a maximum number of users accommodated, a recommended audio rendering mode and a network transmission strategy.
3. The implementation method of the virtual reality-based interactive teaching system according to claim 2, characterized in that, The collection of multi-modal teaching behavior data comprises: The head-mounted display device is provided with a posture sensor for continuously acquiring the pitch angle, yaw angle and roll angle of the user's head at a sampling frequency of not less than 90 Hz, for determining the line-of-sight direction and attention focus; The hand controller is used for tracking the positions and gesture actions of the user's hands in the virtual teaching space in real time, and recognizing teaching interaction behaviors including hand raising, pointing, writing and operating virtual objects; The audio signal of the user is collected by the voice input device, and the speaking and silent states are distinguished, and the speaking duration, volume and speech speed features are extracted; The moving speed, stopping position and relative distance to other users of the virtual avatar are recorded, for representing the spatial movement and interaction behaviors of the user; The head posture data, hand action data, voice data and spatial behavior data are time-stamped, aligned and fused to generate a comprehensive behavior vector, for reflecting the teaching participation degree and interaction intention.
4. The implementation method of the virtual reality-based interactive teaching system according to claim 3, characterized in that, The identification of the teaching scene state includes the following steps: A feature rule set of the teaching scene state is defined in advance, for describing the spatial, posture, voice and group behavior features of different teaching scene states, wherein: the teaching state rule includes that the user located in the podium area maintains a standing posture and the voice activity is more than 0.7, while other users are in the seat area and the proportion of the head direction to the user is greater than 60%; the discussion state rule includes that a group of three to five people are formed in a semantic area, the members in the group are less than 1.5 meters apart and multiple people speak at the same time, the average spatial distance between groups is greater than 4 meters and the audio cross degree between groups is less than 15%; Based on the feature rules, the current multi-modal user behavior data is preliminarily determined to generate a candidate teaching scene state; A multi-modal user behavior vector sequence collected in the past ten seconds is input into a lightweight time series convolution network model, which includes three one-dimensional convolution layers with 32 convolution kernels each, for extracting time series features and outputting the probability distribution of each scene state; According to the probability distribution output by the model, the scene state with the highest probability is selected as the machine learning determination result; The rule determination result and the machine learning determination result are fused to generate the final candidate teaching scene state; Only when the confidence of the newly identified state exceeds the preset threshold of 0.8 for three consecutive seconds, the state conversion is performed, to ensure the stability and accuracy of the scene recognition; The confirmed scene state is used as the current teaching scene recognition result, for driving the interaction strategy, audio rendering and resource scheduling of the virtual teaching system.
5. The implementation method of the virtual reality-based interactive teaching system according to claim 4, characterized in that, The smooth switching of the communication mode includes the following steps: Whether the current communication mode needs to be switched is determined according to the teaching scene state or user demand; When switching to the group discussion mode, the volume of non-group users is linearly reduced to 15% within a 0.5-second transition time, while the volume of group users remains unchanged; When switching to the adjacent communication mode, a spatial range expansion strategy is adopted, and the audio propagation effective radius is linearly expanded from the initial 1 meter to 3 meters at an expansion speed controlled at 0.8 meters per second; According to the new communication mode, the audio domain matrix is recalculated, and independent audio domain identifiers and weight parameters are assigned to each discussion group or spatial range; During the audio domain matrix update, a timestamp is added to the audio data packet being transmitted to ensure that the audio frames before and after the switching are played in the correct timing, avoiding audio skipping or repetition; During the entire switching process, the continuity of the audio stream is maintained to ensure that the teaching communication does not interrupt and a smooth transition is achieved.
6. The implementation method of an interactive teaching system based on virtual reality according to claim 5, characterized in that, The audio domain matrix is stored and calculated using a dynamic sparse matrix structure: the rows of the matrix represent all users in the virtual teaching space, and the columns represent the current active communication domains. The value of each matrix element ranges from 0 to 1, representing the participation intensity of the corresponding user in the communication domain. When a user fully participates in a certain communication domain, the element value is 1, and when the user does not participate at all, the value is 0. When the user participates partially, the value is an intermediate value. The matrix elements are dynamically updated based on the user's semantic area position, current interactive object, and historical communication records, with an update frequency of five times per second. When performing audio mixing calculation, for each target user, its corresponding matrix row vector is extracted, and the vector is dot multiplied with the audio stream of all communication domains. The domains with zero participation intensity are automatically filtered out, and only the audio of non-zero domains is weighted and summed. At the same time, when a certain communication domain has no user participation or audio activity for five consecutive seconds, the domain is automatically removed from the matrix, releasing computing resources and reducing the matrix dimension, maintaining a lightweight running state.
7. The implementation method of the interactive teaching system based on virtual reality according to any one of claims 1-6, characterized in that, The interactive teaching system includes a resource scheduling and performance optimization module and an edge collaboration and network optimization module. The implementation of the interactive teaching system also includes the following steps: During the audio rendering process, the resource scheduling and performance optimization module monitors the frame rate and computing load of the device in real time, and dynamically allocates audio processing resources based on the pre-set computing budget pool; If the frame rate of the device is lower than the pre-set frame rate threshold, the LOD level of part of the audio stream is reduced or the number of audio streams participating in mixing is reduced, and an adaptive code rate control strategy is started to compress the audio data; If the frame rate of the device is higher than the pre-set frame rate threshold and there is a margin in the computing budget, the LOD level of part of the audio stream is increased or the accuracy of spatial audio rendering is increased; The edge collaboration and network optimization module is used to determine the spatial distance between each audio source and the target user. Audio streams with a distance less than a pre-set close distance threshold are marked as close distance audio, and audio streams with a distance greater than the pre-set close distance threshold are marked as far distance audio. The close distance audio is rendered locally in real time on the client, and the far distance audio is pre-mixed on the edge server and then delivered to the client; According to the current teaching scene state and user distribution density, the network topology is dynamically switched. If the user distribution is dense and in a group discussion state, a P2P network topology is used to realize direct connection within the group. If the user distribution is scattered or in a full-broadcast state, a server relay network topology is used to realize unified distribution. The teaching scene state recognition, communication mode switching, audio mixing processing, spatial audio rendering, and resource dynamic scheduling steps are repeatedly executed until the teaching activity in the virtual teaching space ends.
8. The implementation method of an interactive teaching system based on virtual reality according to claim 7, characterized in that, According to the current teaching scene state and user distribution density, the network topology is dynamically switched, if the user distribution is dense and in the group discussion state, P2P network topology is used to realize direct connection within the group; If the user distribution is scattered or in the full broadcast state, the server relay network topology is used to realize unified distribution, including the following steps: According to the three-dimensional coordinates and distribution density information of the users in the current virtual teaching space, the user aggregation of each semantic area is calculated in real time, providing basic data for network topology switching; Combined with the current teaching scene state, it is judged whether the user is in the group discussion, teaching or full broadcast state, and the corresponding network topology switching strategy is determined; When the user distribution is dense and in the group discussion state, P2P direct connection channel is established between users in the same group to realize direct transmission of audio and data within the group, improve the real-time interaction; When the user distribution is scattered or in the full broadcast state, the server relay network topology is enabled, and the audio stream is mixed by the regional access server and the main server before being distributed to each client, ensuring global consistency; In the process of topology switching, the connection state of each client with P2P node or server is updated in real time, maintaining the continuity of data transmission and the synchronous playback of audio stream; Continuously monitor network delay, bandwidth occupation and packet loss rate, if abnormal or high load occurs, automatically adjust network topology parameters or select auxiliary relay node, ensure the stability and communication quality of virtual teaching activities.
Citation Information
Patent Citations
Online teaching video pushing management system
CN110996128A
Online and offline combined multi-dimensional teaching AI (Artificial Intelligence) classroom system
CN112766226A