VR Large-Space Multi-User Collaborative Distributed Audio Interaction System and Method
By collecting and analyzing micro-motion data, constructing an action-audio dataset and calculating the dynamic correlation strength, the problem of insufficient micro-motion capture in existing VR audio interaction systems is solved, achieving more accurate and stable audio demand prediction and improving the naturalness and response efficiency of VR interaction.
Patent Information
- Application Number
- CN202511470236.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Existing VR audio interaction systems struggle to capture the subtle body movements of users during real-time interactions, resulting in insufficient dynamic correlation analysis capabilities for audio needs. This makes it impossible to match users' instantaneous interaction intentions, affecting the accuracy and naturalness of the interaction.
Collect micro-motion data and audio demand data, calculate the direction coefficient of micro-motion data, identify valid micro-motion data and integrate audio demand data, construct an action-audio dataset, extract micro-motion features and audio demand features, calculate dynamic correlation strength, construct a temporal correlation model and trigger an audio prediction signal, and dynamically adjust the parameters of the audio prediction model.
It improves the accuracy and responsiveness of audio interaction, ensures the accuracy and stability of audio demand prediction, and enhances the naturalness and reliability of VR immersive experiences.
Smart Images

Figure CN120949944B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of distributed virtual reality technology, specifically a distributed audio interaction system and method for large-space multi-user collaboration in VR. Background Technology
[0002] With the rapid development of virtual reality (VR) technology, immersive audio interaction has become a core element in enhancing the VR experience, widely used in games, education and training, remote collaboration, and other scenarios. Existing VR audio interaction systems primarily rely on the user's static positional information (such as 3D coordinates) or coarse motion data (such as head orientation angle) to achieve audio control. While these solutions can meet basic interaction needs, they struggle to capture the subtle body movements (such as head rotation angular velocity) generated by the user during real-time interaction. This results in insufficient dynamic correlation analysis capabilities of the system to understand the user's audio needs, making it unable to match the user's instantaneous interactive intentions.
[0003] Existing technologies neglect the directional consistency and stability of user micro-motion data, making it difficult to effectively distinguish between conscious micro-movements (such as actively moving closer to a sound source to increase volume) and unconscious random jerking (such as natural head shaking). For example, when a user makes a random head turn due to slight body shaking, the system may misinterpret it as an intention to adjust the volume, resulting in abnormal volume jumps; while when a user performs slow and stable head micro-movements, the system may experience response delays because it fails to capture the dynamic characteristics of the movement. These problems directly reduce the accuracy and naturalness of VR audio interaction, hindering further improvements in immersive experiences.
[0004] To this end, the present invention provides a distributed audio interaction system and method for large-space multi-user collaboration in VR. Summary of the Invention
[0005] In order to overcome the shortcomings of the prior art, at least one technical problem raised in the background art is solved.
[0006] The technical solution adopted by this invention to solve its technical problem is: a distributed audio interaction method for large-space multi-user collaboration in VR, comprising:
[0007] Collect micro-motion data and audio demand data, calculate the directional pointing coefficient of the micro-motion data; identify valid micro-motion data based on the directional pointing coefficient, and integrate the valid micro-motion data and the corresponding audio demand data to construct an action-audio dataset;
[0008] Based on the constructed action-audio dataset, micro-action features and audio demand features are extracted and the dynamic correlation strength is calculated. Based on the dynamic correlation strength, it is determined whether there is a strong dynamic correlation between the micro-action features and the audio demand features. If there is a strong dynamic correlation, a temporal correlation model between the micro-action features and the audio demand features is constructed and the audio prediction signal is triggered.
[0009] If an audio prediction signal is triggered, an audio prediction model is constructed to predict audio demand features; the predicted audio demand features are obtained and feature matching analysis is performed to obtain the prediction matching degree; based on the prediction matching degree, it is verified whether the predicted audio demand features meet expectations.
[0010] If the predicted audio demand characteristics meet expectations, invalid prediction index is obtained by selecting the prediction matching degree of multiple time windows for invalidity analysis. The parameters of the audio prediction model are dynamically adjusted based on the invalid prediction index, and the model update efficiency is calculated. The dynamic adjustment of the audio prediction model is judged based on the model update efficiency.
[0011] Furthermore, the direction pointing coefficient is calculated as follows:
[0012] The direction vector is obtained by normalizing the micro-motion data of each frame within the time window.
[0013] The directional matching value is obtained by performing a matching analysis on the direction vector of each frame of micro-motion data.
[0014] Calculate the angle between the real-time direction vector and the initial direction vector for each frame within the time window, and perform angle stability analysis to obtain the stable angle value;
[0015] The directional pointing coefficient is obtained by multiplying the directional matching value and the angle stability value.
[0016] Furthermore, the process of performing the fit analysis is as follows:
[0017] The direction vector of the first frame of micro-motion data within the time window is marked as the initial direction vector, and the direction vectors of the remaining frames within the time window are marked as the real-time direction vectors.
[0018] The angle between the real-time direction vector and the initial direction vector in each frame within the time window is calculated using the angle formula.
[0019] Obtain a preset angle threshold, count the number of frames with an angle less than or equal to the preset angle threshold within the time window, and calculate the ratio with the total number of frames within the time window to obtain the direction matching value.
[0020] Furthermore, the process of performing the included angle stability analysis is as follows:
[0021] Calculate the standard deviation of the angle between the real-time direction vector and the initial direction vector for each frame within the time window;
[0022] A formula for calculating the stable value of the angle is constructed by using the standard deviation of the angle between the real-time direction vector and the initial direction vector.
[0023] Obtain the preset maximum fluctuation threshold, input the standard deviation of the angle between the real-time direction vector and the initial direction vector of each frame within the time window, and the preset maximum fluctuation threshold into the angle stability value calculation formula to obtain the angle stability value.
[0024] Furthermore, the process of extracting the micro-motion features and audio demand features is as follows:
[0025] Obtain the three-dimensional angular velocity of the user's head within different time windows;
[0026] The angular displacement data of each frame in different directions is obtained by multiplying the three-dimensional angular velocity of the user's head in different directions in each frame with the time interval within the time window.
[0027] The cumulative angular displacement data in different directions are then accumulated to obtain the cumulative angular displacement in different directions;
[0028] By inputting the cumulative angular displacement in different directions into the modulus formula, the total displacement modulus of the user's head can be obtained;
[0029] The ratio of the total displacement modulus of the user's head within the time window to the preset maximum displacement modulus is processed to obtain the displacement modulus ratio and marked as a micro-motion feature.
[0030] The absolute difference between the volume at the last frame and the volume at the first frame within the time window is used to obtain the volume change amplitude.
[0031] Obtain the preset maximum volume adjustment level, and calculate the volume intensity ratio by comparing the volume change within the time window with the preset maximum volume adjustment level. This ratio is then marked as an audio requirement feature.
[0032] Furthermore, the dynamic correlation strength is calculated as follows:
[0033] The difference between the micro-motion features of adjacent time windows is calculated to obtain the change in micro-motion features between adjacent time windows;
[0034] The change in audio demand characteristics between adjacent time windows is calculated by subtracting the audio demand characteristics between adjacent time windows.
[0035] The number of windows whose changes in micro-motion features and audio demand features have the same sign is counted, and the trend consistency rate is calculated by proportionally dividing the number of windows by the total number of windows.
[0036] The feature ratio of each time window is obtained by comparing the micro-motion features with the audio demand features.
[0037] Calculate the standard deviation of the feature ratios of all time windows, construct an amplitude stability equation based on the standard deviation, and obtain the preset maximum permissible standard deviation;
[0038] The amplitude stability value is obtained by inputting the standard deviation and the preset maximum permissible standard deviation into the amplitude stability equation;
[0039] The dynamic correlation strength is obtained by multiplying the trend consistency rate and the amplitude stability value.
[0040] Furthermore, the method for obtaining the invalid prediction index is as follows:
[0041] Obtain the deviation window, perform deviation analysis based on the deviation window to obtain the predicted deviation value, and perform deviation mean processing on the deviation window to obtain the average deviation degree value.
[0042] The ineffective prediction index is obtained by multiplying the prediction deviation value with the average deviation value.
[0043] Furthermore, the process of performing the aforementioned deviation analysis is as follows:
[0044] The number of deviation windows is counted, and the ratio of the number of deviation windows to the total number of time windows is calculated to obtain the prediction deviation value.
[0045] Furthermore, the process of performing the aforementioned deviation mean processing is as follows:
[0046] Obtain the actual audio demand features, calculate the difference between the predicted audio demand features of the deviation window and the actual audio demand features, and obtain the deviation value of a single deviation window.
[0047] The average deviation value is obtained by summing all deviation values from the window.
[0048] A distributed audio interaction system for large-space, multi-user collaboration in VR, including the following modules:
[0049] Data integration module: Collects micro-motion data and audio demand data, calculates the direction pointing coefficient of micro-motion data; identifies valid micro-motion data based on the direction pointing coefficient, and integrates valid micro-motion data and corresponding audio demand data to construct an action-audio dataset;
[0050] Feature extraction module: Based on the constructed action-audio dataset, extract micro-action features and audio demand features and calculate the dynamic correlation strength; determine whether there is a strong dynamic correlation between micro-action features and audio demand features based on the dynamic correlation strength; if there is a strong dynamic correlation, construct a temporal correlation model between micro-action features and audio demand features and trigger the audio prediction signal.
[0051] Prediction and verification module: If an audio prediction signal is triggered, an audio prediction model is constructed to predict the audio demand features; the predicted audio demand features are obtained and feature matching analysis is performed to obtain the prediction matching degree; based on the prediction matching degree, the predicted audio demand features are verified to see if they meet expectations.
[0052] Model adjustment module: If the predicted audio demand characteristics meet expectations, the prediction matching degree of multiple time windows is selected for invalid analysis to obtain the invalid prediction index; the parameters of the audio prediction model are dynamically adjusted based on the invalid prediction index and the model update efficiency is calculated; the dynamic adjustment of the audio prediction model is judged based on the model update efficiency.
[0053] The beneficial effects of this invention are as follows:
[0054] 1. Collect micro-motion data and audio requirement data, and calculate the direction pointing coefficient of the micro-motion data; identify valid micro-motion data based on the direction pointing coefficient, and integrate the valid micro-motion data and the corresponding audio requirement data to construct an action-audio dataset; by filtering valid micro-motions and associating them with the corresponding audio requirement data, construct a high-quality action-audio dataset to improve the accuracy and relevance of subsequent model training or interactive analysis.
[0055] 2. Based on the constructed action-audio dataset, extract micro-action features and audio demand features and calculate the dynamic correlation strength; determine whether there is a strong dynamic correlation between micro-action features and audio demand features based on the dynamic correlation strength. If there is a strong dynamic correlation, construct a temporal correlation model between micro-action features and audio demand features and trigger an audio prediction signal; by extracting micro-action and audio features, calculating the dynamic correlation strength and constructing a temporal model, real-time prediction of audio demand can be achieved, improving the effectiveness of interactive response.
[0056] 3. If an audio prediction signal is triggered, an audio prediction model is constructed to predict audio demand features; the predicted audio demand features are obtained and feature matching analysis is performed to obtain the prediction matching degree; the predicted audio demand features are verified based on the prediction matching degree to ensure that they meet expectations; by constructing an audio prediction model to predict audio demand features in advance and verifying them through matching degree to ensure that they meet expectations, the accuracy and reliability of audio demand can be improved.
[0057] 4. If the predicted audio demand characteristics meet expectations, select multiple time windows for prediction matching degree invalid analysis to obtain invalid prediction index; dynamically adjust the parameters of audio prediction model based on invalid prediction index and calculate model update efficiency, and judge whether the dynamic adjustment of audio prediction model is qualified based on model update efficiency; by identifying invalid predictions and dynamically optimizing model parameters, and verifying the effectiveness of adjustment, the prediction accuracy and operation stability of audio prediction model can be continuously improved. Attached Figure Description
[0058] The invention will now be further described with reference to the accompanying drawings.
[0059] Figure 1 This is a flowchart illustrating the steps of the distributed audio interaction method for large-space multi-user collaboration in VR as described in an embodiment of the present invention;
[0060] Figure 2 This is a logic judgment diagram for triggering the audio prediction signal as described in an embodiment of the present invention;
[0061] Figure 3 This is a block diagram of the distributed audio interaction system for large-space multi-user collaboration in VR, as described in an embodiment of the present invention. Detailed Implementation
[0062] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0063] Example 1: Please refer to Figure 1 As shown in the embodiment of the present invention, the distributed audio interaction method for large-space multi-user collaboration in VR includes:
[0064] Step 1: Collect micro-motion data and audio requirement data, and calculate the direction pointing coefficient of the micro-motion data; identify valid micro-motion data based on the direction pointing coefficient, and integrate the valid micro-motion data and the corresponding audio requirement data to construct an action-audio dataset;
[0065] In step one, the process of collecting micro-motion data and audio demand data, and calculating the directional pointing coefficient of the micro-motion data is as follows:
[0066] Integrating a miniature inertial measurement unit (MIMU) with a built-in microphone array into VR (virtual reality) devices;
[0067] The miniature inertial measurement unit collects the user's three-dimensional head angular velocity (ω1: pitch, ω2: yaw, ω3: roll) at a sampling frequency of 200Hz, marks it as micro-motion data, and records the acquisition timing of each frame of data;
[0068] Simultaneously use the built-in microphone array to collect users' audio needs data in VR usage scenarios, and align the collected user audio needs data with micro-motion data in time sequence;
[0069] Among them, audio demand data includes audio interaction intent, such as volume adjustment (moving closer to / away from the audio source causes the volume to increase or decrease).
[0070] The collected micro-motion data is divided into different time windows n according to the length of the time window. Then the direction pointing coefficient of the micro-motion data in different time windows is calculated.
[0071] Where n represents the number of time windows;
[0072] Understandably, the time window is 100ms long, which is 20 frames of micro-motion data. If the window is too short, it cannot reflect the direction and trend of the motion. If the window is too long, multiple motions may be mixed in, making it impossible to accurately judge a single motion. The 100ms window covers 20 frames of micro-motion data, which can capture a complete micro-motion trend while avoiding the mixing in of irrelevant motions.
[0073] The method for calculating the direction pointing coefficient of micro-motion data within different time windows is as follows:
[0074] The initial direction vector is obtained by normalizing the micro-motion data of the first frame within the time window.
[0075] For example, the normalization process is as follows: First, calculate the magnitude of the angular velocity of the first frame of micro-motion data. Then, use the angular velocities in different directions of the first frame of micro-motion data to make ratios with the magnitude of the angular velocity to complete the normalization process and obtain the initial direction vector.
[0076] The micro-motion data of frames t (t is 2-20) within the time window are processed using the same normalization method as the micro-motion data of the first frame to obtain the direction vector of frames 2-20, and the real-time direction vector Vt of each frame is obtained.
[0077] The angle φt between the real-time direction vector and the initial direction vector in each frame within the time window is calculated using the included angle formula.
[0078] Within the statistical time window The number of frames less than or equal to the preset angle threshold is compared with the total number of frames within the time window to obtain the direction matching value;
[0079] Preferably, the preset included angle threshold is 25°;
[0080] It is understandable that the 25° angle threshold was obtained by those skilled in the art through experiments. The directional deviation of conscious actions usually does not exceed 25°, while random shaking will frequently exceed 25°.
[0081] Calculate the angle within the time window Standard deviation Through standard deviation Construct a formula for calculating the stable value of the included angle, where the formula is as follows:
[0082] ;
[0083] in, The preset maximum fluctuation threshold;
[0084] Preferably, the maximum fluctuation threshold is 30°;
[0085] It is understandable that the maximum fluctuation threshold of 30° was obtained by those skilled in the art through analysis of industry standards and experimental data;
[0086] The standard deviation of the angle within the time window The preset maximum fluctuation threshold is input into the angle stability value calculation formula to obtain the angle stability value;
[0087] The direction pointing coefficient is obtained by multiplying the direction matching value and the angle stability value.
[0088] In step one, the process of identifying valid micro-motion data based on directional pointing coefficients and integrating the valid micro-motion data with the corresponding audio requirement data to construct an action-audio dataset is as follows:
[0089] Compare the direction pointing coefficient with a preset coefficient threshold;
[0090] It is understandable that the preset coefficient thresholds were obtained by those skilled in the art through quantitative analysis of historical data;
[0091] If the direction pointing coefficient is greater than or equal to the preset coefficient threshold, the micro-motion data within the time window is determined to be conscious head micro-motion and marked as valid micro-motion data.
[0092] If the direction pointing coefficient is less than the preset coefficient threshold, the micro-motion data within this time window is judged as unconscious random head shaking and marked as invalid micro-motion data.
[0093] It should be noted that the physical meaning of the direction pointing coefficient is as follows: The direction pointing coefficient is calculated from the direction matching value and the angle stability value. The direction matching value represents the proportion of frames in which the head movement direction is consistent with the initial direction within the time window, reflecting the overall directional consistency of the user's head micro-movements. The angle stability value represents the degree of fluctuation of the angle between the user's head micro-movements and the initial direction within the time window, reflecting the stability and smoothness of the head micro-movements.
[0094] It should also be noted that the benefits of obtaining the directional pointing coefficient are: through the directional pointing coefficient, the system can automatically identify and remove meaningless random head shaking data of users; it helps to reduce noise interference in subsequent data processing and model training, and improves the accuracy of action-audio correlation analysis;
[0095] Retain the micro-motion data within the time window corresponding to the valid micro-motion data, as well as the corresponding audio requirement data;
[0096] The effective micro-motion data and audio demand data are matched one-to-one according to the time window order to construct an action-audio dataset;
[0097] Step 2: Based on the constructed action-audio dataset, extract micro-action features and audio demand features and calculate the dynamic correlation strength; determine whether there is a strong dynamic correlation between micro-action features and audio demand features based on the dynamic correlation strength. If there is a strong dynamic correlation, construct a temporal correlation model between micro-action features and audio demand features and trigger the audio prediction signal.
[0098] In step two, the process of extracting micro-motion features and audio demand features based on the constructed action-audio dataset and calculating the dynamic correlation strength is as follows:
[0099] Micro-motion features and audio demand features were extracted based on the constructed action-audio dataset;
[0100] The process of extracting micro-motion features is as follows:
[0101] Obtain the three-dimensional angular velocity of the user's head within different time windows, and calculate the cumulative angular velocity displacement in each dimension based on the three-dimensional angular velocity of the user's head;
[0102] The method for calculating the cumulative angular velocity displacement in each dimension is as follows:
[0103] The angular displacement data of each frame in different directions is obtained by multiplying the three-dimensional angular velocity of the user's head in different directions in each frame with the time interval within the time window.
[0104] The cumulative angle displacement in different directions is calculated by accumulating the angle displacement data in each frame within the time window.
[0105] The cumulative angular displacement in different directions is input into the modulus formula, and the cumulative angular displacement modulus of the user's head is calculated and marked as the total displacement modulus.
[0106] The ratio of the total displacement modulus of the user's head within the time window to the preset maximum displacement modulus is obtained and marked as the micro-motion feature M. n ;
[0107] Where n represents the number of time windows;
[0108] Preferably, the maximum displacement modulus is preset to 3 rad;
[0109] For example, assuming a time window of 100ms, a time interval of 0.005 seconds, and 20 frames of micro-motion data, the user's head performs a "slight head tilt + right turn" motion, and the angular velocity data in different directions are as follows:
[0110] Pitch direction (head tilt): The three-dimensional angular velocity of the user's head in each frame is ω1 = 10 rad / s (lasting 20 frames), then the cumulative angular displacement in the pitch direction is... =10 × 0.005 × 20 = 1 rad;
[0111] Yaw direction (right turn): The three-dimensional angular velocity of the user's head in each frame is ω2 = 15 rad / s (lasting 20 frames), then the cumulative angular displacement in the yaw direction is... =15×0.005×20=1.5rad;
[0112] Roll direction (no motion): The 3D angular velocity of the user's head is ω3=0 in each frame, and the cumulative angular displacement in the roll direction is... =0;
[0113] The cumulative angular displacement of the user's head in different directions is input into the modulus formula to obtain the total displacement modulus of the user's head within the time window;
[0114] The formula for the modulus length is:
[0115] ;
[0116] The ratio of the total displacement modulus of the user's head within the time window to the preset maximum displacement modulus is obtained and marked as the micro-motion feature M. n ;
[0117] The process of extracting audio demand features is as follows:
[0118] Based on the volume adjustment event, the absolute difference between the volume of the VR scene in the last frame and the volume of the VR scene in the first frame within the time window is processed to obtain the volume change amplitude.
[0119] The volume intensity ratio is obtained by comparing the volume change within the time window with the preset maximum volume adjustment, and this ratio is marked as the audio demand feature C. n ;
[0120] Preferably, the preset maximum volume adjustment is 85dB;
[0121] For example, the preset maximum volume adjustment level is set by those skilled in the art according to human tolerance;
[0122] The extracted micro-motion features and audio demand features are matched according to the time window order to form feature data pairs;
[0123] The change in micro-motion characteristics between adjacent time windows is obtained by calculating the difference between the micro-motion characteristics of adjacent time windows. The change in audio demand characteristics between adjacent time windows is calculated by subtracting the audio demand characteristics between adjacent time windows. ;
[0124] statistics and The number of windows with the same sign (i.e., both positive or both negative, indicating a consistent trend) is proportionally calculated to the total number of windows to obtain the trend consistency rate.
[0125] The feature ratio of each time window is obtained by comparing the micro-motion features with the audio demand features.
[0126] Calculate the standard deviation S of the feature ratios for all time windows, and construct an amplitude stability equation based on the standard deviation of the feature ratios for all time windows, where the amplitude stability equation is:
[0127] ;
[0128] Among them, S 预 E is the preset maximum permissible standard deviation, and E is the amplitude stability value;
[0129] It is understandable that the preset maximum permissible standard deviation is obtained by those skilled in the art through analysis of historical data;
[0130] The standard deviation of all time window feature ratios and the preset maximum permissible standard deviation are input into the amplitude stability equation to obtain the amplitude stability value;
[0131] The dynamic correlation strength Q is obtained by multiplying the trend consistency rate and the amplitude stability value.
[0132] In step two, the process of determining whether there is a strong dynamic correlation between micro-motion features and audio demand features based on the strength of the dynamic correlation is as follows: If a strong dynamic correlation exists, the process of constructing a temporal correlation model between micro-motion features and audio demand features and triggering the audio prediction signal is as follows:
[0133] The dynamic correlation strength is compared with a preset correlation threshold;
[0134] It should be noted that the preset association threshold was calculated by those skilled in the art using statistical optimization methods;
[0135] If the dynamic correlation strength is less than the preset correlation threshold, it indicates that the micro-motion features and audio demand features have a weak dynamic correlation.
[0136] If the dynamic correlation strength is greater than or equal to the preset correlation threshold, it indicates that there is a strong dynamic correlation between the micro-motion features and the audio demand features.
[0137] It should be noted that the physical meaning of dynamic correlation strength is as follows: dynamic correlation strength is calculated from the trend consistency rate and amplitude stability value. The trend consistency rate reflects the degree of synchronization between the micro-motion characteristics (such as the cumulative displacement modulus of head rotation angular velocity) and the audio demand characteristics (such as volume adjustment intensity) in adjacent time windows (e.g., whether the audio demand increases synchronously when the micro-motion increases). The amplitude stability value reflects the matching stability of the amplitude of the two (e.g., whether the audio demand amplitude changes proportionally when the micro-motion amplitude is large). Specifically, the greater the dynamic correlation strength, the stronger the real-time causal relationship between the user's head micro-motion and audio interaction intention.
[0138] It should also be noted that the benefits of obtaining dynamic correlation strength are: it can more intuitively perceive the correlation between the user's head micro-movements and audio demand events, filter high-quality data for time-series correlation models, optimize the audio demand prediction effect, and achieve more natural VR audio interaction.
[0139] like Figure 2 As shown, if there is a strong dynamic correlation between micro-motion features and audio demand features, a temporal correlation model is constructed by matching the micro-motion features and audio demand features in different time windows according to their temporal correspondence, and the audio prediction signal is triggered synchronously.
[0140] The technical solution of this embodiment is as follows: Collect micro-motion data and audio demand data, and calculate the directional pointing coefficient of the micro-motion data; identify valid micro-motion data based on the directional pointing coefficient, and integrate the valid micro-motion data and corresponding audio demand data to construct an action-audio dataset; construct a high-quality action-audio dataset by filtering valid micro-motions and associating them with corresponding audio demand data, thereby improving the accuracy and relevance of subsequent model training or interactive analysis; extract micro-motion features and audio demand features based on the constructed action-audio dataset and calculate the dynamic correlation strength; determine whether there is a strong dynamic correlation between the micro-motion features and audio demand features based on the dynamic correlation strength; if a strong dynamic correlation exists, construct a temporal correlation model of the micro-motion features and audio demand features and trigger an audio prediction signal; by extracting micro-motion and audio features, calculating the dynamic correlation strength, and constructing a temporal model, real-time prediction of audio demand is achieved, improving the effectiveness of interactive response.
[0141] Example 2: Please refer to Figure 1 As shown in the embodiment of the present invention, the distributed audio interaction method for large-space multi-user collaboration in VR further includes:
[0142] Step 3: If the audio prediction signal is triggered, an audio prediction model is constructed to predict the audio demand features; the audio demand prediction features are obtained and feature matching analysis is performed to obtain the prediction matching degree; based on the prediction matching degree, it is verified whether the predicted audio demand features meet expectations.
[0143] In step three, if an audio prediction signal is triggered, the process of constructing an audio prediction model to predict audio demand features is as follows:
[0144] If an audio prediction signal is triggered, an audio prediction model is constructed and audio demand characteristics are predicted.
[0145] The process of building an audio prediction model and predicting audio demand characteristics is as follows:
[0146] Predict audio demand features based on micro-motion features collected within the time window;
[0147] The micro-motion features within the collected time window are summed and then proportionally calculated with the summed audio demand features to obtain a ratio. This ratio is then proportionally calculated with the total number of windows to obtain the basic correlation ratio TY.
[0148] Understandably, the basic correlation ratio is based on the average ratio calculated from the collected micro-motion features and audio demand features, reflecting the long-term stable correlation trend between micro-motion features and audio demand features, that is, how many units of audio demand features correspond to each unit of micro-motion feature on average, providing a basic reference ratio for audio demand prediction.
[0149] The dynamic correlation strength between micro-motion features and audio demand features is used as a dynamic adjustment coefficient Q, which is correlated with the basic correlation ratio and the real-time acquired micro-motion features M. new Together, we construct an audio prediction formula;
[0150] The audio prediction formula is as follows:
[0151] ;
[0152] The dynamic adjustment coefficient, the basic correlation ratio, and the micro-motion features acquired in real time are input into the audio prediction formula to obtain the predicted audio demand features.
[0153] In step three, the audio demand prediction features are obtained and feature matching analysis is performed to obtain the prediction matching degree. The process of verifying whether the predicted audio demand features meet expectations based on the prediction matching degree is as follows:
[0154] Obtain and predict the actual audio demand features corresponding to the audio demand features;
[0155] The absolute difference between the predicted audio demand characteristics and the actual audio demand characteristics is denoted as the absolute deviation.
[0156] The ratio of absolute deviation to maximum permissible deviation is used to calculate the error proportion, which is then used as the predicted matching degree between the predicted audio demand features and the actual audio demand features.
[0157] It is understandable that the maximum permissible deviation is set by those skilled in the art through experiments;
[0158] The predicted match rate is compared with a preset match rate threshold;
[0159] For example, the preset matching threshold is obtained by those skilled in the art through statistical analysis and error distribution calculation of historical datasets;
[0160] If the predicted matching degree is greater than or equal to the matching degree threshold, it means that the prediction error exceeds the maximum allowable range, the matching fails, and the predicted audio requirement features do not meet expectations.
[0161] If the predicted matching degree is less than the matching degree threshold, it means that the match is successful and the predicted audio demand features meet expectations.
[0162] It should be noted that the purpose of obtaining the predicted matching degree is:
[0163] Function 1: Transforms abstract prediction accuracy into quantifiable numerical indicators, thereby accurately describing the correspondence between predicted audio demand features and actual audio demand features, avoiding the fuzzy interaction problems caused by traditional reliance on subjective judgment or rough thresholds.
[0164] Function 2: Effectively filters out interaction misjudgments caused by prediction errors exceeding the acceptable range, thereby preventing abnormal volume jumps and ensuring that the audio response is highly consistent with the user's actual interaction intent;
[0165] Function 3: Obtaining the predicted matching degree can identify abnormal data containing the predicted audio demand features and the actual audio demand features by identifying the predicted matching degree below the threshold. This gradually improves the system's adaptability to complex interaction scenarios and its long-term robustness.
[0166] Step 4: If the predicted audio demand characteristics meet expectations, select multiple time windows for prediction matching degree invalid analysis to obtain invalid prediction index; dynamically adjust the parameters of the audio prediction model based on the invalid prediction index and calculate the model update efficiency; judge whether the dynamic adjustment of the audio prediction model is qualified based on the model update efficiency.
[0167] In step four, if the predicted audio demand characteristics meet expectations, the process of selecting multiple time windows for prediction matching to perform invalidity analysis and obtain the invalid prediction index is as follows:
[0168] Select the prediction matching degree of multiple consecutive time windows;
[0169] Based on the comparison between the matching degree and the preset matching degree threshold, the time window in which the predicted audio demand features do not meet expectations is marked as the deviation window;
[0170] The number of deviation windows is counted, and the ratio of the number of deviation windows to the total number of selected time windows is used to obtain the prediction deviation value.
[0171] The difference between the predicted audio demand features of the deviation window and the actual audio demand features is calculated to obtain the deviation value of a single deviation window.
[0172] The average deviation value is obtained by summing all deviation values from the window.
[0173] The ineffective prediction index is obtained by multiplying the prediction deviation value with the average deviation value.
[0174] In step four, the parameters of the audio prediction model are dynamically adjusted based on the ineffective prediction index, and the model update efficiency is calculated. The process of judging whether the dynamic adjustment of the audio prediction model is qualified based on the model update efficiency is as follows:
[0175] If the invalid prediction index is lower than or equal to the preset warning threshold, it means that the prediction stability of the audio prediction model is within an acceptable range.
[0176] Fine-tune the basic correlation ratio; calculate the difference between 1 and the invalid prediction index to obtain the fine-tuning coefficient, and multiply the basic correlation ratio with the fine-tuning coefficient to obtain the fine-tuned basic correlation ratio.
[0177] If the invalid prediction index is higher than the preset warning threshold, it means that the prediction stability deviation of the audio prediction model has deviated from the normal range.
[0178] Significantly correct the basic association ratio: Based on recent effective historical micro-motion features and audio demand features, calculate the historical basic association ratio, and calculate the average basic association ratio with the basic association ratio to obtain the average basic association ratio as the corrected basic association ratio.
[0179] For example, recent effective historical micro-motion features and audio demand features are micro-motion features and audio demand features with a dynamic correlation strength greater than or equal to a preset correlation threshold within 24 hours.
[0180] By selecting the time window after adjusting the basic correlation ratio, the adjusted ineffective prediction index is calculated;
[0181] The difference between the ineffective prediction indices before and after adjustment is calculated to obtain the index difference of ineffective predictions;
[0182] The model update efficiency of the audio prediction model is obtained by comparing the difference in invalid prediction index with the invalid prediction index before adjustment.
[0183] The model update efficiency is compared with a preset model update efficiency threshold;
[0184] It should be noted that the preset model update efficiency threshold was obtained by those skilled in the art through multi-dimensional analysis and quantitative comprehensive verification;
[0185] If the model update efficiency is higher than the preset model update efficiency threshold, it means that the dynamic adjustment of the audio prediction model is not up to standard and the parameter adjustment process needs to be repeated until the model update efficiency meets the standard.
[0186] If the model update efficiency is lower than or equal to the preset model update efficiency threshold, it means that the dynamic adjustment of the audio prediction model is qualified.
[0187] For example, the preset model update efficiency threshold is obtained by those skilled in the art through analysis of historical update data;
[0188] It should be noted that the physical meaning of model update efficiency is: quantifying the degree of improvement of parameter adjustment in reducing invalid predictions, reflecting the optimization effect of the adjustment strategy on the model prediction stability. The higher the efficiency, the more effectively the current adjustment can address the prediction bias problem.
[0189] It should also be noted that the benefits of obtaining model update efficiency are as follows: obtaining model update efficiency can intuitively quantify the actual value of adjusting audio prediction model parameters, avoiding ineffective adjustments that consume system resources; it can quickly determine whether the adjustment has specifically solved the prediction bias problem, ensuring that the audio prediction model adapts to the changes in the correlation between user micro-actions and audio needs in real time, and always maintains high prediction accuracy.
[0190] The technical solution of this embodiment is as follows: If an audio prediction signal is triggered, an audio prediction model is constructed to predict audio demand features; the predicted audio demand features are obtained and feature matching analysis is performed to obtain the prediction matching degree; the predicted audio demand features are verified based on the prediction matching degree to ensure that they meet expectations; by constructing an audio prediction model to predict audio demand features in advance and verifying them through matching degree to ensure that they meet expectations, the accuracy and reliability of audio demand can be improved; if the predicted audio demand features meet expectations, the prediction matching degree of multiple time windows is selected for invalid analysis to obtain an invalid prediction index; the parameters of the audio prediction model are dynamically adjusted based on the invalid prediction index and the model update efficiency is calculated; the dynamic adjustment of the audio prediction model is judged to be qualified based on the model update efficiency; by identifying invalid predictions and dynamically optimizing model parameters, while verifying the effectiveness of the adjustment, the prediction accuracy and operational stability of the audio prediction model can be continuously improved.
[0191] Example 3: Please refer to Figure 3 As shown in the embodiment of the present invention, the distributed audio interaction system for large-space multi-user collaboration in VR includes the following modules:
[0192] Data integration module: Collects micro-motion data and audio demand data, calculates the direction pointing coefficient of micro-motion data; identifies valid micro-motion data based on the direction pointing coefficient, and integrates valid micro-motion data and corresponding audio demand data to construct an action-audio dataset;
[0193] Feature extraction module: Based on the constructed action-audio dataset, extract micro-action features and audio demand features and calculate the dynamic correlation strength; determine whether there is a strong dynamic correlation between micro-action features and audio demand features based on the dynamic correlation strength; if there is a strong dynamic correlation, construct a temporal correlation model between micro-action features and audio demand features and trigger the audio prediction signal.
[0194] Prediction and verification module: If an audio prediction signal is triggered, an audio prediction model is constructed to predict the audio demand features; the predicted audio demand features are obtained and feature matching analysis is performed to obtain the prediction matching degree; based on the prediction matching degree, the predicted audio demand features are verified to see if they meet expectations.
[0195] Model adjustment module: If the predicted audio demand characteristics meet expectations, the prediction matching degree of multiple time windows is selected for invalid analysis to obtain the invalid prediction index; the parameters of the audio prediction model are dynamically adjusted based on the invalid prediction index and the model update efficiency is calculated; the dynamic adjustment of the audio prediction model is judged based on the model update efficiency.
[0196] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for distributed audio interaction of VR large space multi-user collaboration, characterized in that: The method comprises the following steps: Collecting micro-motion data and audio demand data, calculating the direction pointing coefficient of the micro-motion data; Identifying effective micro-motion data based on the direction pointing coefficient, integrating the effective micro-motion data and the corresponding audio demand data to construct an action-audio data set; The way to calculate the direction pointing coefficient is: Obtain the direction vector of each frame of micro-motion data in the time window after normalization processing; Based on the direction vector of each frame of micro-motion data, the fitting analysis is carried out to obtain the direction fitting value, Calculate the angle between the real-time direction vector of each frame in the time window and the initial direction vector and carry out the angle stability analysis to obtain the angle stability value; The direction pointing coefficient is obtained by multiplying the direction fitting value and the angle stability value; Based on the constructed action-audio data set, extract the micro-motion feature and the audio demand feature and calculate the dynamic correlation strength; According to the dynamic correlation strength, it is judged whether there is strong dynamic correlation between the micro-motion feature and the audio demand feature, if there is strong dynamic correlation, a time sequence correlation model of the micro-motion feature and the audio demand feature is constructed and an audio prediction signal is triggered; If the audio prediction signal is triggered, an audio prediction model is constructed to predict the audio demand feature; Obtain the audio demand prediction feature and carry out feature matching analysis to obtain the prediction matching degree, and verify whether the predicted audio demand feature meets the expectation based on the prediction matching degree; If the predicted audio demand feature meets the expectation, the prediction matching degrees of multiple time windows are selected for invalidity analysis to obtain an invalid prediction index; Based on the invalid prediction index, the parameters of the audio prediction model are dynamically adjusted and the model update efficiency is calculated, and whether the dynamic adjustment of the audio prediction model is qualified is judged according to the model update efficiency.
2. The method of claim 1, wherein: The process of fitting analysis is: Obtain the direction vector of the first frame of micro-motion data in the time window, mark it as the initial direction vector, and mark the direction vectors of the remaining frames in the time window as real-time direction vectors; Use the angle formula to calculate the angle between the real-time direction vector of each frame in the time window and the initial direction vector; Obtain a preset angle threshold, count the number of frames whose angle is less than or equal to the preset angle threshold in the time window, and perform ratio processing with the total number of frames in the time window to obtain the direction fitting value.
3. The method of claim 1, wherein: The process of angle stability analysis is: Calculate the standard deviation of the angle between the real-time direction vector of each frame in the time window and the initial direction vector; Construct an angle stability value calculation formula based on the standard deviation of the angle between the real-time direction vector and the initial direction vector; Obtain a preset maximum fluctuation threshold, input the standard deviation of the angle between the real-time direction vector of each frame in the time window and the initial direction vector, and the preset maximum fluctuation threshold into the angle stability value calculation formula to obtain the angle stability value.
4. The method of claim 1, wherein: The process of extracting the micro-motion feature and the audio demand feature is: Obtain the three-dimensional angular velocity of the user's head in different time windows; Multiply the three-dimensional angular velocity of the user's head in each direction with the time interval in the time window to obtain the angular displacement data in each direction; And accumulate the angular displacement data in different directions to obtain the cumulative angular displacement in different directions; Input the cumulative angular displacement in different directions into the module length formula to obtain the total displacement module length of the user's head; The total displacement module length of the user's head in the time window is compared with the preset maximum displacement module length to obtain a displacement module length ratio and marked as a micro-motion feature; The absolute difference between the volume at the last frame in the time window and the volume at the first frame is obtained to obtain a volume change amplitude; The preset maximum volume adjustment degree is obtained, and the volume change amplitude in the time window is compared with the preset maximum volume adjustment degree to obtain a volume intensity ratio and marked as an audio demand feature.
5. The method of claim 1, wherein: The calculation of the dynamic correlation strength is as follows: The micro-motion features of adjacent time windows are calculated by difference to obtain the micro-motion feature change amount of adjacent time windows; The audio demand features of adjacent time windows are calculated by difference to obtain the audio demand feature change amount of adjacent time windows; The number of windows with the same sign of the micro-motion feature change amount and the audio demand feature change amount is counted, and the trend consistency rate is calculated by the total number of windows; The micro-motion feature and the audio demand feature of each time window are compared by ratio to obtain the feature ratio of each time window; The standard deviation of the feature ratio of all time windows is calculated, and the amplitude stability equation is constructed based on the standard deviation to obtain the preset maximum allowed standard deviation; The standard deviation and the preset maximum allowed standard deviation are input into the amplitude stability equation to obtain the amplitude stability value; The trend consistency rate and the amplitude stability value are multiplied to obtain the dynamic correlation strength.
6. The method of claim 1, wherein: The invalid prediction index is obtained as follows: The deviation window is obtained, and the prediction deviation value is obtained by deviation analysis based on the deviation window. The deviation average value is obtained by deviation mean processing of the deviation window. The prediction deviation value and the deviation average value are multiplied to obtain the invalid prediction index.
7. The method of claim 6, wherein: The deviation analysis process is as follows: The number of deviation windows is counted, and the prediction deviation value is obtained by comparing the number of deviation windows with the total number of time windows.
8. The method of claim 6, wherein: The deviation mean processing process is as follows: The actual audio demand feature is obtained, and the prediction audio demand feature of the deviation window and the actual audio demand feature are calculated by difference to obtain the deviation degree value of a single deviation window; The deviation degree values of all deviation windows are summarized and averaged to obtain the deviation average degree value.
9. A distributed audio interaction system for VR large space multi-user collaboration, for implementing the method of any one of claims 1-8, characterized in that, It includes the following modules: Data integration module: collect micro-motion data and audio demand data, calculate the direction pointing coefficient of micro-motion data; identify effective micro-motion data based on the direction pointing coefficient, integrate effective micro-motion data and corresponding audio demand data to construct action-audio data set; Feature extraction module: based on the constructed action-audio data set, extract micro-motion features and audio demand features and calculate dynamic correlation strength; determine whether there is strong dynamic correlation between micro-motion features and audio demand features according to the dynamic correlation strength, if there is strong dynamic correlation, construct a time sequence correlation model of micro-motion features and audio demand features and trigger an audio prediction signal; Prediction verification module: if the audio prediction signal is triggered, an audio prediction model is constructed to predict the audio demand feature; obtain the audio demand prediction feature and perform feature matching analysis to obtain the prediction matching degree, and verify whether the predicted audio demand feature meets the expectation based on the prediction matching degree; The model adjustment module: if the predicted audio demand feature meets the expectation, the prediction matching degrees of multiple time windows are selected for invalid analysis to obtain an invalid prediction index; parameters of the audio prediction model are dynamically adjusted based on the invalid prediction index, and a model updating efficiency is calculated, and whether the dynamic adjustment of the audio prediction model is qualified is judged according to the model updating efficiency.
Citation Information
Patent Citations
Portable interactive desktop-level virtual reality system
CN105159450A
Human body figure action identifying method based on intelligent mobile phone
CN106919958A