Intelligent analysis system for ar voice practical training integrated with cross-cultural adaptation
Patent Information
- Application Number
- CN202610953527.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]针对现有技术的不足,本发明提供了融入跨文化适应性的AR语音实训智能分析系统,解决了现有增强现实语音实训系统中多传感器采样频率不同导致的数据时间异步、头部独立运动引起手部相对位移测量误差,以及无法量化评估语音声学能量与手部动能同步相位差、缺乏结合特定文化空间约束的实时视觉纠偏反馈的问题
[0047]1、本发明通过坐标系重构模块结合人体颈部拓扑常数与对齐头部平移向量逆向推算胸骨柄定点绝对坐标,并将虚拟躯干偏航角对应的朝向矢量作为前向轴构建虚拟躯干参考坐标系,消除了头部独立旋转运动对位姿测量造成的误差干扰,提高了手部相对位移数据计算的准确度。
Smart Images

Figure CN122821641A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of training and analysis technology, specifically to an AR voice training intelligent analysis system that incorporates cross-cultural adaptability. Background Technology
[0002] In cross-cultural communication training, augmented reality technology is often used to simulate real language communication scenarios. Complete cross-cultural communication not only includes the accurate expression of language information, but also non-verbal behaviors such as gestures, interpersonal distance, and synchronized action rhythm. Existing augmented reality voice training systems usually acquire and evaluate users' multimodal action data through head-mounted devices.
[0003] However, existing hardware devices suffer from differences in the sampling frequency of different sensors when acquiring raw speech sequences, head pose matrices, and hand spatial coordinates. This can easily cause asynchronous data generation across multiple sources on the time axis, leading to fundamental data deviations in subsequent joint motion and speech analysis. Furthermore, existing training and analysis systems often rely on a single head coordinate system or global world coordinate system as the reference benchmark for motion capture. Since users often experience independent yaw and rotational movements of the head during cross-cultural communication training, directly calculating the relative displacement of the hands based on the head coordinate system introduces measurement errors and fails to accurately reflect the true relative spatial state of the human body.
[0004] In terms of behavioral characteristic assessment, cross-cultural communication has specific cultural requirements for interpersonal spatial distance and the synchronization of gestures. Existing systems can usually only record basic positional trajectories and lack physical quantification methods for calculating the phase difference between speech acoustic energy and hand kinetic energy. This makes it impossible to accurately assess the synchronization state of action beats and speech stress. Furthermore, because existing systems do not incorporate a three-dimensional proximal constraint model that incorporates specific cultural backgrounds, it is difficult to accurately calculate the normal spatial boundary distance of user actions. Therefore, it is impossible to integrate and render visual correction feedback regarding spatial position violations and action misalignment in real time within the user's physical field of vision, thus limiting the guidance effect of cross-cultural training. Therefore, this invention proposes an AR voice training intelligent analysis system that incorporates cross-cultural adaptability to address the shortcomings of existing technologies. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides an AR voice training intelligent analysis system that incorporates cross-cultural adaptability. It solves the problems in existing augmented reality voice training systems, such as asynchronous data time caused by different sampling frequencies of multiple sensors, measurement errors in relative hand displacement caused by independent head movements, inability to quantitatively evaluate the synchronous phase difference between voice acoustic energy and hand kinetic energy, and lack of real-time visual correction feedback that incorporates specific cultural spatial constraints.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an AR voice training intelligent analysis system incorporating cross-cultural adaptability, wherein the AR voice training intelligent analysis system incorporating cross-cultural adaptability includes:
[0007] The data alignment module is configured to acquire the original speech sequence, head pose matrix sequence, and hand spatial coordinate sequence, and map the original speech sequence, head pose matrix sequence, and hand spatial coordinate sequence to a discrete timestamp sequence.
[0008] The coordinate system reconstruction module is set to establish a virtual torso reference coordinate system based on the head posture matrix sequence, and convert the hand spatial coordinate sequence into relative hand displacement data under the virtual torso reference coordinate system.
[0009] The envelope extraction module is configured to extract the acoustic energy envelope and the hand kinetic energy envelope from the original speech sequence and the relative hand displacement data.
[0010] The phase difference calculation module is configured to extract the communication synchronization phase difference by using the acoustic energy envelope and the hand kinetic energy envelope.
[0011] The graphics rendering module is configured to create a three-dimensional near-body constraint geometry. It substitutes the relative displacement data of the hand into the mathematical equations corresponding to the three-dimensional near-body constraint geometry to calculate the normal space boundary distance value, and combines the communication synchronization phase difference to calculate the visual display parameters. The rendered output three-dimensional near-body constraint geometry provides visual feedback for cross-cultural training analysis.
[0012] Preferably, the data alignment module is configured to acquire the original speech sequence, head pose matrix sequence, and hand spatial coordinate sequence, and map the original speech sequence, head pose matrix sequence, and hand spatial coordinate sequence to a discrete timestamp sequence, specifically including:
[0013] The data alignment module selects the screen refresh rate of the augmented reality head-mounted display device or the sampling frequency of the inertial measurement unit as the system reference clock frequency, and generates a discrete timestamp sequence based on the system reference clock frequency;
[0014] The data alignment module applies interpolation and filtering algorithms to output the aligned head rotation matrix, aligned head translation vector, aligned hand spatial coordinates, and aligned tracking confidence scalar mapped to the time axis of the discrete timestamp sequence. It also divides the original speech sequence into high-frequency data segments mapped to the time axis of the discrete timestamp sequence.
[0015] Preferably, establishing a virtual torso reference coordinate system based on the head pose matrix sequence specifically includes:
[0016] The coordinate system reconstruction module extracts the head yaw angle component from the aligned head rotation matrix and calculates the head yaw angular velocity values at consecutive time nodes.
[0017] The dynamic filter damping coefficient is set based on the comparison between the head yaw rate value and the angular velocity dead zone threshold.
[0018] Substitute the dynamic filter damping coefficient into the first-order hysteresis low-pass filter algorithm to calculate the virtual torso yaw angle.
[0019] The coordinate system reconstruction module combines the topological constant of the human neck with the aligned head translation vector to inversely calculate the absolute coordinates of the sternal manubrium. The absolute coordinates of the sternal manubrium are used as the origin of the three-dimensional coordinate system and the orientation vector corresponding to the virtual torso yaw angle is used as the forward axis to construct a virtual torso reference coordinate system.
[0020] Preferably, converting the hand spatial coordinate sequence into relative hand displacement data in a virtual torso reference coordinate system specifically includes:
[0021] The coordinate system reconstruction module extracts the global rotation and global translation parameters of the virtual torso reference coordinate system relative to the system's global world coordinate system, and constructs the inverse matrix of the spatial affine transformation.
[0022] Perform matrix multiplication on the aligned hand spatial coordinates and the inverse of the spatial affine transformation matrix to output hand relative displacement data containing lateral relative displacement components, longitudinal relative displacement components, and normal relative displacement components.
[0023] Preferably, extracting the acoustic energy envelope from the original speech sequence and the relative hand displacement data specifically includes:
[0024] The envelope extraction module extracts the absolute difference between the head yaw angle component and the virtual torso yaw angle as the head line-of-sight yaw angle parameter; and calculates the audio gain compensation value based on the head line-of-sight yaw angle parameter.
[0025] Within the sliding observation time window, short-time root mean square (RMS) calculation is performed on the high-frequency data segments corresponding to the original speech sequence. The RMS calculation result is then multiplied with the audio gain compensation value. The result of the multiplication operation is input into a low-pass filter function for discrete convolution processing to generate a smooth acoustic energy envelope.
[0026] Preferably, extracting the hand kinetic energy envelope from the original speech sequence and the relative hand displacement data specifically includes:
[0027] The envelope extraction module determines whether the alignment tracking confidence scalar is greater than or equal to the stable tracking confidence threshold;
[0028] When the alignment tracking confidence scalar is greater than or equal to the stable tracking confidence threshold, the first-order discrete derivative operation is performed on the lateral relative displacement component, the longitudinal relative displacement component, and the normal relative displacement component to extract the three-dimensional spatial velocity component, and the sum of squares of the three-dimensional spatial velocity components is calculated to generate the hand kinetic energy envelope.
[0029] When the alignment tracking confidence scalar is less than the stable tracking confidence threshold, the hand kinetic energy envelope of the previous time node is calculated using the decay constant to generate the hand kinetic energy envelope of the current time node.
[0030] Preferably, the extraction of the communicative synchronization phase difference using acoustic energy envelope and hand kinetic energy envelope specifically includes:
[0031] The phase difference calculation module extracts the corresponding acoustic energy envelope sequence and hand kinetic energy envelope sequence within the sliding time analysis period;
[0032] The acoustic energy envelope sequence and the hand kinetic energy envelope sequence are respectively subjected to mean removal and normalization preprocessing to output standardized acoustic energy envelope sequence and standardized hand kinetic energy envelope sequence;
[0033] Perform a one-dimensional time-domain cross-correlation function operation on the standardized acoustic energy envelope sequence and the standardized hand kinetic energy envelope sequence;
[0034] Extract the time offset corresponding to the global maximum value of the cross-correlation function and record the time offset as the communication synchronization phase difference.
[0035] Preferably, substituting the relative displacement data of the hand into the mathematical equations corresponding to the three-dimensional near-body constraint geometry to calculate the normal space boundary distance specifically includes:
[0036] The graphics rendering module extracts the horizontal, vertical, and normal geometries from the preset target cultural baseline matrix.
[0037] Based on the virtual torso reference coordinate system, the lateral near-body threshold parameters, longitudinal near-body threshold parameters, and normal near-body threshold parameters are spliced together to form an asymmetric three-dimensional near-body constraint geometry;
[0038] Substitute each coordinate component of the relative displacement data of the hand into the mathematical equation corresponding to the three-dimensional near-body constraint geometry to calculate the normalized radius, and combine the boundary truncation function to calculate and extract the normal space over-boundary distance value.
[0039] Preferably, the calculation of visual display parameters based on the communicative synchronization phase difference specifically includes:
[0040] Visual display parameters include surface rendering color vectors;
[0041] The graphics rendering module uses the absolute value of the communication synchronization phase difference to perform linear interpolation between the preset standard synchronization color vector and the severely out-of-synchronization color vector, and combines the interpolation result with the upper bound truncation function to form the surface rendering color vector.
[0042] Preferably, the rendered output of a 3D near-gravity constrained geometry provides visual feedback for cross-cultural training analysis, specifically including:
[0043] Visual display parameters include surface opacity components;
[0044] When the normal space out-of-bounds distance value is greater than zero, the graphics rendering module uses the normal space out-of-bounds distance value, the opacity gain coefficient, and the maximum opacity limit constant, combined with the upper bound cutoff function, to calculate the surface opacity component.
[0045] The graphics rendering module integrates the surface rendering color vector and the surface opacity component into a rendering data structure. With the transparency blending mode and depth buffer write suppression enabled, it renders and outputs a three-dimensional near-bodily constrained geometry based on the rendering data structure.
[0046] This invention provides an AR voice training intelligent analysis system that incorporates cross-cultural adaptability. It has the following beneficial effects:
[0047] 1. This invention uses a coordinate system reconstruction module to inversely calculate the absolute coordinates of the sternal manubrium fixed point by combining the topological constant of the human neck with the aligned head translation vector, and uses the orientation vector corresponding to the virtual torso yaw angle as the forward axis to construct a virtual torso reference coordinate system. This eliminates the error interference caused by the independent rotation of the head to pose measurement and improves the accuracy of the calculation of the relative displacement data of the hand.
[0048] 2. This invention utilizes an envelope extraction module to perform attenuation calculations using an attenuation constant when the alignment tracking confidence scalar is less than the stable tracking confidence threshold, thus maintaining the continuity of the hand kinetic energy envelope in the time domain. Furthermore, the phase difference calculation module performs one-dimensional time-domain cross-correlation function operations on the standardized acoustic energy envelope sequence and the standardized hand kinetic energy envelope sequence to extract the communicative synchronization phase difference, thereby realizing the physical quantification of the time difference between gesture kinetic energy beats and speech stress beats in cross-cultural communication behaviors.
[0049] 3. This invention reads the target cultural baseline matrix through the graphics rendering module to establish an asymmetric three-dimensional near-body constraint geometry, calculates the surface opacity component using the normal space boundary distance, and calculates the surface rendering color vector using the linear interpolation result of the communication synchronization phase difference. It directly converts the spatial compliance indicators and action beat indicators in cross-cultural behavior into visual display parameters and renders them in the physical field of view, providing direct augmented reality behavior correction feedback. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of the system architecture of the present invention;
[0051] Figure 2 This is a schematic diagram of the overall method flow of the present invention;
[0052] Figure 3 This is a schematic diagram illustrating the decreasing trend of the normal space boundary crossing rate in this invention;
[0053] Figure 4 This is a schematic diagram of the actual distribution of the communication synchronization phase difference convergence band error bars of the present invention;
[0054] Figure 5 This is a schematic diagram of the dynamic response curve of the out-of-bounds distance in a single action according to the present invention. Detailed Implementation
[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] Please see Figure 1 The present invention provides an AR voice training intelligent analysis system with cross-cultural adaptability, including a data alignment module, a coordinate system reconstruction module, an envelope extraction module, a phase difference calculation module, and a graphics rendering module.
[0057] An AR voice training intelligent analysis system with cross-cultural adaptability is deployed in an augmented reality head-mounted display device, which is equipped with a microphone array, an inertial measurement unit, an outward tracking camera array, and a graphics rendering unit.
[0058] The data alignment module connects to the microphone array, the inertial measurement unit, and the outward tracking camera array. The data alignment module is configured to acquire the original speech sequence, absolute rotation matrix, translation vector, and three-dimensional spatial coordinates of key nodes, and output a discretized reference data stream based on a unified timestamp.
[0059] The coordinate system reconstruction module is connected to the data alignment module. The coordinate system reconstruction module is set to establish a virtual torso reference coordinate system and map the absolute coordinates in three-dimensional space to the relative displacement data under the virtual torso reference coordinate system.
[0060] The envelope extraction module is connected to the coordinate system reconstruction module, and the envelope extraction module is set to extract the acoustic energy envelope sequence and the kinetic energy envelope sequence.
[0061] The phase difference calculation module is connected to the envelope extraction module. The phase difference calculation module is set to calculate the communication synchronization phase difference parameter and extract the tolerance threshold from the cultural baseline matrix.
[0062] The graphics rendering module connects the phase difference calculation module and the graphics rendering unit of the augmented reality head-mounted display device. The graphics rendering module is configured to output three-dimensional visual constraints in the physical field of view of the augmented reality head-mounted display device.
[0063] See attached document Figure 2 This invention provides an AR voice training intelligent analysis method that incorporates cross-cultural adaptability, comprising the following steps:
[0064] S100, the data alignment module generates a discrete timestamp sequence at the system reference clock frequency, resamples the original speech sequence, hand spatial coordinate sequence and head posture matrix sequence, and maps the original speech sequence, hand spatial coordinate sequence and head posture matrix sequence to the time axis where the discrete timestamp sequence is located.
[0065] S200, the coordinate system reconstruction module performs yaw dead zone filtering on the head rotation matrix after the timestamp is aligned, calculates and establishes a virtual torso reference coordinate system, and performs affine transformation on the absolute spatial coordinates of the hands to convert them into relative displacement data of the hands in the virtual torso reference coordinate system.
[0066] S300, the envelope extraction module extracts the acoustic energy envelope and the hand kinetic energy envelope within the observation time window, calculates the audio gain compensation value by combining the head line of sight angle parameter to generate a smooth acoustic energy envelope, and performs a first-order derivative operation on the hand relative displacement data by combining the tracking confidence parameter output by the outward tracking camera array to generate a continuous hand kinetic energy envelope.
[0067] S400, the phase difference calculation module performs time-domain cross-correlation function calculation on the acoustic energy envelope sequence and the hand kinetic energy envelope sequence within the analysis period, extracts the time offset corresponding to the global maximum value of the cross-correlation function, and records the time offset as the communication synchronization phase difference;
[0068] S500, the graphics rendering module reads the preset target cultural baseline parameters to establish a three-dimensional near-body constraint geometry with the origin of the virtual torso reference coordinate system as the fixed point, and substitutes the relative displacement data of the hand into the mathematical equation corresponding to the three-dimensional near-body constraint geometry to calculate the normal space boundary distance value.
[0069] In S600, the graphics rendering module inputs the communication synchronization phase difference and the normal space over-boundary distance values into the system's preset joint driving operator to calculate the material transparency parameters of the visual mesh. When the communication synchronization phase difference value deviates from the tolerance range corresponding to the target cultural baseline parameter and the normal space over-boundary distance value is greater than zero, the graphics rendering engine is triggered to render the three-dimensional near-body constraint geometry in the first-person real view, and then loops back to execute S100.
[0070] See attached document Figure 2 Step S100 may include the following sub-steps:
[0071] In step S101, the data alignment module establishes the system reference clock frequency and generates a unified discrete timestamp sequence based on the system reference clock frequency.
[0072] The microphone array, inertial measurement unit, and outward tracking camera array inside the augmented reality head-mounted display device each have their own independent hardware sampling clocks. These different hardware sampling clocks have different operating frequencies. To ensure a unified time reference for subsequent phase calculations, the data alignment module selects either the screen refresh rate of the augmented reality head-mounted display device or the sampling frequency of the inertial measurement unit as the target resampling frequency. This target resampling frequency is defined as the system reference clock frequency. Based on the system reference clock frequency, the data alignment module divides continuous discrete-time indices and generates corresponding discrete timestamp sequences. The mathematical expression for the discrete timestamp sequence is as follows:
[0073] ;
[0074] in, Represents a time node in a discrete timestamp sequence; Represents a continuous discrete-time index. The value of is a sequence of natural numbers that increment from zero; This represents the system's basic sampling period. The value is equal to the reciprocal of the system's reference clock frequency.
[0075] In step S102, the data alignment module acquires the raw data streams output by the microphone array, the inertial measurement unit, and the outward tracking camera array in parallel.
[0076] The data alignment module concurrently calls the application programming interfaces of the microphone array, inertial measurement unit, and outward tracking camera array at the system hardware abstraction layer. The data alignment module reads the raw speech sequence output by the microphone array and simultaneously extracts the raw hardware timestamp of each frame of speech signal assigned by the underlying system. The data alignment module reads the head posture matrix sequence output by the inertial measurement unit and simultaneously extracts the corresponding inertial hardware timestamp. The head posture matrix sequence contains the absolute rotation matrix of the head and the head translation vector on the continuous time axis. The data alignment module simultaneously reads the hand spatial coordinate sequence output by the outward tracking camera array and extracts the hardware timestamp of the visual image frame.
[0077] Considering the blind spots in the field of view and the occlusion phenomenon of the arm itself during optical tracking, the data alignment module reads the wrist node tracking confidence scalar output by the underlying computer vision algorithm while acquiring the hand spatial coordinate sequence. The numerical range of the wrist node tracking confidence scalar is between 0 and 1. As for how the computer vision algorithm outputs the wrist node tracking confidence scalar, those skilled in the art can use existing deep learning hand keypoint detection models to perform probabilistic regression calculations. Deep learning hand keypoint detection and probabilistic regression calculations are well-known technologies in this field and will not be elaborated here.
[0078] In step S103, the data alignment module applies interpolation and filtering algorithms to uniformly map the original speech sequence, hand spatial coordinate sequence, and head pose matrix sequence to the time axis where the discrete timestamp sequence is located.
[0079] The original sampling frequency of the raw speech sequence deviates from the system reference clock frequency. To preserve the high-frequency energy details of the speech signal for subsequent envelope extraction, the data alignment module divides the raw speech sequence into segments corresponding to time nodes based on the mapping relationship between the original hardware timestamps and the system discrete timestamp sequence. Strictly corresponding high-frequency data segments are output as a high-frequency audio buffer sequence aligned with the time axis, replacing the direct downsampling operation that would cause loss of high-frequency features.
[0080] Due to hardware communication bus transmission delays and hardware sampling clock drift, the original hardware timestamps and discrete timestamps of the head pose matrix sequence and hand spatial coordinate sequence differ in time node format. There is a time skew. Since directly interpolating the rotation matrix would destroy the orthogonality of the matrix, the data alignment module first converts the absolute head rotation matrix in the head pose matrix sequence into a quaternion representation, and then uses a spherical linear interpolation algorithm to calculate the time nodes. The corresponding alignment quaternions are then converted back into alignment head rotation matrices. For the head translation vector, hand spatial coordinates, and tracking confidence scalar, the data alignment module uses a cubic spline interpolation algorithm to calculate time nodes. For the corresponding numerical points, the data alignment module outputs the alignment head rotation matrix, alignment head translation vector, alignment hand spatial coordinates, and alignment tracking confidence scalar mapped to the time axis of the discrete timestamp sequence through the above interpolation algorithm. The alignment head rotation matrix, alignment head translation vector, alignment hand spatial coordinates, and alignment tracking confidence scalar are synchronized with the high-frequency audio buffer sequence in the time domain, and together constitute the discretized reference data stream to enter the subsequent feature extraction stage.
[0081] See attached document Figure 2 Step S200 may include the following sub-steps:
[0082] In step S201, the coordinate system reconstruction module extracts the head yaw angle component from the aligned head rotation matrix and calculates the head yaw angular velocity values at consecutive time nodes.
[0083] Augmented reality head-mounted displays can only directly output head posture data. However, the spatial distribution of human communicative gestures is based on the torso as a physical reference. During communication, trainees frequently make rapid head-turning observation movements without rotating their torso. Directly using head posture as a reference will cause the gesture reference frame to drift. To eliminate the coupling error between head posture and torso posture, the coordinate system reconstruction module converts the aligned head rotation matrix into Euler angle data format. The coordinate system reconstruction module separates the Euler angle values of rotation around the spatial vertical axis and defines these Euler angle values as the head yaw angle components. For the algorithm of converting rotation matrices to Euler angles, those skilled in the art can use existing three-dimensional spatial geometry basic operation rules. The algorithm of converting rotation matrices to Euler angles is a well-known technology in this field and will not be described in detail here.
[0084] After extracting the head yaw angle component, the coordinate system reconstruction module uses the time difference between adjacent discrete timestamp sequences to perform first-order discrete difference calculation on the head yaw angle component, obtaining the head yaw angular velocity value at the corresponding time node. The mathematical expression corresponding to the head yaw angular velocity value is as follows:
[0085] ;
[0086] in, This indicates the head yaw rate value at the current time point; This represents the head yaw angle component extracted at the current time point; This represents the head yaw angle component recorded at the previous time point; This indicates the system's basic sampling period.
[0087] In step S202, the coordinate system reconstruction module sets the dynamic filtering damping coefficient based on the head yaw rate value, and substitutes the dynamic filtering damping coefficient into the first-order hysteresis low-pass filtering algorithm to calculate the smooth virtual torso yaw angle.
[0088] The coordinate system reconstruction module presets an angular velocity dead zone threshold representing the physiological torsional limit rate of the human neck. The value range of the angular velocity dead zone threshold is set between 0.5 radians per second and 1.0 radians per second. The coordinate system reconstruction module determines the yaw angular velocity value of the head. If the head yaw rate exceeds the dead zone threshold, the coordinate system reconstruction module determines that the current movement is an independent high-frequency head rotation. In this case, the module forcibly sets the dynamic filtering damping coefficient to 0, preventing the torso orientation from updating in accordance with the head yaw rate component. If the head yaw rate is less than or equal to the dead zone threshold, the module determines that the body is slowly and synchronously rotating in sync with the head. In this case, the module restores the dynamic filtering damping coefficient to the system's preset conventional damping constant. The conventional damping constant is configured as a real number greater than 0 and less than 1, preferably between 0.1 and 0.2, to smooth the torso's following motion.
[0089] After determining the dynamic filter damping coefficient, the coordinate system reconstruction module uses a first-order lag low-pass filter algorithm to calculate the virtual torso yaw angle. The mathematical expression corresponding to the first-order lag low-pass filter algorithm is as follows:
[0090] ;
[0091] in, This represents the virtual torso yaw angle calculated at the current time point; This represents the virtual torso yaw angle recorded at the previous time point; This represents the dynamic filter damping coefficient corresponding to the current time point; This indicates the head yaw angle component extracted at the current time point.
[0092] In step S203, the coordinate system reconstruction module combines the topological constant of the human neck with the aligned head translation vector to inversely calculate the absolute coordinates of the fixed point of the sternal manubrium and establish a relatively static virtual torso reference coordinate system.
[0093] The topological constants of the human neck include the standard physical vertical distance and standard physical horizontal offset distance from the geometric center of the human head to the manubrium of the sternum. The coordinate system reconstruction module reads the standard physical vertical distance and standard physical horizontal offset distance from the system's pre-set human bioscale database, or extracts the above constants by reading the posture calibration data of the trainee when standing still during the system initialization phase. Based on the three-dimensional coordinates aligned with the head translation vector, the coordinate system reconstruction module moves the standard physical vertical distance along the vertically downward direction in space, and moves the standard physical horizontal offset distance along the opposite horizontal vector direction corresponding to the virtual torso yaw angle, to calculate the absolute coordinates of the manubrium of the sternum. The coordinate system reconstruction module uses the absolute coordinates of the manubrium of the sternum as the origin of the three-dimensional coordinate system and the orientation vector corresponding to the virtual torso yaw angle as the forward axis of the three-dimensional coordinate system, to construct a virtual torso reference coordinate system independent of head pitch and roll movements.
[0094] In step S204, the coordinate system reconstruction module generates the inverse spatial affine transformation matrix, and uses the inverse spatial affine transformation matrix to convert the spatial coordinates of the aligned hand into relative displacement data of the hand.
[0095] The aligned hand spatial coordinates output by the augmented reality head-mounted display device are three-dimensional position points expressed based on the system's global world coordinate system. In order to accurately measure the range of motion of the trainee's hand gestures relative to the trainee's own body, the coordinate system reconstruction module extracts the global rotation parameters and global translation parameters of the virtual torso reference coordinate system relative to the system's global world coordinate system. The coordinate system reconstruction module uses the global rotation parameters and global translation parameters to construct a fourth-order homogeneous spatial affine transformation inverse matrix. The coordinate system reconstruction module performs matrix multiplication operations on the aligned hand spatial coordinates and the spatial affine transformation inverse matrix to eliminate the global displacement deviation caused by the trainee's overall movement in physical space. After completing the matrix multiplication operation, the coordinate system reconstruction module outputs hand relative displacement data based on the virtual torso reference coordinate system. The hand relative displacement data includes lateral relative displacement components, longitudinal relative displacement components, and normal relative displacement components distributed along the internal coordinate axes of the virtual torso reference coordinate system.
[0096] See attached document Figure 2 Step S300 may include the following sub-steps:
[0097] In step S301, the envelope extraction module calculates the audio gain compensation value by combining the head gaze angle parameter, and generates a smooth acoustic energy envelope.
[0098] When trainees turn their heads during communication, the sound source from their mouths deviates from the main pickup area of the microphone array directly in front of the augmented reality head-mounted display, causing a physical attenuation of the pickup volume. To eliminate the interference of this attenuation on subsequent acoustic energy calculations, the envelope extraction module extracts the absolute difference between the head yaw angle component and the virtual torso yaw angle. This absolute difference is defined as the head gaze angle parameter. Based on this parameter, the envelope extraction module calculates the audio gain compensation value. The mathematical expression for the audio gain compensation value is as follows:
[0099] ;
[0100] in, This represents the audio gain compensation value calculated at the current time point; Indicates the environmental compensation coefficient; This represents the head yaw angle component extracted at the current time point; This represents the virtual torso yaw angle calculated at the current time point; environmental compensation coefficient. The value is set to a range of 0.5 to 1.5, and the specific value is determined by the pre-calibrated microphone array directivity attenuation curve.
[0101] After determining the audio gain compensation value, the envelope extraction module performs an analysis of the corresponding time nodes in the high-frequency audio buffer sequence within a set sliding observation time window. The high-frequency data segments undergo short-time root mean square (SMS) calculation, and the SMS calculation result is multiplied by the audio gain compensation value. The envelope extraction module inputs the product result into a preset low-pass filter function for discrete convolution processing, outputting the acoustic energy envelope. The SMS calculation reflects the energy fluctuation of the speech signal within a local time period, while the low-pass filter function filters out high-frequency pronunciation details, retaining only the low-frequency envelope contour related to the rhythm of communicative actions. The mathematical expression corresponding to the acoustic energy envelope is as follows:
[0102] ;
[0103] in, Represents the acoustic energy envelope; This represents the audio gain compensation value calculated at the current time point; This indicates the total number of high-frequency audio sampling points included within the sliding observation time window; Indicates the first high-frequency data segment in the current data segment. The square of the amplitude of each high-frequency audio sampling point; This represents the low-pass filter function; This represents the discrete convolution operator.
[0104] The length of the sliding observation time window is typically set between 20 and 50 milliseconds. The envelope extraction module divides the length of the sliding observation time window by the system's basic sampling period to obtain the total number of audio sampling points. Low-pass filter function The cutoff frequency range is set between 2 Hz and 5 Hz, which corresponds to the syllable change frequency under normal human speech rate. For the short-time root mean square calculation and discrete convolution processing of discrete audio signals, those skilled in the art can use existing digital signal time-domain analysis methods. The short-time root mean square calculation and discrete convolution processing of discrete audio signals are well-known technologies in this field and will not be described in detail here.
[0105] In step S302, the envelope extraction module constructs a state machine anti-shake control logic based on the alignment tracking confidence scalar output by the outward tracking camera array.
[0106] Outward-tracking camera arrays have a fixed blind spot in the field of view. When the trainee's hand moves outside the blind spot or when arms obstruct each other, the hand's spatial coordinate sequence will experience abrupt changes or even data loss. These abrupt changes can cause velocity jump distortion in subsequent spatial derivative calculations, producing false kinetic energy peaks. To prevent velocity jump distortion, the envelope extraction module sets a stable tracking confidence threshold. The stable tracking confidence threshold is set between 0.6 and 0.8. The envelope extraction module continuously checks whether the alignment tracking confidence scalar is greater than or equal to the stable tracking confidence threshold, and switches the operating state of the system state machine based on the judgment result.
[0107] In step S303, the envelope extraction module performs calculations on the relative displacement data of the hand under the state machine anti-shake control logic to generate a continuous hand kinetic energy envelope.
[0108] When the alignment tracking confidence scalar is greater than or equal to the stable tracking confidence threshold, the system state machine enters the stable tracking state. The envelope extraction module extracts the lateral, longitudinal, and normal relative displacement components from the hand's relative displacement data. The envelope extraction module performs first-order discrete derivative operations on the lateral, longitudinal, and normal relative displacement components to extract the three-dimensional spatial velocity components. The object's kinetic energy is proportional to the square of its velocity. To simplify the computational load and focus on the relative fluctuation trend of energy, the envelope extraction module normalizes the equivalent mass of the hand and omits constant coefficients. The envelope extraction module calculates the sum of squares of the three-dimensional spatial velocity components to generate the hand kinetic energy envelope at the current time point. The mathematical expression corresponding to the hand kinetic energy envelope in the stable tracking state is as follows:
[0109] ;
[0110] in, This represents the kinetic energy envelope of the hand at the current time point; This represents the lateral relative displacement component at the current time point; This represents the longitudinal relative displacement component at the current time point; This represents the normal relative displacement component at the current time point; This represents the lateral relative displacement component at the previous time point; This represents the longitudinal relative displacement component at the previous time point; This represents the normal relative displacement component of the previous time point; This indicates the system's basic sampling period.
[0111] When the alignment tracking confidence scalar is less than the stable tracking confidence threshold, the system state machine enters the lost maintenance state. In the lost maintenance state, since the spatial coordinates no longer have reference value, the envelope extraction module suspends the first-order discrete derivative operation and starts the exponential decay maintenance logic. The envelope extraction module uses the decay constant to calculate the decay of the hand kinetic energy envelope at the previous time node to generate the hand kinetic energy envelope at the current time node. The decay calculation can avoid a cliff drop in the hand kinetic energy envelope signal and maintain the signal continuity of subsequent cross-correlation calculations. The mathematical expression corresponding to the hand kinetic energy envelope in the lost maintenance state is as follows:
[0112] ;
[0113] in, This represents the envelope of hand kinetic energy recorded at the previous time point; Represents the base of the natural logarithm; Represents the attenuation constant. The value is set between 5.0 and 10.0 to control the smooth rate of descent of the kinetic energy signal during visual tracking loss.
[0114] See attached document Figure 2 Step S400 may include the following sub-steps:
[0115] Step S401: The phase difference calculation module sets the sliding time analysis period, extracts the corresponding acoustic energy envelope sequence and hand kinetic energy envelope sequence within the sliding time analysis period, and performs mean removal and normalization preprocessing on the sequences.
[0116] The length of the sliding time analysis cycle is determined based on the average duration of typical short sentence communication segments in humans, with a specific value range set between 1.5 seconds and 3.0 seconds to ensure that the cycle includes complete acoustic accents and hand gesture beats.
[0117] Since the acoustic energy envelope and the hand kinetic energy envelope belong to completely different physical dimensions, and the signal values of both the acoustic energy envelope and the hand kinetic energy envelope are always greater than or equal to 0, exhibiting significant DC components, directly performing time-domain cross-correlation calculations would lead to a basis shift in the calculation results, masking the true synchronization characteristic peaks. Therefore, the phase difference calculation module first calculates the arithmetic mean and standard deviation of the acoustic energy envelope sequence and the hand kinetic energy envelope sequence within the sliding time analysis period, respectively. Subsequently, the phase difference calculation module subtracts the corresponding arithmetic mean from the original sequence to eliminate the DC component, and divides by the corresponding standard deviation for amplitude normalization, outputting the standardized acoustic energy envelope sequence and the standardized hand kinetic energy envelope sequence. The mathematical expression corresponding to the standardized acoustic energy envelope sequence is as follows:
[0118] ;
[0119] in, This represents the standardized acoustic energy envelope value corresponding to the current time point; This represents the original acoustic energy envelope value; This represents the arithmetic mean of the acoustic energy envelope sequence within the sliding time analysis period; This represents the standard deviation of the acoustic energy envelope sequence within the sliding time analysis period. Similarly, the standardized hand kinetic energy envelope sequence is calculated using the same mathematical logic.
[0120] In step S402, the phase difference calculation module performs a one-dimensional time-domain cross-correlation function operation on the standardized acoustic energy envelope sequence and the standardized hand kinetic energy envelope sequence.
[0121] Time-domain cross-correlation is a mathematical method for measuring the similarity of the relative time shifts of two discrete signals in the time domain. When two sequences are aligned at a specific time offset, their waveform fluctuations are most consistent, and the cumulative sum of their products reaches its maximum. The phase difference calculation module uses a discrete-time offset sliding mechanism to map the joint response distribution matrix of the acoustic energy and kinetic energy sequences at different time differences. The mathematical expression corresponding to the one-dimensional time-domain cross-correlation function is as follows:
[0122] ;
[0123] in, The discrete bias parameter is Calculate the numerical value of the cross-correlation function at time; Indicates the discrete bias parameter; This indicates the total number of sampling points included within the sliding time analysis period; This represents the standardized acoustic energy envelope value corresponding to the current time point; This represents the standardized hand kinetic energy envelope value after adding discrete bias parameters.
[0124] Discrete bias parameter The range of values is .in, This represents the maximum discrete bias index. Based on the physiological perception limit of audio-visual asynchrony in human communication, the maximum physical bias time constant is set to 500 milliseconds. The phase difference calculation module divides 500 milliseconds by the system's basic sampling period. The result is then rounded down to obtain the maximum discrete bias index. For the frequency domain acceleration solution of the underlying Fast Fourier Transform of the cross-correlation function, those skilled in the art can directly call existing digital signal processing acceleration libraries. The frequency domain product and inverse transform acceleration solution are well-known technologies in this field and will not be elaborated here.
[0125] In step S403, the phase difference calculation module extracts the time offset corresponding to the global maximum value of the cross-correlation function and records the time offset as the communication synchronization phase difference.
[0126] The phase difference calculation module calculates discrete bias parameters. Iterate through all cross-correlation functions within the range of values to calculate the numerical values. Searching for an envoy The phase difference calculation module converts the specific offset index that yields the maximum value into a physical time difference, thus completing parameter extraction. The mathematical expression corresponding to the communication synchronization phase difference is as follows:
[0127] ;
[0128] in, This represents the communication synchronization phase difference in the final solution output; Indicates the system's basic sampling period; This indicates that the cross-correlation function is calculated numerically. Get the specific bias index corresponding to the global maximum value.
[0129] Communication synchronization phase difference The physical time difference between the peak of the trainee's gesture kinetic energy and the peak of the speech stress during cross-cultural communication was quantified. A communication synchronization phase difference of less than zero indicates that the gesture action lags behind the speech stress, while a communication synchronization phase difference of more than zero indicates that the gesture action precedes the speech stress.
[0130] See attached document Figure 2 Step S500 finally outputs a normal space boundary crossing distance value representing the degree of hand space boundary crossing. Step S500 may include the following sub-steps:
[0131] In step S501, the graphics rendering module reads the preset target culture baseline matrix and extracts the three-dimensional spatial threshold corresponding to the target culture background from the target culture baseline matrix.
[0132] Cross-cultural proximity studies primarily investigate spatial distance preferences in human communication across different cultural backgrounds. The graphics rendering module loads the corresponding target culture baseline matrix from the system database based on the target communication culture category currently selected by the trainee. The graphics rendering module then extracts horizontal proximity threshold parameters, vertical proximity threshold parameters, and normal proximity threshold parameters from the target culture baseline matrix. The horizontal proximity threshold parameters correspond to the permissible arm extension distance on the left and right sides of the human body, the vertical proximity threshold parameters correspond to the permissible arm extension distance in the vertical direction of the human body, and the normal proximity threshold parameters correspond to the permissible interaction distance directly in front of the human body. The normal proximity threshold parameters corresponding to different cultural backgrounds are typically set between 0.45 meters and 1.2 meters to distinguish between intimate distance, personal distance, and social distance.
[0133] In step S502, the graphics rendering module establishes an asymmetric three-dimensional near-body constraint geometry based on the origin of the three-dimensional coordinate system of the completed virtual torso reference coordinate system.
[0134] Because human limb movements have asymmetrical physiological characteristics, and cross-cultural communication has a much lower tolerance for forward spatial intrusion than for lateral or backward intrusion, the graphics rendering module must construct an asymmetrical geometric space. Centered on the origin of the three-dimensional coordinate system of the virtual torso reference coordinate system, the graphics rendering module divides the three-dimensional space into eight independent octaves. Based on the positive and negative directions of the spatial coordinate axes, the graphics rendering module assigns different coefficient values to the lateral, longitudinal, and normal near-body threshold parameters, splicing them together to form an asymmetrical three-dimensional near-body constrained geometry. For the vertex mesh topology distribution at the bottom layer of the spatial geometry generation algorithm, those skilled in the art can instantiate it using existing computer graphics standard libraries. The spatial geometry generation algorithm is a well-known technology in this field and will not be elaborated upon here.
[0135] In step S503, the graphics rendering module substitutes the relative displacement data of the hand into the mathematical equation corresponding to the three-dimensional near-body constraint geometry, calculates and extracts the final normal space boundary distance value.
[0136] The graphics rendering module determines the sign of each coordinate component in the relative displacement data of the hand and dynamically selects the threshold parameter of the corresponding octave for algebraic operations. In the mathematical equation corresponding to the three-dimensional near-body constraint geometry, each coordinate component is divided by the corresponding threshold parameter and the square root of the sum of squares is performed. Essentially, this calculates the normalized radius of the hand's current spatial position relative to the surface of the asymmetric three-dimensional near-body constraint geometry. When the normalized radius is equal to 1, it means that the hand is precisely located on the surface boundary of the asymmetric three-dimensional near-body constraint geometry. Therefore, subtracting the constant 1 from the normalized radius yields the dimensionless relative value of the distance exceeding the boundary. The mathematical expression corresponding to the normal space boundary distance is as follows:
[0137] ;
[0138] in, This represents the value of the out-of-bounds distance in normal space calculated at the current time point; This represents the lateral relative displacement component at the current time point; This represents the longitudinal relative displacement component at the current time point; This represents the normal relative displacement component at the current time point; This represents the lateral near-body threshold parameter dynamically selected based on the positive and negative signs of the lateral relative displacement components; This represents the longitudinal near-body threshold parameter dynamically selected based on the positive and negative signs of the longitudinal relative displacement components; This represents the normal near-body threshold parameter dynamically selected based on the sign of the normal relative displacement components; This represents the boundary truncation function that extracts the maximum value.
[0139] After truncation using the boundary truncation function, a normal space boundary distance value greater than zero indicates that the gesture has substantially exceeded the close-range safety distance constraint in the corresponding cultural context, while a normal space boundary distance value equal to zero indicates that the gesture is within a reasonable communicative space range.
[0140] See attached document Figure 2 Step S600 may include the following sub-steps:
[0141] In step S601, the graphics rendering module obtains the communication synchronization phase difference and the normal space over-boundary distance values, and constructs the material dynamic rendering logic for asymmetric three-dimensional near-body constrained geometry.
[0142] Cross-cultural communication training requires providing trainees with intuitive visual feedback. The communication synchronization phase difference reflects the degree of matching in communication timing, and the normal space boundary crossing distance reflects the compliance of communication space. Traditional training systems usually use separate data panels to display timing and spatial indicators, which will occupy a large amount of the limited field of view of augmented reality display devices. In order to integrate multi-dimensional evaluation indicators into the physical communication space, the graphics rendering module inputs the communication synchronization phase difference and normal space boundary crossing distance into the material dynamic rendering equation, and jointly calculates the visual display parameters of the surface mesh of the asymmetric three-dimensional near-body constraint geometry. The above joint driving mode uses the color change of the constraint geometry surface to indicate the deviation of the action rhythm, and uses the transparency change of the constraint geometry surface to indicate the degree of spatial boundary crossing, realizing the integrated visual mapping of spatiotemporal parameters.
[0143] In step S602, the graphics rendering module uses the communication synchronization phase difference to determine the surface rendering color vector of the asymmetric three-dimensional near-body constraint geometry.
[0144] When a trainee's gestures and vocal emphasis deviate in time, visual feedback is provided through color gradation. The graphics rendering module pre-defines a standard synchronization color vector and a severe synchronization error color vector. Both vectors are defined using a red-green-blue three-channel color space data structure. The standard synchronization color vector is configured with numerical coordinates representing green in the red-green-blue color space, while the severe synchronization error color vector is configured with numerical coordinates representing red. The graphics rendering module uses the absolute value of the communicative synchronization phase difference to perform linear interpolation between the standard synchronization color vector and the severe synchronization error color vector, calculating the interpolation results for the red, green, and blue channels respectively. The mathematical expression corresponding to the surface rendering color vector is as follows:
[0145] ;
[0146] in, This represents the surface rendering color vector calculated at the current time point; This represents the preset standard synchronized color vector; This represents the preset severely out-of-sync color vector; This represents the absolute value of the communication synchronization phase difference in the final solution output; This indicates the preset phase difference tolerance threshold. This indicates a cutoff function that restricts the values within the parentheses to an upper bound of 1.0, representing the phase difference tolerance threshold. The numerical range is set between 100 milliseconds and 300 milliseconds. The above numerical range corresponds to the physiological limit of human visual and auditory rhythm difference perception. When the absolute value of the communication synchronization phase difference exceeds the phase difference tolerance threshold, the upper bound truncation function will force the interpolation ratio to be locked to a constant 1, so that the surface rendering color vector is completely equivalent to the severely out-of-sync color vector.
[0147] Step S603: The graphics rendering module calculates the surface opacity component of the asymmetric three-dimensional near-body constraint geometry by combining the numerical value of the out-of-bounds distance in normal space.
[0148] To avoid obstructing the trainee's view within the normal communication space, when the normal space over-limit distance is zero, the graphics rendering module maintains the asymmetric 3D near-body constraint geometry in a completely transparent state. When the normal space over-limit distance is greater than zero, the graphics rendering module increases the surface opacity component proportionally to the normal space over-limit distance. The mathematical expression for the surface opacity component is as follows:
[0149] ;
[0150] in, This represents the surface opacity component calculated at the current time point; This represents the preset maximum opacity limit constant; This represents the opacity gain coefficient. This represents the value of the out-of-bounds distance in normal space calculated at the current time point; This indicates that the value within the parentheses is limited to a maximum value equal to the maximum opacity limit constant. The upper bound cutoff function.
[0151] Opacity gain factor The opacity gain coefficient characterizes the sensitivity of visual feedback to spatial boundary violations. Its numerical range is determined by the warning gradient set by the system. For example, if the system is set to reach the maximum warning state when a hand exceeds the boundary by 0.5 meters, then the opacity gain coefficient... Configured to 2.0, maximum opacity limit constant. The numerical range is set between 0.6 and 0.8 to ensure that trainees can still see the physical environment background even when their gestures are seriously out of bounds.
[0152] In step S604, the graphics rendering module integrates the surface rendering color vector and the surface opacity component into a rendering data structure, and sends it to the underlying graphics rendering pipeline of the augmented reality head-mounted display device for rendering output.
[0153] The asymmetric 3D near-body constraint geometry is a virtual barrier surrounding the trainee's body. If a traditional solid occlusion rendering pipeline is used, it will block the perspective effect of the augmented reality head-mounted display on the real physical space. The graphics rendering module forces the alpha blending mode and depth buffer write suppression function to be enabled in the underlying graphics rendering pipeline. The alpha blending mode can weight and blend the pixel values of the surface rendering color vector with the perspective pixel values of the current real environment according to the surface opacity component. In the fragment shader of the underlying graphics rendering pipeline, the graphics rendering module performs color blending and alpha overlay operations on the surface mesh vertex pixels of the asymmetric 3D near-body constraint geometry. The pixel data after color blending and alpha overlay operations is output to the frame buffer of the augmented reality head-mounted display to complete the dynamic visual representation of the spatial physical boundary.
[0154] Specific application examples:
[0155] To better understand the technical solution of the AR voice training intelligent analysis system that incorporates cross-cultural adaptability, the following specific application example illustrates the invention further by having trainees (experimental group members) in the fourth week of training wear augmented reality head-mounted display devices to conduct a simulated report.
[0156] The system hardware is set to a screen refresh rate of 100Hz, which means the system base clock frequency is 100Hz. This derivation is excerpted from the appendix. Figure 5 The timeline just reached This node demonstrates how the system calculates and triggers visual intervention, causing the solid line (experimental group) in the graph to fall back after reaching a peak of approximately 0.07.
[0157] Discrete timestamp generation: System basic sampling period The value is equal to the reciprocal of the system's reference clock frequency: Current discrete-time index The corresponding time node in the discrete timestamp sequence is .
[0158] Coordinate system reconstruction calculation:
[0159] The trainee maintains a forward-looking posture, and the head yaw angle component is extracted at the current time point. The head yaw angle component recorded at the previous time point The numerical formula for the head yaw rate is derived as follows:
[0160] ;
[0161] Since the angular velocity dead zone threshold is less than 0.6 rad / s, the system's preset conventional damping constant is invoked. The current virtual torso yaw angle is calculated. It stabilizes at 0.10 rad.
[0162] Envelope extraction calculation:
[0163] After the inverse affine transformation, the trainee extends their arm forward to emphasize the point. The coordinates of the hand in the virtual torso reference coordinate system are as follows: (Previous time step) The coordinates are (0.190, 0.190, 0.710) meters; at the current time... Lateral relative displacement components Longitudinal relative displacement component Normal relative displacement component The mathematical expression for the kinetic energy envelope of the hand is derived as follows:
[0164] ;
[0165] Calculation of phase difference in communication synchronization:
[0166] A one-dimensional time-domain cross-correlation function calculation is performed within a sliding-time analysis period containing 200 sampling points. Since the trainees are in their fourth week of training and their actions are highly coordinated, the system's underlying calculations determine the discrete bias parameter. At that time, the cross-correlation function is calculated numerically. The mathematical expression for the phase difference in communication synchronization, obtained at the global maximum, is derived as follows:
[0167] ;
[0168] The calculated absolute value is 50 milliseconds, falling within the range of... Figure 4 The experimental group's mean and standard deviation (54.2 ± 8.2 ms) in the mid-term final test (week 4) were within the true distribution range, demonstrating that the trainees had achieved a high degree of communicative synchronization at this time.
[0169] Extraction of normal space out-of-bounds distance values:
[0170] The target culture is set as North American business spaces, and lateral proximity threshold parameters are extracted. Longitudinal near-body threshold parameter Normal near-body threshold parameter Substitute the hand coordinates (0.200, 0.200, 0.727) into the mathematical expression corresponding to the out-of-bounds distance in normal space:
[0171] ;
[0172] The calculated result of 0.070 corresponds to the attached... Figure 5 middle x-axis The solid line (experimental group) reaches a small peak value at 1 second.
[0173] Visual mesh material transparency rendering and motion correction:
[0174] Set the maximum opacity limit constant Opacity gain coefficient The mathematical expression for the surface opacity component is derived as follows:
[0175] ;
[0176] Practical training feedback performance: At a split second, a warning grid with 14% transparency was instantly generated in the trainee's AR field of vision. This triggered a conditioned reflex of muscle memory, causing the trainee to immediately retract their arm backward. Figure 5 It can be clearly observed that after the experimental group (solid line) triggered a peak of 0.07 at 3.0 seconds, the curve showed a steep drop and returned to zero; while at this time, the control group (dashed line), which lacked visual intervention, continued to extend its hand forward, and the distance of the hand that crossed the boundary soared to the boundary state of 0.13, and remained there for as long as 0.5 seconds.
[0177] Experimental verification and effect comparison:
[0178] To verify the effectiveness of the AR voice training intelligent analysis system with cross-cultural adaptability, the invention team conducted a 4-week comparative experiment using the system (control group of 20 people and experimental group of 20 people).
[0179] Verification of the macro-evolution of spatial compliance by appendix Figure 3 It can be known that:
[0180] During the baseline period, due to the close learning habits of their native language cultures, both groups of trainees had a boundary crossing rate of approximately 16 times per 10 minutes (15.6 times in the control group and 16.1 times in the experimental group).
[0181] After 4 weeks of training, the control group, which used traditional video playback training, showed a very slow decline, with 9.4 serious violations still remaining in the 4th week (because they could not obtain spatial awareness at the moment the action occurred).
[0182] The experimental group that received the AR dynamic visual intervention of this invention showed a significant inflection point in the second week, and the out-of-bounds occurrence rate dropped to 0.8 times by the fourth week. This proves that the out-of-bounds transparency rendering mechanism driven by the truncation function can reshape the trainees' cross-cultural communication space sense from the level of neural reflexes.
[0183] The tempo coordination verification is attached. Figure 4 (The phase difference error diagram for communication synchronization) shows that:
[0184] This figure statistically analyzes the absolute values of the communicative synchronization phase difference calculated within the sliding time analysis period. .
[0185] The control group still had a mean of 225.6 milliseconds at the end of the test (week 4) and a long error bar (standard deviation ± 20.5 milliseconds), indicating that the trainees' actions were still lagging behind their speech and that there were significant individual differences within the group.
[0186] During the final test, the mean absolute phase difference of the experimental group had converged significantly to 54.2 milliseconds (i.e., the interval where 50 milliseconds is located), and the error bar was short (standard deviation was only ±8.2 milliseconds). This proves that the phase difference extracted by the one-dimensional time-domain cross-correlation function operation of this invention, through the color vector gradual feedback, corrected the problem of audio-visual asynchrony in the group, so that the rhythm of different trainees was pulled back to the vicinity of the unified ideal synchronization baseline (0 milliseconds).
[0187] Micro-motion control verification conclusions:
[0188] Combined with appendix Figure 5 The single-action acquisition timeline shown in the figure demonstrates that, compared with existing video scoring systems, this invention has the technical advantages of millisecond-level capture, zero-latency rendering, and muscle-level correction, proving the feasibility of using joint driving operators to transform abstract spatiotemporal parameters into physical visual constraints.
Claims
1. An AR voice training intelligent analysis system incorporating cross-cultural adaptability, characterized in that: include: The data alignment module is configured to acquire the original speech sequence, head pose matrix sequence, and hand spatial coordinate sequence, and map them to a discrete timestamp sequence. The coordinate system reconstruction module is configured to establish a virtual torso reference coordinate system based on the head posture matrix sequence, and convert the hand spatial coordinate sequence into relative hand displacement data under the virtual torso reference coordinate system; The envelope extraction module is configured to extract the acoustic energy envelope and the hand kinetic energy envelope from the original speech sequence and the relative hand displacement data; The phase difference calculation module is configured to calculate and extract the communication synchronization phase difference using the acoustic energy envelope and the hand kinetic energy envelope; The graphics rendering module is configured to create a three-dimensional near-body constraint geometry, substitute the relative displacement data of the hand into the mathematical equation corresponding to the three-dimensional near-body constraint geometry to calculate the normal space boundary distance value, and combine the communication synchronization phase difference to calculate the visual display parameters, and render the three-dimensional near-body constraint geometry to provide visual feedback for cross-cultural training analysis.
2. The AR voice training intelligent analysis system with cross-cultural adaptability as described in claim 1, characterized in that, The data alignment module is configured to acquire the original speech sequence, head pose matrix sequence, and hand spatial coordinate sequence, and map them to a discrete timestamp sequence, specifically including: The data alignment module selects the screen refresh rate of the augmented reality head-mounted display device or the sampling frequency of the inertial measurement unit as the system reference clock frequency, and generates the discrete timestamp sequence based on the system reference clock frequency; The data alignment module applies interpolation and filtering algorithms to output the alignment head rotation matrix, alignment head translation vector, alignment hand spatial coordinates, and alignment tracking confidence scalar mapped to the time axis of the discrete timestamp sequence, and divides the original speech sequence into high-frequency data segments mapped to the time axis of the discrete timestamp sequence.
3. The AR voice training intelligent analysis system incorporating cross-cultural adaptability according to claim 2, characterized in that, Establishing a virtual torso reference coordinate system based on the head pose matrix sequence specifically includes: The coordinate system reconstruction module extracts the head yaw angle component from the aligned head rotation matrix and calculates the head yaw angle velocity values at consecutive time nodes. The dynamic filter damping coefficient is set based on the comparison result between the head yaw rate value and the angular velocity dead zone threshold. Substitute the dynamic filter damping coefficient into the first-order hysteresis low-pass filter algorithm to calculate the virtual torso yaw angle. The coordinate system reconstruction module combines the topological constant of the human neck with the aligned head translation vector to inversely calculate the absolute coordinates of the sternal manubrium. The absolute coordinates of the sternal manubrium are used as the origin of the three-dimensional coordinate system and the orientation vector corresponding to the virtual torso yaw angle is used as the forward axis to construct the virtual torso reference coordinate system.
4. The AR voice training intelligent analysis system incorporating cross-cultural adaptability according to claim 3, characterized in that, Converting the hand spatial coordinate sequence into relative hand displacement data in the virtual torso reference coordinate system specifically includes: The coordinate system reconstruction module extracts the global rotation parameters and global translation parameters of the virtual torso reference coordinate system relative to the system's global world coordinate system, and constructs the inverse spatial affine transformation matrix; Perform matrix multiplication on the aligned hand spatial coordinates and the inverse of the spatial affine transformation matrix to output the relative displacement data of the hand, which includes lateral relative displacement components, longitudinal relative displacement components, and normal relative displacement components.
5. The AR voice training intelligent analysis system incorporating cross-cultural adaptability according to claim 4, characterized in that, Extracting the acoustic energy envelope from the original speech sequence and the relative hand displacement data specifically includes: The envelope extraction module extracts the absolute difference between the head yaw angle component and the virtual torso yaw angle as the head line of sight yaw angle parameter. Calculate the audio gain compensation value based on the head gaze angle parameter; Within the sliding observation time window, short-time root mean square (RMS) calculation is performed on the high-frequency data segment corresponding to the original speech sequence. The RMS calculation result is then multiplied with the audio gain compensation value. The product result is input into a low-pass filter function for discrete convolution processing to generate a smooth acoustic energy envelope.
6. The AR voice training intelligent analysis system incorporating cross-cultural adaptability according to claim 5, characterized in that, Extracting the hand kinetic energy envelope from the original speech sequence and the relative hand displacement data specifically includes: The envelope extraction module determines whether the alignment tracking confidence scalar is greater than or equal to the stable tracking confidence threshold. When the alignment tracking confidence scalar is greater than or equal to the stable tracking confidence threshold, the first-order discrete derivative operation is performed on the lateral relative displacement component, the longitudinal relative displacement component, and the normal relative displacement component to extract the three-dimensional spatial velocity components, and the sum of squares of the three-dimensional spatial velocity components is calculated to generate the hand kinetic energy envelope. When the alignment tracking confidence scalar is less than the stable tracking confidence threshold, the hand kinetic energy envelope of the previous time node is calculated using the decay constant to generate the hand kinetic energy envelope of the current time node.
7. The AR voice training intelligent analysis system incorporating cross-cultural adaptability according to claim 1, characterized in that, The extraction of the communication synchronization phase difference using the acoustic energy envelope and the hand kinetic energy envelope specifically includes: The phase difference calculation module extracts the corresponding acoustic energy envelope sequence and hand kinetic energy envelope sequence within the sliding time analysis period; The acoustic energy envelope sequence and the hand kinetic energy envelope sequence are respectively subjected to mean removal and normalization preprocessing to output standardized acoustic energy envelope sequences and standardized hand kinetic energy envelope sequences; Perform a one-dimensional time-domain cross-correlation function operation on the standardized acoustic energy envelope sequence and the standardized hand kinetic energy envelope sequence; Extract the time offset corresponding to the global maximum value of the cross-correlation function, and record the time offset as the communication synchronization phase difference.
8. The AR voice training intelligent analysis system incorporating cross-cultural adaptability according to claim 4, characterized in that, Substituting the relative displacement data of the hand into the mathematical equation corresponding to the three-dimensional near-body constraint geometry to calculate the normal space boundary distance specifically includes: The graphics rendering module extracts the horizontal mystic threshold parameters, the vertical mystic threshold parameters, and the normal mystic threshold parameters from the preset target cultural baseline matrix. Based on the virtual torso reference coordinate system, the lateral near-body threshold parameters, the longitudinal near-body threshold parameters, and the normal near-body threshold parameters are spliced together to form the asymmetric three-dimensional near-body constraint geometry; Substitute each coordinate component of the relative displacement data of the hand into the mathematical equation corresponding to the three-dimensional near-body constraint geometry to calculate the normalized radius, and combine the boundary truncation function to calculate and extract the normal space over-boundary distance value.
9. The AR voice training intelligent analysis system incorporating cross-cultural adaptability according to claim 1, characterized in that, The calculation of visual display parameters based on the aforementioned communication synchronization phase difference specifically includes: The visual display parameters include a surface rendering color vector; The graphics rendering module uses the absolute value of the communication synchronization phase difference to perform linear interpolation calculation between the preset standard synchronization color vector and the severely out-of-synchronization color vector, and combines the interpolation result with the upper bound truncation function to form the surface rendering color vector.
10. The AR voice training intelligent analysis system incorporating cross-cultural adaptability according to claim 9, characterized in that, The rendered output of the 3D near-gravity constrained geometry provides visual feedback for cross-cultural training analysis, specifically including: The visual display parameters include a surface opacity component; When the normal space out-of-bounds distance value is greater than zero, the graphics rendering module uses the normal space out-of-bounds distance value, the opacity gain coefficient, and the maximum opacity limit constant, combined with the upper bound truncation function, to calculate the surface opacity component. The graphics rendering module integrates the surface rendering color vector and the surface opacity component into a rendering data structure. With the transparency blending mode and depth buffer write suppression enabled, the module renders and outputs the three-dimensional near-bodily constrained geometry based on the rendering data structure.