Spatial audio signal generation method and headphone
By fusing game audio signals with multi-dimensional posture signals and rendering them to generate spatial audio signals, the problem of sound field switching delay and low accuracy when the head is turned in headphones is solved, and higher accuracy audio orientation perception is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GEER TECH CO LTD
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-05
AI Technical Summary
Existing spatial audio technology for headphones suffers from delays in sound field switching when the head rotates, and has low orientation mapping accuracy, making it unsuitable for the rapidly changing audio requirements of game genres.
By acquiring audio signals within the game genre and multi-dimensional posture signals of the user, the audio signals are converted into spectral features, and the multi-dimensional posture signals are mapped into multi-dimensional feature vectors, which are then embedded into the spectral features to form a fused feature matrix. The spatial audio signals are then generated using a sound field rendering model for feature rendering.
Reduce sound field switching latency and improve positional accuracy when the user's head moves, matching the rapidly changing audio requirements of game genres.
Smart Images

Figure CN121985283A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of headphone technology, and in particular to a method for generating spatial audio signals and a pair of headphones. Background Technology
[0002] In games, a player's accurate perception of audio location directly affects the gaming experience, especially in first-person shooter (FPS) games and multiplayer online battle arena (MOBA) games. The location recognition of audio such as footsteps, explosions, and teammates' voices is a key basis for players to judge the battle situation and make decisions.
[0003] Existing spatial audio technologies for headphones mostly employ fixed sound field rendering or simple head posture mapping. In this case, head posture data is input as an independent parameter into the sound field rendering model for rendering, resulting in a delay in sound field switching when the head rotates, and low orientation mapping accuracy, which cannot match the rapidly changing audio requirements in game genres. Summary of the Invention
[0004] The main objective of this invention is to provide a spatial audio signal generation method and headphones, which aim to solve the problems of sound field switching delay and low orientation mapping accuracy when the head rotates, which cannot match the rapidly changing audio requirements in game genres.
[0005] To achieve the above objectives, the present invention proposes a method for generating spatial audio signals, the method comprising: Acquire audio signals within the game genre and multi-dimensional posture signals of the user; The audio signal is converted into spectral features, and the multi-dimensional attitude signal is mapped into a corresponding multi-dimensional feature vector; The multi-dimensional feature vectors are embedded into the spectral features to obtain a fused feature matrix; Spatial audio signals are generated by performing feature rendering on the fused feature matrix using a sound field rendering model.
[0006] Optionally, embedding the multi-dimensional feature vector into the spectral features to obtain a fused feature matrix includes: Get the game type of the currently running game; Determine the correlation weights between the spectral features and the feature vectors of each dimension within the multi-dimensional feature vector based on the game type; The fused feature matrix is obtained by fusing the dimensional feature vectors within the multi-dimensional feature vectors into the spectral feature based on the correlation weights.
[0007] Optionally, the multi-dimensional attitude signal is a three-dimensional attitude parameter signal; the acquisition of the audio signal within the game type and the multi-dimensional attitude signal within the game type includes: User attitude data is collected using inertial sensors; The attitude data is denoised using a Kalman filter algorithm. The denoised attitude data is decomposed into three-dimensional attitude parameter signals by quaternion transformation; Obtain the initial audio signal within the game type; The high-frequency components in the initial audio signal are pre-emphasized to obtain the audio signal.
[0008] Optionally, the noise reduction of the attitude data using the Kalman filter algorithm includes: Get the game type of the currently running game; Adjust the filtering coefficients of the Kalman filter algorithm according to the game type; The attitude data is denoised using a Kalman filter algorithm with adjusted filter coefficients.
[0009] Optionally, the spatial audio signal generation method further includes: A unified timestamp is assigned to each frame of the multi-dimensional attitude signal and each frame of the audio signal through a timestamp alignment mechanism. Using the signal frequency of the audio signal as a reference, the multi-dimensional attitude signal after timestamp allocation is linearly interpolated using a linear interpolation algorithm; The step of converting the audio signal into spectral features and mapping the multi-dimensional attitude signal into a corresponding multi-dimensional feature vector includes: The audio signal after timestamp allocation is converted into spectral features using short-time Fourier transform, and the multi-dimensional attitude signal after linear interpolation is mapped into the corresponding multi-dimensional feature vector.
[0010] Optionally, the spatial audio signal generation method further includes: Obtain the initial sound field rendering model for the execution feature rendering; Retrieve audio feature templates for different game genres from the storage; Set a scene embedding layer within the initial sound field rendering model; The scene embedding layer is used to embed the target audio feature template that matches the fusion feature matrix into the initial sound field rendering model; The sound field rendering model is obtained by adjusting the rendering parameters using the target audio feature template.
[0011] Optionally, after obtaining the initial sound field rendering model for performing feature rendering, the method further includes: The input layer and decoder of the initial sound field rendering model are adjusted according to the fusion feature matrix; The encoder in the adjusted initial sound field rendering model is lightweighted to obtain an optimized initial sound field rendering model.
[0012] Optionally, adjusting the input layer and decoder of the initial sound field rendering model according to the fused feature matrix includes: Based on the audio dynamic features within the fused feature matrix, an audio feature extraction branch is added to the input layer of the initial sound field rendering model. The audio feature extraction branch includes a first convolutional kernel, a second convolutional kernel, and a third convolutional kernel. The first convolutional kernel, the second convolutional kernel, and the third convolutional kernel respectively extract high-frequency features, mid-frequency features, and low-frequency features within the audio signal, and fuse them with the corresponding dimension of the attitude parameter signals within the multi-dimensional attitude signal. The optimized initial sound field rendering model is obtained by adjusting the number of decoding layers in the decoder of the initial sound field rendering model according to the audio feature extraction branch, wherein the number of decoding layers is the same as the number of audio feature extraction branches.
[0013] Optionally, the step of lightweighting the encoder within the adjusted initial sound field rendering model to obtain an optimized initial sound field rendering model includes: The number of first self-attention mechanisms and the number of second self-attention mechanisms in each layer of the encoder of the initial sound field rendering model are configured according to the game type. The first self-attention mechanism is used to focus on the frequency correlation of the audio signal, and the second self-attention mechanism is used to focus on the correlation between the audio signal and the multi-dimensional pose signal. By reducing the convolution dimension of the in-encoder feedforward neural network after configuring the self-attention mechanism, an optimized initial sound field rendering model is obtained. In addition, to achieve the above objectives, the present invention also provides a headset for performing any of the spatial audio signal generation methods described above; The headset includes: a posture sensor, an audio input module, a clock module, a processor, a fusion chip, and a sound generation module; The attitude sensor, the audio input module, and the clock module are all connected to the fusion chip, the fusion chip is also connected to the processor, and the processor is also connected to the sound generation module. The processor is equipped with a sound field rendering model.
[0014] This invention provides a method for generating spatial audio signals and a headset. The method includes: acquiring audio signals within a game genre and multi-dimensional posture signals of a user; converting the audio signals into spectral features and mapping the multi-dimensional posture signals into corresponding multi-dimensional feature vectors; embedding the multi-dimensional feature vectors into the spectral features to obtain a fused feature matrix; and performing feature rendering on the fused feature matrix using a sound field rendering model to generate spatial audio signals. This invention, by acquiring audio signals within a game genre and multi-dimensional posture signals of a user, and embedding the multi-dimensional feature vectors mapped from the posture signals into the spectral features of the audio signals, fuses the audio signals with the multi-dimensional posture signals. Furthermore, by performing feature rendering on the fused feature matrix, it can reduce the latency of sound field switching and improve the accuracy of orientation when the user's head moves, thereby matching the rapidly changing audio requirements of game genres. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating the first embodiment of the spatial audio signal generation method proposed in this invention. Figure 2 This is a schematic diagram of the first process of a second embodiment of the spatial audio signal generation method proposed in this invention; Figure 3 This is a second flowchart illustrating a second embodiment of the spatial audio signal generation method proposed in this invention. Figure 4 This is a schematic diagram of the third process of a second embodiment of the spatial audio signal generation method proposed in this invention; Figure 5 This is a schematic diagram of the first process of the third embodiment of the spatial audio signal generation method proposed in this invention; Figure 6 This is a schematic diagram of the second process of the third embodiment of the spatial audio signal generation method proposed in this invention; Figure 7 This is a schematic diagram of the structure of the optimized sound field rendering model of the present invention; Figure 8 This is a schematic diagram of the structure of the headphones proposed in this invention.
[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a specific posture. If the specific posture changes, the directional indicators will also change accordingly.
[0020] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the use of "and / or" or "and / or" throughout the text includes three parallel solutions. For example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0021] It should be understood that during audio signal output, different head positions result in different sound fields in the user's headphones, allowing the user to more accurately experience audio changes within the game. For example, when playing FPS games, if an audio signal is generated at a fixed location, the audio might be directly in front of the user if their head is not turned; however, if the user's head is turned 90 degrees, the audio might be located to their left or right. In this case, the audio sound field input to the user's ears needs to be adjusted by 90 degrees as well, so that the user perceives the audio's location as being to their left or right.
[0022] For the sound field adjustment process, fixed sound field rendering or simple head posture mapping is usually used. In this case, the head posture data is input as an independent parameter into the sound field rendering model for rendering, which causes a delay in sound field switching when the head rotates, and the orientation mapping accuracy is low, which cannot match the rapidly changing audio requirements in game genres.
[0023] Reference Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the spatial audio signal generation method proposed in this invention. Based on Figure 1 This invention proposes a method for generating spatial audio signals.
[0024] In this embodiment, the spatial audio signal generation method includes: Step S10: Acquire audio signals within the game genre and multi-dimensional posture signals of the user; It is understood that the execution entity of the spatial audio signal generation method can be a headset or a controller within a headset. In this embodiment and the following embodiments, the controller within a headset is used as the execution entity to describe the spatial audio signal generation method.
[0025] It should be noted that the audio signals within the game category refer to the audio signals generated within the game during gameplay. These audio signals can be in-game voice communication signals, signals emitted by in-game characters, or signals generated by changes in in-game objects. The multi-dimensional posture signals refer to the signals corresponding to the head posture of the user wearing the headphones. These multi-dimensional posture signals can include posture signals composed of head rotation angles, lateral tilt angles, and pitch angles, such as a 15-degree head rotation, a 5-degree lateral tilt, or a 10-degree head tilt.
[0026] In the specific acquisition process, audio signals within the game genre can be acquired through the audio input module on the headset. For example, the sampling frequency of the audio input module can be set to 96kHz, the bit depth to 24bit, and the audio format to multi-channel Pulse Code Modulation (PCM) signal. Then, the audio signal within the game genre can be directly acquired using the parameter-set audio input module. For the user's multi-dimensional posture signals, an inertial measurement unit (IMU) installed in the headset can be used to acquire the user's multi-dimensional posture signals during gameplay.
[0027] Step S20: Convert the audio signal into spectral features and map the multi-dimensional attitude signal into a corresponding multi-dimensional feature vector.
[0028] It should be understood that spectral characteristics refer to the features of an audio signal in the frequency domain. These spectral characteristics can visually demonstrate the intensity characteristics of an audio signal at different frequencies. Multidimensional feature vectors are a representation of the detailed features that make up a multidimensional attitude signal; for example, a multidimensional attitude signal is composed of different features such as rotation angle, yaw angle, and pitch angle.
[0029] In practical implementation, audio signals within a game genre can be converted into spectral features using a short-time Fourier transform, for example, into a 2048*128 spectral graph, where 2048 represents the number of time frames and 128 represents the number of frequency units. For multi-dimensional attitude signals, interference data can be filtered out or standardized first. Then, multiple features of the standardized multi-dimensional attitude signal can be extracted, and each feature can be vectorized to obtain the corresponding multi-dimensional feature vector.
[0030] Step S30: Embed the multi-dimensional feature vector into the spectral feature to obtain the fused feature matrix.
[0031] It should be understood that the fused feature matrix is a feature matrix that includes the spectral features of the audio signal and the spectral features of the multi-dimensional feature vectors. This fused feature matrix includes a weight matrix for the spectral features corresponding to the audio signal and a weight matrix for each dimensional feature vector.
[0032] In practice, given the multi-dimensional feature vector and the spectral features of the audio signal, each feature within the multi-dimensional feature vector can be adjusted according to the weight matrix of the spectral features to generate a feature weight matrix with the same dimension as the weight matrix of the spectral features. Then, the elements of the two weight matrices are multiplied to obtain the feature spectrum corresponding to the feature weight matrix. Finally, the feature spectrum corresponding to each feature is embedded into the feature spectrum of the audio signal in a weighted manner to obtain the fused feature matrix.
[0033] Step S40: Generate spatial audio signals by performing feature rendering on the fused feature matrix using the sound field rendering model.
[0034] It should be noted that the sound field rendering model can be a real-time multimodal sound field rendering system based on deep learning. This system generates a three-dimensional audio signal by deeply fusing spectral features and multi-dimensional pose data. The input to this rendering model can be a fused feature matrix, and the output is a spatial audio signal. The spatial audio signal is an audio signal input to the user's ear that possesses spatial awareness. When the user receives the spatial audio signal, they can determine its specific source based on this signal.
[0035] In practice, the fusion feature matrix can be input into the sound field rendering model. The sound field rendering model can directly analyze the spectral features and multi-dimensional posture data within the fusion feature matrix, thereby determining and outputting the spatial audio signal that the user needs to receive under the current posture.
[0036] This embodiment provides a method for generating spatial audio signals. The method includes: acquiring audio signals within a game genre and multi-dimensional posture signals of a user; converting the audio signals into spectral features and mapping the multi-dimensional posture signals into corresponding multi-dimensional feature vectors; embedding the multi-dimensional feature vectors into the spectral features to obtain a fused feature matrix; and performing feature rendering on the fused feature matrix using a sound field rendering model to generate spatial audio signals. This embodiment, by acquiring audio signals within a game genre and multi-dimensional posture signals of a user, and embedding the multi-dimensional feature vectors mapped from the multi-dimensional posture signals into the spectral features of the audio signals, fuses the audio signals with the multi-dimensional posture signals. Furthermore, it performs feature rendering on the fused feature matrix. This reduces the latency of sound field switching and improves the accuracy of orientation when the user's head moves, thereby matching the rapidly changing audio requirements of game genres.
[0037] Based on the first embodiment of the spatial audio signal generation method described above, a second embodiment of the spatial audio signal generation method of the present invention is proposed. (Refer to...) Figure 2 , Figure 2 This is a schematic diagram of the first process of a second embodiment of the spatial audio signal generation method proposed in this invention.
[0038] In this embodiment, step S30 includes: Step S301: Obtain the game type of the currently running game.
[0039] It should be noted that the game type can be role-playing game, action game, adventure game, shooting game, etc. In practice, the type of game the user is currently playing can be determined by comparing the game name and specific content.
[0040] Step S302: Determine the correlation weights between the spectral features and the feature vectors of each dimension within the multi-dimensional feature vector according to the game type.
[0041] It should be noted that the correlation weight refers to the proportion of correlation between a single-dimensional feature vector and the spectral features of the audio signal. The sum of the correlation weights between feature vectors of all dimensions and the spectral features is 1.
[0042] It should be understood that the correlation weights between the various dimensions within the multi-dimensional feature vector and the audio signal differ for different game genres. For example, in shooting games, feature vectors for dimensions such as head pitch angle, yaw angle, and roll angle have relatively high correlation weights with the audio signal, while feature vectors for dimensions such as footsteps and voice communication have relatively low correlation weights. In role-playing games (RPGs), feature vectors for dimensions such as character movements and voice have relatively high correlation weights with the audio signal.
[0043] Therefore, given a specific game genre, the correlation weights between the feature vectors of each dimension and the corresponding spectral features of the audio signal can be determined based on the game genre. For example, in FPS games, the audio signal is the enemy's gunshot, so the correlation weights for dimensions such as the user's head yaw angle and roll angle are relatively large to accurately determine the enemy's position.
[0044] Step S303: Based on the correlation weight, fuse each of the dimensional feature vectors in the multi-dimensional feature vector into the spectral feature to obtain the fused feature matrix.
[0045] It should be understood that, given the correlation weights between each dimension of features and the spectral features, the parameter values and weight values corresponding to the feature vectors of each dimension can be calculated. Then, the calculated parameter values are used to generate a weight matrix of the feature with the same dimension as the weight matrix of the spectral features. The elements of the two weight matrices are then multiplied to obtain the feature spectrum corresponding to the feature weight matrix. Finally, the feature spectrum corresponding to each feature is embedded into the feature spectrum of the audio signal by weighting to obtain the fused feature matrix.
[0046] Reference Figure 3 , Figure 3 This is a second flowchart illustrating a second embodiment of the spatial audio signal generation method proposed in this invention.
[0047] In this embodiment, step S10 includes: Step S11: Collect the user's attitude data using an inertial sensor.
[0048] It should be noted that the user's posture data refers to the characteristics of the user's head while playing games while wearing headphones. This posture data may include features such as head tilt and head rotation.
[0049] An inertial sensor can be a sensor that includes a gyroscope and an accelerometer. The gyroscope is mainly used to determine the angle of change of the user's head, and the accelerometer is used to determine the rate of change of the user's head. The sampling rate of this inertial sensor is set to 1kHz, and the data dimensions include x / y / z axis angular velocity and x / y / z axis acceleration, etc.
[0050] In practice, the gyroscope within an inertial sensor can be used to collect the amplitude of the angle change of the user's head, and an accelerometer can be used to detect the acceleration data during the head change process. Based on the amplitude of the angle change and the acceleration, the user's head posture data can be determined. For example, an inertial sensor can collect one frame of posture data every 1ms, including angular velocity (range: -2000° / s~2000° / s) and acceleration (range: -16g~16g).
[0051] Step S12: Denoise the attitude data using the Kalman filter algorithm.
[0052] It should be noted that the Kalman filter algorithm is an algorithm that uses the state equations of a linear system to make an optimal estimate of the system state using the system's input and output observation data. In this process, because the observed data includes the influence of noise in the system, the process of processing the data using the Kalman filter algorithm can also be regarded as a noise reduction process.
[0053] It should be understood that the collected user posture data may contain noise due to factors such as the user's hair and sweat. In this embodiment, the noise in the posture data can be filtered out using a Kalman filter algorithm, thereby preventing the noise data from affecting the spatial audio signal.
[0054] Step S13: Decompose the denoised attitude data into three-dimensional attitude parameter signals through quaternion transformation.
[0055] It should be understood that the three-dimensional attitude parameter signal refers to a signal presenting an attitude including three dimensions. In this embodiment, the three-dimensional attitude parameter signal can specifically be the pitch angle, yaw angle, and roll angle of the user's head; these pitch angle, yaw angle, and roll angle can accurately determine the user's head attitude data when combined. Quaternion conversion transforms a composite parameter into a superposition of three different parameters. For example, simultaneously adjusting the pitch angle, yaw angle, and roll angle, as well as adjusting the head pose data, may yield the same adjustment result. For example, if the user's head is tilted to the upper left, the corresponding pitch angle is upward, the yaw angle is upward, and the roll angle remains unchanged.
[0056] In practical implementation, the pitch, yaw, and roll angles can be determined from the acquired attitude data using quaternion transformation. The specific value of each dimension in the three-dimensional attitude parameter signal corresponds to the position point in the spatial coordinate system of the attitude data. The quaternion transformation formulas are: q0 = cos(θ / 2), q1 = sin(θ / 2)*nx, q2 = sin(θ / 2)*ny, q3 = sin(θ / 2)*nz, where θ is the rotation angle, and nx, ny, and nz are the unit vectors of rotation along the x-axis, y-axis, and z-axis, respectively.
[0057] Step S14: Obtain the initial audio signal within the game type.
[0058] It should be understood that the initial audio signal is the audio signal directly output from the game being run by the user. This audio signal has not undergone any processing, such as filtering or amplification. In practice, the initial audio signal within the game type can be directly acquired through the audio input module on the headset. The audio input module acquires one audio sample every 10.4μs, and every 128 samples constitute one frame of audio data.
[0059] Step S15: Pre-emphasize the high-frequency components in the initial audio signal to obtain the audio signal.
[0060] It should be understood that an initial audio signal typically includes low-frequency, mid-frequency, and high-frequency components. However, high-frequency signals usually have lower energy; for example, the audio signal in a game is mainly human voices or music, and its energy is concentrated in the low-frequency components, while the high-frequency components themselves are relatively weak. Furthermore, when noise is present in the audio signal, quantization applies uniform noise to every frequency range. For high-frequency components, this noise can significantly affect the signal-to-noise ratio, and may even cause the high-frequency components to be completely masked by the noise.
[0061] Pre-emphasis is used to increase the amplitude of high-frequency components before quantization of an audio signal. During quantization, this can effectively reduce the signal-to-noise ratio. Of course, after quantization, in order to ensure the originality of the audio signal, the high-frequency part of the pre-emphasized audio signal needs to be de-emphasized.
[0062] In practice, the emphasis coefficient of the high-frequency components can be determined first. This coefficient ranges from 0 to 1; the larger the coefficient, the stronger the high-frequency boost. In this embodiment, an emphasis coefficient of 0.97 can be used. Then, a high-pass filter is set based on this coefficient, and the high-frequency components of the audio signal are emphasized using this filter.
[0063] By pre-emphasizing the high-frequency components of the audio signal, the signal-to-noise ratio of the audio signal can be effectively reduced, thereby obtaining a more accurate audio signal.
[0064] Step S12 includes: Step S121: Obtain the game type of the currently running game.
[0065] It should be noted that the game type can be role-playing game, action game, adventure game, shooting game, etc. In practice, the type of game the user is currently playing can be determined by comparing the game name and specific content.
[0066] Step S122: Adjust the filtering coefficients of the Kalman filter algorithm according to the game type.
[0067] It should be understood that Kalman filtering does not have fixed filter coefficients; its core is the dynamically calculated Kalman gain. This gain is the optimal weight automatically calculated by the algorithm based on prediction and observation uncertainties. By setting the process noise and observation noise during the audio signal acquisition process, we indirectly adjust the behavior of the Kalman gain, thereby controlling the filter's response speed and smoothness.
[0068] The process noise is referred to as prediction noise. In determining the noise included in an audio signal, it can be obtained through both prediction and observation. However, the filtering effect of different acquisition methods on the noise in the audio signal varies. For example, when the proportion of observation noise is large, the filter responds more sensitively to new measurement data, but may be more affected by observation noise; when the proportion of pre-defined noise is large, the filter output is smoother, but there may be a delay in tracking the true state of the system. Therefore, in determining the noise in an audio signal, a balance between the two methods needs to be considered.
[0069] It's important to note that the filter coefficients in the Kalman filter algorithm refer to the optimal weight matrix used to weight the new pose data in the state update equation. The filter coefficients for the user's head movements differ across game genres. For example, in shooting games, the audio signal has relatively few key feature vectors, so the filter coefficient can be set to 0.1. Conversely, in role-playing games, the audio signal has more key feature vectors, so the filter coefficient can be set to 0.3.
[0070] Step S123: Use the Kalman filter algorithm with adjusted filter coefficients to reduce noise in the attitude data.
[0071] It should be understood that, given the filter coefficients of the Kalman filter algorithm, these coefficients can be directly set within the Kalman filter algorithm. This allows the Kalman filter algorithm to filter out noise in the attitude data, thereby preventing noise data from affecting the spatial audio signal.
[0072] It is understandable that the state equation of the Kalman filter is set as X(k)=A*X(k-1)+B*U(k)+W(k), and the observation equation is set as Z(k)=H*X(k)+V(k), where A is the state transition matrix and H is the observation matrix.
[0073] Reference Figure 4 , Figure 4 This is a schematic diagram of the third process of the second embodiment of the spatial audio signal generation method proposed in this invention.
[0074] In this embodiment, the spatial audio signal generation method further includes: Step S21: Assign a unified timestamp to each frame of the multi-dimensional attitude signal and each frame of the audio signal through a timestamp alignment mechanism.
[0075] It should be understood that during data acquisition, a sampling frequency of 96kHz is typically used for audio signals, while a lower sampling frequency is required for multi-dimensional attitude signals, which are usually sampled at 1kHz. During data fusion, it is crucial to ensure that each audio signal corresponds to the multi-dimensional attitude signal at the same time. If the sampling frequencies are different, their timestamps will not align, resulting in time misalignment and making it impossible to obtain an accurate fused feature matrix.
[0076] Therefore, in this embodiment, a timestamp alignment mechanism can also be used to uniformly allocate timestamps to each frame of the multi-dimensional attitude signal and each frame of the audio signal. For example, during the time synchronization phase, the clock module can generate a timestamp every 1ms. The signal frames of the multi-dimensional attitude signal are directly associated with the timestamps, and the audio frames of the audio signal have their timestamps calculated based on the acquisition start time and frame length. Alternatively, a high-precision clock module with an accuracy of ±1ms can be built into the earphone. The clock signal output by this module can be used to set the time of each frame of the audio signal, and the time of each frame of the multi-dimensional attitude signal can also be set according to this clock signal. In this case, the multi-dimensional attitude signal and the audio signal are aligned in time.
[0077] Step S22: Using the signal frequency of the audio signal as a reference, perform linear interpolation on the multi-dimensional attitude signal after timestamp allocation using a linear interpolation algorithm.
[0078] It should be understood that after time alignment of the audio signal and the multi-dimensional attitude signal, there are many blank frames in the multi-dimensional attitude signal because the sampling frequency of the multi-dimensional attitude signal is lower than that of the audio signal. In the process of mapping and converting the multi-dimensional attitude signal, to avoid the problem of unrecognized blank frames leading to time inconsistencies, linear interpolation can also be performed on the multi-dimensional attitude signal in this embodiment.
[0079] It should be noted that the linear interpolation algorithm is an algorithm that interpolates multi-dimensional attitude signals based on the linear relationship between the frames of the multi-dimensional attitude signal. The linear interpolation algorithm can maintain the regularity of the multi-dimensional attitude signal.
[0080] In practical implementation, the frequency of an audio signal can be used as a reference, and a linear interpolation algorithm can be used to linearly interpolate the multi-dimensional attitude signal after timestamp allocation. For example, if the frequency of the audio signal is 96kHz, a linear interpolation algorithm can be used to adjust the frequency of the multi-dimensional attitude signal from 1kHz to 96kHz to ensure time dimension alignment between the audio signal and the multi-dimensional attitude signal.
[0081] Accordingly, step S20 is: Step S20': The audio signal after the timestamp allocation is converted into spectral features using short-time Fourier transform, and the multi-dimensional attitude signal after linear interpolation is mapped into the corresponding multi-dimensional feature vector.
[0082] In practical implementation, the audio signal after timestamp allocation can be converted into spectral features using short-time Fourier transform. For multi-dimensional attitude signals, interference data in the linearly interpolated multi-dimensional attitude signal can be filtered out or the data can be standardized. Then, multiple features of the standardized multi-dimensional attitude signal can be extracted, and each feature can be vectorized to obtain the multi-dimensional feature vector corresponding to the multi-dimensional feature vector.
[0083] Based on the first or second embodiment of the spatial audio signal generation method described above, a third embodiment of the spatial audio signal generation method of the present invention is proposed. (Refer to...) Figure 5 , Figure 5 This is a schematic diagram of the first process of the third embodiment of the spatial audio signal generation method proposed in this invention.
[0084] In this embodiment, the spatial audio signal generation method further includes: Step S41: Obtain the initial sound field rendering model for performing feature rendering.
[0085] It should be noted that the initial sound field rendering model is used to render individual head pose data. This initial sound field rendering model, based on a convolutional model of Head Related Transfer Functions (HRTF), relies on a fixed HRTF and cannot improve rendering quality by learning the audio patterns of different game genres.
[0086] In practice, an initial sound field rendering model can be obtained by directly rendering individual head pose data. Alternatively, a preliminary sound field rendering model can be established through the mapping relationship between head pose data and spatial audio signals.
[0087] Step S42: Obtain the stored audio feature templates for different game types.
[0088] It should be understood that while this initial sound field rendering model can accurately render individual head pose data, it may have some bias when rendering the fused feature matrix because it does not consider the specific game genre. For example, it might use rendering parameters for a shooting game to render the fused feature matrix for a role-playing game.
[0089] Therefore, in this embodiment, it is also necessary to combine audio feature templates of different game types into the rendering process during the initial sound field rendering model data rendering process, so as to obtain more accurate spatial audio signals.
[0090] It should be noted that the audio feature templates are templates representing user audio features across different game genres. These audio feature templates can include those for FPS, MOBA, RPG, etc., such as FPS game footstep sound spectrum templates and MOBA game skill sound effect templates. The audio feature templates can be pre-stored in the headset's memory. Matching different audio feature templates assigns dynamic weights to different game genres, thereby adjusting the rendering process according to the specific game genre.
[0091] Step S43: Set a scene embedding layer within the initial sound field rendering model.
[0092] It should be noted that the scene embedding layer is a structure that embeds audio feature templates into the initial sound field rendering model. This scene embedding layer is a virtual structure and can be implemented through software programming. This scene embedding layer can be placed between the encoder and decoder of the model, thereby using audio feature templates to adjust to the fused feature matrix.
[0093] Step S44: Using the scene embedding layer, embed the target audio feature template that matches the fusion feature matrix into the initial sound field rendering model.
[0094] It should be understood that during the rendering of the fusion feature matrix, the scene embedding layer can embed the target audio feature template corresponding to the current game type into the initial sound field rendering model, and then use the currently running game type to combine with the rendering process of the fusion feature matrix.
[0095] In specific implementation, the scene embedding layer can determine the game type by obtaining the game name or game content, and then determine the target audio feature template corresponding to the game type. The game type matches the fusion feature matrix, that is, the target audio feature template matches the fusion feature matrix; then the obtained target audio template is embedded into the initial sound field rendering model.
[0096] Step S45: Adjust the rendering parameters using the target audio feature template to obtain the sound field rendering model.
[0097] It should be understood that different audio feature templates correspond to different dynamic weights. The target audio feature template corresponds to the dynamic weight of the current game type. The rendering parameters of the initial sound field rendering model can be adjusted according to the dynamic weight to obtain the sound field rendering model.
[0098] The sound field rendering model is one that can accurately render the fusion feature matrix in conjunction with the current game type. The rendering parameters of this sound field rendering model are related to the target audio feature template set for the current game type.
[0099] Furthermore, in this embodiment, step S41 is followed by: Step S46: Adjust the input layer and decoder of the initial sound field rendering model according to the fusion feature matrix.
[0100] It should be noted that the input layer of the initial sound field rendering model is a structure used to receive parameters of different dimensions of the fused feature matrix, and the decoder is a device used to decode the parameters of one dimension to obtain the original information.
[0101] It should be understood that the input to the initial sound field rendering model is typically head pose data, which is not optimized for the dynamic characteristics of game audio, such as the broadband characteristics of sudden explosions and the periodic characteristics of footsteps. However, the fused feature matrix includes not only multi-dimensional pose signals but also audio signals. Therefore, the fused feature matrix contains data in multiple dimensions, especially the spectral characteristics of the audio signal. In order for the feature data within the spectral characteristics to be rendered by the model, the input layer and decoder of the initial sound field rendering model need to be adjusted.
[0102] In the specific adjustment process, the input layer and encoder of the initial sound field rendering model can be adjusted according to the feature dimension of the spectral features specifically included in the fusion feature matrix. For example, the dimension of the data that the input layer can receive and the dimension of the data that the decoder can decode can be increased.
[0103] Step S47: Perform lightweight processing on the encoder in the adjusted initial sound field rendering model to obtain an optimized initial sound field rendering model.
[0104] In addition, a deep learning-based sound field rendering model is used for some initial sound field rendering models. Although this model can effectively improve rendering accuracy by learning a large number of parameters and performing rendering, the parameters of this part of the model are redundant. When running on the embedded device at the headphone end, the frame rate is ≤20FPS, which cannot meet the real-time requirement of ≥60FPS in the game scene, resulting in extremely poor game screen smoothness.
[0105] Therefore, in this embodiment, the initial sound field rendering model also needs to be lightweighted to appropriately reduce its execution parameters and effectively improve the smoothness of the game screen. Considering that the encoder in the initial sound field rendering model has a large number of layers, and each layer contains a multi-head self-attention mechanism and a feedforward neural network, in specific implementation, the encoder within the initial sound field rendering model can be lightweighted, for example, by reducing the number of encoder layers and the number of self-attention mechanisms within each layer. The self-attention mechanism is a mechanism that allows users to focus on the local correlation features of audio signals within the game genre. These local correlation features include the signal characteristics of the audio signal and the correlation features between the audio signal and multi-dimensional pose signals.
[0106] Reference Figure 6 , Figure 6 This is a schematic diagram of the second process of the third embodiment of the spatial audio signal generation method proposed in this invention. Step S46 specifically includes: Step S461: Add an audio feature extraction branch to the input layer of the initial sound field rendering model based on the audio dynamic features in the fusion feature matrix.
[0107] It should be noted that the audio feature extraction branch is used to extract audio dynamic features of different frequencies within the fusion feature matrix. This branch includes a first convolutional kernel, a second convolutional kernel, and a third convolutional kernel. These kernels extract high-frequency, mid-frequency, and low-frequency features from the audio signal, respectively. These high-frequency, mid-frequency, and low-frequency features from the audio signal within the fusion feature matrix are then fused with the corresponding dimension's attitude parameter signals from the multi-dimensional attitude signal. The first, second, and third convolutional kernels use three kernels of different sizes, such as 3×3, 5×5, and 7×7, to convolve the spectral features, extracting high-frequency features such as footsteps, mid-frequency features such as teammates' voices, and low-frequency impact features such as explosions.
[0108] In practical implementation, considering that the dynamic features of audio mainly include low-frequency features, mid-frequency features and high-frequency features in audio, the audio feature extraction branches of the input layer can be increased to three audio feature extraction branches during the process of adding audio feature extraction branches, so that features of different frequencies of audio signals can be extracted.
[0109] Step S462: Adjust the number of decoding layers in the decoder of the initial sound field rendering model according to the audio feature extraction branch to obtain the optimized initial sound field rendering model.
[0110] It should be understood that a single decoder can decode parameters of a frequency feature. In this embodiment, the spectral features in the fused feature matrix include three frequency features: low-frequency features, mid-frequency features, and high-frequency features. Therefore, in this embodiment, it is also necessary to determine the number of decoding layers in the decoder of the initial sound field rendering model. The number of decoding layers is the same as the number of audio feature extraction branches, both corresponding to the low-frequency, mid-frequency, and high-frequency features of the spectral features.
[0111] It should be noted that the optimized initial sound field rendering model can directly receive and render the low-frequency, mid-frequency, and high-frequency features within the fused feature matrix.
[0112] Step S47 includes: Step S471: Configure the number of first self-attention mechanisms and the number of second self-attention mechanisms in each layer of the encoder of the initial sound field rendering model according to the game type.
[0113] It should be understood that, considering the redundancy in data processing within the encoder, lightweighting of the encoder is necessary. Given that different game genres have different spectral characteristics—for example, for shooter games, which require encoding fewer spectral features, the encoder can be significantly lightweighted; while for role-playing games, which require encoding more spectral features, only minor lightweighting of the encoder is possible.
[0114] It should be noted that, in order to focus on the spectral characteristics of the audio signal and the feature correlation between the audio signal and the multi-dimensional attitude signal, a certain number of self-attention mechanisms need to be set in each layer of the encoder to ensure that the frequency correlation of the spectral characteristics and the correlation between the audio signal and the multi-dimensional attitude signal will not fail to be encoded correctly.
[0115] The first self-attention mechanism is used to focus on the frequency correlation of the audio signal, and the second self-attention mechanism is used to focus on the correlation between the audio signal and the multi-dimensional attitude signal.
[0116] In practice, the self-attention mechanism of each layer within the encoder can be configured according to different game types. For example, the self-attention mechanism can be set to 8 heads, with the first self-attention mechanism having 4 heads and the second self-attention mechanism also having 4 heads. Of course, the specific number of the first and second self-attention mechanisms needs to be adjusted according to the game type. For example, for role-playing games, the first self-attention mechanism can be set to 3 heads and the second self-attention mechanism to 5 heads.
[0117] Step S472: Reduce the convolution dimension of the feedforward neural network within the encoder after configuring the self-attention mechanism to obtain the optimized initial sound field rendering model.
[0118] Each layer of the encoder includes a self-attention mechanism and a feedforward neural network. After limiting the number of self-attention mechanisms, the feedforward neural network can be lightweighted to further improve the encoder's lightweightness.
[0119] In practical implementation, the feedforward neural network can be lightweighted by reducing the convolutional dimension. For example, replacing the core convolution of the feedforward neural network with a 1×1 convolution can reduce the number of parameters by 60%, ensuring improved efficiency of the sound field rendering model.
[0120] Reference Figure 7 , Figure 7 This is a schematic diagram of the optimized sound field rendering model of this invention. Figure 7In this model, the input layer has three convolutional kernels, and the lightweight encoder is illustrated using an 8-head self-attention mechanism. The scene embedding layer embeds the target audio feature template corresponding to the current game type into the sound field rendering model; the audio feature template can be pre-stored within the scene embedding layer. The decoder is a three-layer decoder used to decode the low-frequency, mid-frequency, and high-frequency features in the fused feature matrix. The output layer converts the decoded spectral features into a two-channel spatial audio signal using inverse Fourier transform.
[0121] In the specific rendering process, the input layer can receive low-frequency, mid-frequency, and high-frequency features from the fused feature matrix. The encoder uses a first self-attention mechanism to focus on the frequency correlation of the audio signal based on the audio feature spectrum, and a second self-attention mechanism to focus on the correlation between the audio signal and the multi-dimensional pose signal. The low-frequency, mid-frequency, and high-frequency features are encoded through a convolutional neural network with reduced dimensionality. The scene embedding layer can embed the target audio feature template corresponding to the current game type, thereby adjusting the encoded data. The adjusted data is decoded by the corresponding layer of the decoder to obtain the rendered audio signal. Finally, the output layer converts the decoded spectral features into a two-channel spatial audio signal through inverse Fourier transform and outputs it to the headphone speaker unit for sound generation.
[0122] Furthermore, to achieve the above objectives, the present invention also provides a headset, referring to... Figure 8 , Figure 8 This is a schematic diagram of the structure of the headphones proposed in this invention.
[0123] exist Figure 8 The headset includes: a posture sensor 10, an audio input module 20, a clock module 30, a processor 40, a fusion chip 50, and a sound generation module 60; The attitude sensor 10, the audio input module 20, and the clock module 30 are all connected to the fusion chip 50. The fusion chip 50 is also connected to the processor 40, and the processor 40 is also connected to the sound generation module 60. The processor 40 is equipped with a sound field rendering model.
[0124] The headset is used to execute any of the spatial audio signal generation methods described above, and to implement the functions of any of the above embodiments. The specific implementation process can be referred to the spatial audio signal generation method described above, and will not be repeated here.
[0125] The above description is merely an exemplary embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention specification and drawings under the technical concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.
Claims
1. A method for generating spatial audio signals, characterized in that, The method includes: Acquire audio signals within the game genre and multi-dimensional posture signals of the user; The audio signal is converted into spectral features, and the multi-dimensional attitude signal is mapped into a corresponding multi-dimensional feature vector; The multi-dimensional feature vectors are embedded into the spectral features to obtain a fused feature matrix; Spatial audio signals are generated by performing feature rendering on the fused feature matrix using a sound field rendering model.
2. The spatial audio signal generation method as described in claim 1, characterized in that, The step of embedding the multi-dimensional feature vector into the spectral features to obtain a fused feature matrix includes: Get the game type of the currently running game; Determine the correlation weights between the spectral features and the feature vectors of each dimension within the multi-dimensional feature vector based on the game type; The fused feature matrix is obtained by fusing the dimensional feature vectors within the multi-dimensional feature vectors into the spectral feature based on the correlation weights.
3. The spatial audio signal generation method as described in claim 1, characterized in that, The multi-dimensional attitude signal is a three-dimensional attitude parameter signal; the acquisition of the audio signal and the multi-dimensional attitude signal within the game type includes: User attitude data is collected using inertial sensors; The attitude data is denoised using a Kalman filter algorithm; The denoised attitude data is decomposed into three-dimensional attitude parameter signals by quaternion transformation; Obtain the initial audio signal within the game type; The high-frequency components in the initial audio signal are pre-emphasized to obtain the audio signal.
4. The spatial audio signal generation method as described in claim 3, characterized in that, The noise reduction of the attitude data using the Kalman filter algorithm includes: Get the game type of the currently running game; Adjust the filtering coefficients of the Kalman filter algorithm according to the game type; The attitude data is denoised using a Kalman filter algorithm with adjusted filter coefficients.
5. The spatial audio signal generation method as described in claim 1, characterized in that, The spatial audio signal generation method further includes: A unified timestamp is assigned to each frame of the multi-dimensional attitude signal and each frame of the audio signal through a timestamp alignment mechanism. Using the signal frequency of the audio signal as a reference, the multi-dimensional attitude signal after timestamp allocation is linearly interpolated using a linear interpolation algorithm; The step of converting the audio signal into spectral features and mapping the multi-dimensional attitude signal into a corresponding multi-dimensional feature vector includes: The audio signal after timestamp allocation is converted into spectral features using short-time Fourier transform, and the multi-dimensional attitude signal after linear interpolation is mapped into the corresponding multi-dimensional feature vector.
6. The spatial audio signal generation method as described in claim 1, characterized in that, The spatial audio signal generation method further includes: Obtain the initial sound field rendering model for the execution feature rendering; Retrieve audio feature templates for different game genres from the storage; Set a scene embedding layer within the initial sound field rendering model; The scene embedding layer is used to embed the target audio feature template that matches the fusion feature matrix into the initial sound field rendering model; The sound field rendering model is obtained by adjusting the rendering parameters using the target audio feature template.
7. The spatial audio signal generation method as described in claim 6, characterized in that, After obtaining the initial sound field rendering model for performing feature rendering, the process further includes: The input layer and decoder of the initial sound field rendering model are adjusted according to the fusion feature matrix; The encoder in the adjusted initial sound field rendering model is lightweighted to obtain an optimized initial sound field rendering model.
8. The spatial audio signal generation method as described in claim 7, characterized in that, The step of adjusting the input layer and decoder of the initial sound field rendering model according to the fused feature matrix includes: Based on the audio dynamic features within the fusion feature matrix, an audio feature extraction branch is added to the input layer of the initial sound field rendering model; The optimized initial sound field rendering model is obtained by adjusting the number of decoding layers in the decoder of the initial sound field rendering model according to the audio feature extraction branch, wherein the number of decoding layers is the same as the number of audio feature extraction branches.
9. The spatial audio signal generation method as described in claim 7, characterized in that, The step of lightweighting the encoder within the adjusted initial sound field rendering model to obtain an optimized initial sound field rendering model includes: The number of first self-attention mechanisms and the number of second self-attention mechanisms in each layer of the encoder of the initial sound field rendering model are configured according to the game type. The first self-attention mechanism is used to focus on the frequency correlation of the audio signal, and the second self-attention mechanism is used to focus on the correlation between the audio signal and the multi-dimensional pose signal. By reducing the convolution dimension of the feedforward neural network within the encoder after configuring the self-attention mechanism, an optimized initial sound field rendering model is obtained.
10. A type of headset, characterized in that, Used to perform the spatial audio signal generation method according to any one of claims 1 to 9; The headset includes: a posture sensor, an audio input module, a clock module, a processor, a fusion chip, and a sound generation module; The attitude sensor, the audio input module, and the clock module are all connected to the fusion chip, the fusion chip is also connected to the processor, and the processor is also connected to the sound generation module. The processor is equipped with a sound field rendering model.