Game audio processing methods, devices, and headphones

By dividing the game audio signal into frequency bands and extracting transient feature parameters, and dynamically adjusting the suppression threshold, the harshness problem of headphones when processing high-frequency transient signals is solved, achieving a combination of hearing protection and sound fidelity.

CN122124458APending Publication Date: 2026-06-02DONGGUAN DESHENG IND CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGGUAN DESHENG IND CO LTD
Filing Date
2026-01-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing headphones are prone to producing sharp and piercing sounds when processing high-frequency transient signals, which can damage users' hearing and affect the detail and spatial positioning of game sound effects.

Method used

By dividing the game audio signal into frequency bands, extracting transient feature parameters of the high-frequency band, performing multi-dimensional fusion judgment to identify transient high-frequency harsh sounds, and dynamically adjusting the suppression threshold according to the occurrence status in the time dimension to implement dynamic suppression processing, while compensating for the mid-to-high frequency bands.

Benefits of technology

It effectively reduces the risk of hearing stimulation from high-frequency harsh sounds, preserves the details and positioning characteristics of game audio to the maximum extent, and achieves real-time recognition and distortion-free suppression of high-frequency harsh sounds, taking into account both hearing protection and game audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122124458A_ABST
    Figure CN122124458A_ABST
Patent Text Reader

Abstract

This invention relates to a game audio processing method, apparatus, and headset. The method includes: dividing the acquired game audio signal into frequency bands; extracting transient feature parameters to characterize the transient changes in the audio signal within the high-frequency band; performing multi-dimensional fusion judgment based on the transient feature parameters to identify whether there are transient audio events with auditory stimulation risks in the high-frequency band audio signal, and determining the audio signal as transient high-frequency harsh sound when a preset judgment condition is met; determining a dynamic suppression threshold based on the occurrence status information of the transient high-frequency harsh sound in the time dimension; implementing dynamic suppression processing on the target high-frequency band when the loudness of the transient high-frequency harsh sound exceeds the dynamic suppression threshold, and compensating for key positioning sound effects in the mid-high frequency band. Using the above scheme, the risk of high-frequency harsh sound stimulating the user's hearing can be effectively reduced while preserving the detailed information and positioning characteristics of the game audio to the maximum extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a method, apparatus and headphones for processing game audio. Background Technology

[0002] With the continuous development of modern video game technology, audio effects in games have gradually become a key factor in enhancing player immersion and experience. Audio processing plays an important role in game sound effects, background music, and special effects audio, especially in high-intensity and complex scenes, where dynamic adjustment of audio is particularly important.

[0003] For example, first-person shooter (FPS) games have extremely high requirements for the audio detail performance of headphones. In order to achieve accurate sound field positioning and environmental sound effect reproduction, headphones usually adopt an acoustic design that enhances high-frequency response.

[0004] However, this design can produce instantaneous high-intensity, high-frequency signals in certain game scenarios (such as close-range gunshots, explosions, grenade detonations, etc.), resulting in sharp and piercing sound output. Summary of the Invention

[0005] Based on this, embodiments of the present invention provide a game audio processing method, device, and headphones, which can effectively reduce the risk of high-frequency harsh sounds stimulating the user's hearing while preserving the detailed information and positioning characteristics of game audio to the maximum extent.

[0006] In a first aspect, the present invention provides a game audio processing method, comprising:

[0007] The acquired game audio signal is divided into frequency bands, at least low-frequency audio signal, mid-high frequency audio signal and high-frequency audio signal;

[0008] Within the high-frequency audio signal, transient feature parameters are extracted to characterize the transient change characteristics of the audio signal. These transient feature parameters are used to reflect at least the abrupt change characteristics of the high-frequency audio signal in the time and frequency dimensions.

[0009] Based on the transient feature parameters, a multi-dimensional fusion judgment is performed to identify whether there are transient audio events with auditory stimulation risks in the high-frequency audio signal, and when the preset judgment conditions are met, the corresponding audio signal is judged as transient high-frequency harsh sound.

[0010] Based on the occurrence status information of the transient high-frequency harsh sound in the time dimension, the corresponding dynamic suppression threshold is determined;

[0011] When the loudness of the transient high-frequency harsh sound exceeds the dynamic suppression threshold, dynamic suppression processing is performed on the target high-frequency band, and compensation processing is performed on the key positioning sound effects in the mid-high frequency band.

[0012] In one embodiment, extracting transient feature parameters for characterizing the transient change characteristics of the audio signal within the high-frequency audio signal includes:

[0013] Within the high-frequency audio signal, the audio signal is divided into short-time frame segments, discretizing the continuous audio signal into multiple audio frames;

[0014] For each audio frame, the energy change rate and spectral change amplitude of the audio signal are determined based on the comparison relationship between the audio frame and its adjacent audio frames.

[0015] Based on the energy change rate and the spectral change amplitude, audio frames in which the energy and / or spectral change are significantly altered are marked as candidate transient audio frames.

[0016] For the candidate transient audio frame, the corresponding peak duration is determined to obtain the transient feature parameters.

[0017] In one embodiment, determining the energy change rate and spectral change amplitude of the audio signal for each audio frame, based on the comparison relationship between the audio frame and adjacent audio frames, includes:

[0018] For each audio frame, a short-time Fourier transform is performed to calculate the energy distribution of each frequency band and obtain a spectrum.

[0019] Based on the spectrum, the energy increment of each frequency band is extracted and the instantaneous frequency change is calculated;

[0020] The amplitude of the spectral change is calculated based on the spectral changes of each audio frame and its adjacent frames, and the rate of change characteristic of the audio frame is obtained using a sliding window algorithm based on the signal rate of change.

[0021] Calculate the rate of energy change of the audio frame relative to the preceding and following audio frames, and perform a weighted average of the energy change of each frequency band based on a weighting function to obtain the weighted change amplitude of the frequency characteristics;

[0022] By performing multi-channel spectral analysis, the number of frequency abrupt changes and their relative distribution within each frequency band are calculated. Combined with the temporal characteristics of the sound, the energy change rate and spectral change amplitude of the audio signal are determined.

[0023] In one embodiment, the multi-dimensional fusion judgment based on the transient features, when a preset judgment condition is met, determines the corresponding audio signal as a transient high-frequency harsh sound, including:

[0024] Based on the transient feature parameters, an auditory stimulus risk index for transient audio events is calculated, and the auditory stimulus risk index is compared with a preset risk threshold.

[0025] If the auditory stimulus risk index exceeds the risk threshold, obtain the spatial orientation information or sound image localization parameters corresponding to the transient audio event.

[0026] Based on the spatial orientation information or sound image positioning parameters, determine whether the transient audio event conforms to the positioning sound effect feature pattern used to prompt the player's spatial orientation;

[0027] If it is determined that the transient audio event does not conform to the positioning sound effect feature pattern, the transient audio event is judged as a transient high-frequency piercing sound;

[0028] If the location sound effect characteristic pattern is determined to be consistent with the specified pattern, the determination of the harsh sound of the transient audio event is abandoned.

[0029] In one embodiment, determining the corresponding dynamic suppression threshold based on the occurrence status information of the transient high-frequency harsh sound in the time dimension includes:

[0030] The time interval and number of occurrences of the transient high-frequency piercing sound are recorded in multiple consecutive game frames.

[0031] Based on the recorded occurrence time intervals and frequency, a time distribution description is formed to characterize the distribution characteristics of transient high-frequency harsh sounds in the time dimension;

[0032] Based on the time distribution description results, determine whether the transient high-frequency harsh sound in the current audio scene shows a trend of increasing density, frequency, or intermittent change.

[0033] When the trend indicates that the piercing sound events are becoming more frequent or intense, control parameters for adjusting the dynamic suppression threshold are generated.

[0034] The dynamic suppression threshold is adaptively adjusted according to the control parameters to obtain a dynamic suppression threshold that matches the current audio scene.

[0035] In one embodiment, determining whether the transient high-frequency harsh sound exhibits a trend of increasing density, frequency, or intermittent change in the current audio scene based on the time distribution description result includes:

[0036] Based on the time distribution description results, the occurrence frequency of each transient high-frequency piercing sound is calculated, and the frequency change curve is determined;

[0037] The frequency variation curve is smoothed in the time domain by using a sliding window algorithm to eliminate high-frequency noise and extract frequency fluctuation characteristics within a time period.

[0038] The frequency fluctuation characteristics are classified using time-based clustering analysis to identify frequency fluctuation patterns within a time period, including dense, frequent, or intermittent change patterns.

[0039] Based on the frequency fluctuation pattern, its stability in historical audio data is analyzed, and an audio fluctuation characteristic curve is generated by a fitting algorithm.

[0040] Based on the audio fluctuation characteristic curve and the changing trend of the frequency fluctuation pattern, the probability of transient high-frequency harsh sounds appearing in future audio frames is predicted and compared with a preset threshold, and the occurrence trend is output.

[0041] In one embodiment, the dynamic suppression processing of the target high-frequency band includes:

[0042] When a transient high-frequency harsh sound is detected and its loudness exceeds the dynamic suppression threshold, the dynamic suppression control process is triggered.

[0043] Based on the current dynamic suppression threshold, an initial suppression parameter is applied to the high-frequency band corresponding to the transient high-frequency harsh sound.

[0044] After applying the initial suppression parameters, the transient feature parameters in the suppressed audio signal are extracted again;

[0045] Based on the changes in the transient characteristic parameters before and after the dynamic suppression treatment, an effective evaluation result is formed for adjusting the dynamic suppression treatment.

[0046] If the effective evaluation results indicate that the inhibition effect has not reached the preset target, the inhibition parameters are adaptively adjusted and the inhibition is reapplied.

[0047] When the effective evaluation results indicate that the inhibition effect has reached the preset target, the current inhibition parameters are maintained.

[0048] When the transient high-frequency harsh sound is detected to disappear or the risk is reduced, the suppression intensity is gradually released based on the historical change trajectory of the suppression parameter.

[0049] In one embodiment, the compensation processing for key positioning sound effects in the mid-to-high frequency band includes:

[0050] After completing the dynamic suppression process, the energy distribution of each channel in the mid-to-high frequency band is detected.

[0051] Based on the detected channel energy distribution, determine whether dynamic suppression processing causes a shift in the relative energy relationship between key positioning sound effects in the channels;

[0052] The degree of impact of dynamic suppression processing on spatial positioning consistency is determined based on the degree of deviation in the relative energy relationship.

[0053] Based on the degree of influence, a compensation strategy is determined to correct the channel energy relationship offset for key positioning sound effects; wherein, while implementing the compensation strategy, the compensation magnitude is constrained not to exceed the minimum range required to offset the degree of influence;

[0054] After the compensation process is completed, the relative energy relationship between the channels is re-detected, and if it is not restored to the preset spatial positioning consistency range, the compensation process is iteratively executed.

[0055] In a second aspect, the present invention provides a game audio processing device, comprising:

[0056] The segmentation unit is configured to divide the acquired game audio signal into frequency bands, at least dividing it into low-frequency audio signal, mid-high frequency audio signal and high-frequency audio signal;

[0057] The extraction unit is configured to extract transient feature parameters within the high-frequency audio signal to characterize the transient change characteristics of the audio signal, wherein the transient feature parameters are at least used to reflect the abrupt change characteristics of the high-frequency audio signal in the time and frequency dimensions.

[0058] The judgment unit is configured to perform multi-dimensional fusion judgment based on the transient feature parameters to identify whether there is a transient audio event with auditory stimulation risk in the high-frequency audio signal, and to judge the corresponding audio signal as transient high-frequency harsh sound when the preset judgment conditions are met.

[0059] The determining unit is configured to determine the corresponding dynamic suppression threshold based on the occurrence status information of the transient high-frequency harsh sound in the time dimension;

[0060] The processing unit is configured to perform dynamic suppression processing on the target high-frequency band and compensation processing on key positioning sound effects in the mid-high frequency band when the loudness of the transient high-frequency harsh sound exceeds the dynamic suppression threshold.

[0061] Thirdly, the present invention provides an earphone, including the game audio processing device as described in the foregoing embodiments.

[0062] The above approach involves dividing the game audio signal into frequency bands, decoupling the low-frequency, mid-high-frequency, and high-frequency audio signals. This allows for differentiated processing strategies for different spectral components, avoiding distortion caused by uniform suppression of the entire audio signal. Furthermore, for transient audio components in the high-frequency band that are prone to causing auditory stimulation, transient feature parameters are extracted to characterize their abrupt changes in time and frequency dimensions. Through multi-dimensional fusion judgment, real-time and accurate identification of transient high-frequency harsh sounds posing an auditory stimulation risk is achieved. Thus, the suppression threshold is dynamically determined based on the occurrence status of transient high-frequency harsh sounds in the time dimension, allowing the suppression intensity to adaptively adjust with the frequency and duration of the harsh sound. When the loudness of the harsh sound exceeds the dynamic suppression threshold, dynamic suppression is applied only to the target high-frequency band, while compensation is provided for key positioning sound effects in the mid-high frequency band, thereby avoiding the adverse effects of high-frequency suppression on sound effect positioning and spatial perception. Therefore, this solution can effectively reduce the risk of high-frequency harsh sounds stimulating users' hearing while preserving the detailed information and positioning characteristics of game audio to the maximum extent, achieving real-time identification and distortion-free suppression of high-frequency harsh sounds, thus balancing hearing protection and game audio experience. Attached Figure Description

[0063] Figure 1 A flowchart of a game audio processing method provided in an embodiment of the present invention;

[0064] Figure 2 A flowchart for extracting transient feature parameters is provided in an embodiment of the present invention;

[0065] Figure 3 A flowchart for determining a dynamic suppression threshold is provided in an embodiment of the present invention;

[0066] Figure 4 This is a schematic diagram of a game audio processing device provided in an embodiment of the present invention. Detailed Implementation

[0067] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the specific details described below are only a part of the embodiments of the present invention, and the present invention can be implemented in many other embodiments different from those described herein. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0069] As mentioned in the background technology, in first-person shooter games, instantaneous high-intensity high-frequency signals are generated, resulting in sharp and piercing sound output.

[0070] This linear approach lacks targeted dynamic protection mechanisms. For example, some products use fixed EQ filters, which reduce high-frequency intensity but sacrifice game audio details; a few devices with volume limiting functions achieve protection through hard attenuation, resulting in sound fragmentation and distortion, disrupting the game's immersion. Prolonged use of such headphones for gaming exposes users to the risk of irreversible noise-induced hearing damage.

[0071] Therefore, in order to address the contradiction between high-frequency protection and sound fidelity in existing headphones, this invention provides a game audio processing solution that enables real-time identification and distortion-free suppression of high-frequency harsh sounds, thus protecting hearing while preserving game audio details to the maximum extent.

[0072] Specifically, this solution divides the game audio signal into frequency bands, decoupling the low-frequency, mid-high-frequency, and high-frequency audio signals. This allows for differentiated processing strategies for different spectral components, avoiding distortion introduced by uniformly suppressing the entire audio signal. Based on this, transient feature parameters are extracted from transient audio components in the high-frequency band that are prone to causing auditory stimulation. These parameters characterize the abrupt changes in time and frequency dimensions. Through multi-dimensional fusion judgment, real-time and accurate identification of transient high-frequency harsh sounds posing an auditory stimulation risk is achieved. Thus, the suppression threshold is dynamically determined based on the occurrence status of the transient high-frequency harsh sound in the time dimension, allowing the suppression intensity to adaptively adjust according to the frequency and duration of the harsh sound. When the loudness of the harsh sound exceeds the dynamic suppression threshold, dynamic suppression is applied only to the target high-frequency band, while compensation is provided for key positioning sound effects in the mid-high frequency band, thereby avoiding the adverse effects of high-frequency suppression on sound effect positioning and spatial perception. Therefore, this solution can effectively reduce the risk of high-frequency harsh sounds stimulating users' hearing while preserving the detailed information and positioning characteristics of game audio to the maximum extent, achieving real-time identification and distortion-free suppression of high-frequency harsh sounds, thus balancing hearing protection and game audio experience.

[0073] To enable those skilled in the art to better understand and implement the present invention, the following detailed description of the specific solutions, principles, advantages, and effects of the present invention is provided with reference to the accompanying drawings and specific embodiments.

[0074] See Figure 1 The flowchart shown in this embodiment of the invention provides a game audio processing method, as follows: Figure 1 As shown, the following steps can be performed:

[0075] S101, divide the acquired game audio signal into frequency bands, and at least obtain low-frequency audio signal, mid-high frequency audio signal and high-frequency audio signal.

[0076] Specifically, due to the significant differences in the spectral distribution, temporal variation characteristics, and auditory perception of different sound sources in game audio, especially the transient characteristics of sudden energy increases and extremely short durations in high-frequency bands such as explosions, metal collisions, and skill triggers, if the overall suppression is not differentiated, it can easily lead to sound quality degradation, loss of positioning information, and even affect the player's immersion.

[0077] Based on this, this solution first divides the acquired game audio signal into frequency bands, obtaining at least low-frequency, mid-to-high-frequency, and high-frequency audio signals. This decouples the audio components of different frequency bands in the processing path, providing a foundation for subsequent targeted analysis and processing. Frequency band division avoids unnecessary processing of non-risk frequency bands, reducing the overall risk of audio distortion.

[0078] In one example, the low-frequency audio signal is in the range of 20Hz-500Hz, where a low-pass filter is used to preserve the low-frequency vibration of the explosion sound (which is not involved in the detection of harsh sounds).

[0079] The mid-to-high frequency audio signal ranges from 500Hz to 4kHz. Bandpass filtering is used to preserve key positioning sound effects such as footsteps (1kHz-2kHz) and reloading sounds (2kHz-3kHz). A dynamic protection threshold is set for this frequency band (10dB higher than the harsh sound threshold) to avoid false suppression.

[0080] The frequency range of high-frequency audio signals is between 4kHz and 10kHz. Adaptive filtering is used to retain only the instantaneous peak signal (harsh sound characteristics) in this frequency range, while attenuating continuous and stable high-frequency signals (such as ambient wind sounds) to reduce invalid detection.

[0081] S102, In the high-frequency audio signal, extract transient feature parameters to characterize the transient change characteristics of the audio signal. The transient feature parameters are used to reflect the abrupt change characteristics of the high-frequency audio signal in the time and frequency dimensions.

[0082] Specifically, after completing the frequency band division of the game audio signal, this solution further conducts a special analysis on the high-frequency audio signal.

[0083] High-frequency audio signals typically contain a large number of transient components, which manifest as extremely short-duration, steep-rise energy abrupt changes on the time axis, and as high-energy concentration and dramatic spectral variations on the frequency axis. While these transient high-frequency components can enhance the impact of sound effects, when their amplitude or frequency is too high, they can easily cause stimulation or even potential damage to the user's hearing.

[0084] Thus, this scheme extracts transient feature parameters from high-frequency audio signals to characterize their transient changes. These transient feature parameters include at least those reflecting abrupt changes in the time dimension and those reflecting abrupt changes in the frequency dimension, such as transient energy change rate, peak envelope steepness, spectral flux, and transient duration. By extracting these transient feature parameters from high-frequency audio signals, the suddenness and stimulating nature of the audio signal can be characterized from multiple dimensions, providing a quantitative basis for subsequent identification.

[0085] In some embodiments, see Figure 2 The flowchart shown in the embodiment of the present invention provides a method for extracting transient feature parameters, such as... Figure 2 As shown, it includes:

[0086] S201, within the high-frequency audio signal, performs short-time frame segmentation on the audio signal, discretizing the continuous audio signal into multiple audio frames.

[0087] Specifically, when processing the input high-frequency audio signal, parameters such as the sampling rate, analysis window length, and frame shift size of the audio signal can be predetermined, thereby dividing the continuous time-domain audio signal into multiple audio frames that are continuous in time and may overlap or not overlap.

[0088] In practical applications, different combinations of frame segmentation parameters will form a variety of audio frame processing schemes. Each audio frame processing scheme can achieve frame-by-frame processing of the same high-frequency audio signal. The difference lies in the different time resolution and frequency resolution of each audio frame, which results in differences in the ability to characterize transient audio features.

[0089] In this embodiment, the audio frame processing scheme may include: the frame length of the audio frame (e.g., 5ms, 10ms or 20ms); the frame shift size between adjacent audio frames; the window function type (e.g., Hamming window, Hanning window or rectangular window); and the arrangement order of each audio frame on the time axis.

[0090] As a non-limiting example, for a high-frequency audio signal, the following audio frame processing schemes can be formed: Scheme 1 is an audio frame sequence with a frame length of 10ms and a frame shift of 5ms; Scheme 2 is an audio frame sequence with a frame length of 5ms and a frame shift of 2.5ms.

[0091] It should be noted that the above audio frame processing scheme is only an example to illustrate that the same high-frequency audio signal can be discretized by different combinations of short-time frame parameters. This embodiment of the invention does not limit this.

[0092] S202, for each audio frame, based on the comparison relationship between the audio frame and adjacent audio frames, determine the energy change rate and spectral change amplitude of the audio signal.

[0093] Specifically, after completing the frame segmentation of the high-frequency audio signal, the frame energy and spectral characteristics of each audio frame can be calculated separately.

[0094] Frame energy can be obtained by summing or averaging the squares of the amplitudes of the sampling points within the audio frame; spectral characteristics can be characterized by the amplitude spectrum or power spectrum obtained by performing a fast Fourier transform on the audio frame.

[0095] Based on this, by comparing the current audio frame with its adjacent audio frames (such as the previous frame and / or the next frame), the energy change rate and spectral change amplitude between adjacent audio frames can be calculated. The energy change rate characterizes the degree of abrupt change in the audio signal over time, while the spectral change amplitude characterizes the degree of change in the frequency distribution of the audio signal.

[0096] In one embodiment, step S202 may include: performing a short-time Fourier transform on each audio frame to calculate the energy distribution of each frequency band and obtain a spectrum; extracting the energy increment of each frequency band and calculating the instantaneous frequency change based on the spectrum; calculating the amplitude of the spectrum change based on the spectrum change between each audio frame and its adjacent frames, and using a sliding window algorithm based on the signal rate of change to determine the rate of change characteristic of the audio frame; calculating the energy rate of change of the audio frame relative to the preceding and following audio frames, and performing a weighted average of the energy change of each frequency band based on a weighting function to obtain the weighted amplitude of the frequency characteristic; and calculating the number of frequency abrupt changes and their relative distribution within each frequency band through multi-channel spectrum analysis, and determining the energy rate of change and the amplitude of the spectrum change of the audio signal by combining the time-domain characteristics of the sound.

[0097] S203, based on the rate of energy change and the magnitude of spectral change, mark audio frames with significant changes in energy and / or spectrum as candidate transient audio frames.

[0098] Specifically, after obtaining the energy change rate and spectral change amplitude corresponding to each audio frame, they can be compared with a preset threshold.

[0099] When the rate of change of energy and / or the magnitude of change of spectrum of an audio frame exceeds the corresponding threshold, it indicates that there is a significant transient change in the audio frame, and the audio frame can be marked as a candidate transient audio frame.

[0100] After identifying candidate transient audio frames, the transient events can be further aggregated and analyzed by combining adjacent candidate frames on the timeline, thereby determining the start and end positions of the corresponding transient.

[0101] Based on the time interval between the start and end positions of the transient event, the peak duration corresponding to the transient event can be determined as one of the transient characteristic parameters.

[0102] It should be noted that the peak durations corresponding to different candidate transient audio frames may be different, and under different audio content, there may be multiple transient events with the same or similar transient characteristic parameters.

[0103] S204. For candidate transient audio frames, determine the corresponding peak duration to obtain transient feature parameters.

[0104] Specifically, for the obtained candidate transient audio frames, the peak duration is determined to obtain the corresponding transient feature parameters.

[0105] Specifically, energy analysis is performed on candidate transient audio frames in the time or time-frequency domain to determine the location of transient peaks within the audio frame. A transient peak can be the point in time when the audio signal amplitude, energy, or spectral amplitude reaches a local maximum.

[0106] After determining the peak position, the signal amplitude or energy is searched for points before and after the peak position along the time axis where the signal amplitude or energy falls below a preset threshold. The threshold can be determined based on the background noise level, the average signal energy, or a preset scaling factor. When the signal amplitude or energy drops below the threshold, the start and end times of the peak are determined.

[0107] The peak duration is calculated based on the time interval between the start and end times. Peak duration is used to characterize the duration of transient audio events in the time dimension.

[0108] Furthermore, the peak duration is used as a transient feature parameter, or combined with other transient-related feature parameters (such as peak amplitude, rise time, decay time, etc.) to form a set of transient feature parameters for describing candidate transient audio frames.

[0109] Thus, by performing short-time frame segmentation on high-frequency audio signals and calculating the energy change rate and spectral change amplitude based on the comparison relationship between adjacent audio frames, this invention can effectively characterize the abrupt changes in the time and frequency dimensions of audio signals, thereby improving the accuracy of transient audio feature detection. At the same time, by marking audio frames with significant changes in energy and / or spectrum as candidate transient audio frames and further determining their corresponding peak duration as transient feature parameters, the invalid processing of non-transient audio frames is reduced, the computational complexity is lowered, and the processing efficiency is improved, resulting in better stability and reliability of the extracted transient features.

[0110] S103 performs multi-dimensional fusion judgment based on transient feature parameters to identify whether there are transient audio events with auditory stimulation risks in high-frequency audio signals, and when the preset judgment conditions are met, the corresponding audio signal is judged as transient high-frequency harsh sound.

[0111] Specifically, after obtaining transient feature parameters, a multi-dimensional fusion judgment mechanism is introduced to improve the accuracy and robustness of transient high-frequency harsh sound recognition.

[0112] Specifically, this scheme performs multi-dimensional fusion judgment based on transient feature parameters. By comprehensively analyzing the synergistic relationship of multiple transient features in the time domain, frequency domain, and energy change level, it determines whether there are transient audio events with auditory stimulation risks in high-frequency audio signals. When the fusion judgment result meets the preset judgment conditions, the corresponding audio signal is judged as transient high-frequency harsh sound.

[0113] By using a multi-dimensional fusion judgment method, normal high-frequency detail sound effects can be effectively distinguished from abnormal transient high-frequency sounds with potential stimulation risks, thereby avoiding false suppression of normal sound effects and improving the overall intelligence level of audio processing.

[0114] In one embodiment, step S103 may include: calculating an auditory stimulus risk index for a transient audio event based on transient feature parameters, and comparing the auditory stimulus risk index with a preset risk threshold; if the auditory stimulus risk index exceeds the risk threshold, obtaining spatial orientation information or sound image positioning parameters corresponding to the transient audio event; determining whether the transient audio event conforms to a positioning sound effect feature pattern used to prompt the player's spatial orientation based on the spatial orientation information or sound image positioning parameters; if it is determined that the transient audio event does not conform to the positioning sound effect feature pattern, classifying the transient audio event as a transient high-frequency piercing sound; if it is determined that the transient audio event conforms to the positioning sound effect feature pattern, abandoning the piercing sound judgment of the transient audio event.

[0115] Specifically, based on the extracted transient feature parameters, transient audio events are analyzed to calculate corresponding auditory stimulus risk indices. These indices characterize the potential discomfort or risk that the transient audio event may cause to the auditory system in terms of spectral energy distribution, transient rise characteristics, and high-frequency stimulation intensity. Subsequently, the auditory stimulus risk indices are compared with pre-set risk thresholds.

[0116] When the auditory stimulus risk index does not exceed the risk threshold, the transient audio event is considered to have a low auditory stimulus risk and no further processing is performed; when the auditory stimulus risk index exceeds the risk threshold, a further discrimination process is triggered to distinguish whether the transient audio event belongs to a functionally significant location sound effect.

[0117] When the auditory stimulus risk index exceeds the risk threshold, the spatial orientation information or acoustic image localization parameters corresponding to the transient audio event are obtained. The spatial orientation information or acoustic image localization parameters may include, but are not limited to, the azimuth angle of the sound source, the energy difference between the left and right channels, the phase difference, the time delay difference, or the position vector in the three-dimensional sound field, which are used to reflect the perceived spatial location characteristics of the transient audio event.

[0118] Based on spatial orientation information or sound image positioning parameters, it is determined whether transient audio events conform to the positioning sound effect characteristic pattern used to prompt players' spatial orientation. The positioning sound effect characteristic pattern is a predefined pattern rule used to describe the typical characteristics of functional sound effects used to guide or prompt players of direction, target location, or environmental changes in game or interactive scenarios in terms of spatial positioning properties.

[0119] If a transient audio event is determined not to conform to the location sound effect characteristic pattern, it is considered that the transient audio event does not have a clear spatial cue function and its auditory stimulus risk index has exceeded the risk threshold. Therefore, the transient audio event is judged as a transient high-frequency piercing sound.

[0120] Conversely, if a transient audio event is determined to conform to the characteristic pattern of location sound effects, even if its auditory stimulus risk index exceeds the risk threshold, the transient audio event is still considered to be a location sound effect with functional significance, used to convey spatial orientation information to the player. Therefore, the harsh sound judgment of the transient audio event is abandoned to avoid misjudging normal location prompt sound effects.

[0121] In one embodiment, typical harsh sound frequencies within the 4kHz-10kHz range in the database can be matched (e.g., gunshot characteristic frequency of 6.2kHz, grenade explosion characteristic frequency of 8.5kHz); the rising edge slope of the detected signal can be detected (the rising edge slope of the harsh sound is ≥1.5V / ms, significantly higher than the 0.5V / ms of footsteps); and the peak duration of the detected signal can be detected (the peak duration of the harsh sound is ≤100ms to avoid misjudging continuous high-frequency sound effects).

[0122] Furthermore, when two or more of the three conditions are met, the signal is determined to be a harsh sound signal that needs to be suppressed.

[0123] S104. Determine the corresponding dynamic suppression threshold based on the occurrence status information of transient high-frequency harsh sounds in the time dimension.

[0124] Specifically, after identifying transient high-frequency harsh sounds, this solution further considers their occurrence status information in the time dimension. Among them, different transient high-frequency harsh sounds have significant differences in occurrence frequency, duration and density. If a fixed suppression threshold is used, it is easy to fail to suppress them enough in transient dense scenes or to oversuppress them in occasional scenes.

[0125] Therefore, this solution dynamically determines the corresponding suppression threshold based on the occurrence status information of transient high-frequency piercing sounds in the time dimension. The occurrence status information includes at least the occurrence frequency, adjacent occurrence interval, and duration of the transient high-frequency piercing sounds. By incorporating the above time dimension information into the threshold calculation process, the suppression threshold can adaptively adjust according to the actual occurrence of transient piercing sounds, thereby achieving more reasonable suppression intensity control in different game scenarios.

[0126] In some embodiments, see Figure 3 The flowchart shown in the embodiment of the present invention provides a method for determining a dynamic suppression threshold, as follows: Figure 3 As shown, it includes:

[0127] S301 records the time interval and number of occurrences of transient high-frequency piercing sounds in multiple consecutive game frames.

[0128] Specifically, in order to more precisely depict the changing patterns of transient high-frequency piercing sounds in the time dimension, timestamps are recorded for each detected transient high-frequency piercing sound event in the audio signals corresponding to multiple consecutive game frames. Using a sliding time window consisting of a preset number of frames or a preset duration as the statistical unit, the time interval between adjacent piercing sound events and the number of times they occur within the sliding time window are counted respectively.

[0129] S302, based on the recorded occurrence time intervals and frequency, forms a time distribution description result to characterize the distribution characteristics of transient high-frequency harsh sounds in the time dimension.

[0130] Specifically, based on the occurrence time interval and frequency, a time distribution description is generated to characterize the temporal distribution properties of transient high-frequency piercing sounds. This description includes not only the average occurrence time interval and frequency per unit time of the piercing sound events, but also an index of the dispersion of the time intervals to reflect the concentration and stability of the piercing sound events along the time axis. In this way, the time distribution description can simultaneously reflect both the frequency of piercing sound events and the regularity of their temporal distribution.

[0131] S303, based on the time distribution description results, determine whether the transient high-frequency harsh sound shows a trend of increasing density, frequency, or intermittent change in the current audio scene.

[0132] Specifically, based on the time distribution description results, the occurrence trend of transient high-frequency harsh sounds in the current audio scene is determined.

[0133] Specifically, when the average time interval between occurrences of a piercing sound event shows a continuous decreasing trend over multiple consecutive sliding time windows, and the number of occurrences per unit time exceeds a preset frequency threshold, it is determined that the transient high-frequency piercing sound is showing a trend of increasing density or frequency. When the time interval between occurrences of piercing sound events varies significantly but the number of occurrences does not reach the frequency threshold, it is determined that the transient high-frequency piercing sound is showing a trend of intermittent change. By introducing a trend judgment mechanism based on continuous time windows, misjudgments caused by a single abnormal event can be avoided.

[0134] In one example, step S303 includes: calculating the frequency of each transient high-frequency harsh sound based on the time distribution description results, and determining the frequency change curve; smoothing the frequency change curve in the time domain using a sliding window algorithm to eliminate high-frequency noise and extracting frequency fluctuation features within the time period; classifying the frequency fluctuation features using time-based clustering analysis to identify frequency fluctuation patterns within the time period, including dense, frequent, or intermittent change patterns; analyzing the stability of the frequency fluctuation patterns in historical audio data and generating an audio fluctuation feature curve using a fitting algorithm; predicting the probability of transient high-frequency harsh sounds appearing in future audio frames based on the audio fluctuation feature curve and the changing trend of the frequency fluctuation patterns, comparing it with a preset threshold, and outputting the occurrence trend.

[0135] Specifically, based on the time distribution description results, the frequency of occurrence of each transient high-frequency piercing sound within a preset time unit is statistically analyzed to calculate the frequency of occurrence of the transient high-frequency piercing sound. The time unit can be an audio frame, a millisecond-level time slice, or a second-level time window, which can be set according to the application scenario.

[0136] Furthermore, by plotting the time axis on the horizontal axis and the frequency of transient high-frequency piercing sounds on the vertical axis, the calculated frequencies within each time unit are continuously mapped to form a frequency variation curve for the transient high-frequency piercing sounds. This frequency variation curve characterizes the temporal evolution of the transient high-frequency piercing sounds in the audio signal and reflects the trend of its intensity change over time.

[0137] Because real-world audio signals may contain environmental noise, detection errors, or random interference, frequency variation curves often contain high-frequency jitter components. To reduce the impact of these unstable factors on subsequent analysis, this invention introduces a sliding window algorithm to smooth the frequency variation curves in the time domain.

[0138] A sliding time window with a preset window width is set on the frequency variation curve and slides along the time axis according to a preset step size. Statistical operations are performed on the frequency values ​​within each sliding window, including but not limited to mean calculation, weighted average, or median filtering. Through the above processing, high-frequency noise components in the frequency variation curve are effectively eliminated, resulting in a smoothed frequency variation.

[0139] Based on this, frequency fluctuation characteristics within a time period are extracted from the smoothed frequency change curve. These frequency fluctuation characteristics include, but are not limited to, the average frequency, the amplitude of frequency change, the fluctuation period, and the rate of rise or fall, which are used to quantitatively describe the change state of transient high-frequency harsh sound within a specified time range.

[0140] To classify the extracted frequency fluctuation features, this invention further employs a time-based clustering analysis method. By comparing the similarity between frequency fluctuation features within different time periods, frequency fluctuation features with similar trends are grouped into the same category, thereby identifying the frequency fluctuation patterns of transient high-frequency harsh sounds within a time period.

[0141] Frequency fluctuation patterns include at least the following types: Intensive pattern: characterized by the concentrated occurrence of transient high-frequency harsh sounds within a short period, with its frequency change curve exhibiting a clear peak concentration or sudden increase; Frequent pattern: characterized by the continuous occurrence of transient high-frequency harsh sounds at a relatively high and stable frequency over a longer period; Intermittent change pattern: characterized by the periodic or irregular intermittent occurrence of transient high-frequency harsh sounds on the time axis, with its frequency change curve showing obvious fluctuation intervals. Identifying frequency fluctuation patterns can effectively characterize the temporal behavior of transient high-frequency harsh sounds.

[0142] After obtaining the frequency fluctuation pattern, this invention further analyzes the stability of the frequency fluctuation pattern by combining it with historical audio data. Specifically, the currently identified frequency fluctuation pattern is compared with the frequency fluctuation pattern of the corresponding time period in the historical audio data to evaluate its recurrence and persistence characteristics at different time scales, in order to determine whether the frequency fluctuation pattern has a stable evolutionary pattern.

[0143] Based on the stability analysis results above, a fitting algorithm is used to model the historical frequency fluctuation characteristics and generate audio fluctuation characteristic curves. The fitting algorithm may include linear fitting, polynomial fitting, exponential fitting, or time-series-based prediction models to functionally describe the frequency fluctuation trend over time, thereby obtaining audio fluctuation characteristic curves that reflect the long-term variation characteristics of transient high-frequency harsh sounds.

[0144] Furthermore, by combining the audio fluctuation characteristic curve and the changing trends of the identified frequency fluctuation patterns, the probability of transient high-frequency harsh sounds occurring in future audio frames is predicted. The prediction process is based on the current frequency change state, historical stability analysis results, and the evolution direction of the frequency fluctuation patterns, calculating the likelihood of transient high-frequency harsh sounds occurring in subsequent time periods.

[0145] Subsequently, the predicted occurrence probability is compared with a preset threshold. When the occurrence probability is greater than or equal to the preset threshold, the occurrence trend information of the transient high-frequency harsh sound is output. The trend information is used to indicate the possible enhancement, continuation, or deterioration trend of the transient high-frequency harsh sound in future audio signals. When the occurrence probability is lower than the preset threshold, it is determined that the occurrence trend of the transient high-frequency harsh sound is within a controllable range.

[0146] S304 generates control parameters for dynamically adjusting the suppression threshold when a trend indicates that the piercing sound events are becoming more frequent or concentrated.

[0147] Specifically, when the trend analysis indicates that the piercing sound events are becoming more frequent or concentrated, control parameters for dynamically adjusting the suppression threshold are generated based on the time distribution description. These control parameters may include the adjustment direction, magnitude, or rate of the suppression threshold. The adjustment magnitude is positively correlated with the frequency of piercing sounds and the degree of shortening of the time interval, thus resulting in higher suppression intensity as the piercing sounds become more concentrated.

[0148] S305 adaptively adjusts the dynamic suppression threshold according to the control parameters to obtain a dynamic suppression threshold that matches the current audio scene.

[0149] Specifically, the dynamic suppression threshold is adaptively adjusted according to the control parameters, so that the dynamic suppression threshold gradually decreases when the harsh sound is dense or frequent, so as to enhance the suppression effect on transient high-frequency harsh sound; when the harsh sound shows intermittent changes or the frequency decreases, the dynamic suppression threshold is maintained or adjusted back to avoid excessive suppression of normal high-frequency sound effects or speech details.

[0150] In one example, in a typical combat scenario: the frequency of a piercing sound is determined at an interval greater than 500ms, with a suppression threshold of 99dB, resulting in slight suppression; in an intense combat scenario: the frequency of a piercing sound is determined at an interval less than 100ms, with a suppression threshold of 95dB, resulting in severe suppression; and in a stealthy, infiltrative scenario: the frequency of a piercing sound is determined at an interval greater than 5s, with a suppression threshold of 102dB, maximizing the preservation of detail.

[0151] The above scheme records the time intervals and frequency of transient high-frequency piercing sounds across multiple game frames, and uses this data to create a temporal distribution description of the piercing sounds, thus accurately determining the trend of piercing sounds in the current audio scene. When piercing sounds show a trend of increasing density or frequency, corresponding control parameters are generated, and the dynamic suppression threshold is adaptively adjusted. This allows the suppression intensity to dynamically match the changes in the temporal distribution characteristics of the piercing sounds, effectively enhancing the suppression effect in scenarios with high piercing sound incidence, while avoiding over-suppression when piercing sounds are intermittent or decreasing. This balances piercing sound suppression effectiveness with maintaining audio quality, ultimately improving overall auditory comfort.

[0152] S105: When the loudness of transient high-frequency harsh sound exceeds the dynamic suppression threshold, dynamic suppression processing is applied to the target high-frequency band, and compensation processing is applied to the key positioning sound effects in the mid-high frequency band.

[0153] Specifically, after determining the dynamic suppression threshold, the high-frequency audio signal can be dynamically processed based on this threshold.

[0154] Specifically, when the loudness of transient high-frequency harsh sound exceeds the dynamic suppression threshold, dynamic suppression processing is applied to the target high-frequency band to reduce the loudness level of transient high-frequency harsh sound, thereby reducing the stimulation to the user's hearing.

[0155] Meanwhile, considering that mid-to-high frequency audio signals usually carry key positioning sound information in games, such as the direction of footsteps, the location of skill releases, and environmental reflections, suppressing only the high frequency band may cause an energy imbalance in auditory perception, thereby affecting the accuracy of spatial positioning.

[0156] Based on this, this solution dynamically suppresses the target high-frequency band while compensating for key positioning sound effects in the mid-to-high frequency band. By appropriately compensating for the mid-to-high frequency band, it is possible to suppress harsh high-frequency stimuli while maintaining or even enhancing the spatial positioning and layering of game audio, thereby ensuring the player's immersive experience and game performance while protecting auditory health.

[0157] In some embodiments, implementing dynamic suppression processing on a target high-frequency band may include: triggering a dynamic suppression control process when a transient high-frequency harsh sound is detected and its loudness exceeds a dynamic suppression threshold; applying initial suppression parameters to the high-frequency band corresponding to the transient high-frequency harsh sound according to the current dynamic suppression threshold; after applying the initial suppression parameters, re-extracting transient feature parameters from the suppressed audio signal; forming an effective evaluation result for adjusting the dynamic suppression processing based on the changes in transient feature parameters before and after the dynamic suppression processing; adaptively adjusting the suppression parameters and reapplying suppression when the effective evaluation result indicates that the suppression effect has not reached the preset target; maintaining the current suppression parameters when the effective evaluation result indicates that the suppression effect has reached the preset target; and gradually releasing the suppression intensity based on the historical change trajectory of the suppression parameters when the transient high-frequency harsh sound is detected to have disappeared or the risk has decreased.

[0158] Specifically, when a transient high-frequency harsh sound is detected in the audio signal and its loudness exceeds the dynamic suppression threshold, the dynamic suppression control process is triggered. The dynamic suppression threshold can be adaptively set based on the overall audio energy level, environmental noise characteristics, or user perception model to distinguish between high-risk transient high-frequency harsh sounds that need to be suppressed and acceptable normal high-frequency components. For example, a pre-set database of high-frequency harsh sound features from games can be used to determine the current dynamic suppression threshold.

[0159] After the dynamic suppression control process is triggered, initial suppression parameters are applied to the high-frequency band corresponding to the transient high-frequency harsh sound, based on the current dynamic suppression threshold. The initial suppression parameters include, but are not limited to, suppression gain, suppression bandwidth, duration of action, and initial response speed, which are used to initially reduce the target high-frequency harsh sound without significantly damaging the overall audio listening experience.

[0160] After applying the initial suppression parameters, the suppressed audio signal is re-analyzed to extract the suppressed transient characteristic parameters. These transient characteristic parameters may include transient energy peak value, rise time slope, high-frequency energy proportion, and duration, and are used to characterize the actual effect of the suppression process on transient high-frequency harshness.

[0161] Furthermore, the transient characteristic parameters before and after dynamic suppression processing are compared and analyzed, and effective evaluation results are generated based on their changes to adjust the dynamic suppression processing. These effective evaluation results characterize the degree to which the current suppression parameters attenuate transient high-frequency harshness and their impact on the overall characteristics of the audio signal.

[0162] When the effective evaluation results indicate that the current suppression effect has not reached the preset target, the suppression parameters are adaptively adjusted based on the evaluation results, and the suppression process is reapplied. The adaptive adjustment process can use an iterative approach to gradually optimize the suppression intensity and range of action to avoid excessive suppression that could degrade sound quality.

[0163] When the effective evaluation results show that the suppression effect has reached the preset target, the current suppression parameters are kept unchanged, and the changes in transient high-frequency harsh sounds in the audio signal are continuously monitored.

[0164] Furthermore, when the transient high-frequency harshness is detected to disappear or its risk level is reduced to a preset safe range, the suppression intensity is gradually released based on the historical change trajectory of the suppression parameters. By gradually restoring high-frequency energy, new auditory discomfort caused by abrupt changes in suppression parameters is avoided, thereby ensuring the continuity and naturalness of the audio signal during dynamic changes.

[0165] For example, when the loudness exceeds a threshold, within a preset duration (e.g., 5ms), the EQ parameter adjustment value is determined based on the game's high-frequency harsh sound feature database, and then the EQ equalization curve is modified. Specifically, for pulsed harsh sounds (such as gunshots): an exponential curvature adjustment is used to quickly suppress the peak and then slowly recover, avoiding sound gaps; for continuous harsh sounds (such as explosion aftershocks): a composite curvature of linear and logarithmic curvature is used to smoothly suppress high-frequency energy.

[0166] Accordingly, compensation processing is performed on key positioning sound effects in the mid-to-high frequency band, including: after completing dynamic suppression processing, detecting the energy distribution of each channel in the mid-to-high frequency band; based on the detected channel energy distribution, determining whether the dynamic suppression processing has caused a shift in the relative energy relationship between the key positioning sound effects and the channels; determining the degree of impact of the dynamic suppression processing on spatial positioning consistency based on the degree of shift in the relative energy relationship; determining a compensation strategy to correct the shift in the channel energy relationship of the key positioning sound effects based on the degree of impact; wherein, while implementing the compensation strategy, the compensation amplitude is constrained not to exceed the minimum range required to offset the degree of impact; after the compensation processing is completed, the relative energy relationship between the channels is re-detected, and if it has not recovered to the preset spatial positioning consistency range, the compensation processing is iteratively executed.

[0167] Specifically, after completing the dynamic suppression process, the energy value of each channel in the target frequency band is obtained based on short-time energy, frequency band integral energy, or other equivalent energy characterization methods.

[0168] Based on the detected channel energy distribution, it is determined whether dynamic suppression processing has caused a shift in the relative energy relationship between key localization effects in each channel. This determination can be achieved by comparing the current energy ratio or difference between channels with the reference energy relationship before dynamic suppression processing.

[0169] When a shift in relative energy relationship is detected, the degree of impact of dynamic suppression processing on spatial positioning consistency is determined based on the extent of the shift. This degree of impact characterizes how much the current channel energy relationship deviates from the preset spatial positioning consistency range.

[0170] Subsequently, based on the degree of impact, a compensation strategy was determined to correct the channel energy relationship shift for key positioning sound effects. The compensation strategy includes adjusting the gain of at least one channel to restore the relative energy relationship between the channels.

[0171] Among these measures, while implementing the compensation strategy, the compensation range is constrained to ensure that it does not exceed the minimum range required to offset the impact, so as to avoid overcompensation introducing new audio-visual distortion or sound quality degradation.

[0172] After the compensation process is completed, the relative energy relationship between each channel is re-detected, and it is determined whether it has been restored to the preset spatial positioning consistency range. If it has not been restored to the preset range, the compensation process is iteratively executed based on the new detection results until the relative energy relationship between the channels meets the requirements of spatial positioning consistency or the preset iteration termination condition is reached.

[0173] For example, after dynamic suppression processing, the energy distribution of key positioning sound effects in each channel within the mid-to-high frequency band is detected, and the detected channel energy distribution is compared with the reference energy distribution before dynamic suppression processing. By analyzing the changes in the energy proportion of each channel in the key frequency band, it is determined whether the dynamic suppression processing has caused a change in the relative energy relationship between channels; the degree of change in the relative energy relationship is used to characterize the degree of deviation of the spatial directivity of the key positioning sound effects relative to the reference state.

[0174] Based on the degree of offset, the impact of dynamic suppression processing on spatial positioning consistency is divided into different levels. The impact level is used to reflect whether the stability of key positioning sound effects in spatial perception is affected, and the severity of the impact.

[0175] Thus, after determining the impact of dynamic suppression processing on spatial positioning consistency, a compensation strategy is adaptively determined based on the degree of impact to correct the energy relationship of key positioning sound effects across channels. The compensation strategy includes: appropriately increasing the energy proportion of channels with relative energy lower than the reference state in the key frequency band; and appropriately decreasing the energy proportion of channels with relative energy higher than the reference state in the key frequency band, thereby bringing the relative energy relationship between channels closer to the reference state. The intensity of the compensation corresponds to the degree of impact; a smaller energy correction is used when the impact is small, and a relatively larger energy correction is used when the impact is large.

[0176] In one example, while suppressing harsh sounds, amplitude compensation (≤3dB) is applied to key localization sound effects in the mid-to-high frequency range (e.g., 500Hz-4kHz) to counteract the slight impact of the suppression algorithm on adjacent frequency bands, ensuring that the sound field localization accuracy is not compromised.

[0177] The game audio processing method has been described in detail above through some embodiments. In order to enable those skilled in the art to better understand and implement it, the corresponding device is also described in detail below through some embodiments.

[0178] See Figure 4 The diagram shown is a structural schematic of a game audio processing device provided in an embodiment of the present invention. Figure 4 As shown, the game audio processing device 400 may include:

[0179] The segmentation unit 410 is configured to perform frequency band segmentation on the acquired game audio signal, and to segment at least the low-frequency audio signal, the mid-high frequency audio signal and the high-frequency audio signal.

[0180] Extraction unit 420 is configured to extract transient feature parameters in high-frequency audio signals to characterize the transient change characteristics of audio signals. The transient feature parameters are used to reflect at least the abrupt change characteristics of high-frequency audio signals in the time and frequency dimensions.

[0181] The judgment unit 430 is configured to perform multi-dimensional fusion judgment based on transient feature parameters to identify whether there are transient audio events with auditory stimulation risks in high-frequency audio signals, and to judge the corresponding audio signal as transient high-frequency harsh sound when the preset judgment conditions are met.

[0182] The determining unit 440 is configured to determine the corresponding dynamic suppression threshold based on the occurrence status information of transient high-frequency harsh sound in the time dimension;

[0183] The processing unit 450 is configured to perform dynamic suppression processing on the target high-frequency band when the loudness of the transient high-frequency harsh sound exceeds the dynamic suppression threshold, and to perform compensation processing on key positioning sound effects in the mid-high frequency band.

[0184] For further details regarding the division unit 410, extraction unit 420, judgment unit 430, determination unit 440, and processing unit 450, please refer to the aforementioned examples.

[0185] It is understandable that the above division of units is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the above units can be implemented by the processor calling software.

[0186] The present invention also provides an earphone that includes the game audio processing device as described in the foregoing embodiments.

[0187] In one embodiment, the headphones may include:

[0188] Signal acquisition module: Employs a 24-bit or higher high-precision ADC (analog-to-digital converter) to sample the analog electrical signal input from the headphones, with a sampling frequency of no less than 48kHz, ensuring complete capture of high-frequency signals.

[0189] DSP processing unit: It adopts a high-performance digital signal processor, which can complete parameter updates and signal processing within 5ms, and the core frequency is not less than 300MHz, providing real-time guarantee for operation.

[0190] Signal output module: Includes digital-to-analog conversion and power amplification circuits, receives the signal processed by the DSP and drives the headphones to produce sound.

[0191] Storage module: Dual database structure, including: FPS game high-frequency piercing sound feature database (including multi-dimensional features such as frequency, peak duration, and rise slope) and game sound effect protection database (including frequency range and amplitude threshold of key sound effects such as footsteps and reloading sounds).

[0192] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0193] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications, substitutions, and improvements without departing from the concept of the present invention, and these should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of this invention should be determined by the claims.

Claims

1. A method for processing game audio, characterized in that, include: The acquired game audio signal is divided into frequency bands, at least low-frequency audio signal, mid-high frequency audio signal and high-frequency audio signal; Within the high-frequency audio signal, transient feature parameters are extracted to characterize the transient change characteristics of the audio signal. These transient feature parameters are used to reflect at least the abrupt change characteristics of the high-frequency audio signal in the time and frequency dimensions. Based on the transient feature parameters, a multi-dimensional fusion judgment is performed to identify whether there are transient audio events with auditory stimulation risks in the high-frequency audio signal, and when the preset judgment conditions are met, the corresponding audio signal is judged as transient high-frequency harsh sound. Based on the occurrence status information of the transient high-frequency harsh sound in the time dimension, the corresponding dynamic suppression threshold is determined; When the loudness of the transient high-frequency harsh sound exceeds the dynamic suppression threshold, dynamic suppression processing is applied to the target high-frequency band, and compensation processing is applied to the key positioning sound effects in the mid-high frequency band (in order to reduce the impact of suppression processing on the sound field positioning accuracy).

2. The game audio processing method according to claim 1, characterized in that, The step of extracting transient feature parameters to characterize the transient change characteristics of the audio signal within the high-frequency band includes: Within the high-frequency audio signal, the audio signal is divided into short-time frame segments, discretizing the continuous audio signal into multiple audio frames; For each audio frame, the energy change rate and spectral change amplitude of the audio signal are determined based on the comparison relationship between the audio frame and its adjacent audio frames. Based on the energy change rate and the spectral change amplitude, audio frames in which the energy and / or spectral change are significantly altered are marked as candidate transient audio frames. For the candidate transient audio frame, the corresponding peak duration is determined to obtain the transient feature parameters.

3. The game audio processing method according to claim 2, characterized in that, For each audio frame, determining the energy change rate and spectral change amplitude of the audio signal based on the comparison relationship between that audio frame and adjacent audio frames includes: For each audio frame, a short-time Fourier transform is performed to calculate the energy distribution of each frequency band and obtain a spectrum. Based on the spectrum, the energy increment of each frequency band is extracted and the instantaneous frequency change is calculated; The amplitude of the spectral change is calculated based on the spectral changes of each audio frame and its adjacent frames, and the rate of change characteristic of the audio frame is obtained using a sliding window algorithm based on the signal rate of change. Calculate the rate of energy change of the audio frame relative to the preceding and following audio frames, and perform a weighted average of the energy change of each frequency band based on a weighting function to obtain the weighted change amplitude of the frequency characteristics; By performing multi-channel spectral analysis, the number of frequency abrupt changes and their relative distribution within each frequency band are calculated. Combined with the temporal characteristics of the sound, the energy change rate and spectral change amplitude of the audio signal are determined.

4. The game audio processing method according to claim 1 or 2, characterized in that, The multi-dimensional fusion judgment based on the transient features, when a preset judgment condition is met, determines the corresponding audio signal as a transient high-frequency harsh sound, including: Based on the transient feature parameters, an auditory stimulus risk index for transient audio events is calculated, and the auditory stimulus risk index is compared with a preset risk threshold. If the auditory stimulus risk index exceeds the risk threshold, obtain the spatial orientation information or sound image localization parameters corresponding to the transient audio event. Based on the spatial orientation information or sound image positioning parameters, determine whether the transient audio event conforms to the positioning sound effect feature pattern used to prompt the player's spatial orientation; If it is determined that the transient audio event does not conform to the positioning sound effect feature pattern, the transient audio event is judged as a transient high-frequency piercing sound; If the location sound effect characteristic pattern is determined to be consistent with the specified pattern, the determination of the harsh sound of the transient audio event is abandoned.

5. The game audio processing method according to claim 1, characterized in that, The step of determining the corresponding dynamic suppression threshold based on the occurrence status information of the transient high-frequency harsh sound in the time dimension includes: The time interval and number of occurrences of the transient high-frequency piercing sound are recorded in multiple consecutive game frames. Based on the recorded occurrence time intervals and frequency, a time distribution description is formed to characterize the distribution characteristics of transient high-frequency harsh sounds in the time dimension; Based on the time distribution description results, determine whether the transient high-frequency harsh sound in the current audio scene shows a trend of increasing density, frequency, or intermittent change. When the trend indicates that the piercing sound events are becoming more frequent or intense, control parameters for adjusting the dynamic suppression threshold are generated. The dynamic suppression threshold is adaptively adjusted according to the control parameters to obtain a dynamic suppression threshold that matches the current audio scene.

6. The game audio processing method according to claim 5, characterized in that, The step of determining whether the transient high-frequency harsh sound exhibits a trend of increasing density, frequency, or intermittent change in the current audio scene based on the time distribution description results includes: Based on the time distribution description results, the occurrence frequency of each transient high-frequency piercing sound is calculated, and the frequency change curve is determined; The frequency variation curve is smoothed in the time domain by using a sliding window algorithm to eliminate high-frequency noise and extract frequency fluctuation characteristics within a time period. The frequency fluctuation characteristics are classified using time-based clustering analysis to identify frequency fluctuation patterns within a time period, including dense, frequent, or intermittent change patterns. Based on the frequency fluctuation pattern, its stability in historical audio data is analyzed, and an audio fluctuation characteristic curve is generated by a fitting algorithm. Based on the audio fluctuation characteristic curve and the changing trend of the frequency fluctuation pattern, the probability of transient high-frequency harsh sounds appearing in future audio frames is predicted, and compared with a preset threshold, and the occurrence trend is output.

7. The game audio processing method according to claim 1, characterized in that, The dynamic suppression processing of the target high-frequency band includes: When a transient high-frequency harsh sound is detected and its loudness exceeds the dynamic suppression threshold, the dynamic suppression control process is triggered. Based on the current dynamic suppression threshold, an initial suppression parameter is applied to the high-frequency band corresponding to the transient high-frequency harsh sound. After applying the initial suppression parameters, the transient feature parameters in the suppressed audio signal are extracted again; Based on the changes in the transient characteristic parameters before and after the dynamic suppression treatment, an effective evaluation result is formed for adjusting the dynamic suppression treatment. If the effective evaluation results indicate that the inhibition effect has not reached the preset target, the inhibition parameters are adaptively adjusted and the inhibition is reapplied. When the effective evaluation results indicate that the inhibition effect has reached the preset target, the current inhibition parameters are maintained. When the transient high-frequency harsh sound is detected to disappear or the risk is reduced, the suppression intensity is gradually released based on the historical change trajectory of the suppression parameter.

8. The game audio processing method according to claim 7, characterized in that, The compensation processing for key positioning sound effects in the mid-to-high frequency band includes: After completing the dynamic suppression process, the energy distribution of each channel in the mid-to-high frequency band is detected. Based on the detected channel energy distribution, determine whether dynamic suppression processing causes a shift in the relative energy relationship between key positioning sound effects in the channels; The degree of impact of dynamic suppression processing on spatial positioning consistency is determined based on the degree of deviation in the relative energy relationship. Based on the degree of influence, a compensation strategy is determined to correct the channel energy relationship offset for key positioning sound effects; wherein, while implementing the compensation strategy, the compensation magnitude is constrained not to exceed the minimum range required to offset the degree of influence; After the compensation process is completed, the relative energy relationship between the channels is re-detected, and if it is not restored to the preset spatial positioning consistency range, the compensation process is iteratively executed.

9. A game audio processing device, characterized in that, include: The segmentation unit is configured to divide the acquired game audio signal into frequency bands, at least dividing it into low-frequency audio signal, mid-high frequency audio signal and high-frequency audio signal; The extraction unit is configured to extract transient feature parameters within the high-frequency audio signal to characterize the transient change characteristics of the audio signal, wherein the transient feature parameters are at least used to reflect the abrupt change characteristics of the high-frequency audio signal in the time and frequency dimensions. The judgment unit is configured to perform multi-dimensional fusion judgment based on the transient feature parameters to identify whether there is a transient audio event with auditory stimulation risk in the high-frequency audio signal, and to judge the corresponding audio signal as transient high-frequency harsh sound when the preset judgment conditions are met. The determining unit is configured to determine the corresponding dynamic suppression threshold based on the occurrence status information of the transient high-frequency harsh sound in the time dimension; The processing unit is configured to perform dynamic suppression processing on the target high-frequency band and compensation processing on key positioning sound effects in the mid-high frequency band when the loudness of the transient high-frequency harsh sound exceeds the dynamic suppression threshold.

10. An earphone, characterized in that, Includes the game audio processing device as described in claim 9.