An audio processing method, system, computing device, and storage medium

CN122511271APending Publication Date: 2026-08-04ZHUHAI KINGSOFT ONLINE GAME TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHUHAI KINGSOFT ONLINE GAME TECH CO LTD
Filing Date
2026-05-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

然而,上述方式会导致听者在场景中移动(如回头、快速靠近/远离声源)时,音频信号突然出现或消失,产生听觉断层,破坏听觉沉浸感

Benefits of technology

[0011] An audio processing method provided in one embodiment of this specification acquires the interaction behavior data of a virtual character in response to the interaction behavior between a virtual character and multiple sound source objects in a virtual scene. Each sound source object has object audio data, which is pre-set with multiple levels of detail. Based on the interaction behavior data, a target level of detail for each object's audio data is determined from the multiple levels of detail. Based on the playback parameters corresponding to the target level of detail, the audio data of each object is played in the virtual scene. This method can dynamically adjust the rendering quality of the object audio data of each sound source according to the real-time interaction state between the virtual character and the sound source objects, avoiding the waste of computing power or loss of key sounds caused by fixed levels. This effectively reduces the computational load and ensures the overall frame rate stability of the virtual scene. Especially in virtual scenes with a large number of concurrent sound sources or frequent interactions, it can balance the on-demand allocation of computing resources, the importance of sound sources, and the continuity of auditory perception, ensuring the user's auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511271A_ABST
    Figure CN122511271A_ABST
Patent Text Reader

Abstract

This specification provides an audio processing method, system, computing device, and storage medium. The method includes: in response to the interaction between a virtual character and multiple sound source objects in a virtual scene, acquiring interaction behavior data of the virtual character, wherein each sound source object has object audio data, and the object audio data is pre-set with multiple levels of detail; based on the interaction behavior data, determining the target level of detail for each object audio data from the multiple levels of detail; and playing the object audio data in the virtual scene based on the playback parameters corresponding to the target level of detail. This method can balance the on-demand allocation of computing resources, the importance of sound sources, and the continuity of auditory perception, ensuring the user's auditory experience. It can be widely applied in the fields of digital cultural and creative product production software and digital cultural and creative software in the next-generation information technology field (such as artificial intelligence, virtual reality, etc.).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of audio processing, and in particular to an audio processing method, system, computing device, and storage medium. Background Technology

[0002] With the development of digital cultural products, factors such as the number of sound sources, the scale of scene space, and the frequency of environmental interaction in digital cultural products have increased significantly, leading to increasingly higher complexity in audio rendering. This has hindered the stable operation of digital cultural products, resulting in phenomena such as screen stuttering and audio rendering stuttering in the scene.

[0003] Existing methods for reducing the complexity of audio rendering in digital cultural products mainly optimize based on the distance between the sound source and the listener. For example, when the sound source is far from the listener, high-frequency components of the audio are directly cut off or the sound source is muted; when the sound source is close to the listener, the complete audio output is restored. However, the above methods can cause the audio signal to suddenly appear or disappear when the listener moves in the scene (such as turning around, quickly approaching / moving away from the sound source), creating auditory gaps and disrupting the sense of auditory immersion.

[0004] In summary, existing solutions for reducing audio rendering complexity suffer from auditory discontinuities, thus diminishing the user's auditory experience. Summary of the Invention

[0005] In view of this, embodiments of this specification provide an audio processing method. One or more embodiments of this specification also relate to an audio processing system, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0006] According to a first aspect of the embodiments of this specification, an audio processing method is provided, comprising: In response to the interaction between the virtual character and multiple sound source objects in the virtual scene, the interaction behavior data of the virtual character is obtained. Each sound source object has object audio data, and the object audio data has multiple levels of detail pre-set. Based on interactive behavior data, the target detail level of each object's audio data is determined from multiple detail levels; Based on the playback parameters corresponding to the target detail level, the audio data of each object is played in the virtual scene.

[0007] According to a second aspect of the embodiments of this specification, an audio processing system is provided, comprising: The load detection module is configured to respond to the interaction behavior between the virtual character and multiple sound source objects in the virtual scene, and to obtain the interaction behavior data of the virtual character. Each sound source object has object audio data, and the object audio data has multiple levels of detail pre-set. The sound source classification module is configured to determine the target level of detail for each object's audio data from multiple levels of detail based on interactive behavior data; The audio signal output module is configured to play the audio data of each object in the virtual scene based on the playback parameters corresponding to the target detail level.

[0008] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which implement the steps of the above method when executed by the processor.

[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0010] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0011] An audio processing method provided in one embodiment of this specification acquires the interaction behavior data of a virtual character in response to the interaction behavior between a virtual character and multiple sound source objects in a virtual scene. Each sound source object has object audio data, which is pre-set with multiple levels of detail. Based on the interaction behavior data, a target level of detail for each object's audio data is determined from the multiple levels of detail. Based on the playback parameters corresponding to the target level of detail, the audio data of each object is played in the virtual scene. This method can dynamically adjust the rendering quality of the object audio data of each sound source according to the real-time interaction state between the virtual character and the sound source objects, avoiding the waste of computing power or loss of key sounds caused by fixed levels. This effectively reduces the computational load and ensures the overall frame rate stability of the virtual scene. Especially in virtual scenes with a large number of concurrent sound sources or frequent interactions, it can balance the on-demand allocation of computing resources, the importance of sound sources, and the continuity of auditory perception, ensuring the user's auditory experience. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating an audio processing method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of a three-layer detail hierarchy of object audio data provided in one embodiment of this specification; Figure 3 This is a schematic diagram illustrating the detail level switching triggering and decision logic of object audio data according to one embodiment of this specification; Figure 4 This is a flowchart illustrating the switching of detail levels of object audio data according to one embodiment of this specification; Figure 5 This is a schematic diagram of the architecture of an audio processing system provided in one embodiment of this specification; Figure 6 This is a flowchart illustrating the operation of an audio processing system according to one embodiment of this specification. Figure 7 This is a schematic diagram of the structure of an audio processing system provided in one embodiment of this specification; Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0017] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0018] Virtual scenes: A virtual scene is a data representation of a spatial environment generated by computing devices, used to carry virtual characters, sound source objects, and interactive behaviors in digital cultural products. For example, in role-playing digital cultural products, virtual scenes can include different terrains and buildings such as forests, cities, and dungeons; in virtual reality applications, virtual scenes can be simulated conference rooms, classrooms, or factory workshops.

[0019] Virtual characters: Virtual characters are objects controlled by users or systems that can move and interact in a virtual environment. For example, in role-playing digital cultural products, virtual characters can be hero characters controlled by players; in simulation management digital cultural products, virtual characters can be workers or customers controlled by the system.

[0020] Detail level: The detail level is a set of rendering parameters that are pre-set for a piece of audio data and correspond to different levels of computational complexity and auditory quality. Each detail level is associated with a specific set of playback parameters.

[0021] Computing devices: Computer devices designed to perform one or more specific tasks. They are less powerful than personal mainframes, but have significant advantages in terms of size and power consumption. They are commonly used in various electronic and mechanical control devices.

[0022] In digital cultural products that include large-scale complex virtual scenes (such as open-world interactive applications and large-scale virtual simulation scenes), the complexity of audio rendering increases exponentially with the increase of the number of sound sources, the scale of the scene space, and the frequency of environmental interaction. Therefore, it is necessary to effectively control the complexity of audio rendering in order to ensure the user's auditory experience.

[0023] Currently, methods for reducing audio rendering control are relatively simple, mostly optimizing only based on the distance between the sound source and the listener. Their core flaws are as follows: 1. Obvious "Popping Effect": When the sound source is far from the listener, high-frequency components are directly cut off or the sound source is muted; when the sound source is close, the full audio output is restored. This crude switching method causes the audio signal to suddenly appear or disappear when the user controls the virtual character to move in the virtual scene (e.g., turning around, quickly approaching / moving away from the sound source), creating a noticeable auditory gap, known in the industry as the "Popping Effect," which severely damages the auditory immersion. 2. Unreasonable allocation of computing power and serious waste of resources: Traditional methods do not consider the real-time CPU load and the importance of the sound source itself. In extreme scenarios (such as large-scale battles or multiple concurrent sound sources), excessive CPU load can easily lead to frame rate drops and audio stuttering; while maintaining high-detail audio rendering even when the sound source is of low importance results in a significant waste of computing resources. 3. Discontinuous listening experience: Traditional switching processes lack a smooth transition mechanism. Even without obvious popping effects, sudden changes in audio parameters can lead to abrupt and discontinuous listening experiences, further affecting the user experience.

[0024] Furthermore, among existing related technologies, audio cross-fade is often used for audio transitions, but it is mostly applied to offline audio editing scenarios. It uses transcendental functions to achieve volume gradation, which has the problem of high real-time calculation latency and cannot adapt to the real-time audio rendering needs of large-scale complex scenarios. Although audio parameter interpolation technology can be used for signal smoothing, it is not deeply integrated with the audio detail level layer control, and cannot solve the parameter connection problem when switching between different detail levels, thus failing to fundamentally solve the above-mentioned defects.

[0025] In view of the above problems, this specification provides an audio processing method, and also relates to an audio processing system, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0026] See Figure 1 , Figure 1 A flowchart of an audio processing method according to an embodiment of this specification is shown, including the following specific steps: Step 102: In response to the interaction between the virtual character and multiple sound source objects in the virtual scene, obtain the interaction behavior data of the virtual character, wherein each sound source object has object audio data, and the object audio data has multiple levels of detail pre-set.

[0027] The embodiments in this specification apply to scenarios involving the processing of audio data in digital cultural products. For example, scenarios involving the processing of audio data in digital cultural products such as virtual reality or role-playing.

[0028] The sound source object is an entity in the virtual scene that can emit sound and has independent audio data. For example, non-player characters, interactive objects, environmental elements (such as flowing water, wind sounds), or the virtual character itself in the virtual scene.

[0029] Interactive behavior refers to spatial or logical actions that occur between a virtual character and a sound source object and can be detected by the system. Examples include a virtual character moving closer to or away from a sound source object, colliding with a sound source object, triggering dialogue, performing an attack, or using an item.

[0030] Interactive behavior data is a set of numerical values ​​or labels that describe the specific attributes of interactive behavior. For example, the real-time distance between the virtual character and the sound source object, their relative movement speed, the type of interactive action (such as touch or dialogue), and the identification of the scene area where the interaction occurs.

[0031] Object audio data refers to audio resources that are bound to a single sound source object and need to be rendered and played in a virtual scene. Examples include greetings from non-player characters, gun firing sounds, chest opening notifications, or background ambient sounds like flowing water.

[0032] Multiple detail levels are at least two rendering levels pre-defined for the same audio data object, ranging from low to high, corresponding to different listening quality and computational complexity. For example, the first detail level is full quality, the second detail level is medium quality (reduced sampling rate or fewer channels), and the third detail level is low-fidelity quality that retains only the identifiable core frequencies.

[0033] Optionally, one way to obtain the interaction behavior data of a virtual character in response to the interaction between the virtual character and multiple sound source objects in the virtual scene is to: calculate and record the interaction behavior data in real time by monitoring the position changes of the virtual character in the virtual scene and collision events with the bounding boxes of the sound source objects. For example, in each frame rendering loop, all sound source objects are traversed, the Euclidean distance between the virtual character and each sound source object is calculated, and whether a near-field trigger occurs. Another way to implement this is to generate interaction behavior data by parsing the operation signals of user input devices (such as keyboards, mice, and gamepads) and combining them with the current task state of the virtual character. For example, when the user presses an interaction key and the virtual character faces a non-player character, the system records the interaction behavior type as "dialogue" and marks the non-player character as the interaction object. This specification does not limit this aspect in the embodiments.

[0034] Optionally, one way to pre-define multiple levels of detail in the object audio data is as follows: During the development phase of digital cultural products, audio designers manually specify multiple levels of detail and their corresponding playback parameters for each segment of object audio data. For example, for an explosion sound effect, the first level of detail could be specified as a 44,100 Hz sampling rate, stereo, with reverb; the second level of detail as a 22,050 Hz sampling rate, mono, without reverb; and the third level of detail as a 11,025 Hz sampling rate, mono, with a frequency range limited to 300-3000 Hz. Another approach is to automatically generate multiple levels of detail using an audio analysis algorithm, retaining different proportions of the dominant frequency components as output levels based on the spectral energy distribution of the original audio. For example, for a long duration of environmental wind sound, the algorithm automatically extracts the frequency band with energy concentrated between 100 and 2000 Hz as the third level of detail, while retaining the entire frequency band as the first level of detail. This specification does not limit the scope of this approach in the embodiments.

[0035] For example, in a shooting simulation digital culture product, a virtual character walks through a warehouse scene with a gun. The surrounding sound sources include the virtual character's own footsteps, distant enemy gunfire, nearby teammates reloading, and the sound of a fan running. As the virtual character quickly runs towards the warehouse door, the system acquires real-time data on the distance changes between the virtual character and each sound source (e.g., the distance for footsteps remains constant at 0.5 meters, while the distance for enemy gunfire gradually decreases from 50 meters to 10 meters), the virtual character's movement speed, and whether the character is currently in combat.

[0036] Step 102 obtains the interaction behavior data of the virtual character by responding to the interaction behavior between the virtual character and multiple sound source objects in the virtual scene, which provides a data foundation for subsequently determining the target detail level of the audio data of each object.

[0037] Step 104: Based on the interaction behavior data, determine the target detail level of the audio data of each object from multiple detail levels.

[0038] The target level of detail is the specific level of detail selected for the audio data of a certain object in the current interactive state. This level determines the set of playback parameters when rendering the audio data subsequently. For example, when the virtual character is very close to the sound source, the target level of detail is the first level of detail; when the distance is far and the sound source is part of the background atmosphere, the target level of detail is the third level of detail.

[0039] Optionally, one way to determine the target detail level of each object's audio data from multiple detail levels based on interaction behavior data is to directly determine the target detail level by looking up a preset distance-detail level mapping table based on the distance values ​​in the interaction behavior data. For example, the first detail level is selected when the distance is less than 5 meters, the second detail level is selected when the distance is between 5 and 15 meters, and the third detail level is selected when the distance is greater than 15 meters. Another approach is to calculate a comprehensive score using multiple dimensions in the interaction behavior data (such as distance, sound source type, and plot importance), and then divide the target detail level according to the score threshold. For example, the first detail level is selected when the comprehensive score is greater than 0.7, the second detail level is selected when it is between 0.3 and 0.7, and the third detail level is selected when it is less than 0.3. This specification does not limit this approach in the embodiments.

[0040] For example, following the above gunfight simulation example, the system obtains that the distance between the virtual character and the enemy's gunshot rapidly decreases from fifty meters to ten meters, and the virtual character is currently in combat. Based on internal rules (distance less than fifteen meters and sound source is combat type), the system determines the target detail level of the audio data of the gunshot object as the first detail level; while for the distant fan spinning sound, its distance is always greater than thirty meters and the sound source type is environmental background, the system determines its target detail level as the third detail level.

[0041] Step 104, based on interactive behavior data, determines the target detail level of each object's audio data from multiple detail levels. This enables the level of detail in audio rendering to dynamically adapt to the real-time interactive context of the virtual character, avoiding the waste of computing resources or the loss of important sound information caused by fixed levels.

[0042] Step 106: Based on the playback parameters corresponding to the target detail level, play the audio data of each object in the virtual scene.

[0043] Playback parameters are the set of parameters used during the rendering and output of object audio data. For example, playback parameters can be categorized according to audio output parameter standards, such as sampling rate, bit rate, number of audio channels, and audio frequency range. They can also be categorized according to rendering logic, such as core data, physical sound effects, and environmental sound effects. Furthermore, they can be categorized according to the percentage of computational overhead (CPU performance overhead), such as 5%-10%, 30%-40%, and 100%. The aforementioned core data can be data that retains only the core content that allows the user's hearing to identify within the object audio data; that is, it retains the core frequency components of the object audio data (such as character footsteps at 100-800Hz, basic NPC dialogue at 300-3400Hz), maintaining the basic perceptibility of the sound source without adding any additional effects such as distance attenuation or reverberation.

[0044] Indicatively, Figure 2This specification illustrates a three-layer detail hierarchy structure of object audio data according to an embodiment. Table 1 below (showing playback parameters corresponding to the three detail levels) further clarifies this structure. Figure 2 The content will be introduced in detail below: Table 1. Playback parameters corresponding to the three levels of detail.

[0045] It should be noted that the computational overhead percentages for each detail level in Table 1 are based on 100% of the 7.1 channel full detail rendering of a single sound source. In real-world scenarios with multiple sound sources running concurrently, the system can allocate the rendering computational power of each sound source's detail level in a cumulative manner, prioritizing the computational power supply to the core sound source and avoiding any loss of auditory quality from the core sound source.

[0046] Optionally, one way to play audio data of each object in a virtual scene based on playback parameters corresponding to the target detail level is to call the dynamic parameter setting interface provided by the audio engine to apply the playback parameters corresponding to the target detail level to the currently playing audio stream in real time. For example, obtain the audio component object in Unreal Engine and call functions to set the sampling rate, set the number of channels, etc. Another way is to maintain an independent renderer for each object's audio data, and after determining the target detail level, reinitialize the renderer and load the audio resources of the corresponding level for playback. For example, stop the current renderer, release the original resources, and then create a new renderer with new parameters and start playback from the beginning. The embodiments in this specification do not limit this approach.

[0047] For example, for the aforementioned gunshot audio data, the target detail level is the first detail level. The system plays the gunshot according to the sampling rate of 44,100 Hz, stereo, and full-band parameters, so that the virtual character hears a realistic and clearly positioned gunshot. For the fan sound, it is played according to the third detail level parameters, and only mono low-frequency sound is output. Although high-frequency details are lost, computing resources are greatly saved.

[0048] Step 106 plays the audio data of each object in the virtual scene based on the playback parameters corresponding to the target detail level. This allows the computational cost of audio rendering to match the current interaction requirements, ensuring the auditory quality of important sound sources while reducing the overall audio rendering complexity, thereby reducing screen and audio stuttering.

[0049] In the embodiments of this specification, interaction behavior data of the virtual character is obtained in response to the interaction behavior between the virtual character and multiple sound source objects in the virtual scene. Each sound source object has object audio data, which is pre-set with multiple levels of detail. Based on the interaction behavior data, the target level of detail for each object audio data is determined from the multiple levels of detail. Based on the playback parameters corresponding to the target level of detail, the audio data of each object is played in the virtual scene. This method can dynamically adjust the rendering quality of the object audio data of each sound source according to the real-time interaction state between the virtual character and the sound source objects, avoiding the waste of computing power or loss of key sounds caused by fixed levels. This effectively reduces the computational load and ensures the overall frame rate stability of the virtual scene. Especially in virtual scenes with a large number of concurrent sound sources or frequent interactions, it can balance the on-demand allocation of computing resources, the importance of sound sources, and the continuity of auditory perception, ensuring the user's auditory experience.

[0050] In one optional embodiment of this specification, step 104 includes the following specific steps: Based on interaction behavior data, determine the priority parameters corresponding to the audio data of each object; Based on priority parameters, the target level of detail for each object's audio data is determined from multiple levels of detail.

[0051] The priority parameter is a quantitative indicator that reflects the importance of each object's audio data to the virtual character at the current moment. For example, the priority parameter can range from 0 to 1, where 1 represents the most important (such as key plot dialogue, the sound of an enemy approaching), and 0 represents the least important (such as distant, irrelevant background birdsong).

[0052] Optionally, one way to determine the priority parameters corresponding to the audio data of each object based on interaction behavior data is to look up a preset priority mapping table based on the distance value and sound source type in the interaction behavior data to directly obtain the priority value. For example, when the distance is less than two meters and the sound source type is "warning", the priority parameter is set to 0.95; when the distance is greater than thirty meters and the sound source type is "environment", the priority parameter is set to 0.05. Another way to implement this is to use a deep learning model, taking the multi-dimensional features of the interaction behavior data as input and outputting a priority score. For example, constructing a fully connected network with two hidden layers, taking distance, relative speed, sound source type encoding, and plot progress as input, and outputting a priority parameter between 0 and 1. Yet another way to implement this is to use a fuzzy inference system to fuzzify and defuzzify the interaction distance, sound source type, and user operation frequency to obtain the priority parameter. For example, outputting a high priority when the distance is close and the operation frequency is high, and outputting a low priority when the distance is far and the operation frequency is low. This specification does not limit this approach in the embodiments.

[0053] Optionally, one implementation for determining the target detail level of each object's audio data from multiple detail levels based on the priority parameter is to compare the priority parameter with two thresholds and directly select the corresponding detail level based on the comparison result. For example, if the priority is greater than 0.7, the first detail level (highest quality) is selected; if the priority is between 0.3 and 0.7, the second detail level (medium quality) is selected; and if the priority is less than 0.3, the third detail level (lowest quality) is selected. Another implementation is to sort the priority parameters of all object audio data and allocate detail levels according to a preset ratio. For example, the top 20% with the highest priority are allocated to the first detail level, the middle 60% to the second detail level, and the bottom 20% to the third detail level. Yet another implementation is to treat the priority parameter as a continuous variable and calculate the index value of the detail level using a fitting function (such as rounding or mapping). For example, the target detail level = rounded up (priority parameter × 3), resulting in 1, 2, or 3. This specification does not limit the implementation in this way.

[0054] Optionally, in one embodiment, multiple levels of detail audio data can be pre-stored for each object's audio data. After determining the target level of detail for each object's audio data from multiple levels of detail based on interaction behavior data, the audio data corresponding to the target level of detail is played in the virtual scene. For example, developers pre-store three versions of audio files for the sound of a moving car engine: a high-fidelity stereo version, a medium-quality version, and a low-quality version. When a virtual character stands next to the car and observes it closely, the system determines the target level of detail based on interaction behavior data (distance 0.5 meters, observing the vehicle) and then plays the high-fidelity version. When the virtual character runs away to a distance of 50 meters and turns its back to the car, the system determines the target level of detail and plays the low-quality version. When the virtual character is about 20 meters away and talking to someone else, the system determines the target level of detail and plays the medium-quality version. By pre-storing multiple levels of detail object audio data, the computational overhead of downsampling or effects processing the original audio at runtime is avoided, further reducing the real-time CPU load.

[0055] For example, in a role-playing digital cultural product, a virtual character is currently conversing with a key non-player character. Interaction data shows that the sound source of this dialogue is classified as "story dialogue," and the virtual character is only 1.2 meters away from the character. Based on this data, the system determines the priority parameter of this dialogue sound source to be 0.92, which is greater than 0.7. Therefore, the target detail level of this dialogue audio data is set to the first detail level, ensuring that the user can clearly hear the emotional details and lip-sync of each line. Simultaneously, a non-player-controlled pet cat occasionally meows in the distance behind the virtual character. Its priority parameter is only 0.1, less than 0.3. Therefore, the system sets the target detail level of this cat's meow to the third detail level, retaining only the recognizable outline of the "meow" sound, significantly reducing its rendering overhead.

[0056] In the embodiments described in this specification, priority parameters corresponding to the audio data of each object are determined based on interactive behavior data; the target level of detail for the audio data of each object is determined from multiple levels of detail based on the priority parameters; and playback is performed based on the playback parameters corresponding to the target level of detail. This approach can reasonably distinguish the importance of sound sources according to the priority parameters, allocating high-priority sound sources to high-quality rendering and reducing the detail of low-priority sound sources to save computing resources, thereby reducing the overall computational load, ensuring stable frame rates in the virtual scene, and improving the user's auditory experience.

[0057] In one optional embodiment of this specification, the interactive behavior data includes at least one dimension, and the interactive behavior data of at least one dimension includes at least one of the following: the sound source type of each object's audio data, the plot tag of the virtual scene, the interaction distance between each sound source object and the virtual character, and the correlation between each object's audio data and the user's operation. Based on interaction behavior data, priority parameters corresponding to the audio data of each object are determined, including: The priority parameters corresponding to the audio data of each object are obtained by weighted summation of the interaction behavior data of at least one dimension.

[0058] Sound source type is a data identifier that represents the functional attributes of an object's audio data. For example, sound source types can be divided into ambient sounds (wind sounds, water sounds), character action sounds (footsteps, sword swinging sounds), dialogue voice (non-player character greetings, quest prompts), combat sound effects (gunshots, explosions), interface feedback sounds (clicks, prompts), etc.

[0059] The plot tags of a virtual scene are data identifiers that identify the current narrative stage of the virtual scene. For example, plot tags may include "main plot climax", "side quest calm period", "new player tutorial area", "dungeon battle preparation", etc.

[0060] The interaction distance between each sound source object and the virtual character represents the spatial distance between each sound source object and the virtual character in the virtual scene. For example, the interaction distance can be the Euclidean distance between each sound source object and the virtual character in the virtual scene.

[0061] The correlation between audio data of each object and user operation is a quantitative representation of the degree to which the audio data of each object affects the user's current operation (such as key press, click, movement, attack). For example, the correlation of the gunshot sound produced after the user presses the fire button is set to 1.0; the correlation of the background music that plays automatically when the user does not perform any operation is set to 0.

[0062] Optionally, one way to obtain the priority parameters corresponding to the audio data of each object by weighted summation of the interaction behavior data of at least one dimension can be to pre-assign a weight coefficient to each dimension, and then sum the normalized scores of each dimension by their corresponding weights. Another way can be to calculate the weights of each dimension using the analytic hierarchy process (AHP), and then perform a weighted summation. For example, a pairwise comparison matrix can be constructed, and the eigenvectors can be calculated to obtain the weight values ​​of each dimension. This specification does not limit the embodiments described herein.

[0063] For example, in a post-apocalyptic survival digital cultural product, a virtual character is performing a main storyline mission: "Infiltrating the enemy base." The system calculates priority parameters for a background alarm sound: the sound source type is "environmental alarm," with a score of 0.6; the storyline tag is "main storyline infiltration," with a score of 0.9; the interaction distance is 25 meters, with a normalized distance score of 0.3; and the user operation relevance is 0. Setting the weights for each dimension to 0.2, 0.3, 0.3, and 0.2 respectively, the priority parameter = 0.2 × 0.6 + 0.3 × 0.9 + 0.3 × 0.3 + 0.2 × 0 = 0.48. This value falls between 0.3 and 0.7, and the system determines the target detail level of this alarm sound as the second detail level to balance auditory atmosphere and computational power consumption. Meanwhile, the footsteps of an enemy five meters in front of the virtual character have a sound source type score of 0.9, a plot tag score of 0.9, a distance score of 0.9, and a relevance score of 0.8. After weighted summation, the priority parameter is as high as 0.85, and it is assigned to the first level of detail to ensure that the user can clearly locate the enemy's position.

[0064] Priority parameters are quantitative indicators that characterize the importance of each object's audio data to the virtual character in the virtual scene. They reflect the perceived importance of each object's audio data to the user in the virtual scene. For example, semantic classification and quantitative assignment can be performed on the audio data of each type of object in the virtual scene. The sound source levels can be divided according to the function, interaction relationship, and plot weight of the object's audio data in the scene, and a weight value can be assigned to each level. Priority parameters can be divided into core sound sources, secondary sound sources, auxiliary sound sources, etc. Optionally, the classification of priority parameters can include both user manual preset and system automatic classification, and dynamic adjustment during scene operation is supported (such as the importance of NPC sound sources increasing as the plot progresses).

[0065] For example, Table 2 schematically shows the sound source level (i.e., priority parameter) classification table of the object audio data, which can divide the sound source level of the object audio data into three categories: core sound source, secondary sound source, and auxiliary sound source. The weight assignment, definition, and typical examples of each category are shown in Table 2: Table 2. Sound Source Classification Table for Object Audio Data

[0066] Optionally, the system can automatically classify the priority parameters of object audio data based on four dimensions: sound source type, interaction distance between each sound source object and the virtual character, plot tags of the virtual scene, and the relevance of each object's audio data to user operations, using a weighted scoring algorithm. For example, the weights of the four dimensions are 0.3, 0.2, 0.3, and 0.2, respectively, and the total score is the weight value W (priority parameter) of the object audio data. The sound source level is automatically classified according to the range of W's value. Illustratively, the interaction distance dimension scores highest when the interaction distance is ≤5m, and lowest when the interaction distance is >20m; plot tags can be pre-labeled by scene developers (e.g., "key plot points" or "background atmosphere").

[0067] In the embodiments of this specification, interactive behavior data is collected from multiple dimensions such as sound source type, plot tag, interaction distance, and user operation relevance. The priority parameters are obtained by weighted summation of the data from each dimension. The target detail level is determined and played based on the priority parameters. This can more accurately assess the importance of each sound source, avoid the waste of computing resources or the degradation of important sounds caused by misjudgment from a single dimension, thereby achieving stable frame rate in virtual scene operation and ensuring that the auditory experience of important sound sources is not lost, thus improving the user experience.

[0068] In one optional embodiment of this specification, before determining the target level of detail for each object's audio data from multiple levels of detail based on a priority parameter, the method further includes: Obtain the computational resource usage of audio data for each object; Based on priority parameters, the target level of detail for each object's audio data is determined from multiple levels of detail, including: Based on priority parameters and computational resource utilization, the target detail level of each object's audio data is determined from multiple detail levels.

[0069] The computational resource utilization rate of each object's audio data is the ratio between the computational resources consumed by the computing unit (such as the central processing unit) in rendering each object's audio data and the total computing resources. It reflects the workload of the computing unit in rendering each object's audio data.

[0070] Indicatively, Figure 3 This diagram illustrates a detail-level switching triggering and decision-making logic for object audio data according to an embodiment of this specification. Figure 3 As shown, the CPU audio processing utilization rate of the system CPU can be collected at a sampling frequency of 10ms / sample through the CPU load acquisition interface of the driver layer (only the CPU utilization related to audio rendering is counted, excluding the utilization of other processes). A 50ms sliding window averaging filtering algorithm is used to smooth the collected data to eliminate instantaneous load fluctuations (such as sudden load increases and decreases caused by sudden sound sources), and output a stable CPU audio processing utilization rate value (denoted as C, with a value range of 0-100%). Illustratively, two configurable CPU load thresholds can be preset: a high load threshold (denoted as C1, configurable range 60%-80%, default 70%) and a low load threshold (denoted as C2, configurable range 20%-40%, default 30%), and C2 must strictly satisfy C2 < C1 to avoid frequent switching caused by threshold overlap. The specific switching trigger logic is as follows: (1) When C > C1, the detail level of the object audio data is switched downward (e.g., environment layer → detail layer, detail layer → core layer), which can prioritize reducing the detail level of the low-priority parameter object audio data to save computing power; (2) When C < C2, the audio detail level is switched upward (e.g., core layer → detail layer, detail layer → environment layer), which prioritizes increasing the detail level of the low-priority parameter object audio data (e.g., the detail level of the core sound source) to restore the ultimate listening experience; (3) When C2 ≤ C ≤ C1, the current detail level of the object audio data remains unchanged to avoid fluctuations in listening experience caused by frequent switching.

[0071] Optionally, dedicated computing resources for audio processing can be reserved in advance for the audio data of the rendered object, so that these dedicated computing resources can be isolated from the computing load unrelated to audio processing.

[0072] Optionally, one way to obtain the computational resource utilization rate of each object's audio data is to read the CPU time percentage of the audio processing thread through the operating system's performance monitoring application programming interface. For example, in a Windows operating system, a performance counter query function can be called to obtain the percentage of CPU time used by the audio rendering thread in the most recent frame period. Another approach is to embed time stubs in the audio rendering pipeline, count the number of microseconds consumed in rendering all object audio data per frame, and then divide by the frame period to obtain the utilization rate. This specification does not limit the implementation of this method.

[0073] Optionally, one implementation for determining the target detail level of each object's audio data from multiple detail levels based on priority parameters and computational resource utilization can be as follows: when the computational resource utilization is higher than a certain load threshold, the target detail level of one or more object audio data with lower priority parameters is determined first; when the computational resource utilization is lower than a certain load threshold, the target detail level of one or more object audio data with higher priority parameters is determined first. Another implementation can be: a weighted combination of computational resource utilization and priority parameters; when this weighted combination exceeds a certain upper limit, the detail level is switched downwards; when it is lower than a certain lower limit, it is switched upwards. This specification does not limit the implementation in this way.

[0074] For example, in a massively multiplayer online role-playing game (MMORPG), a virtual character is in the main city square, surrounded by thirty non-player characters simultaneously emitting dialogue, and five ambient sound effects (fountain, pigeons, bells). The system monitors the computational resource utilization of audio processing in real time; the current value is 73%, exceeding the preset high-load threshold of 70%. At this point, the system iterates through all object audio data, switching the target detail level of the ten background ambient sound effects with the lowest priority parameters (such as distant bells and wheel sounds) from the second detail level to the third detail level, causing the computational resource utilization to drop to 65%. Subsequently, the virtual character is teleported to an open wilderness scene, where the number of surrounding sound sources is reduced to five, and the computational resource utilization drops to 25%, below the low-load threshold of 30%. The system then raises the target detail level of the non-player character dialogue with the highest priority parameter from the second detail level back to the first detail level, and gradually upgrades other secondary sound sources to restore the auditory quality to its optimal state.

[0075] In the embodiments of this specification, by obtaining the computing resource occupancy rate of each object's audio data, and based on the priority parameter and computing resource occupancy rate, the target detail level of each object's audio data is determined from multiple detail levels. This allows for real-time perception of the computing load status. When the load is high, the detail level of low-priority sound sources is automatically reduced to release computing power, and when the load is low, the detail level of high-priority sound sources is automatically increased to optimize the listening experience. This makes the dynamic allocation of computing resources more reasonable, effectively controls the peak CPU load, keeps the overall frame rate stable, and ensures that the listening continuity of important sound sources is not affected.

[0076] In one optional embodiment of this specification, the plurality of detail levels include a first detail level, a second detail level, and a third detail level, in which the audio detail is sequentially reduced; Based on priority parameters and computational resource utilization, the target detail level of each object's audio data is determined from multiple detail levels, including: Based on priority parameters, determine the switching order of playback parameters corresponding to the audio data of each object; If the calculated resource utilization rate is greater than the first threshold, the second or third detail level will be determined as the target detail level based on the playback parameter switching order. If the calculated resource utilization rate is less than the second threshold, the second detail level or the first detail level is determined as the target detail level based on the playback parameter switching order.

[0077] Audio details are parameters that reflect the fineness of an object's audio data. For example, audio details can include frequency resolution, spatial positioning accuracy, and dynamic range.

[0078] The first level of detail is the highest quality level of detail, which can be used for the most critical or closest sound sources to ensure a complete auditory experience. For example, the first level of detail can be the core layer mentioned above, or it can be an ultra-high fidelity layer, used for the voices of key plot characters, the player's own action sound effects, and the main interactive feedback sounds.

[0079] The second level of detail is a medium-quality level of detail, which can be used for sound sources of general importance or at medium distances, striking a balance between auditory quality and computational power. For example, the second level of detail can be the aforementioned level of detail, or it can be a balanced level, used for dialogue of ordinary non-player characters, non-core combat sound effects, and ambient sounds at medium distances.

[0080] The third level of detail is the lowest quality level of detail, which can be used for unimportant or distant sound sources, prioritizing the conservation of computational resources. For example, the third level of detail can be the aforementioned environment layer, or it can be a placeholder layer, used for distant background noise, low-priority ambient sound effects, and a large number of concurrent, negligible sound sources.

[0081] The playback parameter switching order is a data representation that indicates the order in which playback parameters are switched for each audio data object. For example, when the computational resource utilization rate is greater than a certain preset threshold, it indicates that the computational load is high, and playback parameters corresponding to the detail levels of audio data objects with lower priority parameters (such as auxiliary sound sources) are switched first, from high to low. When the computational resource utilization rate is less than a certain preset threshold, it indicates that the computational load is low, and playback parameters corresponding to the detail levels of audio data objects with higher priority parameters (such as the core sound source) are switched first, from low to high.

[0082] For example, the switching order of playback parameters for audio data of each object can be as follows: (1) For core sound sources: even if the CPU load is slightly higher than C1 (exceeding the threshold ≤ 5%), it will not easily switch to the low detail layer; when the CPU load drops back to below C1, it will switch upward first; the system reserves 10%-15% of dedicated computing power for core sound sources to ensure that the core listening experience is not damaged; (2) For secondary sound sources: strictly follow the switching logic of CPU load threshold, dynamically adjust the detail level according to the change of C, do not enjoy computing power reservation, and take into account both listening experience and computing power; (3) For auxiliary sound sources: when the CPU load is close to C1 (C≥C1-5%), it will switch downward first, and maintain core layer rendering by default. It will only switch upward when the CPU load is much lower than C2 (C≤C2-10%), which is the core object of computing power saving.

[0083] Optionally, one implementation for determining the playback parameter switching order corresponding to the audio data of each object based on priority parameters is as follows: the order of downward switching is based on the order of priority parameters from smallest to largest, and the order of upward switching is based on the order of priority parameters from largest to smallest. For example, for three sound sources with priorities of 0.9, 0.5, and 0.2, the downward switching order is 0.2, 0.5, and 0.9, and the upward switching order is 0.9, 0.5, and 0.2. Another implementation is to bind the priority parameters to preset levels (core, secondary, auxiliary), and then switch according to distance or random order within the same level. For example, the switching order of the core sound source is always later than that of the secondary and auxiliary sound sources, and the auxiliary sound source is always the first to be downgraded. This specification does not limit this aspect in the embodiments.

[0084] The first threshold is an upper limit used to determine whether the computing resource utilization is too high. When this value is exceeded, the detail level is switched down. For example, the first threshold can be set to 70%. When the computing resource utilization is greater than 70%, the system needs to reduce the detail level of some object audio data to release computing power.

[0085] The second threshold is a lower limit used to determine whether the computing resource utilization rate is too low. When it falls below this value, the level of detail is switched upwards. For example, the second threshold can be set to 30%. When the computing resource utilization rate is less than 30%, the system can increase the level of detail of some audio data to optimize the listening experience.

[0086] Optionally, if the calculated resource occupancy rate is greater than the first threshold, one implementation method for determining the second or third detail level as the target detail level based on the playback parameter switching order is as follows: Following the downward switching order, sound sources currently at the first detail level are sequentially switched to the second detail level. If the resource occupancy rate is still greater than the first threshold, then sound sources currently at the second detail level are further switched to the third detail level. For example, the resource occupancy rate is recalculated after each sound source switch until it drops below the first threshold. Another implementation method is to set the target detail level of all sound sources with a priority lower than a certain threshold to the third detail level at once, and set the rest to the second detail level. For example, all sound sources with a priority less than 0.4 are directly downgraded to the third detail level. This specification does not limit this implementation.

[0087] If the calculated resource utilization rate is less than the second threshold, the implementation method of determining the second or first detail level as the target detail level based on the playback parameter switching order can be referred to the above. If the calculated resource utilization rate is greater than the first threshold, the implementation method of determining the second or third detail level as the target detail level based on the playback parameter switching order will not be elaborated here.

[0088] For example, in an open-world game, the current computational resource utilization is 78%, exceeding the first threshold of 70%. The system first switches the three lowest-priority auxiliary sound sources (such as distant river sounds, rustling leaves, and insect chirps) from the second level of detail to the third level of detail, reducing the utilization to 72%, still above 70%. The system then switches the next two lowest-priority secondary sound sources (such as distant non-player character voices) from the first level of detail to the second level of detail, reducing the utilization to 66%, which is below the first threshold, and stops switching. Subsequently, the user leaves the area, reducing the computational resource utilization to 28%, below the second threshold of 30%. The system then maintains the highest-priority core sound source (player footsteps) at the first level of detail, while restoring the downgraded secondary sound sources from the second level of detail back to the first level of detail, increasing the utilization to 35%, and stops restoring.

[0089] In the embodiments of this specification, by pre-setting three audio detail levels that decrease sequentially, the playback parameter switching order is determined based on priority parameters. When the computing resource utilization rate is higher than the first threshold, the detail level is switched downwards in sequence, and when it is lower than the second threshold, it is switched upwards in sequence. This enables systematic and step-by-step batch resource adjustment, prioritizing the listening experience of important sound sources and quickly releasing the computing power of unnecessary sound sources, thereby significantly improving frame rate stability. Furthermore, throughout the adjustment process, both the continuity of listening experience and the importance of sound sources are taken into account.

[0090] In one optional embodiment of this specification, after determining the playback parameter switching order corresponding to the audio data of each object based on the priority parameter, the method further includes: Based on priority parameters and computational resource utilization, the switching decision factors are determined; If the calculated resource utilization rate is greater than the first threshold, the second or third detail level is determined as the target detail level based on the playback parameter switching order, including: If the calculated resource utilization rate is greater than the first threshold, and / or the switching decision factor is greater than the third threshold, the second or third detail level will be determined as the target detail level based on the switching order of the playback parameters. If the calculated resource utilization rate is less than the second threshold, the second level of detail or the first level of detail is determined as the target level of detail based on the playback parameter switching order, including: If the calculated resource utilization rate is less than the second threshold, and / or the switching decision factor is less than the fourth threshold, the second detail level or the first detail level is determined as the target detail level based on the switching order of the playback parameters.

[0091] The switching decision factor is a composite indicator representing the priority parameter of the target audio data and the computing resource utilization rate, used to more accurately determine whether a switching detail level is needed. For example, the switching decision factor can be defined as the product of the priority parameter and the computing resource utilization rate, or a weighted sum of the two, to reflect the combined effect of "importance and load pressure," avoiding misjudgments triggered by a single parameter at the detail level. Illustratively, the computing resource utilization rate and the priority parameter of the target audio data can be jointly quantified to construct a switching decision factor F = W × (C / 100), where W represents the priority parameter (such as the sound source importance weight W), and C represents the computing resource utilization rate. The system presets decision factor thresholds (F1=0.6, F2=0.3) and combines them with the computational resource utilization rate and the priority parameters of the object's audio data to form a complete joint decision logic. For example, when F>F1, it triggers a downward switch of the detail level, and the switching priority is: auxiliary sound source> secondary sound source> core sound source; when F<F2, it triggers an upward switch of the detail level, and the switching priority is: core sound source> secondary sound source> auxiliary sound source; when F2≤F≤F1, it maintains the current detail level; if the core sound source's F>F1, but the computational resource utilization rate exceeds the threshold ≤5%, it still maintains the original level.

[0092] The third threshold is an upper limit used to determine whether the switching decision factor is too high. For example, the third threshold can be set to 0.6. When the switching decision factor is greater than 0.6, it means that under the combined effect of the current priority parameter and the computing resource utilization rate, the system should reduce the detail level of the sound source to release computing power.

[0093] The fourth threshold is a lower limit used to determine whether the switching decision factor is too low. For example, the fourth threshold can be set to 0.3. When the switching decision factor is less than 0.3, it means that under the combined effect of the current priority parameter and the computing resource utilization rate, the system should improve the detail level of the sound source to optimize the listening experience. Optionally, if the computing resource utilization rate is greater than the first threshold and / or the switching decision factor is greater than the third threshold, one implementation method for determining the second or third detail level as the target detail level based on the switching order of playback parameters can be: triggering a downgrade when both conditions are met, otherwise maintaining the original level. For example, the first threshold is 70%, the third threshold is 0.6, and a downward switch is only performed when the utilization rate is >70% and the decision factor is >0.6. Another implementation method can be: triggering a downgrade when either of the two conditions is met. For example, if the utilization rate is >70% or the decision factor is >0.6, a downgrade begins when either condition is met.

[0094] Similarly, if the calculated resource utilization rate is less than the second threshold, and / or the switching decision factor is less than the fourth threshold, the implementation method of determining the second detail level or the first detail level as the target detail level based on the switching order of playback parameters can be referred to the above implementation method of determining the second detail level or the third detail level as the target detail level based on the switching order of playback parameters if the calculated resource utilization rate is greater than the first threshold, and / or the switching decision factor is greater than the third threshold. It will not be elaborated here.

[0095] For example, in a highly dynamic tactical shooting game, the current computational resource utilization rate is 65%, which is less than the first threshold of 70%. However, because the decision factor F = priority 0.9 × 0.65 = 0.585 for the gunshot sound source with the highest priority parameter, which is slightly less than the third threshold of 0.6, the system does not trigger a downgrade and maintains the first level of detail for the gunshot sound, ensuring the sound quality in combat. In another scenario, the computational resource utilization rate is 68%, which is close to but not yet at the threshold. However, the priority of a certain secondary dialogue sound source is low (0.4), and its decision factor F = 0.4 × 0.68 = 0.272, which is less than the fourth threshold of 0.3. The system determines that although the overall load is not high at this time, the importance of this sound source is very low and the load is close to the upper limit. Therefore, it downgrades it from the second level of detail to the third level of detail, releasing computing power in advance to reserve space for possible sudden sound sources. Conversely, when the occupancy rate is 32%, slightly higher than the second threshold of 30%, but the decision factor of the core sound source F=0.9×0.32=0.288, which is less than the fourth threshold of 0.3, the system still triggers an upward switch, upgrading the core sound source from the second level of detail to the first level of detail, so that the user can obtain the best listening experience as soon as possible. This specification does not limit this aspect in the embodiments.

[0096] In the embodiments described in this specification, a switching decision factor is constructed based on priority parameters and computing resource utilization. When the load is higher than a first threshold and / or the decision factor is greater than a third threshold, the system switches downwards; when the load is lower than a second threshold and / or the decision factor is less than a fourth threshold, the system switches upwards. Playback is then performed based on the playback parameters corresponding to the target detail level after the switch. This approach comprehensively assesses the dual impact of sound source importance and computing resource load, avoiding frequent erroneous switching caused by fluctuations in a single threshold. This allows for precise allocation of computing resources, achieving a stable frame rate while ensuring a seamless auditory experience.

[0097] In one optional embodiment of this specification, audio data of each object is played in the virtual scene based on playback parameters corresponding to the initial level of detail; Before step 106, the following is also included: Switch the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level.

[0098] The playback parameters corresponding to the initial detail level are the set of playback parameters associated with the audio data of each object currently in use in the virtual scene before the playback parameters are switched.

[0099] Optionally, one implementation method for playing audio data of each object in a virtual scene based on playback parameters corresponding to the initial level of detail can be: when the digital cultural product starts, a default level of detail is assigned to each sound source according to the initial position of the virtual character and the preset quality configuration, and playback begins with the corresponding playback parameters applied. Another implementation method can be: the user manually selects the global initial level of detail through the settings menu; for example, in "Performance Mode," the initial level of detail for all object audio data is set to the third level of detail (lowest quality), and in "Quality Mode," all sound sources are initially set to the first level of detail (highest quality). This specification does not limit the implementation of this method in the embodiments.

[0100] Optionally, one way to switch the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level is to immediately stop the audio rendering of the initial detail level, release the computing resources it occupies, and restart playing the same audio resource using the playback parameters of the target detail level. Another way is to render the audio of both levels simultaneously during a brief overlap period, gradually changing the volume ratio to smoothly transition to the target level. This specification does not limit the implementation of this method in the embodiments.

[0101] For example, the virtual character was initially standing at a distance of 30 meters from the fountain, and the fountain sound was played at the initial detail level 3 (mono, narrowband). When the virtual character quickly ran to the fountain (one meter away), the system determined the target detail level of the fountain sound to be the first level.

[0102] This embodiment switches the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level, and plays the audio based on the switched playback parameters. This avoids audio interruptions or pops caused by direct hard switching, thereby ensuring that the continuity of the listening experience is not compromised by the switching of detail levels while reducing the computational load.

[0103] In one optional embodiment of this specification, switching the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level includes: Configure transition durations for the audio data of each object separately; During the transition period, the playback parameters corresponding to the initial detail level are switched to the playback parameters corresponding to the target detail level based on the volume gradient curve. The volume gradient curve is pre-constructed based on a first-order rational function.

[0104] The transition duration is the time interval from the start of the transition to its completion. Within this interval, the audio signals of the initial detail level and the target detail level are rendered simultaneously and mixed and output according to the gradient curve. For example, the overlap transition duration T can be set (configurable range 50-200ms, default 100ms). T can be adaptively adjusted according to the sound source level, balancing auditory continuity and computational consumption: core sound source T=100-200ms (prioritizing auditory quality), secondary sound source T=50-100ms, auxiliary sound source T=50ms (prioritizing computational efficiency).

[0105] The volume transition curve is a function curve describing the change in the volume ratio of the initial detail level and the target detail level over time during the transition duration. For example, the volume transition curve can be an exponential curve, a linear curve, or a rational function curve. This embodiment uses a first-order rational function curve to balance smoothness and computational efficiency.

[0106] For example, the function formula for the volume button curve is as follows: (1) Current detail level fade-out curve: Vs(t)=t / (k+t), where t is the transition time (0≤t≤T), k is the adjustment coefficient (default k=50, which can be adjusted according to the scene), and Vs(t) is the current level volume ratio (from 1 to 0). (2) Target level fade-in curve: Vt(t)=1-Vs(t), where Vt(t) is the target level volume ratio (from 0 to 1).

[0107] The rational function curves described above are smooth and without abrupt changes, and the computational cost is only 1 / 5 of that of traditional transcendental functions, which greatly reduces real-time computation latency and ensures that the switching process is seamless.

[0108] To address the issue of high real-time computational latency in traditional crossfade-in / fade-out algorithms, a rational function is used to construct the volume gradient curve, replacing traditional transcendental functions (such as sine and exponential functions). This improves computational efficiency while ensuring a smooth transition. Additionally, a configurable overlapping transition interval is set to achieve smooth superposition of the audio signals of the current and target layers.

[0109] Optionally, one way to configure the transition duration for each object's audio data is to dynamically calculate the transition duration based on the priority parameter of the object's audio data; the higher the priority, the longer the transition duration. Another way is for the audio designer to preset a fixed transition duration for each sound source type during the development phase and store it in a configuration table. For example, 200 milliseconds for dialogue, 50 milliseconds for gunshots, and 100 milliseconds for ambient sounds. This specification does not limit this approach in the embodiments.

[0110] Alternatively, one way to construct a volume gradient curve based on a first-order rational function is to pre-generate a volume lookup table with a length equal to the transition duration, where each discrete time point corresponds to an initial level volume ratio and a target level volume ratio, and the table is directly looked up at runtime.

[0111] Optionally, during the transition period, one way to switch the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level based on the volume gradient curve is to simultaneously start two audio renderers, outputting audio streams according to the initial detail level parameters and the target detail level parameters respectively, and then weighting and mixing the two audio streams according to the volume gradient curve before outputting them to the audio device. Another approach is to first mute the initial detail level, and then render only the target detail level and gradually increase its volume from 0 to 1 according to the volume gradient curve. This specification does not limit the implementation of this approach.

[0112] For example, within a 150-millisecond transition period, the system gradually decreases the gain of the initial detail level according to the volume gradient curve, while simultaneously increasing the gain of the target detail level. Because a first-order rational function is used, the first half of the curve changes rapidly, while the second half changes more slowly, which aligns with the human ear's perception of volume changes (more sensitive to changes in high sound pressure levels, less sensitive to changes in low sound pressure levels). After the switch is complete, rendering of the initial detail level stops, and the system plays entirely using the target detail level.

[0113] In the embodiments described in this specification, by configuring a transition duration for the audio data of each object, and within the transition duration, switching the playback parameters of the initial detail level to the playback parameters of the target detail level based on the volume gradient curve, it is possible to achieve a smooth fade-out of the old level and a smooth fade-in of the new level, avoiding popping sounds or auditory breaks caused by abrupt changes in detail levels. Furthermore, the gradient curve constructed based on a first-order rational function has a small computational load and does not increase the CPU burden, thus ensuring stable frame rate while effectively maintaining auditory continuity.

[0114] In one optional embodiment of this specification, during the transition period, switching the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level based on the volume gradient curve includes: During the transition period, based on the volume gradient curve, determine the volume ratio of each object's audio data played according to the playback parameters corresponding to the initial detail level and the playback parameters corresponding to the target detail level. Based on the volume ratio, switch the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level.

[0115] The volume ratio refers to the relative magnitude of the output audio data volumes of the initial detail level and the target detail level at any given moment within the transition interval. For example, within the transition interval T, the system simultaneously renders the object audio data of the current detail level and the target detail level, adjusting the volume ratio of the two detail levels in real time according to the aforementioned curve to achieve smooth overlay of the object audio data and avoid popping or abrupt changes in volume. After the transition, rendering of the current detail level is stopped, and only the object audio data output of the target detail level is retained. Optionally, the playback parameters of the target detail level can be calibrated after the transition.

[0116] Optionally, within the transition duration, one approach to determining the volume ratio of each object's audio data played according to the playback parameters corresponding to the initial detail level and the playback parameters corresponding to the target detail level, based on the volume gradation curve, could be: Calculate the initial level volume ratio using the volume gradation curve formula based on the current time t, where the target level volume ratio = 1 - the initial level volume ratio. Another approach could be: pre-discretize the volume gradation curve into multiple time slices, maintaining a fixed volume ratio within each time slice, and directly look up the ratio in a table at runtime based on the current time slice. For example, divide 100 milliseconds into 10 segments of 10 milliseconds each, with each segment having a preset ratio value.

[0117] Optionally, one way to switch the playback parameters corresponding to the initial detail level to the target detail level based on the volume ratio is to multiply each sample point of the initial level audio stream by the initial level volume ratio, and each sample point of the target level audio stream by the target level volume ratio, then add the two products to obtain the output sample point. For example, for the nth sample point, the output = initial ratio × initial sample value + target ratio × target sample value. Another way to implement this is to use the audio engine's mixer interface to set the volume ratio for each of the two audio sources, and then have the mixer automatically mix them.

[0118] For example, within a 100-millisecond transition duration, the system calculates the volume percentage every frame (e.g., every 10 milliseconds). At millisecond 0 of the transition, the initial level percentage is 1.0, the target level percentage is 0.0, and the output is entirely the initial level sound; at millisecond 25, the initial percentage is 0.75, and the target percentage is 0.25; at millisecond 50, each accounts for 0.5; at millisecond 75, the initial percentage is 0.25, and the target percentage is 0.75; at millisecond 100, the initial percentage is 0.0, and the target percentage is 1.0. Within each frame, the system renders audio data from both levels simultaneously and mixes them according to the aforementioned weights before feeding them into the audio buffer.

[0119] In the embodiments of this specification, the volume ratio of the initial detail level and the target detail level is calculated in real time during the transition period, and the audio signals of the two object audio data are weighted and mixed according to the ratio before output. This makes the volume ratio of the object audio data corresponding to the initial detail level and the target detail level change continuously with time, thereby making the output waveform of the object audio data continuous and eliminating the sound discontinuity problem caused by sudden silence or sudden recovery.

[0120] In one optional embodiment of this specification, prior to step 106, the method further includes: Perform parameter interpolation on the playback parameters corresponding to the target detail level to obtain the playback parameters corresponding to the target detail level after parameter interpolation; Step 106 includes the following specific steps: Based on the playback parameters corresponding to the target detail level after parameter interpolation, the audio data of each object is played in the virtual scene.

[0121] Parametric interpolation is the process of gradually changing playback parameters from a reference value corresponding to one detail level to a reference value corresponding to another detail level over a transition period. For example, when the reverberation intensity changes from 20% to 80%, interpolation will cause the reverberation intensity to rise uniformly (or along a curve) from 20% to 80% during the transition period.

[0122] Optionally, the playback parameters corresponding to the target detail level can be obtained by performing parameter interpolation on the playback parameters. The implementation method can be as follows: for each parameter that needs to be interpolated, the current parameter value is calculated by linear interpolation or shifted linear interpolation based on the current time t and the transition duration T.

[0123] For example, when switching the audio data of a sound source object from the second level of detail to the first level of detail, the reverberation intensity parameter needs to gradually increase from 20% to 80% over a transition period of 100 milliseconds. During the transition, the system adjusts the reverberation intensity every 20 milliseconds: initially 20%; after 20 milliseconds, it becomes 32%; after 40 milliseconds, it becomes 44%; after 60 milliseconds, it becomes 56%; after 80 milliseconds, it becomes 68%; and at the end of the 100-millisecond transition, it reaches exactly 80%.

[0124] In the embodiments of this specification, playback parameters corresponding to the target detail level are interpolated to obtain the interpolated playback parameters. Based on the interpolated playback parameters, the audio data of each object is played in the virtual scene, which enables acoustic parameters such as reverberation intensity, sampling rate, and spatial azimuth angle to change continuously during the transition period rather than abruptly. This ensures that the physical characteristics and spatial positioning of the audio remain consistent even when the detail level changes due to the adjustment of the computational load, further improving the continuity of the user's listening experience.

[0125] In one optional embodiment of this specification, audio data of each object is played in the virtual scene based on playback parameters corresponding to the initial level of detail; The playback parameters corresponding to the target detail level are interpolated to obtain the interpolated playback parameters corresponding to the target detail level, including: Determine the amount of change between the playback parameters corresponding to the initial level of detail and the playback parameters corresponding to the target level of detail; If the change is greater than the preset change, linear interpolation is performed on the playback parameters corresponding to the target detail level to obtain the linearly interpolated playback parameters corresponding to the target detail level. If the change is not greater than the preset change, the playback parameters corresponding to the target detail level are subjected to shifted linear interpolation to obtain the playback parameters corresponding to the target detail level after shifted linear interpolation.

[0126] The preset change amount is a threshold used to determine whether a change in playback parameters is significant. Different preset change amounts can be set for different types of parameters. For example, the preset change amount is set to 10000Hz for sampling rate and 20% for reverberation intensity. If the change amount exceeds this threshold, the playback parameter is considered to have changed significantly; otherwise, the playback parameter is considered to have changed gradually.

[0127] Optionally, one way to determine the change between the playback parameters corresponding to the initial detail level and the playback parameters corresponding to the target detail level is to calculate the absolute difference and compare it with a preset change. Another way is to calculate the relative percentage change; if it exceeds a percentage threshold, it is considered significant. For example, if the reverberation intensity changes from 50% to 55%, the relative change is 10%, and the preset threshold is 15%, then it is not significant. This specification does not limit the scope of this embodiment.

[0128] To address the abrupt changes in the physical and spatial parameters of audio data at different detail levels, this specification employs a combination of linear interpolation and shifted linear interpolation in its embodiments to achieve a smooth transition of all relevant playback parameters within the transition interval T. The shifted linear interpolation algorithm, balancing computational efficiency and interpolation accuracy, is the system's default interpolation algorithm and is suitable for real-time audio rendering scenarios. Specifically: (1) Interpolation parameter range: The parameters to be interpolated include two main categories: physical parameters and spatial parameters of the target audio data. Specifically, these can be: distance attenuation coefficient, Doppler frequency shift, occlusion attenuation coefficient, reverberation intensity, reverberation time, spatial azimuth angle, spatial elevation angle, etc. The interpolation range of each parameter is strictly matched with the reference value of the parameter at the target detail level to ensure the continuity of the parameters after switching and to prevent abrupt changes.

[0129] (2) Selection of interpolation algorithm: 1) Linear interpolation: Suitable for scenarios where parameters change gradually (such as distance attenuation coefficient interpolation from core layer to detail layer). Algorithm formula: P(t)=Ps+(Pt-Ps)×(t / T), where Ps is the initial value of the parameter of the current detail layer, Pt is the baseline value of the parameter of the target detail layer, t is the transition time, and P(t) is the parameter value at time t. 2) Shifted linear interpolation: Suitable for scenarios with large parameter variations (such as reverberation intensity interpolation from detail layer to environment layer). It adds a shift step size (default 0.05, configurable) to linear interpolation, dividing parameter variations into multiple smooth stages to avoid parameter abrupt changes caused by a single interpolation. The algorithm efficiency is consistent with linear interpolation, and the interpolation accuracy is improved by 30%. The specific formula is: P(t)=Ps+(Pt-Ps)×(t / (T×s)), where s is the shift step size.

[0130] 3) Synchronous execution logic: Indicatively, Figure 4 This diagram illustrates a flowchart of a detail level switching process for object audio data according to one embodiment of this specification. Figure 4 As shown, parameter interpolation and crossfade-in / fade-out start and end synchronously. Within the transition interval T, the interpolation of all parameters and the gradual change of volume are coordinated in real time to ensure that the physical characteristics and spatial positioning of the audio after the switch remain consistent with those before the switch, avoiding abrupt listening sensations caused by parameter changes. After the interpolation is completed, the parameters at the target detail level can be calibrated in real time to ensure that they match the environment (such as indoor / outdoor) and sound source location of the current virtual scene.

[0131] Optionally, if the change is greater than a preset change, linear interpolation is performed on the playback parameters corresponding to the target detail level to obtain the linearly interpolated playback parameters corresponding to the target detail level. One implementation method is to apply a linear interpolation formula to calculate the parameter value once for each audio frame or each sampling point. For example, the reverberation intensity is updated once for each audio frame. Another implementation method is to use the smoothing parameter mechanism built into the audio engine, setting the smoothing time to be equal to the transition duration, and the engine automatically performs linear interpolation. This specification does not limit this aspect in the embodiments.

[0132] Optionally, if the change is not greater than a preset change, a shifted linear interpolation process is performed on the playback parameters corresponding to the target detail level to obtain the playback parameters corresponding to the target detail level after shifted linear interpolation. One implementation method is to set the shift step size s=0.1, complete the interpolation in the first 10% of the transition time, and maintain the target value for the remaining 90% of the time. Another implementation method is to dynamically adjust the shift step size according to the change; the smaller the change, the smaller the step size, and the faster the interpolation. This specification does not limit this approach in the embodiments.

[0133] For example, when switching detail levels, the reverberation intensity changes from 45% to 50%, a change of 5%, less than the preset change of 20%. The system determines the change is insignificant and uses shifted linear interpolation with a shift step size s=0.1 and a transition time T=100ms. Thus, within the first 10ms after the switch, the reverberation intensity linearly increases from 45% to 50%; then remains constant at 50% for the next 90ms. The user hardly perceives the change in reverberation intensity because the change is small and rapid, avoiding meaningless, prolonged gradual changes. In another scenario, the sampling rate changes from 10kHz to 44kHz, a change of 34kHz, much greater than the preset change of 10kHz. The system uses linear interpolation to uniformly increase the frequency from 10kHz to 44kHz within 100ms. The user will perceive the sound quality gradually becoming clearer, natural, and not abrupt.

[0134] In the embodiments of this specification, by adaptively selecting between linear interpolation and shifted linear interpolation according to the magnitude of parameter change, linear interpolation is used for significantly changing playback parameters to make the change evenly distributed throughout the transition period, avoiding spectral jumps. For parameters that do not change significantly, shifted linear interpolation is used to quickly complete the change in the first short period of the transition period, avoiding lengthy gradual changes, thus achieving a balance between smoothness and response speed.

[0135] It should be noted that the audio processing methods provided in this manual can be applied to various industries or scenarios, such as virtual reality processing software, home entertainment product software, digital cultural product production software, digital cultural creative software, digital cultural creative design, education, news, cultural content industry software, digital publishing software, digital music development and production, and digital mobile multimedia development and production. In some cases, they can also be applied to fields such as animation and game production engine software and development systems, game and animation software, animation and game digital content services, digital film and television development and production, and digital performance development and production.

[0136] Corresponding to the above method embodiments, this application also provides an audio processing system. Figure 5 A schematic diagram of the architecture of an audio processing system according to an embodiment of this application is shown.

[0137] like Figure 5 As shown, the audio processing system adopts a four-layer modular architecture design, consisting of a hardware layer, a driver layer, a core algorithm layer, and an application layer from bottom to top. Each layer is independently encapsulated and reserves standardized data interaction interfaces, supporting cross-platform and cross-audio rendering engine adaptation. The specific functions and module composition of each layer are as follows: Hardware layer: As the physical foundation for system operation, it includes hardware devices such as CPU (responsible for audio rendering computing power allocation and calculation), audio decoding chip (responsible for audio signal decoding), sound card (responsible for audio signal conversion and output), and speakers / headphones (responsible for final audio playback); its core functions are final audio signal output, CPU load data acquisition, audio hardware driver adaptation, and providing a stable hardware operating environment for the upper layer.

[0138] Driver layer: Includes two core sub-modules: hardware driver and audio engine interface adaptation module; the core function is data transfer between the hardware layer and the core algorithm layer, specifically including: real-time acquisition and transmission of CPU load data, hardware distribution of audio rendering parameters, and standardized interface adaptation of mainstream audio rendering engines (such as OpenAL, Unity Audio, Unreal Audio), ensuring that the system can be quickly deployed based on existing audio rendering engines without reconstructing the underlying architecture.

[0139] Core Algorithm Layer: This is the core execution layer of the system and the core unit for dynamic switching of audio detail levels. It includes five sub-modules: load detection module, sound source classification module, switching decision module, smooth switching module, and parameter calibration module. The functions of each sub-module are as follows: (1) Load detection module: responsible for real-time acquisition of CPU audio processing utilization rate, using sliding window average filtering algorithm to smooth the acquired data, eliminate false triggering caused by instantaneous load fluctuations, and output stable CPU load parameters. (2) Sound source classification module: responsible for real-time detection, importance classification and quantification of all sound sources in the scene, supporting two modes: user manual preset classification standards and system automatic classification, and dynamically updating sound source importance parameters; (3) Switching decision module: Based on the CPU load parameters output by the load detection module and the sound source importance parameters output by the sound source classification module, a joint decision logic is constructed to determine whether it is necessary to switch the audio detail level, the switching direction (up / down) and the switching object; (4) Smooth switching module: It integrates cross-fade-in and fade-out and parameter interpolation mechanisms to perform smooth switching between various audio detail levels, eliminating Popping effect and auditory discontinuity; (5) Parameter calibration module: After the switching is completed, the audio parameters of the target level are calibrated in real time to ensure that the parameters match the current scene environment and sound source status and maintain the stability of the listening experience.

[0140] Application layer: It includes three sub-modules: scene configuration module, parameter preset module, and status monitoring module. It provides users with a visual operation interface. Its core functions are: scene parameter configuration, core threshold preset, system operation status monitoring, and switch record query. System debugging and maintenance can be completed without professional technicians.

[0141] The four-layer architecture achieves data interaction with the local data bus through the standardized TCP / IP protocol, with a data transmission latency of ≤10ms, meeting the real-time audio rendering requirements of large-scale complex scenarios. The interfaces between each layer adopt a standardized design, which facilitates future expansion and maintenance.

[0142] Indicatively, Figure 6 A flowchart of an audio processing system according to an embodiment of this application is shown.

[0143] like Figure 6 As shown, the audio processing system adopts a fully automated, non-human-interventional workflow. From system initialization to switching execution during operation, and then to maintaining the state after switching, all are automatically completed by the core algorithm layer. At the same time, exception handling and stability assurance strategies are added to avoid switching failures or abnormal listening experiences caused by extreme conditions. The system's entire workflow logic is divided into three parts: the initialization phase, the operation phase, and the exception handling phase, detailed as follows: Initialization phase: When the system starts, it completes audio data loading, parameter presets, engine adaptation, and module initialization to prepare for subsequent operation. Specific steps include: 1) Audio data loading: Load the three-layer audio data of all sound sources in the scene, preprocess the audio data (format standardization, frequency component extraction, parameter calibration), and store the audio data of each layer in the local cache (cache size is configurable, default is 100MB) to achieve fast retrieval and avoid audio stuttering when switching. 2) Parameter preset: Through the parameter preset module of the application layer, core parameters such as CPU load threshold (C1, C2), sound source importance weight standard, cross-fade-in and fade-out transition duration T, parameter interpolation algorithm type, switching delay time, and reverberation time range can be configured. Users can save multiple sets of parameter configuration schemes according to scenario requirements. 3) Engine and hardware adaptation: The driver layer completes the interface adaptation with the existing audio rendering engine, initializes the audio output device and CPU load acquisition interface of the hardware layer, establishes the data transmission channel, tests the data transmission latency, and ensures that the latency is ≤10ms; 4) Module initialization: The core algorithm layer initializes each functional module, completes the sampling frequency calibration of the load detection module (ensuring 10ms / sampling), loads the automatic classification rules of the sound source classification module, configures the decision factor threshold of the switching decision module, and initializes the algorithm parameters of the smooth switching module. 5) Self-test: The system completes a self-test to verify the effectiveness of data interaction, parameter transmission, and algorithm execution at each layer, and to check whether the audio rendering at each layer is normal and whether the switching mechanism can be triggered normally. After the self-test is passed, the system enters the running phase; if the test fails, a fault prompt is issued and the faulty module is identified.

[0144] Operation phase: The system enters normal operation, completing data acquisition, decision-making, switchover execution, and status maintenance in real time. Specific steps include: 1) Real-time data acquisition: The load detection module collects the CPU audio processing utilization rate at a frequency of 10ms / time, and outputs a stable C value after 50ms sliding window filtering; the sound source classification module performs real-time detection of sound sources in the scene (addition / disappearance / location change / plot-related change), and dynamically updates the sound source importance weight W at a frequency of 50ms / time. 2) Switching decision judgment: The switching decision module calculates the decision factor F based on C and W, and combines the preset F1 and F2 thresholds and sound source priority to determine whether it is necessary to switch the audio detail level, and at the same time determine the switching direction (up / down) and the switching object (core / secondary / auxiliary sound source). 3) Smooth transition execution: If a transition is required, the smooth transition module activates a dual smoothing mechanism, simultaneously executing cross-fade-in and fade-out and parameter interpolation, completing the level transition within the transition duration T; during the transition, the volume gradation and parameter interpolation progress are monitored in real time to ensure a smooth transition; 4) State Maintenance: After the switch is completed, the system maintains the audio rendering of the target level, and the parameter calibration module calibrates the target level parameters in real time; at the same time, it continuously monitors the changes of C and W until the next switch trigger condition is met. 5) Data recording: The status monitoring module records audio detail level switching records in real time, including switching time, sound source name, level before and after switching, CPU load value, sound source weight value, and switching duration, which facilitates subsequent scene optimization and fault diagnosis. The record retention time is configurable (default 7 days).

[0145] Exception handling phase For extreme anomalies during system operation (such as sudden changes in CPU load, sudden addition / disappearance of sound sources, and loss of audio data), a dedicated anomaly handling strategy is designed to ensure system stability and continuity of audio experience, as detailed below: 1) Sudden CPU load change: When the CPU load suddenly increases by more than 20% within 50ms (e.g., more than 20 concurrent sound sources suddenly appear), the system directly switches all auxiliary sound sources to the core layer, and secondary sound sources to the core layer / detail layer (based on priority). The core sound source maintains its original layer, and at the same time, an emergency smooth switch is initiated (T=50ms) to quickly reduce computing power consumption and avoid audio stuttering; when the CPU load suddenly drops by more than 20%, the upward switch is triggered after a 100ms delay to avoid frequent switching. 2) Sudden addition / disappearance of sound sources: When multiple sound sources (≥10) are suddenly added to the scene, the system will prioritize setting the newly added auxiliary sound sources as the core layer for rendering, and then adjust according to the decision logic after the CPU load stabilizes; when a sound source suddenly disappears, the system will use a fade-in fade-out mechanism (T=50ms) to reduce the volume of the sound source to 0 to avoid the auditory discontinuity caused by the sudden disappearance of the sound. 3) Audio data loss: When audio data at a certain level is lost, the system automatically interpolates from adjacent levels to generate replacement data (e.g., if environment layer data is lost, temporary environment layer data is generated by interpolating from detail layer parameters) to ensure the continuity of audio rendering. At the same time, a data loss prompt is issued to notify the user to replenish the audio data. 4) Algorithm execution error: If the smooth switching algorithm fails, the system will immediately stop switching, maintain the current level of audio rendering, and restart the smooth switching module to re-execute the switching operation. If it fails 3 times in a row, a fault message will be issued.

[0146] The above audio processing system can achieve the following: 1) Completely eliminate the popping effect and improve the continuity of the listening experience: Through the parameterized connection of the three-layer audio performance structure, combined with the dual smoothing mechanism of optimized cross-fade-in and high-precision parameter interpolation, non-destructive switching of audio detail levels is achieved, which completely solves the problems of instantaneous popping and abrupt appearance / disappearance of sound in traditional methods; the switching transition is imperceptible, and the physical characteristics and spatial positioning of the audio are consistent, which significantly improves the user's auditory immersion, especially suitable for scenarios with high requirements for auditory experience, such as games and virtual simulations.

[0147] 2) Fine-grained allocation of computing power to balance system performance and auditory experience: Based on the joint switching decision logic of CPU real-time load and the semantic importance of sound sources, computing resources are allocated on demand and supplied with priority. In extremely complex scenarios (such as multiple sound sources concurrently), the CPU audio processing utilization rate can be reduced by 30%-50%, and the system frame rate can be stabilized at more than 60fps. When the CPU is under low load, the detail level of the core sound source is automatically improved to restore the ultimate auditory experience, avoid the ineffective waste of computing resources, and improve the computing power utilization rate by more than 40%.

[0148] 3) Lightweight design to meet the real-time needs of large-scale complex scenarios: All core algorithms have been optimized for lightweight operation, with real-time calculation latency ≤5ms and switching response time ≤200ms, supporting high-frequency, multi-source layer switching; the system adopts a modular architecture with data transmission latency ≤10ms, which can adapt to the real-time audio rendering needs of large-scale complex scenarios such as open-world games, large-scale virtual simulations, and metaverses, without affecting the user's real-time operation experience.

[0149] 4) Strong compatibility, easy deployment and promotion: The system's four-layer architecture is fully compatible with existing audio rendering engines (OpenAL, Unity Audio, Unreal Audio), requiring no large-scale modifications to existing audio resources. It can be quickly deployed by simply loading three layers of audio data and configuring relevant parameters. It supports cross-platform operation (Windows, Linux, Android, iOS) and can be widely used in various large-scale and complex scenarios such as games, virtual simulation, metaverse, and multiplayer online interaction, with high practicality and market promotion value.

[0150] 5) High configurability and intelligence to adapt to different scenario requirements: The system's core parameters (CPU load threshold, transition time, sound source weight, interpolation algorithm, etc.) all support visual configuration and adaptive adjustment, which can be flexibly optimized according to the needs of different application scenarios; the importance of sound sources supports automatic classification and dynamic updates, eliminating the need for manual configuration of each sound source, realizing intelligent audio detail level control, and reducing user operation and maintenance costs.

[0151] Corresponding to the above method embodiments, this application also provides an audio processing system. Figure 7 A schematic diagram of an audio processing system according to an embodiment of this application is shown. Figure 7 As shown, the system includes: The load detection module 702 is configured to respond to the interaction behavior between the virtual character and multiple sound source objects in the virtual scene, and to acquire the interaction behavior data of the virtual character, wherein each sound source object has object audio data, and the object audio data has multiple levels of detail preset. The sound source classification module 704 is configured to determine the target level of detail for each object's audio data from multiple levels of detail based on interactive behavior data; The audio signal output module 706 is configured to play audio data of each object in the virtual scene based on the playback parameters corresponding to the target detail level.

[0152] Optionally, the sound source classification module 704 is further configured as follows: Based on interaction behavior data, determine the priority parameters corresponding to the audio data of each object; Based on priority parameters, the target level of detail for each object's audio data is determined from multiple levels of detail.

[0153] Optionally, the interaction behavior data includes at least one dimension, and the interaction behavior data of at least one dimension includes at least one of the following: the sound source type of each object's audio data, the plot tag of the virtual scene, the interaction distance between each sound source object and the virtual character, and the correlation between each object's audio data and the user's operation. The sound source classification module 704 is further configured as follows: The priority parameters corresponding to the audio data of each object are obtained by weighted summation of the interaction behavior data of at least one dimension.

[0154] Optionally, the audio processing system also includes a load detection module, configured to obtain the computational resource utilization of each object's audio data before determining the target level of detail for each object's audio data from multiple levels of detail based on priority parameters; The sound source classification module 704 is further configured as follows: Based on priority parameters and computational resource utilization, the target detail level of each object's audio data is determined from multiple detail levels.

[0155] Optionally, the multiple levels of detail include a first level of detail, a second level of detail, and a third level of detail, in which the audio detail decreases sequentially. The sound source classification module 704 is further configured as follows: Based on priority parameters, determine the switching order of playback parameters corresponding to the audio data of each object; If the calculated resource utilization rate is greater than the first threshold, the second or third detail level will be determined as the target detail level based on the playback parameter switching order. If the calculated resource utilization rate is less than the second threshold, the second detail level or the first detail level is determined as the target detail level based on the playback parameter switching order.

[0156] Optionally, the audio processing system also includes a switching decision module, configured to determine switching decision factors based on priority parameters and computing resource utilization. The sound source classification module 704 is further configured as follows: If the calculated resource utilization rate is greater than the first threshold, and / or the switching decision factor is greater than the third threshold, the second or third detail level will be determined as the target detail level based on the switching order of the playback parameters. If the calculated resource utilization rate is less than the second threshold, and / or the switching decision factor is less than the fourth threshold, the second detail level or the first detail level is determined as the target detail level based on the switching order of the playback parameters.

[0157] Optionally, the audio data of each object in the virtual scene is played based on the playback parameters corresponding to the initial level of detail; the audio processing system also includes a smooth switching module, which is configured to switch the playback parameters corresponding to the initial level of detail to the playback parameters corresponding to the target level of detail.

[0158] Optionally, the smooth switching module is further configured as follows: Configure transition durations for the audio data of each object separately; During the transition period, the playback parameters corresponding to the initial detail level are switched to the playback parameters corresponding to the target detail level based on the volume gradient curve. The volume gradient curve is pre-constructed based on a first-order rational function.

[0159] Optionally, the smooth switching module is further configured as follows: During the transition period, based on the volume gradient curve, determine the volume ratio of each object's audio data played according to the playback parameters corresponding to the initial detail level and the playback parameters corresponding to the target detail level. Based on the volume ratio, switch the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level.

[0160] Optionally, before playing the audio data of each object in the virtual scene based on the playback parameters corresponding to the target detail level, the smooth switching module is further configured to: perform parameter interpolation on the playback parameters corresponding to the target detail level to obtain the playback parameters corresponding to the target detail level after parameter interpolation; Based on the playback parameters corresponding to the target detail level after parameter interpolation, the audio data of each object is played in the virtual scene.

[0161] Optionally, the audio data of each object in the virtual scene is played based on the playback parameters corresponding to the initial level of detail; The smooth switching module is further configured to: acquire spatial coordinate data of the virtual narrative scene; Determine the amount of change between the playback parameters corresponding to the initial level of detail and the playback parameters corresponding to the target level of detail; If the change is greater than the preset change, linear interpolation is performed on the playback parameters corresponding to the target detail level to obtain the linearly interpolated playback parameters corresponding to the target detail level. If the change is not greater than the preset change, the playback parameters corresponding to the target detail level are subjected to shifted linear interpolation to obtain the playback parameters corresponding to the target detail level after shifted linear interpolation.

[0162] The audio processing system provided in this application acquires the interaction behavior data of the virtual character by responding to the interaction behavior between the virtual character and multiple sound source objects in the virtual scene. Each sound source object has object audio data, which is pre-set with multiple levels of detail. Based on the interaction behavior data, the system determines the target level of detail for each object's audio data from these multiple levels of detail. Based on the playback parameters corresponding to the target level of detail, the system plays the audio data of each object in the virtual scene. This system can dynamically adjust the rendering quality of the object audio data for each sound source according to the real-time interaction status between the virtual character and the sound source objects, avoiding the waste of computing power or loss of key sounds caused by fixed levels. This effectively reduces the computational load and ensures a stable overall frame rate for the virtual scene. Especially in virtual scenes with a large number of concurrent sound sources or frequent interactions, it can balance the on-demand allocation of computing resources, the importance of sound sources, and the continuity of auditory perception, thus guaranteeing the user's auditory experience.

[0163] The above is an illustrative scheme of an audio processing system according to this embodiment. It should be noted that the technical solution of this audio processing system and the technical solution of the aforementioned audio processing method belong to the same concept. Details not described in detail in the technical solution of the audio processing system can be found in the description of the technical solution of the aforementioned audio processing method. Furthermore, the components in the system embodiment should be understood as functional modules necessary to implement each step of the program flow or each step of the method; these functional modules are not actual functional divisions or separations. The system claims defined by such a set of functional modules should be understood as a functional module architecture that primarily implements the solution through the computer program described in the specification, and not as a physical system that primarily implements the solution through hardware.

[0164] Figure 8 This specification illustrates a structural block diagram of a computing device according to one embodiment. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.

[0165] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0166] In one embodiment of this specification, the above-described components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0167] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.), smart home appliances, multimedia playback devices, intelligent voice interaction devices, or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server. A server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0168] The memory 810 is used to store computer programs / instructions, and the processor 820 is used to execute the following computer programs / instructions, which, when executed by the processor, implement the steps of the above-described audio processing method.

[0169] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the audio processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the audio processing method described above.

[0170] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described audio processing method.

[0171] The above is an illustrative embodiment of a computer-readable storage medium. It should be noted that the technical solution of this storage medium and the technical solution of the aforementioned audio processing method belong to the same concept. Details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the aforementioned audio processing method.

[0172] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described audio processing method.

[0173] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described audio processing method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described audio processing method.

[0174] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0175] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0176] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0177] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0178] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An audio processing method, characterized in that, include: In response to the interaction between a virtual character and multiple sound source objects in a virtual scene, the interaction behavior data of the virtual character is obtained, wherein each sound source object has object audio data, and the object audio data has multiple levels of detail pre-set. Based on the interaction behavior data, the target detail level of each object's audio data is determined from multiple detail levels; Based on the playback parameters corresponding to the target detail level, the audio data of each object is played in the virtual scene.

2. The method according to claim 1, characterized in that, The process of determining the target level of detail for each object's audio data from multiple levels of detail based on the interaction behavior data includes: Based on the interaction behavior data, the priority parameters corresponding to the audio data of each object are determined; Based on the priority parameter, the target detail level of each object's audio data is determined from multiple detail levels.

3. The method according to claim 2, characterized in that, The interactive behavior data includes at least one dimension, and the interactive behavior data of at least one dimension includes at least one of the following: the sound source type of the audio data of each object, the plot tag of the virtual scene, the interaction distance between each sound source object and the virtual character, and the correlation between the audio data of each object and the user operation. The step of determining the priority parameters corresponding to the audio data of each object based on the interaction behavior data includes: The interaction behavior data of at least one dimension are weighted and summed to obtain the priority parameters corresponding to the audio data of each object.

4. The method according to claim 2, characterized in that, Before determining the target level of detail for each object's audio data from multiple levels of detail based on the priority parameter, the method further includes: Obtain the computational resource utilization rate of the audio data of each object; The step of determining the target level of detail for each object's audio data from multiple levels of detail based on the priority parameter includes: Based on the priority parameter and the computing resource utilization rate, the target detail level of each object's audio data is determined from multiple detail levels.

5. The method according to claim 4, characterized in that, The multiple levels of detail include a first level of detail, a second level of detail, and a third level of detail, in which the audio detail decreases sequentially. The process of determining the target detail level of each object's audio data from multiple detail levels based on the priority parameter and the computing resource utilization rate includes: Based on the priority parameter, the playback parameter switching order corresponding to the audio data of each object is determined; If the computing resource occupancy rate is greater than the first threshold, the second detail level or the third detail level is determined as the target detail level based on the playback parameter switching order; If the computational resource utilization rate is less than the second threshold, the second detail level or the first detail level is determined as the target detail level based on the playback parameter switching order.

6. The method according to claim 5, characterized in that, After determining the playback parameter switching order corresponding to the audio data of each object based on the priority parameter, the method further includes: Based on the priority parameters and the computing resource utilization rate, the switching decision factor is determined; If the computational resource occupancy rate is greater than a first threshold, determining the second detail level or the third detail level as the target detail level based on the playback parameter switching order includes: If the computing resource utilization rate is greater than the first threshold, and / or the switching decision factor is greater than the third threshold, the second detail level or the third detail level is determined as the target detail level based on the playback parameter switching order; If the computational resource occupancy rate is less than the second threshold, determining the second detail level or the first detail level as the target detail level based on the playback parameter switching order includes: If the computing resource utilization rate is less than the second threshold, and / or the switching decision factor is less than the fourth threshold, the second detail level or the first detail level is determined as the target detail level based on the playback parameter switching order.

7. The method according to any one of claims 1-6, characterized in that, The virtual scene plays the audio data of each object based on the playback parameters corresponding to the initial level of detail. Before playing the audio data of each object in the virtual scene based on the playback parameters corresponding to the target detail level, the method further includes: Switch the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level.

8. The method according to claim 7, characterized in that, Switching the playback parameters corresponding to the initial level of detail to the playback parameters corresponding to the target level of detail includes: Configure a transition duration for the audio data of each object; During the transition period, the playback parameters corresponding to the initial detail level are switched to the playback parameters corresponding to the target detail level based on the volume gradient curve, wherein the volume gradient curve is pre-constructed based on a first-order rational function.

9. The method according to claim 8, characterized in that, During the transition period, switching the playback parameters corresponding to the initial detail level to the playback parameters corresponding to the target detail level based on the volume gradient curve includes: During the transition period, based on the volume gradient curve, the volume ratio of each object's audio data played according to the playback parameters corresponding to the initial detail level and the playback parameters corresponding to the target detail level is determined; Based on the volume ratio, the playback parameters corresponding to the initial detail level are switched to the playback parameters corresponding to the target detail level.

10. The method according to claim 1, characterized in that, Before playing the audio data of each object in the virtual scene based on the playback parameters corresponding to the target detail level, the method further includes: Perform parameter interpolation on the playback parameters corresponding to the target detail level to obtain the interpolated playback parameters corresponding to the target detail level. The step of playing the audio data of each object in the virtual scene based on the playback parameters corresponding to the target detail level includes: Based on the playback parameters corresponding to the target detail level after parameter interpolation, the audio data of each object is played in the virtual scene.

11. The method according to claim 10, characterized in that, The virtual scene plays the audio data of each object based on the playback parameters corresponding to the initial level of detail. The step of performing parameter interpolation on the playback parameters corresponding to the target detail level to obtain the interpolated playback parameters corresponding to the target detail level includes: Determine the amount of change between the playback parameters corresponding to the initial level of detail and the playback parameters corresponding to the target level of detail; If the change is greater than the preset change, linear interpolation is performed on the playback parameters corresponding to the target detail level to obtain the linearly interpolated playback parameters corresponding to the target detail level. If the change amount is not greater than the preset change amount, the playback parameters corresponding to the target detail level are subjected to shifted linear interpolation to obtain the playback parameters corresponding to the target detail level after shifted linear interpolation.

12. An audio processing system, characterized in that, include: The load detection module is configured to respond to the interaction behavior between the virtual character and multiple sound source objects in the virtual scene, and to acquire the interaction behavior data of the virtual character, wherein each sound source object has object audio data, and the object audio data is pre-set with multiple detail levels; The sound source classification module is configured to determine the target level of detail for each object's audio data from multiple levels of detail based on the interactive behavior data. The audio signal output module is configured to play the audio data of each object in the virtual scene based on the playback parameters corresponding to the target detail level.

13. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, It stores a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.

15. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.