Dynamic sound field interaction control method and device, electronic equipment and storage medium

By dividing the sound source logical distribution areas in the virtual scene and adjusting the spatial audio parameters and emotional response in real time, the problems of insufficient dynamic sound source distribution and single emotional response in traditional sound source systems are solved, and dynamic sound field interactive control with highly realistic sound effects is achieved.

CN120704635APending Publication Date: 2025-09-26SHANGHAI NETEASE CUICAN NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510804942.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional sound source systems find it difficult to achieve dynamic distribution of sound sources, real-time emotional response, and coordinated control of group behavior in dynamic interactive scenarios, resulting in a lack of continuity and multi-dimensional adaptation of the sound field atmosphere.

Method used

The virtual scene is divided into multiple logical sound source distribution areas, and the sound source distribution data is collected in real time through the group sound source control node. The spatial audio parameters are dynamically adjusted, and a dynamic sound field is generated based on the emotional dynamic response parameters to achieve the collaborative expression of layered sound effect behaviors.

Benefits of technology

It achieves precise mapping of the sound field and multi-dimensional dynamic adaptation, improves the immersiveness of the sound field and the coherence of group behavior simulation, and solves the problems of insufficient sound source adaptation and single emotional response in traditional sound source systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704635A_ABST
    Figure CN120704635A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic sound field interaction control method and device, electronic equipment and a storage medium, and the method comprises the steps: dividing a virtual scene into a plurality of sound source logic distribution regions, associating each region with a group sound source control node, collecting sound source distribution data in real time according to the position of a virtual object, and dynamically adjusting node space audio parameters; and configuring an emotion envelope parameter based on a preset interaction event type, generating an emotion dynamic response parameter to drive execution of a layered sound effect behavior, and outputting and generating a dynamic sound field based on the spatial audio parameter, the emotion dynamic response parameter and the layered sound effect behavior. According to the invention, sound source partition and emotion envelope control can be realized, global emotion values are calculated based on event driving, layered sound effects are intelligently switched, the problems of single emotion and rigid positioning of a traditional system are solved, multi-dimensional dynamic adaptation of sound field accompanying and scenes is realized, and emotion level collaborative expression is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for dynamic sound field interactive control based on emotion perception. Background Art

[0002] In the field of sound field construction for dynamic interactive scenes, immersive sound field construction is one of the key technologies to enhance user experience. Real-time sound control systems must take into account the dynamic adaptation of the spatial distribution of sound sources and emotional responses to simulate group behavior characteristics in complex environments. For example, in virtual reality and racing simulation games, the dynamic realism of audience sound sources must simultaneously meet the real-time adaptation of the spatial distribution of the sound field (such as the location of the audience seats in the straightaway area and the finish area) and emotional responses (such as cheering when overtaking and exclaiming when the vehicles collide) to simulate the complexity of audience group behavior in real events.

[0003] Traditional sound source mapping technology is usually based on static or event-triggered point sound sources. For example, fixed point sound sources are pre-set in the audience area next to the track, or temporary sound sources are dynamically generated at the event location in response to specific game events (such as overtaking and vehicle collisions) (such as screams when vehicles collide). Or, based on the heat map algorithm of the track partition, the audience seats are divided into multiple logical areas (such as straight areas, curve areas, and finish areas). The activity weight of each area is calculated based on the real-time data of the game (such as vehicle position, overtaking events, and ranking changes) to optimize the sound field distribution. However, in the process of implementing the technology in dynamic interactive scenarios, it was found that it was still difficult to simultaneously meet the requirements of dynamic distribution of sound sources, real-time response to emotions, and collaborative control of group behavior in highly realistic scenarios. Summary of the Invention

[0004] The embodiments of the present application provide a dynamic sound field interactive control method, device, electronic device and storage medium to solve the technical problems of dynamic distribution of sound sources, real-time emotional response and group behavior coordination in high-fidelity scenes in related technologies.

[0005] In a first aspect, an embodiment of the present application provides a dynamic sound field interactive control method, the method comprising:

[0006] Divide the virtual scene into multiple sound source logical distribution areas, each sound source logical distribution area is associated with a group sound source control node, collect sound source distribution data within each sound source logical distribution area in real time according to the position of the virtual object, and dynamically adjust the spatial audio parameters of the group sound source control node;

[0007] According to a preset interaction event type, generating an emotional dynamic response parameter by dynamically configuring an emotional envelope parameter, and driving the group sound source control node to perform a layered sound effect behavior based on the emotional dynamic response parameter and the spatial audio parameter;

[0008] A dynamic sound field is generated based on the spatial audio parameters, the emotional dynamic response parameters and the output of the layered sound effect behavior.

[0009] In a second aspect, an embodiment of the present application provides a dynamic sound field interactive control device,

[0010] The device comprises:

[0011] A sound source logical partitioning module is used to divide the virtual scene into multiple sound source logical distribution areas, each of which is associated with a group sound source control node. The sound source distribution data within each sound source logical distribution area is collected in real time according to the position of the virtual object in the interaction path, and the spatial audio parameters of the group sound source control node are dynamically adjusted.

[0012] An emotion envelope parameter configuration module is used to generate emotion dynamic response parameters by dynamically configuring emotion envelope parameters according to a preset interaction event type;

[0013] A layered sound effect synthesis module is used to drive the group sound source control node to perform layered sound effect behavior based on the emotional dynamic response parameters and the spatial audio parameters, and to generate a dynamic sound field based on the spatial audio parameters, the emotional dynamic response parameters and the output of performing the layered sound effect behavior.

[0014] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory storing a plurality of instructions; a processor loading instructions from the memory to execute the steps of any one of the dynamic sound field interactive control methods provided in the embodiment of the present application.

[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a plurality of instructions suitable for loading by a processor to execute the steps of any dynamic sound field interactive control method provided in an embodiment of the present application.

[0016] By adopting the solution of the embodiment of the present application, the distribution data of interactive objects in the virtual scene can be collected in real time, combined with the spatial evolution of the logical distribution area of ​​the sound source and the emotional envelope control, to achieve accurate mapping of the dynamic sound field: the global emotional value is calculated based on the event-driven emotional value, and the intelligent switching of layered sound effects is driven, which solves the problems of single emotional response and rigid spatial positioning of traditional sound source systems, and realizes multi-dimensional dynamic adaptation of the sound field as the virtual object behavior and virtual scene changes, and the coordinated expression and dynamic interaction of multi-level emotional sound effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 This is a flow chart of an embodiment of a dynamic sound field interactive control method provided in an embodiment of the present application;

[0019] Figure 2 It is a schematic diagram showing the change in the direction of the manipulation perspective corresponding to the virtual object and the relative position of the interactive object in the virtual space;

[0020] Figure 3 This is a diagram of the audience emotional response curve of a dynamic sound field interactive control method provided in an embodiment of the present application. Figure 1 ;

[0021] Figure 4 This is a diagram of the audience emotional response curve of a dynamic sound field interactive control method provided in an embodiment of the present application. Figure 2 ;

[0022] Figure 5 Schematic diagram of an audio duration compression mechanism of a dynamic sound field interactive control method provided in an embodiment of the present application;

[0023] Figure 6 Schematic diagram of emotional state hierarchical control of a dynamic sound field interactive control method provided in an embodiment of the present application;

[0024] Figure 7 is a schematic diagram of a differentiated volume response curve of a dynamic sound field interactive control method provided in an embodiment of the present application;

[0025] Figure 8-A 1 is a schematic diagram of a mapping curve of speech fundamental frequency versus Doppler frequency shift in a dynamic sound field interactive control method provided in an embodiment of the present application;

[0026] Figure 8-B 1 is a schematic diagram of a response curve of voice volume changing with camera speed in a dynamic sound field interactive control method provided in an embodiment of the present application;

[0027] Figure 9 1 is a schematic diagram of parameter configuration of an emotional envelope in a dynamic sound field interactive control method provided in an embodiment of the present application;

[0028] Figure 10 Schematic diagram showing the configuration of sound effect rule parameters of a semantic feedback layer of a dynamic sound field interactive control method provided in an embodiment of the present application;

[0029] Figure 11 1 is a schematic diagram of a curve showing dynamic changes in audience emotion values ​​during a racing game in a racing game scene using a dynamic sound field interactive control method provided in an embodiment of the present application;

[0030] Figure 12 It is a display diagram of a blueprint type audience;

[0031] Figure 13 It is a display diagram of a patch type audience;

[0032] Figure 14 It is a display diagram of plant type audience;

[0033] Figure 15 Schematic diagram of a semantic feedback resource encapsulation structure of a dynamic sound field interactive control method provided in an embodiment of the present application;

[0034] Figure 16 This is a schematic diagram of multi-language dynamic dialogue resource configuration for a dynamic sound field interactive control method provided in an embodiment of the present application;

[0035] Figure 17 is a structural diagram of a dynamic sound field interactive control device provided in an embodiment of the present application;

[0036] Figure 18 It is a structural diagram of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0037] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application. At the same time, in the description of the embodiments of the present application, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0038] In the description of this application, the word "for example" is used to mean "used as an example, illustration or illustration". Any embodiment described in this application as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.

[0039] Before explaining the embodiments of the present application in detail, some terms involved in the embodiments of the present application are first explained.

[0040] Wwise (WaveWorks Interactive Sound Engine) is a professional audio engine and toolset used for audio design, integration, and management in games, virtual reality (VR), augmented reality (AR), and interactive media. Its core function is to provide developers with an efficient toolchain for dynamic control and optimization of complex audio logic.

[0041] RTPC (Real-Time Parameter Control) is a technology used in audio engines (such as Wwise) to dynamically control audio parameters. It maps real-time variables (such as in-game data, user behavior, and environmental conditions) to audio properties (such as volume, pitch, and reverberation intensity) to achieve real-time dynamic adjustment of sound effects.

[0042] ADSR (Attack-Decay-Sustain-Release), attack-decay-sustain-disappearance, is a four-stage model that describes the sound envelope. It is used to control the complete form of the sound from triggering to disappearing. Specifically, (1) Attack: the time it takes for the sound to reach its peak intensity from zero (such as the instantaneous peak of the audience's cheers); (2) Decay: the time it takes for the sound to decay from the peak to the sustained level (such as the volume of the audience's cheers falling back after reaching the peak); (3) Sustain: the stage of maintaining a stable volume (such as the audience's cheers maintaining a high volume); (4) Release: the time it takes for the sound to decay from the sustained level to zero (such as the audience's cheers gradually weakening and disappearing).

[0043] To facilitate understanding of the embodiment of the present application, a dynamic sound field interactive control method is provided, and the relevant application scenarios of the dynamic sound field interactive control method are first described. Specifically, the dynamic sound field interactive control method provided by the present application is based on emotion and spatial perception, and is applied to the field of sound field construction in dynamic interactive scenes. It can be specifically applied to virtual reality game scenes, virtual meetings, online education, smart homes, smart car systems, live broadcasts / performance interactions, medical / psychological treatment scenes, etc. Specifically, taking a simulated racing game as an example, in a real racing scene, the emotional fluctuations of the audience group will present dynamic envelope characteristics in the time dimension (such as cheers showing an ADSR change curve of Attack onset, Decay attenuation, Sustain continuation, and Release disappearance as the overtaking action occurs) and density gradient distribution in the spatial dimension (such as the audience on the curve side showing a higher sound pressure level and a denser sound source localization due to close observation of the overtaking action). This emotion-driven sound field dynamic characteristic is the core element in creating an immersive event atmosphere. To simulate the interactive experience of a real racing scene in racing games, traditional audience sound source systems in racing games often focus on constructing static point sound sources or event-triggered point sound sources. This is done by pre-setting fixed point sound sources in the audience seating area next to the track, or dynamically generating temporary sound sources at the event location when specific game events (such as overtaking or collisions) occur. For example, a temporary exclamation sound source with a radius of 5m is generated when a vehicle crashes, and the sound and image direction moves with the playback camera. Alternatively, based on a heat map algorithm for track zoning, the audience seating area is divided into multiple logical areas (such as straights, curves, and finish areas). The "activity weight" of each area is calculated based on real-time race data (such as vehicle position, overtaking events, and ranking changes). For example, when the player's vehicle approaches the finish straight, the volume of the audience in the finish area increases, and the sound and image positioning shifts to the front; when the vehicle drifts around a curve, the density of the audience cheering in the curve area increases, and the spatial audio engine adjusts the balance of the left and right channels to match the player's perspective, thereby optimizing the sound field distribution.

[0044] However, in actual racing game scenarios, since player vehicles are usually fast and change direction frequently, in high-speed racing scenes, when the player vehicle is at the boundary of multiple audience point sound source areas, there will be auditory discontinuity and lack of coherence. Based on this situation, the audience sound source system in traditional racing games is constructed in a static point sound source manner, which will have the technical problem of low coupling between the sound source distribution and the dynamic scene, making it difficult to adapt to the high-speed changes in vehicle position and perspective switching; traditional racing games use event-triggered point sound sources, and the binding method of sound source layering elements and track data lacks dynamic adaptability and lacks fine-grained control of the audience's emotional dynamic envelope (such as ADSR characteristics), resulting in a lack of coherence in the temporal changes of the sound field atmosphere. In addition, the game relies on discrete events to trigger sound effects, and the overall control of the audience's emotional trend as the event progresses is slightly monotonous, and the simulation of group behavior is weak.

[0045] Precisely to address the aforementioned technical issues, the present application provides a method, device, electronic device, and storage medium for dynamic sound field interaction control. By dynamically adjusting the spatial distribution of sound sources, implementing real-time parameter control based on emotional envelopes, and triggering layered sound effects, the present application addresses the technical issues of insufficient sound source adaptation, single emotional response, and lack of group behavior hierarchy in traditional sound field technologies. This enables highly realistic sound effects interaction in dynamic interactive scenarios (such as simulated racing games, virtual conferences, smart homes, and medical assistance), significantly enhancing the immersive experience of the sound field. This will be explained in detail below with reference to the embodiments.

[0046] Specifically, this embodiment will be described from the perspective of a dynamic sound field interactive control device. This dynamic sound field interactive control device can be integrated into an electronic device. That is, the dynamic sound field interactive control method of the embodiment of the present application can be executed by an electronic device. Optionally, the electronic device can include a terminal device. The terminal device can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, game console, or personal computer (PC).

[0047] The dynamic sound field interactive control method provided in the embodiments of the present application can be applied, for example, to a racing game system. The racing game system may include a player terminal device and a server. The terminal device may include both receiving and transmitting hardware, i.e., a device capable of performing two-way communication over a two-way communication link. The player terminal device and the server may communicate bidirectionally via a network.

[0048] Optionally, the server may be a standalone server, or a server network or server cluster consisting of servers, including but not limited to a computer, a network host, a single network server, a set of multiple network servers, or a cloud server consisting of multiple servers. A cloud server is composed of a large number of computers or network servers based on cloud computing.

[0049] In an embodiment of the present application, the interactive method applied to the game can be run on a local terminal device or a server. When the interactive method in the game is run on a server, the method can be implemented and executed based on a cloud interactive system, wherein the cloud interactive system includes a server and a client device.

[0050] In an optional embodiment, various cloud applications, such as cloud gaming, can be run under the cloud interaction system. Taking cloud gaming as an example, cloud gaming refers to a gaming method based on cloud computing. In the cloud gaming operating mode, the operating body of the game program and the main body presenting the game screen are separated. The storage and operation of the interactive methods within the game are completed on the cloud gaming server. The role of the client device is to receive and send data and present the game screen. For example, the client device can be a display device with data transmission capabilities close to the user, such as a mobile terminal, television, computer, PDA, etc.; however, it is the cloud gaming server in the cloud that performs information processing. When playing the game, the player operates the client device to send operation instructions to the cloud gaming server. The cloud gaming server runs the game according to the operation instructions, encodes and compresses the game screen and other data, and returns it to the client device via the network. Finally, the client device decodes and outputs the game screen.

[0051] In an optional embodiment, taking a game as an example, a local terminal device stores a game program and is used to present the game screen. The local terminal device is used to interact with the player through a graphical user interface, that is, conventionally downloading and installing the game program through an electronic device and running it. The local terminal device can provide the graphical user interface to the player in a variety of ways, for example, it can be rendered and displayed on the terminal's display screen, or provided to the player through holographic projection. For example, the local terminal device may include a display screen and a processor, the display screen is used to present the graphical user interface, the graphical user interface includes the game screen, and the processor is used to run the game, generate the graphical user interface, and control the display of the graphical user interface on the display screen.

[0052] The following is a detailed description of each step in conjunction with the accompanying drawings. In the embodiments of this application, the execution subject is a terminal device as an example. It should be noted that the order in which the following embodiments are described does not limit the preferred order of the embodiments. Although the flowcharts show a logical order, in some cases, the steps shown or described may be performed in a different order than that shown in the accompanying drawings.

[0053] The embodiment of the present application provides a dynamic sound field interaction control method that can enhance the high-fidelity sound effect interaction in game scenes.

[0054] Please refer to Figure 1 Taking a terminal as an example, the present application embodiment provides a dynamic sound field interactive control method. The specific process of the dynamic sound field interactive control method can be as follows: Steps S101 to 103, wherein:

[0055] Step S101: Divide the virtual scene into multiple sound source logical distribution areas, each sound source logical distribution area is associated with a group sound source control node, and the sound source distribution data in each sound source logical distribution area is collected in real time according to the position of the virtual object in the interaction path, and the spatial audio parameters of the group sound source control node are dynamically adjusted.

[0056] In some embodiments of the present application, a graphical user interface is provided by a terminal device, and the graphical user interface displays a target game screen, and the target game screen at least partially includes a virtual scene and a controlled virtual object located in the virtual scene.

[0057] The graphical user interface may be used to display the content of the target game by the player logging into the game account, and to manipulate the virtual objects in the virtual scene in the target game.

[0058] The virtual scene also includes interactive objects that interact with the virtual objects, and interactive paths of the virtual objects in the virtual scene. Based on the interactive object aggregation nodes generated by the aggregation of the interactive objects, the spatial distribution area in the virtual scene forms the sound source logical distribution area.

[0059] For example, the target game screen includes a virtual scene of a racing game, the virtual object is a racing car controlled by a player logged into their game account within the virtual scene, the interactive object is a spectator watching the race, and the interactive path is the racing track within the virtual scene. The spatial distribution area of ​​the spectator within the virtual scene, such as the auditorium area outside the track, forms the logical distribution area of ​​the sound source.

[0060] The virtual scene may also include a race scene of the surrounding environment of the track.

[0061] In some embodiments of the present application, a virtual scene is divided into a plurality of sound source logical distribution areas by a preset spatial rule, and each sound source logical distribution area is associated with a group sound source control node.

[0062] In some embodiments of the present application, the preset spatial rule is a four-quadrant spatial division rule, that is, the virtual scene is divided into four sound source logical distribution areas of left front, right front, left rear, and right rear according to the four quadrants of the XZ plane, and each sound source logical distribution area is associated with a group sound source control node. For example, in a virtual scene of a racing game, the environment around the track is divided into four quadrant logical areas (left front, right front, left rear, and right rear) according to the spatial position, and each area corresponds to a different block of the audience seats, wherein the left front quadrant corresponds to the audience seats in front of the left side of the vehicle, the right front quadrant corresponds to the high-density audience area outside the curve, and the left rear / right rear quadrants correspond to the audience gathering areas at the ends of the track respectively. The area boundaries are defined by the XZ plane coordinate system and stored in the scene configuration file.

[0063] In some embodiments of the present application, sound source distribution data within each sound source logical distribution area is collected in real time according to the position of the virtual object in the interaction path, and the spatial audio parameters of the group sound source control node are dynamically adjusted, wherein the sound source distribution data includes the aggregated position (such as average coordinates), density (such as the number of entities per unit area) and dynamic behavior parameters (such as emotional value, activity) of the interactive objects within the corresponding sound source logical distribution area.

[0064] Specifically, the reference points for dividing the logical distribution areas of sound sources are determined based on the real-time position of the virtual object in the interaction path, and the logical distribution areas of sound sources are dynamically adjusted based on the reference points; the sound source distribution data within each adjusted logical distribution area of ​​sound sources is collected, and the collected sound source distribution data is mapped to the real-time parameter control (RTPC) of the audio engine to drive the spatial access and density feedback of the group sound source control nodes. The density feedback includes volume and the number of sound effect superposition layers. For example, when a car approaches a turning area on the track, the azimuth focus and volume intensity of the sound source in the corresponding area are enhanced.

[0065] In some embodiments of the present application, before dividing the virtual scene into a plurality of logical distribution areas of sound sources, the method further includes: in the system startup phase, loading and triggering the basic loop layer of group sound effects (such as crowd background sound) through the initialization function to provide underlying support for the dynamic sound field interaction. For example, when a racing game is started, the basic loop layer of group sound effects is loaded and triggered through the initialization function, and the basic cheering loop sound effect of the audience seats is played. When the player's vehicle enters the track scene, the engine loads the audience seat model, sound effect resources and event configuration; after the resource loading is completed, the initialization function is called to trigger the basic loop layer of group sound effects, and start playing the basic audience cheers, and the volume is initialized to the default value (such as 50%).

[0066] In some embodiments of the present application, the sound source distribution data within each sound source logical distribution area is collected in real time according to the position of the virtual object in the interaction path, specifically:

[0067] Determining reference points for dividing logical distribution areas of sound sources according to interaction rules of virtual objects; wherein the interaction rules are dynamically set according to the interaction viewing angle state of the virtual objects;

[0068] Taking the reference point as the center, scan and record all interactive object nodes within a preset perception radius and store them in a node list; wherein the preset perception radius is used to limit the maximum number of the interactive object nodes;

[0069] According to the positions of all interactive object nodes in the node list, each interactive object node is divided into the corresponding sound source logical distribution area according to the sound source logical distribution area (such as four quadrants), and the number and average position of interactive objects in each sound source logical distribution area are counted.

[0070] In some embodiments of the present application, the interaction rule is to follow the virtual object according to the interaction perspective of the virtual object, and the reference point is bound to the center of mass coordinate system of the virtual object. For example, when the virtual object is a racing car, the interaction perspective of the virtual object is the driver's control perspective of the racing car, that is, the reference point is determined by following the driver's control perspective of the racing car.

[0071] For example, the system scans spectator location data for the four quadrants surrounding the player's vehicle, aggregates the audience distribution, and stores it separately for each sub-track. During the race, the system uses this audience data to calculate, in real time, the location of sound sources, crowd size, and sentiment in the four quadrants surrounding the player's vehicle. When the orientation of the player's vehicle (the Listener) changes, the location of the corresponding spectator sound source also changes.

[0072] See also Figure 2 , Figure 2This diagram illustrates the change in the virtual object's control perspective and the relative position of the interactive object in virtual space. For example, in a racing simulation game, the area surrounding the player's vehicle (virtual object) is divided into four quadrants in the XZ coordinate system (e.g., front left, front right, rear left, and rear right). Each quadrant corresponds to a logical sound source distribution area (e.g., areas 0, 1, 2, and 3), representing the audience distribution data for different sub-tracks. After scanning the audience positions, data is stored separately for each sub-track (e.g., area 1 corresponds to sub-track 1, recording parameters such as audience number and emotion). During the race, the sound source properties of the four surrounding quadrants are calculated in real time based on the player's vehicle's position, including sound source location, crowd size, and emotion data. When the player's vehicle's orientation (Listener) changes, the system remaps the sound source position. For example, if the player's vehicle turns 90° left, the "front left" quadrant becomes "front," and the sound source position must be updated in the coordinate system simultaneously. If the vehicle turns and causes a densely populated area of ​​​​audience in a quadrant (such as Z sub-track 3) to move to the rear of the vehicle, the system will reduce the volume of the sound source in that area, while enhancing the sound field effect of the new front quadrant to enhance the immersive feeling.

[0073] Such refined data partition storage and real-time calculation achieve a high degree of synchronization between vehicle movement and audience soundscape.

[0074] In some embodiments of the present application, the interaction rule is that if the interaction perspective of the virtual object does not follow the virtual object, the reference point is generated based on the global focus position. For example, the reference point is generated based on the global camera focus position and does not follow the control perspective of the player's vehicle.

[0075] In some embodiments of the present application, dynamically adjusting the spatial audio parameters of the group sound source control node specifically includes:

[0076] The corresponding weight coefficient is defined according to the sound field density scenario, and the corresponding weight coefficient calculation is performed on the number of interactive objects in the logical distribution area of ​​each sound source according to the corresponding sound field density scenario to obtain the modified density parameter of the effective group interactive objects.

[0077] The sound field density scenarios include low-density scenarios (such as car practice), medium-density scenarios (such as car qualifying), and high-density scenarios (such as official car races). For example, based on the definition of the event type mapping rules for a car racing game:

[0078] For racing practice sessions, the weight coefficient is <1 to reduce the audience density. Through weighted calculation of the weight coefficient, the total volume of cheers is reduced to simulate a sparse audience scene and highlight the sound of vehicle engines.

[0079] For car qualifying, the coefficient is ≈1, maintaining the original density; through weighted calculation of the weight coefficient, the sound field distribution is balanced to reflect the medium audience participation.

[0080] Formal car races: coefficient > 1, amplifying the audience density; through weighted calculation of the weight coefficient, the high-frequency cheering components are enhanced, and the pressure of the sound field increases the tension of the race.

[0081] In some embodiments of the present application, after collecting the sound source distribution data in each sound source logical distribution area in real time according to the position of the virtual object in the interaction path, the method further includes:

[0082] Calculating the distance between the average position of the interactive objects in the logical distribution area of ​​the sound source and the current position of the virtual object;

[0083] determining whether the distance exceeds a preset distance;

[0084] When the distance does not exceed the preset distance, correcting the average position of the group interaction objects using preset parameters to obtain corrected sound source distribution data;

[0085] When the distance exceeds the preset distance, the collected original sound source distribution data is retained.

[0086] The preset parameters include a fixed offset distance and an offset angle. Specifically, the fixed offset distance is offset from the reference point along the XZ plane toward the offset angle to generate corresponding corrected sound source position data.

[0087] For example, in a virtual scene of a racing game, when the average position of the audience in a certain quadrant is only 5 meters away from the vehicle (the preset distance is 8 meters), the system triggers the correction mechanism and generates an offset sound source in a 45° direction (offset angle) and 10 meters (offset distance) to avoid cheers coming from directly under the vehicle (which violates the laws of physics).

[0088] In some embodiments of the present application, after collecting the sound source distribution data in each sound source logical distribution area in real time according to the position of the virtual object in the interaction path, the method further includes:

[0089] According to the real-time change of the position of the virtual object in the interaction path, the first transition variable is used to update the sound source position within the logical distribution area of ​​each sound source in real time.

[0090] Specifically, a first transition variable is preset to represent the per-frame displacement of the group sound source control node from its current coordinates to its target coordinates. As the virtual object's position changes, the average coordinates of the group sound source control node are gradually updated based on the first transition variable. For example, when a vehicle turns left, the sound source originally in the "front left" quadrant needs to transition to the "front" position. The system gradually updates the coordinates using the first transition variable to avoid jumps in the sound image and eliminate the auditory discontinuity caused by sudden changes in the sound source's position.

[0091] In some embodiments of the present application, after collecting the sound source distribution data within the logical distribution area of ​​each sound source in real time according to the position of the virtual object in the interaction path, the average position of the group sound source control node is linearly transitioned from the current position to the target position using the second transition variable.

[0092] Specifically, the second transition variable is preset to represent the per-frame change in the number of group interaction objects from the current value to the target value. When the position of the virtual object changes, the number of group interaction objects is gradually updated according to the second transition variable. For example, when the number of spectators increases from 100 to 200, the volume of cheers increases according to a logarithmic curve to avoid the mechanical feeling caused by linear growth.

[0093] In some embodiments of the present application, before dynamically adjusting the spatial audio parameters of the group sound source control node, the method further includes:

[0094] By dynamically updating the world coordinates and orientation of interactive object nodes, the real-time collected interactive node distribution data is mapped to a 3D sound field engine (such as Wwise).

[0095] For example, by calling the graphics engine API, the final calculated sound source coordinates (x, y, z) and Euler angles (pitch, yaw, roll) are written into the transformation matrix of the spectator GameObject to complete the spatial mapping in the world coordinate system, ensuring that the sound source position is strictly synchronized with vehicle movement, spectator behavior, and event scenes, thereby achieving a high degree of consistency between the spatial orientation of the sound effect and visual perception.

[0096] In some embodiments of the present application, dynamically adjusting the spatial audio parameters of the group sound source control node specifically includes:

[0097] Mapping the number of interactive objects within each sound source logical distribution area to the group size parameter of the audio engine to drive dynamic adjustment of the volume of the group interactive objects;

[0098] Obtain group emotion values ​​(e.g., 0 to 100) from the global emotion manager, map them to emotion intensity parameters, and dynamically adjust global audio mixing parameters (e.g., volume balance, reverberation intensity, and high-frequency attenuation).

[0099] The real-time audio parameters (RTPC) of the group sound source control nodes within the logical distribution area of ​​each sound source are dynamically adjusted according to the movement speed of the manipulation perspective of the virtual object, driving the Doppler effect simulation and enhancing the frequency offset of the relative movement of the group sound source control nodes.

[0100] See also Figure 3 , Figure 3 Schematic diagram of audience emotional response curve of a dynamic sound field interactive control method provided in an embodiment of the present application Figure 1 .

[0101] In some embodiments of the present application, the number of interactive objects within each logical distribution area of ​​a sound source is mapped to the audio engine's crowd size parameter to drive dynamic adjustment of the volume of these group interactive objects. Specifically, the number of audience members in four quadrants (e.g., front left, front right, rear left, and rear right) is obtained through real-time scanning and stored as crowdSize_Q1 through crowdSize_Q4. Four independent RTPC parameters (e.g., CrowdSize_Q1 through CrowdSize_Q4) are created in the audio engine (e.g., Wwise), each corresponding to the number of audience members in a quadrant. The volume of preset audio events in Wwise (e.g., cheers) is bound to the RTPC parameters. When CrowdSize_Q{i} changes, the volume of the corresponding soundtrack is automatically adjusted.

[0102] By mapping the number of interactive objects in the logical distribution area of ​​each sound source to the group size parameter of the audio engine to drive the dynamic adjustment of the volume of the group interactive objects, the number of spectators in each quadrant independently drives the volume of its sound source to avoid the flattening of the sound field caused by global unified adjustment. For example, if the audience in the right front is dense (crowdSize_Q2=300), the volume of the right front sound source is significantly higher than that in the right rear (crowdSize_Q4=50). When the player's vehicle enters the curve, the number of people in the right front audience area surges (crowdSize_Q2 rises from 100 to 400), and the corresponding sound source volume increases from -15dB to +5dB, strengthening the sound field focus in the curve area.

[0103] See also Figures 4 to 7 ,in, Figure 4 Schematic diagram of audience emotional response curve of a dynamic sound field interactive control method provided in an embodiment of the present application Figure 2 ; Figure 5 Schematic diagram of an audio duration compression mechanism of a dynamic sound field interactive control method provided in an embodiment of the present application; Figure 6 Schematic diagram of emotional state hierarchical control of a dynamic sound field interactive control method provided in an embodiment of the present application; Figure 7 Schematic diagram of a differentiated volume response curve of a dynamic sound field interactive control method provided in an embodiment of the present application.

[0104] In some embodiments of the present application, mapping the number of interactive objects within each logical distribution area of ​​a sound source to a group size parameter of an audio engine to drive dynamic adjustment of the volume of the group interactive objects includes:

[0105] Adjusting the basic volume gain of each group sound source according to a preset proportional coefficient based on the group size parameter; and

[0106] When the group size exceeds the threshold, the sound source position random offset algorithm is activated to simulate the difference in group spatial distribution density.

[0107] In some embodiments of the present application, group emotion values ​​are obtained from a global emotion manager, mapped to emotion intensity parameters, and global audio mixing parameters are dynamically adjusted; wherein, the dynamic adjustment of global audio mixing parameters performs three-level response control, including audio duration compression, emotional state layered control, and layered volume dynamic response.

[0108] For example, the current global audience emotion value is read from the global emotion manager (CrowdExcitementManager). This value is dynamically calculated based on the following factors:

[0109] Player behavior (e.g. drifting, overtaking, collision);

[0110] The stage of the race (e.g., last lap, near the finish line);

[0111] Audience group emotional envelope (ADSR envelope controls the intensity and duration of sound effect attenuation).

[0112] Register a global RTPC parameter in Wwise, with a range of 0 to 100, corresponding to the emotion value. Dynamically control the audio properties in Wwise using the RTPC parameter value:

[0113] (1) Audio duration compression (Duration)

[0114] Define the inverse mapping relationship between the emotion value (0-100) and the audio trigger delay duration (Duration):

[0115] Emotion value = 0 → Duration = 2.0 seconds (full delay, such as a long brewing sound effect)

[0116] Emotion value = 70 → Duration = 1.5 seconds (partially compressed)

[0117] Emotion value = 100 → Duration = 0 seconds (zero delay instant response)

[0118] (2) Hierarchical control of emotional states

[0119] Establish an emotional state machine and divide it into three levels of state threshold intervals:

[0120] bored: 0-30, enables low-frequency dynamic compression (<200Hz rolloff)

[0121] excited: 31-70, activates the reverb intensity gradient (0.7→1.0)

[0122] very_excited: 71-100, triggers exclusive sound layer (crowd_group_veryexcited_oneshot)

[0123] (3) Layered volume dynamic response

[0124] Bind independent gain curves to each emotional state:

[0125] very_excited state layer: When the emotion value is 0→100, the volume increases nonlinearly from -20dB to +2dB (breaking the conventional upper limit of 0dB);

[0126] Excited state layer: Use a smooth rising curve (-30dB→0dB)

[0127] For example, when the player's car sprints for the last lap (excitement 70→95), trigger:

[0128] Duration is compressed from 1.5 seconds to 0 seconds (instantly releasing the very_excited layer sound effect);

[0129] Volume increased from -6dB to 0dB;

[0130] The reverberation intensity coefficient increases from 0.7 to 1.0 to simulate the resonance of the entire venue.

[0131] By mapping the audience's emotional values ​​to global mixing parameters, the dynamic adjustment of audio properties is directly driven by changes in emotional values. A single parameter controls multi-dimensional mixing, simplifying the logic complexity. Through fine-tuning of acoustic characteristics (volume, frequency band, reverberation), players' perception of the event atmosphere is deepened. The Duration zero-delay mechanism avoids the problem of ADSR envelope response lag at emotional peaks (in high-real-time scenarios such as game sprints).

[0132] Please also see Figure 8-A and Figure 8-B ,in, Figure 8-A This is a schematic diagram of a mapping curve of voice fundamental frequency versus Doppler shift in a dynamic sound field interactive control method provided in an embodiment of the present application. The horizontal axis represents Doppler shift (doppler) in the range of -200 to 200; the vertical axis represents voice fundamental frequency (voice pitch) in the range of -4800 to 4800. Figure 8-B This is a schematic diagram of a response curve of voice volume changing with camera speed in a dynamic sound field interactive control method provided in an embodiment of the present application. The horizontal axis represents the camera speed camera_speed, ranging from 0 to 150 units; the vertical axis represents the voice volume voice Volume, ranging from -200 to 200.

[0133] For example, obtain the current camera velocity vector through the game engine API, register the camera velocity parameter in Wwise, and bind it to the Doppler effect control interface of the audio event. In Wwise, enable the Doppler effect plug-in for the audience sound source and bind the camera velocity parameter. Adjust high-frequency attenuation and low-frequency enhancement based on the camera velocity to simulate the sound field compression effect caused by high-speed movement.

[0134] For example, when the player's vehicle approaches a curve, the system scans the audience nodes in the surrounding seats, calculates the intensity of cheers in each quadrant, and adjusts the Doppler effect according to the vehicle speed to create a sense of presence when turning at high speed.

[0135] By manipulating the RTPC parameters driven by the viewpoint (camera) speed through virtual objects, dynamic adaptation of the Doppler effect of the four-quadrant audience sound source is achieved. The Doppler effect simulates the frequency change of the sound source when the camera moves at high speed (for example, the pitch increases when approaching the sound source and decreases when moving away). The four quadrants respond independently to the camera speed, enhancing the spatial hierarchy of the sound field and the realism of movement (for example, when the camera moves to the left, the Doppler effect of the sound source on the left is more significant).

[0136] In some embodiments of the present application, the spatial audio parameters of the group sound source control node are dynamically adjusted, and the following is also included: prioritizing the interactive object nodes in specific areas (such as stands) to ensure that the sound effect resources of key sound sources (such as close-range groups) are loaded and played first. By marking high-priority audience nodes (such as VIP seats and finish line audiences), their sound effect resources are ensured to be loaded first; for example, when the player's vehicle approaches the finish line, the cheers of the finish line audience seats are independent of other areas, and the volume and clarity are significantly enhanced.

[0137] In some embodiments of the present application, dynamically adjusting the spatial audio parameters of the group sound source control node also includes: when an abnormal event (such as a collision or exceeding a limit) is detected, triggering corresponding negative feedback sound effects (such as exclamations or alarms), and driving global mixing adjustments through markers. For example, when a vehicle crashes (an abnormal state), the system immediately triggers an exclamation (CallOut sound effect) from the audience and reduces the volume of background looping sound effects to highlight the event feedback.

[0138] Step S102 : generating emotional dynamic response parameters by dynamically configuring emotional envelope parameters according to a preset interactive event type, and driving the group sound source control node to execute layered sound effect behavior based on the emotional dynamic response parameters and the spatial audio parameters.

[0139] The preset interaction event types include user-specific behaviors and system status changes, such as racing drift, overtaking, and achievement unlocking. The sound field includes sound effects and positions.

[0140] In some embodiments of the present application, the emotional dynamic response parameters are configured for each type of interactive event through the emotional envelope parameters, including:

[0141] Attack rate: The rate at which the emotion value rises when an event is triggered (e.g., a drift event ATK = 1.5 / sec means that the excitement value increases by 1.5 points per second);

[0142] Release: The rate at which the emotion value transitions to the sustained value after the Attack phase ends (e.g., Decay = 2 / sec for a drift event).

[0143] Sustain: The stable level of the emotion value during the event (e.g., Sustain = 70 for a drift event);

[0144] Release rate: the rate at which the emotion value decreases after the event ends (e.g., RLS = 8 / sec for a drift event);

[0145] Sentiment Contribution (Cntrb): The baseline impact of an event on the overall sentiment value. Positive events enhance it, while negative events suppress it (e.g., drift contribution +70 points, static contribution -50 points).

[0146] In some embodiments of the present application, according to a preset interaction event type, the emotional dynamic response parameters are generated by dynamically configuring the emotional envelope parameters, including:

[0147] Manage the envelope lifecycle according to the preset interaction event type and calculate the real-time envelope contribution value of each interaction event;

[0148] The global emotion value is calculated by integrating the envelope contribution values ​​of all interaction events as the emotion dynamic response parameter.

[0149] In some embodiments of the present application, driving the group sound source control node to perform a layered sound effect behavior includes:

[0150] selecting a sound effect level and volume based on the spatial audio parameters;

[0151] The corresponding level of sound effects is triggered by the global emotion value and the emotion value threshold to perform layered sound effect behavior; wherein the emotion value threshold is the minimum emotion value that triggers the corresponding level of feedback sound effects.

[0152] In some embodiments of the present application, the step of performing a layered sound effect behavior includes:

[0153] Continuous emotional preparation sound effect level selection is performed according to the spatial audio parameters.

[0154] Convert the global emotion value into a sound effect trigger probability to trigger the localized semantic feedback sound effect as the output for executing the layered sound effect behavior;

[0155] A transient emotional outburst response is performed according to the spatial audio parameter and the global emotional value.

[0156] The envelope processing logic based on the emotion envelope parameters includes:

[0157] Attack phase: When an event is triggered, the emotion value rises according to the attack slope to the upper limit of the contribution. For example, after the drift event is triggered, the excitement level rises at a rate of 1.5 / sec to 70 points.

[0158] Sustain phase: During the duration of the event (if the drift has not ended), the sentiment value maintains the upper limit of the contribution;

[0159] Release phase: After the event ends, the emotion value decreases to the baseline according to the Release slope; for example, after the drift event ends, the excitement decreases at a rate of 8 / sec;

[0160] Termination condition: At the end of the game, all envelopes are forced to be released and the situation value returns to zero.

[0161] See also Figure 9 , is a schematic diagram of parameter configuration of the emotional envelope in a dynamic sound field interactive control method provided in an embodiment of the present application. Figure 9 The preset parameters map the relationship between event events and emotional responses. For example, during a racing game, when a series of parameters such as goingWrongway(…), stayingStill(…), airTime(…), and drifting(…) are triggered, the system enters the envelope Attack state if the conditions are met, otherwise it does not enter the Attack state.

[0162] In some embodiments of the present application, the global emotion value is calculated by integrating the envelope contribution values ​​of all interaction events using the following formula (1).

[0163] Global emotion value = default emotion value + ∑(each envelope contribution value) Formula (1)

[0164] The default emotion value is the initial value of the sound source control node.

[0165] In some embodiments of the present application, the global emotion value is converted into a sound effect triggering probability by the following formula (2).

[0166] P 触发 =(Current emotion value - emotion value threshold) / (100 - emotion value threshold)×maximum probability formula (2)

[0167] Among them, the current emotion value is the global emotion value, the emotion value threshold is the minimum emotion requirement for triggering the corresponding category Callout sound effect, and the maximum probability is the highest trigger probability of this type of sound effect when the emotion value reaches 100.

[0168] See also Figure 10 , Figure 10 This is a schematic diagram showing the configuration of the sound effect rule parameters of the semantic feedback layer of a dynamic sound field interactive control method provided in an embodiment of the present application.

[0169] For example, if the current excitement level is 50, the excitement threshold corresponding to the drift event is 25, and the maximum probability is 80%, then the probability of triggering the callout sound effect is:

[0170] P=(50-25)-(100-25)×80%=26.67%

[0171] If the random number is ≤ 26.67%, a callout sound effect is triggered in the quadrant where the drift event is located (e.g., front left), and the voice message "Drifting is so cool!" is played.

[0172] In some embodiments of the present application, after step S101, i.e., dividing the virtual scene into a plurality of logical sound source distribution areas, each logical sound source distribution area being associated with a group sound source control node, collecting sound source distribution data within each logical sound source distribution area in real time based on the position of a virtual object in an interaction path, and dynamically adjusting the spatial audio parameters of the group sound source control node, the method further includes:

[0173] According to the key interaction event type and the spatial audio parameters of the group sound source control node, the group sound source control node is driven to perform a layered sound effect behavior.

[0174] Specifically, the key interaction event is predefined by the system and is different from the preset interaction event, such as a car crossing the finish line or a car crash. The key interaction event is pre-assigned an event identifier and is bound to a corresponding audio file.

[0175] When a key event is detected (such as a vehicle crossing the finish line), the system immediately generates a sound effect trigger instruction without relying on the global emotion value, and directly calls the sound effect corresponding to the key interaction event. The sound effect playback position is determined according to the spatial quadrant in which the key interaction event occurs (such as the finish line is located on the finish straight, corresponding to the "front right" quadrant); the coordinate information of the reference point (from step S101) is called to drive the 3D sound and image positioning of the audio engine.

[0176] For example, in a virtual scene of a racing game, when a vehicle collision is detected, the following actions are triggered immediately:

[0177] Negative emotional feedback: Reduce global excitement (e.g., from 70 to 30) and reduce the volume of cheers;

[0178] Instantaneous sound effect triggering: playing the audience's exclamation (OneShot sound effect), and the sound effect is positioned in the quadrant closest to the collision point (through the spatial audio parameters of the group sound source control node determined in step S101).

[0179] For key events in virtual scenes that require instant feedback (such as crossing the finish line, crashing, unlocking achievements, etc.), the probability model based on emotional values ​​is bypassed and the corresponding semantic sound effects (such as cheering, exclamation, and prompt voice) are directly triggered to ensure that users receive deterministic and high-priority auditory feedback at critical moments, enhancing the immediacy and immersion experience in the virtual scene.

[0180] Traditional heat map algorithms only statically partition audience activity. The present invention uses an event-driven envelope accumulation mechanism and combines multiple parameters (such as drift duration and remaining game time) to calculate the global emotion value in real time, and triggers the callout sound effect through a probability model, which solves the problem of a single emotional response mechanism and realizes refined timing control of the sound field atmosphere.

[0181] Step S103: generating a dynamic sound field based on the spatial audio parameters, the emotional dynamic response parameters, and the output of the layered sound effect behavior.

[0182] The layered sound effect definition includes:

[0183] Base Loop: This layer serves as the foundational sound effect layer. It consists of continuous group vocals (e.g., crowd whispers, humming), or continuous audience behavior (e.g., prolonged background applause, cheering noise), setting the overall emotional tone (e.g., calm chatter, moderate anticipation, or passionate excitement).

[0184] Event-driven instantaneous response layer (OneShot): This layer responds to specific events on stage / venue (e.g., exciting performance moments, unexpected turns), with the audience providing short, immediate feedback, including sudden bursts of applause, cheers, exclamations, whistles, etc.

[0185] Semantic feedback layer (CallOut): The audience emits group voice or sound feedback with clear semantics or strong emotional orientation, such as slogans (such as "Go Team XX!"), collective responses to the host's call-and-response interaction, neat cheers (such as "Bravo!", "Encore!"), and collective exclamations expressing strong emotions (such as "Wow!", "Oh!"), etc.

[0186] Specifically, based on the emotion value threshold and probability formula, the CallOut sound effect is triggered, and the sound effect resources of the Loop, CallOut and OneShot mechanisms are called in layers to simulate the group behavior of the interactive objects, such as:

[0187] The loop mechanism switches the sound effect layer based on the number of audience members (e.g. small: 5-29 people, group: 30-749 people), and selects positive or negative loop sound effects based on the emotion value;

[0188] The CallOut mechanism triggers localized language feedback (such as "Go!" or "Oops!") based on excitement thresholds, enhancing the regional characteristics of the event;

[0189] The OneShot mechanism triggers instantaneous sound effects during specific events (such as drifting and crashing), staggering the frequency band of the looped sound effects to avoid interference.

[0190] In some embodiments of the present application, the emotional response of the Loop mechanism specifically includes: switching the basic loop sound effect level according to the distribution data of the logical distribution area of ​​each sound source (such as audience density and position), the global emotional value (calculated by comprehensively considering the contribution value of all interactive event envelopes), and the interactive event type. For example, the calm layer (with a population range of 0 to 30 people) corresponds to low-volume background noise (such as crowd whispers), the medium layer (with a population range of 30 to 70 people) corresponds to standard cheers and applause, and the warm layer (with a population range of 70 to 100 people) corresponds to high-frequency enhanced intensive cheers and musical instrument sounds, dynamically adjusting the loop sound effect level and volume, and outputting spatial sound image positioning that matches the distribution of interactive objects (such as increasing the volume of sound sources in the high-density area in the right front quadrant).

[0191] In some embodiments of the present application, the emotional response of the callout mechanism specifically includes: calculating the probability of triggering the sound effect call based on the distribution data of the logical distribution area of ​​each sound source, the global emotion value and emotion value threshold, and the type of interactive event, and triggering the corresponding callout sound effect according to the trigger probability, as well as a short-term emotion value increase (such as excitement +5 points after triggering). For example, loading localized voice resources according to the type of interactive event (such as the Chinese "Overtaking Successfully!"); the playback position is determined by the quadrant where the event is located (such as left front drift triggering the playback of the left front sound source).

[0192] In some embodiments of the present application, the emotional response of the Oneshot mechanism specifically includes: directly triggering a short sound effect (such as playing a "scream" during a car crash) according to the distribution data of the logical distribution area of ​​each sound source and the key event trigger signal, and selecting the sound effect level according to the distribution data of the logical distribution area of ​​the sound source (such as the audience density), which can be that the small layer (with a number of people ranging from 5 to 29 people) selects a short single sound effect, and the group layer (with a number of people ranging from 30 to 749 people) selects a multi-part superimposed sound effect, so as to instantly play the corresponding short-effect sound feedback, and the sound and image positioning matches the quadrant where the event occurs.

[0193] In some embodiments of the present application, the layered sound effects also include: a secondary noisemaker layer, independent of the main loop layer, which dynamically loads resources based on the actual interaction type to be superimposed on the loop layer sound effects, such as adding distinctive sound effects (horn sounds, whistles) to enhance regional cultural characteristics (such as the vuvuzela on the South African racetrack). Based on the distribution data of the logical distribution area of ​​each sound source and the key event trigger signal, the secondary noisemaker layer is called and controlled by RTPC to load sound effects.

[0194] Based on the dynamic adjustment of the spatial audio parameters and emotional dynamic response output of the group sound source control node, the group behavior of interactive objects is dynamically simulated through layered sound effects, and a dynamic sound field based on emotional fluctuations, spatial orientation and semantic feedback is constructed. Specifically, the sound effect level and volume are selected based on the dynamic adjustment of the spatial audio parameters of the group sound source control node (such as audience distribution and density data), and the sound source position is dynamically adjusted to match the control perspective of the virtual object (such as the sound and image offset when the vehicle turns to match the racing player's perspective). Furthermore, the corresponding level of sound effect triggering is driven by the global emotion value and emotion value threshold, and the probability model is covered in the direct trigger (OneShot / CallOut) to ensure the immediacy of key feedback.

[0195] For example, when a car drifts through a corner (positive emotional feedback), it is detected that there are dense spectators in the left front quadrant (group layer), and the excitement level is calculated to rise to 60 (drift event contribution +70). The CallOut mechanism triggers the left front sound source to play "Beautiful drift!" according to probability, and the OneShot mechanism superimposes a high-frequency whistle (noisemaker layer) to enhance the atmosphere.

[0196] See also Figure 11The following is a schematic diagram of the dynamic change curve of the audience's emotional level during a racing game in a racing game scenario, using a dynamic sound field interactive control method provided by an embodiment of the present application. The horizontal axis represents time, and the vertical axis represents the emotional level (0-100). The following describes the starting phase: the emotional level is initially at a medium level (e.g., 50); the elimination countdown triggers: the player enters the danger zone, causing the emotional level to fluctuate; and the impact of the first elimination countdown (key event marker) on the emotional level. Figure 2 The emotional curve visually demonstrates how this design drives the complete evolution of the sound field from calm to warm, from lows to highs, providing players with a coherent and layered immersive experience.

[0197] Based on the emotional value threshold and probability model, the basic loop sound effect (Loop), event-driven instantaneous sound effect (OneShot) and semantic feedback sound effect (CallOut) are triggered in a coordinated manner, which solves the problem of insufficient hierarchy in group behavior simulation and realizes the coordinated expression and dynamic interaction of multi-level sound effects of emotions.

[0198] In some embodiments of the present application, before step S101, i.e., dividing the virtual scene into a plurality of logical sound source distribution areas, each logical sound source distribution area being associated with a group sound source control node, collecting sound source distribution data within each logical sound source distribution area in real time based on the position of the virtual object in the interaction path, and dynamically adjusting the spatial audio parameters of the group sound source control node, the method further includes:

[0199] Identify the types of interactive objects in the virtual scene, aggregate the identified interactive objects to generate corresponding interactive object aggregation nodes, and configure multimodal resources for the interactive object aggregation nodes.

[0200] The multimodal resources include at least a language package and a sound effect library associated with the scene configuration file. The interactive object type is a classification of interactive objects in the virtual scene (such as static models, dynamic agents, and environmental elements).

[0201] In some embodiments of the present application, the interactive object types in the virtual scene are identified by an automated tool, and the interactive object types include audience characters in the virtual scene of a racing game, specifically including:

[0202] Blueprint type audience: Define audience behavior logic (such as applauding, standing) through custom blueprints, such as Figure 12 As shown;

[0203] Patch type audience: Use static mesh to simulate efficient rendering of audience in the stands, such as Figure 13 As shown;

[0204] Plant type audience: audience models that are refreshed in batches through the vegetation system (such as virtual crowds on the grass), such as Figure 14 shown.

[0205] In some embodiments of the present application, the identified interactive objects are aggregated to generate corresponding interactive object aggregation nodes, specifically: the identified interactive objects are aggregated into regular units divided according to the interactive path or logical area to form corresponding interactive object aggregation nodes; wherein each node contains the number and location information of the interactive objects.

[0206] For example, the spectators on both sides of the track are aggregated according to a 10m×10m grid to generate cylindrical marks (the height represents the number of spectators).

[0207] In some embodiments of the present application, the interactive objects identified by aggregation are adjusted by dynamically configuring parameters to generate corresponding interactive object aggregation nodes; wherein the parameters include a spacing threshold that defines the minimum interval of the interactive object distribution, a density threshold that defines the maximum density of the interactive object distribution, and a rule unit for dividing the interactive path or logical area.

[0208] For example, adjusting the grid size used to divide the track can influence the distribution of audience cluster nodes (e.g., an overly large grid can cause spectators to cluster in the center of the track); adjusting the audience spacing threshold (e.g., 70 cm) can determine whether to merge into the same cluster node; and adjusting the density parameter of the patch audience can control the spacing of the spectator models in the stands. In a virtual practice race, increasing the grid size (Grid Cell = 200 cm) and reducing the number of cluster nodes can reduce computational overhead. In a virtual race, reducing the grid size (Grid Cell = 50 cm) can improve audience distribution accuracy and enhance sound field detail.

[0209] In some embodiments of the present application, the multimodal resources also include a multilingual resource library. In a racing simulation game, the Callout mechanism requires the audience to issue short sentences with localized language characteristics (such as "Come on!" in Chinese and "Go!" in English) based on the player's driving behavior (such as drifting, overtaking, and crashing) to enhance the immersion in the regional culture of the track. Multimodal resources are configured for the interactive object aggregation node. Specifically, Callout audio resources in different languages ​​are independently packaged as Wwise Bank files, and independent audio resource paths and trigger logic are defined for each language; Figure 15 As shown in the figure, specify the corresponding language bank to be loaded in the track configuration file to ensure that different tracks (such as loading Chinese for the Shanghai track and French for the France track) call the correct localized voice. Configure multilingual Dynamic Dialogue resources through Wwise and bind them to the track data. Specify the loaded Wwise voice bank (Bank) in the track data file, such as "English_Callouts.bank". Figure 16As shown in the figure, the trigger logic is that when the excitement envelope exceeds a threshold, the system randomly selects a callout audio from the corresponding bank based on the language configured for the current track. For example, when drifting on the Shanghai track, the Chinese callout "Beautiful corner!" is triggered, while on the Monza track, the Italian callout "Fantastico!" is triggered.

[0210] In this way, by identifying the types of interactive objects in the virtual scene, extracting the location, density and distribution data, generating the corresponding interactive object aggregation nodes, mapping the interactive object distribution to computable aggregation nodes (such as the height of a cylinder corresponding to the number of people), and constructing the initial distribution data of the interactive objects in the virtual scene by dynamically configuring parameters and binding multimodal resources.

[0211] This embodiment also provides a dynamic sound field interactive control device, which can be integrated into a terminal device. Figure 17 As shown, the dynamic sound field interactive control device may include:

[0212] The sound source logical partitioning module 201 is used to divide the virtual scene into multiple sound source logical distribution areas, each of which is associated with a group sound source control node. The sound source distribution data within each sound source logical distribution area is collected in real time according to the position of the virtual object in the interaction path, and the spatial audio parameters of the group sound source control node are dynamically adjusted.

[0213] The emotion envelope parameter configuration module 202 is used to generate emotion dynamic response parameters by dynamically configuring emotion envelope parameters according to a preset interaction event type;

[0214] The layered sound effect synthesis module 203 is used to drive the group sound source control node to perform layered sound effect behavior based on the emotional dynamic response parameters and the spatial audio parameters of the group sound source control node, and to generate a dynamic sound field based on the output of the spatial audio parameters, emotional dynamic response parameters and layered sound effect behavior of the sound source control node.

[0215] In some embodiments of the present application, the sound source logical partitioning module 201 divides the virtual scene into multiple sound source logical distribution areas according to preset spatial rules.

[0216] In some embodiments of the present application, the preset space rule is a four-quadrant space division rule.

[0217] In some embodiments of the present application, the sound source logical partitioning module 201 is used to:

[0218] Determining a reference point for dividing a sound source logical distribution area according to the real-time position of the virtual object in the interaction path, and dynamically adjusting the sound source logical distribution area based on the reference point;

[0219] Collect the sound source distribution data within the adjusted logical distribution area of ​​each sound source, and map the collected sound source distribution data to the real-time parameter control of the audio engine to drive the spatial access and density feedback of the group sound source control node.

[0220] In some embodiments of the present application, the sound source logical partitioning module 201 is used to:

[0221] Determining reference points for dividing logical distribution areas of sound sources according to interaction rules of virtual objects; wherein the interaction rules are dynamically set according to the interaction viewing angle state of the virtual objects;

[0222] Taking the reference point as the center, scan and record all interactive object nodes within a preset perception radius and store them in a node list; wherein the preset perception radius is used to limit the maximum number of the interactive object nodes;

[0223] According to the positions of all interactive object nodes in the node list, each interactive object node is divided into a corresponding sound source logical distribution area, and the number and average position of interactive objects in each sound source logical distribution area are counted.

[0224] In some embodiments of the present application, the interaction rule is to follow the virtual object according to the interaction perspective of the virtual object, and the reference point is bound to the center of mass coordinate system of the virtual object.

[0225] In some embodiments of the present application, the interaction rule is that the interaction perspective of the virtual object does not follow the virtual object, and the reference point is generated based on a global focus position.

[0226] In some embodiments of the present application, the sound source logical partitioning module 201 is used to:

[0227] The corresponding weight coefficient is defined according to the sound field density scenario, and the corresponding weight coefficient calculation is performed on the number of interactive objects in the logical distribution area of ​​each sound source according to the corresponding sound field density scenario to obtain the modified density parameter of the effective group interactive objects.

[0228] In some embodiments of the present application, the sound source logical partitioning module 201 is used to:

[0229] Calculating the distance between the average position of the interactive objects in the logical distribution area of ​​the sound source and the current position of the virtual object;

[0230] When it is determined that the distance does not exceed the preset distance, the average position of the group interaction objects is corrected using preset parameters to obtain corrected sound source distribution data; wherein the preset parameters include a fixed offset distance and an offset angle;

[0231] When it is determined that the distance exceeds the preset distance, the collected original sound source distribution data is retained.

[0232] In some embodiments of the present application, the sound source logical partitioning module 201 is used to:

[0233] According to the real-time changes in the position of the virtual object in the interaction path, the sound source position within the logical distribution area of ​​each sound source is updated in real time using the first transition variable; wherein, the first transition variable is preset to represent the displacement of the group sound source control node from the current coordinate to the target coordinate per frame.

[0234] In some embodiments of the present application, the sound source logical partitioning module 201 is used to:

[0235] The average position of the group sound source control node is linearly transitioned from the current position to the target position using a second transition variable; wherein the second transition variable is preset to represent the per-frame change in the number of group interaction objects from the current value to the target value.

[0236] In some embodiments of the present application, the emotion envelope parameter configuration module 202 is configured to:

[0237] Mapping the number of interactive objects within each sound source logical distribution area to the group size parameter of the audio engine to drive dynamic adjustment of the volume of the group interactive objects;

[0238] Obtain group emotion values ​​from the global emotion manager, map them to emotion intensity parameters, and dynamically adjust global audio mixing parameters;

[0239] The real-time audio parameters of the group sound source control nodes within the logical distribution area of ​​each sound source are dynamically adjusted according to the movement speed of the manipulation perspective of the virtual object, driving the Doppler effect simulation and enhancing the frequency offset of the relative movement of the group sound source control nodes.

[0240] In some embodiments of the present application, the emotion envelope parameter configuration module 202 is configured to:

[0241] Prioritize interactive object nodes in a specific area and load sound effect resources by marking priority.

[0242] In some embodiments of the present application, the emotion envelope parameter configuration module 202 is configured to:

[0243] Managing the envelope lifecycle based on the emotional dynamic response parameters to calculate the real-time envelope contribution value of each interactive event, and calculating the global emotional value by integrating the envelope contribution values ​​of all interactive events;

[0244] The global emotion value is converted into a sound effect trigger probability to trigger a semantic feedback sound effect, and the sound field change of the virtual scene is driven according to the spatial audio parameters of the group sound source control node.

[0245] In some embodiments of the present application, the emotion envelope parameter configuration module 202 is used to calculate the global emotion value by integrating all interaction event envelope contribution values ​​through formula (1);

[0246] Global emotion value = default emotion value + ∑(each envelope contribution value) Formula (1)

[0247] The default emotion value is the initial value of each interactive object node.

[0248] In some embodiments of the present application, the emotion envelope parameter configuration module 202 is used to convert the global emotion value into a sound effect triggering probability by using formula (2);

[0249] P 触发 =(Current emotion value - emotion value threshold) / (100 - emotion value threshold)×maximum probability formula (2)

[0250] Among them, the current emotion value is the global emotion value, the emotion value threshold is the minimum emotion requirement for triggering the semantic feedback sound effect of the corresponding category, and the maximum probability is the highest trigger probability of this type of sound effect when the emotion value reaches 100.

[0251] In some embodiments of the present application, the emotion envelope parameter configuration module 202 is configured to:

[0252] According to the key interaction event type and the spatial audio parameters of the group sound source control node, the group sound source control node is driven to perform layered sound effect behavior; wherein, the key interaction event is pre-defined by the system and is different from the preset interaction event.

[0253] In some embodiments of the present application, the emotion envelope parameter configuration module 202 is configured to:

[0254] The identified interactive objects are aggregated into regular units divided according to the interactive paths or logical areas to form corresponding interactive object aggregation nodes; wherein each node contains the number and location information of the interactive objects.

[0255] In some embodiments of the present application, the layered sound effect synthesis module 203 is used to: generate corresponding interactive object aggregation nodes by adjusting the interactive objects identified by aggregation through dynamic configuration parameters; wherein the parameters include a spacing threshold that defines the minimum interval of the interactive object distribution, a density threshold that defines the maximum density of the interactive object distribution, and a rule unit for dividing the interactive path or logical area.

[0256] Accordingly, an embodiment of the present application further provides an electronic device, which may be a terminal, such as a smartphone, a tablet computer, a laptop computer, a touch screen, a game console, a personal computer (PC), a personal digital assistant (PDA), or the like. Alternatively, the electronic device may be a server.

[0257] like Figure 18 As shown, Figure 18 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 is electrically connected to the memory 1102. Those skilled in the art will appreciate that the electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0258] The processor 1101 is the control center of the electronic device 1100. It connects the various parts of the entire electronic device 1100 using various interfaces and lines. By running or loading software programs and / or units stored in the memory 1102 and calling data stored in the memory 1102, it executes various functions of the electronic device 1100 and processes data, thereby monitoring the electronic device 1100 as a whole. The processor 1101 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0259] In the embodiment of the present application, the processor 1101 in the electronic device 1100 loads instructions corresponding to one or more application processes into the memory 1102 according to the following steps, and the processor 1101 runs the application stored in the memory 1102 to implement various functions, such as:

[0260] Divide the virtual scene into multiple sound source logical distribution areas, each sound source logical distribution area is associated with a group sound source control node, collect sound source distribution data in each sound source logical distribution area in real time according to the position of the virtual object in the interaction path, and dynamically adjust the spatial audio parameters of the group sound source control node;

[0261] According to a preset interactive event type, generating emotional dynamic response parameters by dynamically configuring emotional envelope parameters, and driving the group sound source control node to perform layered sound effect behavior based on the emotional dynamic response parameters and the spatial audio parameters of the group sound source control node;

[0262] Based on the dynamic adjustment of the spatial audio parameters of the group sound source control node and the emotional dynamic response output, the group behavior of the interactive objects is dynamically simulated through layered sound effects, and a dynamic sound field based on emotional fluctuations, spatial orientation and semantic feedback is constructed.

[0263] In some embodiments of the present application, the processor executing the application may further implement the following steps: dividing the virtual scene into multiple logical sound source distribution areas, specifically: dividing the virtual scene into multiple logical sound source distribution areas according to preset spatial rules.

[0264] In some embodiments of the present application, the processor executing the application program may further implement the following steps: the preset space rule is a four-quadrant space division rule.

[0265] In some embodiments of the present application, the processor executing the application may further implement the following steps: collecting sound source distribution data within each sound source logical distribution area in real time based on the position of the virtual object in the interaction path, and dynamically adjusting the spatial audio parameters of the group sound source control node, specifically:

[0266] Determining a reference point for dividing a sound source logical distribution area according to the real-time position of the virtual object in the interaction path, and dynamically adjusting the sound source logical distribution area based on the reference point;

[0267] Collect the sound source distribution data within the adjusted logical distribution area of ​​each sound source, and map the collected sound source distribution data to the real-time parameter control of the audio engine to drive the spatial access and density feedback of the group sound source control node.

[0268] In some embodiments of the present application, the processor executing the application program may further implement the following steps: collecting sound source distribution data within each sound source logical distribution area in real time according to the position of the virtual object in the interaction path, specifically:

[0269] Determining reference points for dividing logical distribution areas of sound sources according to interaction rules of virtual objects; wherein the interaction rules are dynamically set according to the interaction viewing angle state of the virtual objects;

[0270] Taking the reference point as the center, scan and record all interactive object nodes within a preset perception radius and store them in a node list; wherein the preset perception radius is used to limit the maximum number of the interactive object nodes;

[0271] According to the positions of all interactive object nodes in the node list, each interactive object node is divided into a corresponding sound source logical distribution area, and the number and average position of interactive objects in each sound source logical distribution area are counted.

[0272] In some embodiments of the present application, the processor executing the application may further implement the following steps: the interaction rule is to follow the virtual object according to the interaction perspective of the virtual object, and the reference point is bound to the center of mass coordinate system of the virtual object.

[0273] In some embodiments of the present application, the processor executing the application may further implement the following steps: the interaction rule is that the interaction perspective of the virtual object does not follow the virtual object, and the reference point is generated based on the global focus position.

[0274] In some embodiments of the present application, the processor executing the application program may further implement the following steps: dynamically adjusting the spatial audio parameters of the group sound source control node, specifically including:

[0275] The corresponding weight coefficient is defined according to the sound field density scenario, and the corresponding weight coefficient calculation is performed on the number of interactive objects in the logical distribution area of ​​each sound source according to the corresponding sound field density scenario to obtain the modified density parameter of the effective group interactive objects.

[0276] In some embodiments of the present application, the processor executing the application may further implement the following steps: after collecting the sound source distribution data within each sound source logical distribution area in real time according to the position of the virtual object in the interaction path, the following steps may also be included:

[0277] Calculating the distance between the average position of the interactive objects in the logical distribution area of ​​the sound source and the current position of the virtual object;

[0278] When it is determined that the distance does not exceed the preset distance, the average position of the group interaction objects is corrected using preset parameters to obtain corrected sound source distribution data; wherein the preset parameters include a fixed offset distance and an offset angle;

[0279] When it is determined that the distance exceeds the preset distance, the collected original sound source distribution data is retained.

[0280] In some embodiments of the present application, the processor executing the application may further implement the following steps: after collecting the sound source distribution data within each sound source logical distribution area in real time according to the position of the virtual object in the interaction path, the following steps may also be included:

[0281] According to the real-time changes in the position of the virtual object in the interaction path, the sound source position within the logical distribution area of ​​each sound source is updated in real time using the first transition variable; wherein, the first transition variable is preset to represent the displacement of the group sound source control node from the current coordinate to the target coordinate per frame.

[0282] In some embodiments of the present application, the processor executing the application may further implement the following steps: after collecting the sound source distribution data within each sound source logical distribution area in real time according to the position of the virtual object in the interaction path, the following steps may also be included:

[0283] The average position of the group sound source control node is linearly transitioned from the current position to the target position using a second transition variable; wherein the second transition variable is preset to represent the per-frame change in the number of group interaction objects from the current value to the target value.

[0284] In some embodiments of the present application, the processor executing the application program may further implement the following steps: dynamically adjusting the spatial audio parameters of the group sound source control node, specifically including:

[0285] Mapping the number of interactive objects within each sound source logical distribution area to the group size parameter of the audio engine to drive dynamic adjustment of the volume of the group interactive objects;

[0286] Obtain group emotion values ​​from the global emotion manager, map them to emotion intensity parameters, and dynamically adjust global audio mixing parameters;

[0287] The real-time audio parameters of the group sound source control nodes within the logical distribution area of ​​each sound source are dynamically adjusted according to the movement speed of the manipulation perspective of the virtual object, driving the Doppler effect simulation and enhancing the frequency offset of the relative movement of the group sound source control nodes.

[0288] In some embodiments of the present application, the processor executing the application program may further implement the following steps: dynamically adjusting the spatial audio parameters of the group sound source control node, further comprising:

[0289] Prioritize interactive object nodes in a specific area and load sound effect resources by marking priority.

[0290] In some embodiments of the present application, the processor executing the application may further implement the following steps: based on the emotional dynamic response parameter and the spatial audio parameter of the group sound source control node, driving the group sound source control node to perform a layered sound effect behavior, specifically including:

[0291] Managing the envelope lifecycle based on the emotional dynamic response parameters to calculate the real-time envelope contribution value of each interactive event, and calculating the global emotional value by integrating the envelope contribution values ​​of all interactive events;

[0292] The global emotion value is converted into a sound effect trigger probability to trigger a semantic feedback sound effect, and the sound field change of the virtual scene is driven according to the spatial audio parameters of the group sound source control node.

[0293] In some embodiments of the present application, the processor executing the application may further implement the following steps: calculating the global emotion value by integrating all interactive event envelope contribution values ​​through formula (1);

[0294] Global emotion value = default emotion value + ∑(each envelope contribution value) Formula (1)

[0295] The default emotion value is the initial value of each interactive object node.

[0296] In some embodiments of the present application, the processor executing the application may further implement the following steps: converting the global emotion value into a sound effect triggering probability by using formula (2);

[0297] P 触发 =(Current emotion value - emotion value threshold) / (100 - emotion value threshold)×maximum probability formula (2)

[0298] Among them, the current emotion value is the global emotion value, the emotion value threshold is the minimum emotion requirement for triggering the semantic feedback sound effect of the corresponding category, and the maximum probability is the highest trigger probability of this type of sound effect when the emotion value reaches 100.

[0299] In some embodiments of the present application, the processor executing the application program may further implement the following steps: dividing the virtual scene into a plurality of logical sound source distribution areas, each logical sound source distribution area being associated with a group sound source control node, collecting sound source distribution data within each logical sound source distribution area in real time based on the position of the virtual object in the interaction path, and dynamically adjusting the spatial audio parameters of the group sound source control node. The method may further include:

[0300] According to the key interaction event type and the spatial audio parameters of the group sound source control node, the group sound source control node is driven to perform layered sound effect behavior; wherein, the key interaction event is pre-defined by the system and is different from the preset interaction event.

[0301] In some embodiments of the present application, the processor executing the application program may further implement the following steps: dividing the virtual scene into a plurality of logical sound source distribution areas, each logical sound source distribution area being associated with a group sound source control node, collecting sound source distribution data within each logical sound source distribution area in real time based on the position of the virtual object in the interaction path, and before dynamically adjusting the spatial audio parameters of the group sound source control node, the method may further include:

[0302] Identify the types of interactive objects in the virtual scene, aggregate the identified interactive objects to generate corresponding interactive object aggregation nodes, and configure multimodal resources for the interactive object aggregation nodes; wherein the multimodal resources include at least a language pack and a sound effect library associated with the scene configuration file; the interactive object type is a classification of interactive objects in the virtual scene.

[0303] In some embodiments of the present application, the processor executing the application may further implement the following steps: aggregating the identified interactive objects to generate corresponding interactive object aggregation nodes, specifically:

[0304] The identified interactive objects are aggregated into regular units divided according to the interactive paths or logical areas to form corresponding interactive object aggregation nodes; wherein each node contains the number and location information of the interactive objects.

[0305] In some embodiments of the present application, the processor executing the application may further implement the following steps: dynamically configuring parameters to adjust the aggregated identified interactive objects to generate corresponding interactive object aggregation nodes; wherein the parameters include a spacing threshold defining the minimum interval of the interactive object distribution, a density threshold defining the maximum density of the interactive object distribution, and rule units for dividing the interactive path or logical area.

[0306] The present invention realizes the precise mapping of dynamic sound fields by collecting interactive object distribution data in virtual scenes in real time, combining the spatial evolution of the logical distribution area of ​​sound sources with the emotional envelope control: the global emotional value is calculated based on the event-driven emotional value, driving the intelligent switching of layered sound effects, solving the problems of single emotional response and rigid spatial positioning of traditional sound source systems, realizing multi-dimensional dynamic adaptation of the sound field as the virtual object behavior and virtual scene change, and the coordinated expression and dynamic interaction of multi-level emotional sound effects.

[0307] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0308] Optional, such as Figure 18 As shown, the electronic device 1100 further includes: a touch screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. Among them, the processor 1101 is electrically connected to the touch screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107 respectively. Those skilled in the art will understand that Figure 18 The electronic device structure shown in the figure does not constitute a limitation to the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0309] The touch display screen 1103 can be used to display a graphical user interface and receive user actions on the operation instructions generated by the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display the information input by the user or the information provided to the user and various graphical user interfaces of the electronic device, and these graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. Optionally, a liquid crystal display (LCD), an organic light emitting diode (OLED, Organic Light-Emitting Diode) and the like can be used to configure the display panel. The touch panel can be used to collect the user's touch operations on or near it (such as the user uses any suitable object or accessory such as a finger, a stylus on the touch panel or near the touch panel) and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 1101, and can receive the command sent by the processor 1101 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event. Then the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 1103 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 1103 can also be used as part of the input unit 1106 to realize the input function.

[0310] The radio frequency circuit 1104 may be used to transmit and receive radio frequency signals, so as to establish wireless communication with a network device or other electronic devices through wireless communication, and to transmit and receive signals with the network device or other electronic devices.

[0311] The audio circuit 1105 can be used to provide an audio interface between the user and the electronic device through a speaker and microphone. The audio circuit 1105 can convert the received audio data into an electrical signal and transmit it to the speaker, which then converts it into a sound signal for output. On the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and converted into audio data. The audio data is then output to the processor 1101 for processing, and then sent to another electronic device through the radio frequency circuit 1104, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earphone jack to provide communication between external headphones and the electronic device.

[0312] The input unit 1106 may be configured to receive input digital, character information, or user feature information (such as fingerprint, iris, or facial information), and to generate keyboard, mouse, joystick, optical, or trackball signal input related to user settings and function control.

[0313] Power supply 1107 is used to supply power to various components of electronic device 1100. Optionally, power supply 1107 can be logically connected to processor 1101 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. Power supply 1107 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0314] although Figure 18 Not shown, the electronic device 1100 may further include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.

[0315] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0316] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0317] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of computer programs, which can be loaded by a processor to execute any of the dynamic sound field interactive control methods provided in the embodiments of the present application. The computer program can execute the following steps of the dynamic sound field interactive control method:

[0318] Divide the virtual scene into multiple sound source logical distribution areas, each sound source logical distribution area is associated with a group sound source control node, collect sound source distribution data in each sound source logical distribution area in real time according to the position of the virtual object in the interaction path, and dynamically adjust the spatial audio parameters of the group sound source control node;

[0319] According to a preset interactive event type, generating emotional dynamic response parameters by dynamically configuring emotional envelope parameters, and driving the group sound source control node to perform layered sound effect behavior based on the emotional dynamic response parameters and the spatial audio parameters of the group sound source control node;

[0320] Based on the dynamic adjustment of the spatial audio parameters of the group sound source control node and the emotional dynamic response output, the group behavior of the interactive objects is dynamically simulated through layered sound effects, and a dynamic sound field based on emotional fluctuations, spatial orientation and semantic feedback is constructed.

[0321] The present invention realizes the precise mapping of dynamic sound fields by collecting interactive object distribution data in virtual scenes in real time, combining the spatial evolution of the logical distribution area of ​​sound sources with the emotional envelope control: the global emotional value is calculated based on the event-driven emotional value, driving the intelligent switching of layered sound effects, solving the problems of single emotional response and rigid spatial positioning of traditional sound source systems, realizing multi-dimensional dynamic adaptation of the sound field as the virtual object behavior and virtual scene change, and the coordinated expression and dynamic interaction of multi-level emotional sound effects.

[0322] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0323] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0324] Since the computer program stored in the computer-readable storage medium can execute any dynamic sound field interactive control method provided in the embodiments of the present application, the beneficial effects that can be achieved by any dynamic sound field interactive control method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0325] According to one aspect of the present application, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in various optional implementations of the above embodiments.

[0326] In the above-mentioned embodiments of the dynamic sound field interactive control device, computer-readable storage medium, electronic device, and computer program product, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes and beneficial effects of the above-described dynamic sound field interactive control device, computer-readable storage medium, computer program product, electronic device, and corresponding units thereof can be referred to the description of the dynamic sound field interactive control method in the above embodiments, and the specific details will not be repeated here.

[0327] The above is a detailed introduction to a dynamic sound field interactive control method, device, electronic device, computer-readable storage medium and computer program product provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and embodiments of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific embodiments and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A dynamic sound field interactive control method, characterized in that: The method comprises: Divide the virtual scene into multiple sound source logical distribution areas, each sound source logical distribution area is associated with a group sound source control node, collect sound source distribution data within each sound source logical distribution area in real time according to the position of the virtual object, and dynamically adjust the spatial audio parameters of the group sound source control node; According to a preset interaction event type, generating an emotional dynamic response parameter by dynamically configuring an emotional envelope parameter, and driving the group sound source control node to perform a layered sound effect behavior based on the emotional dynamic response parameter and the spatial audio parameter; A dynamic sound field is generated based on the spatial audio parameters, the emotional dynamic response parameters and the output of the layered sound effect behavior.

2. The dynamic sound field interactive control method according to claim 1, characterized in that: The virtual scene is divided into multiple logical distribution areas of sound sources, specifically: The virtual scene is divided into multiple logical distribution areas of sound sources through preset spatial rules.

3. The dynamic sound field interactive control method according to claim 2, characterized in that: The preset space rule is a four-quadrant space division rule.

4. The dynamic sound field interactive control method according to claim 1, characterized in that: According to the position of the virtual object in the interaction path, the sound source distribution data within the logical distribution area of ​​each sound source is collected in real time, and the spatial audio parameters of the group sound source control node are dynamically adjusted, specifically: Determining a reference point for dividing a sound source logical distribution area according to the real-time position of the virtual object in the interaction path, and dynamically adjusting the sound source logical distribution area based on the reference point; Collect the sound source distribution data within the adjusted logical distribution area of ​​each sound source, and map the collected sound source distribution data to the real-time parameter control of the audio engine to drive the spatial access and density feedback of the group sound source control node.

5. The dynamic sound field interactive control method according to claim 4, characterized in that: According to the position of the virtual object in the interaction path, the sound source distribution data within the logical distribution area of ​​each sound source is collected in real time, specifically: Determining reference points for dividing logical distribution areas of sound sources according to interaction rules of virtual objects; wherein the interaction rules are dynamically set according to the interaction viewing angle state of the virtual objects; Taking the reference point as the center, scan and record all interactive object nodes within a preset perception radius and store them in a node list; wherein the preset perception radius is used to limit the maximum number of the interactive object nodes; According to the positions of all interactive object nodes in the node list, each interactive object node is divided into a corresponding sound source logical distribution area, and the number and average position of interactive objects in each sound source logical distribution area are counted.

6. The dynamic sound field interactive control method according to claim 5, characterized in that: The interaction rule is to follow the virtual object according to the interaction perspective of the virtual object, and the reference point is bound to the center of mass coordinate system of the virtual object.

7. The dynamic sound field interactive control method according to claim 5, characterized in that: The interaction rule is that the interaction perspective of the virtual object does not follow the virtual object, and the reference point is generated based on a global focus position.

8. The dynamic sound field interactive control method according to claim 1, characterized in that: Dynamically adjusting the spatial audio parameters of the group sound source control node specifically includes: The corresponding weight coefficient is defined according to the sound field density scenario, and the corresponding weight coefficient calculation is performed on the number of interactive objects in the logical distribution area of ​​each sound source according to the corresponding sound field density scenario to obtain the modified density parameter of the effective group interactive objects.

9. The dynamic sound field interactive control method according to claim 1, characterized in that: After collecting the sound source distribution data in the logical distribution area of ​​each sound source in real time according to the position of the virtual object in the interaction path, it also includes: Calculating the distance between the average position of the interactive objects in the logical distribution area of ​​the sound source and the current position of the virtual object; When it is determined that the distance does not exceed the preset distance, the average position of the group interaction objects is corrected using preset parameters to obtain corrected sound source distribution data; wherein the preset parameters include a fixed offset distance and an offset angle; When it is determined that the distance exceeds the preset distance, the collected original sound source distribution data is retained.

10. The dynamic sound field interactive control method according to claim 1, characterized in that: After collecting the sound source distribution data in the logical distribution area of ​​each sound source in real time according to the position of the virtual object in the interaction path, it also includes: According to the real-time changes in the position of the virtual object in the interaction path, the sound source position within the logical distribution area of ​​each sound source is updated in real time using the first transition variable; wherein, the first transition variable is preset to represent the displacement of the group sound source control node from the current coordinate to the target coordinate per frame.

11. The dynamic sound field interactive control method according to claim 1, characterized in that: After collecting the sound source distribution data in the logical distribution area of ​​each sound source in real time according to the position of the virtual object in the interaction path, it also includes: The average position of the group sound source control node is linearly transitioned from the current position to the target position using a second transition variable; wherein the second transition variable is preset to represent the per-frame change in the number of group interaction objects from the current value to the target value.

12. The dynamic sound field interactive control method according to claim 1, characterized in that: Dynamically adjusting the spatial audio parameters of the group sound source control node specifically includes: Mapping the number of interactive objects within each sound source logical distribution area to the group size parameter of the audio engine to drive dynamic adjustment of the volume of the group interactive objects; Obtain group emotion values ​​from the global emotion manager, map them to emotion intensity parameters, and dynamically adjust global audio mixing parameters; The real-time audio parameters of the group sound source control nodes within the logical distribution area of ​​each sound source are dynamically adjusted according to the movement speed of the manipulation perspective of the virtual object, driving the Doppler effect simulation and enhancing the frequency offset of the relative movement of the group sound source control nodes.

13. The dynamic sound field interactive control method according to claim 1, characterized in that: Dynamically adjusting the spatial audio parameters of the group sound source control node also includes: Prioritize interactive object nodes in a specific area and load sound effect resources by marking priority.

14. The dynamic sound field interactive control method according to claim 1, characterized in that: According to the preset interaction event type, the emotional dynamic response parameters are generated by dynamically configuring the emotional envelope parameters, including: Manage the envelope lifecycle according to the preset interaction event type and calculate the real-time envelope contribution value of each interaction event; The global emotion value is calculated by integrating the envelope contribution values ​​of all interaction events as the emotion dynamic response parameter.

15. The dynamic sound field interactive control method according to claim 14, characterized in that: Driving the group sound source control node to perform layered sound effect behaviors includes: selecting a sound effect level and volume based on the spatial audio parameters; The corresponding level of sound effects is triggered by the global emotion value and the emotion value threshold to perform layered sound effect behavior; wherein the emotion value threshold is the minimum emotion value that triggers the corresponding level of feedback sound effects.

16. The dynamic sound field interactive control method according to claim 15, characterized in that: The steps to implement the layered sound behavior include: Performing continuous emotional preparation sound effect level selection according to the spatial audio parameters; Convert the global emotion value into a sound effect trigger probability to trigger the localized semantic feedback sound effect as the output for executing the layered sound effect behavior; A transient emotional outburst response is performed according to the spatial audio parameter and the global emotional value.

17. The dynamic sound field interactive control method according to claim 16, characterized in that: The global emotion value is calculated by integrating the envelope contribution values ​​of all interactive events through formula (1); Global emotion value = default emotion value + ∑(contribution value of each envelope) Formula (1) Wherein, the default emotion value is the initial value of the sound source control node.

18. The dynamic sound field interactive control method according to claim 17, characterized in that: The global emotion value is converted into a sound effect triggering probability by formula (2); P 触发 =(Current emotion value - emotion value threshold) / (100 - emotion value threshold) × Maximum probability formula (2) Among them, the current emotion value is the global emotion value, and the maximum probability is the highest triggering probability of the corresponding level sound effect when the emotion value reaches 100.

19. The dynamic sound field interactive control method according to claim 1, characterized in that: The method further comprises dividing the virtual scene into a plurality of logical sound source distribution areas, each logical sound source distribution area being associated with a group sound source control node, collecting sound source distribution data within each logical sound source distribution area in real time according to the position of the virtual object in the interaction path, and dynamically adjusting the spatial audio parameters of the group sound source control node. According to the key interaction event type and the spatial audio parameters of the group sound source control node, the group sound source control node is driven to perform layered sound effect behavior; wherein, the key interaction event is pre-defined by the system and is different from the preset interaction event.

20. The dynamic sound field interactive control method according to claim 1, characterized in that: The method further comprises dividing the virtual scene into a plurality of logical sound source distribution areas, each logical sound source distribution area being associated with a group sound source control node, collecting sound source distribution data within each logical sound source distribution area in real time according to the position of the virtual object in the interaction path, and dynamically adjusting the spatial audio parameters of the group sound source control node. Identify the types of interactive objects in the virtual scene, aggregate the identified interactive objects to generate corresponding interactive object aggregation nodes, and configure multimodal resources for the interactive object aggregation nodes; wherein the multimodal resources include at least a language pack and a sound effect library associated with the scene configuration file; the interactive object type is a classification of interactive objects in the virtual scene.

21. The dynamic sound field interactive control method according to claim 20, characterized in that: The identified interactive objects are aggregated to generate corresponding interactive object aggregation nodes, specifically: The identified interactive objects are aggregated into regular units divided according to the interactive paths or logical areas to form corresponding interactive object aggregation nodes; wherein each node contains the number and location information of the interactive objects.

22. The dynamic sound field interactive control method according to claim 21, characterized in that: The interactive objects identified by aggregation are adjusted by dynamically configuring parameters to generate corresponding interactive object aggregation nodes; wherein the parameters include a spacing threshold defining the minimum interval of interactive object distribution, a density threshold defining the maximum density of interactive object distribution, and rule units for dividing interactive paths or logical areas.

23. A dynamic sound field interactive control device, characterized in that: The device comprises: A sound source logical partitioning module is used to divide the virtual scene into multiple sound source logical distribution areas, each of which is associated with a group sound source control node. The sound source distribution data within each sound source logical distribution area is collected in real time according to the position of the virtual object in the interaction path, and the spatial audio parameters of the group sound source control node are dynamically adjusted. An emotion envelope parameter configuration module is used to generate emotion dynamic response parameters by dynamically configuring emotion envelope parameters according to a preset interaction event type; A layered sound effect synthesis module is used to drive the group sound source control node to perform layered sound effect behavior based on the emotional dynamic response parameters and the spatial audio parameters, and to generate a dynamic sound field based on the spatial audio parameters, the emotional dynamic response parameters and the output of performing the layered sound effect behavior.

24. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of the dynamic sound field interactive control method according to any one of claims 1 to 22.

25. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the dynamic sound field interactive control method according to any one of claims 1 to 22.

Citation Information

Cited By

  • Audio adjustment processing method and device, equipment and medium

    CN121490382A