Immersive sound field rendering method and system for acoustic loudspeaker

By identifying focal points and ambient sound sources in the virtual space, and employing differentiated rendering strategies and processor load adjustments, the problem of imbalance between spatial resolution and computational efficiency in acoustic speaker sound field rendering was solved, achieving high-resolution sound source localization and a low-latency immersive sound field experience.

CN121985284APending Publication Date: 2026-05-05深圳市云科实业有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
深圳市云科实业有限公司
Filing Date
2026-02-06
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing acoustic speakers suffer from an imbalance between spatial resolution and computational efficiency during sound field rendering, making it difficult for users to accurately determine the source of sound, overloading the system's computing unit, and causing fuzzy sound source localization, slow system response, and disruption of audiovisual synchronization in multi-user collaborative training scenarios.

Method used

By identifying focal sound sources and ambient sound sources in the virtual space, a differentiated rendering strategy is adopted. The focal sound sources are rendered at high resolution, while the ambient sound sources are rendered with low resources. The rendering strategy is adjusted in combination with processor utilization and temperature to ensure the computing resources of the focal sound sources. The sound is then played using an acoustic speaker array.

Benefits of technology

It achieves computational efficiency at high spatial resolution, ensuring accurate perception of key sounds and an immersive experience for users, reducing system computational load, and improving the immersion and interactivity of virtual reality training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985284A_ABST
    Figure CN121985284A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an acoustic horn immersion type sound field rendering method and system, and relates to the technical field of sound field rendering, and the method comprises the steps: obtaining a user position and a head posture in a virtual space; acquiring the position and intensity of a virtual space sound source; identifying a focus sound source currently concerned by the user according to the user position, the head posture, the sound source position and the intensity; identifying sound sources except the focus sound source as environment sound sources; performing sound field rendering on the environmental sound source to obtain a shared environmental sound field signal, the computing resource of the environmental sound source being lower than the computing resource of the focus sound source; performing convolution operation on the original audio signal of the focus sound source and the head acoustic transmission function to obtain a personalized high-resolution signal; and superposing the shared environment sound field signal and the personalized high-resolution signal to generate a binaural signal, and playing the binaural signal through an acoustic horn array. According to the method, the immersion and the actual effect of virtual reality training can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sound field rendering technology, and in particular to an immersive sound field rendering method and system for acoustic speakers. Background Technology

[0002] In related technologies, acoustic speakers often face a thorny problem when rendering sound fields: how to ensure sufficiently accurate spatial sound localization while avoiding system sluggishness due to excessive computation. This imbalance between spatial resolution and computational efficiency often makes it difficult for users to accurately determine the source of sound in the virtual world, and also overwhelms the computing units supporting the system. Summary of the Invention

[0003] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes an immersive sound field rendering method and system for acoustic speakers, which aims to solve the problems of imbalance between spatial resolution and computational efficiency when performing sound field rendering with existing acoustic speakers, which makes it difficult for users to accurately determine the sound source, causes excessive load on the system's computing unit, and results in problems such as blurred sound source localization, slow system response, and disruption of audiovisual synchronization in multi-user collaborative training scenarios.

[0004] In a first aspect, embodiments of this application provide an immersive sound field rendering method for acoustic speakers, including: Obtain the user's location and head pose in the virtual space; Obtain the location and intensity of a sound source in virtual space; Identify the user's current focus sound source based on user location, head posture, sound source location, and intensity; Identify sound sources other than the focal sound source as ambient sound sources; Perform sound field rendering on the ambient sound source to obtain a shared ambient sound field signal. The computational resources of this ambient sound source are lower than those of the focal sound source. A personalized high-resolution signal is obtained by convolving the original audio signal of the focal sound source with the head acoustic transfer function. The shared ambient sound field signal is superimposed with the personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array.

[0005] Furthermore, based on the above method, the steps of convolving the original audio signal of the focal sound source with the head acoustic transfer function to obtain a personalized high-resolution signal include: Obtain the user's head orientation based on head posture; Based on the head orientation, retrieve the corresponding head acoustic transfer function from the pre-stored head transfer function database. A personalized high-resolution signal is obtained by convolving the original audio signal of the focal sound source with the retrieved head acoustic transfer function.

[0006] In some preferred embodiments, the step of superimposing a shared ambient sound field signal with a personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array, includes: Obtain the position and pose information of virtual objects in virtual space, as well as the geometric information of virtual objects; For each focal sound source, calculate the sound propagation path from the focal sound source to the user's ears; Based on the preset operation path of the current training task, and according to the virtual object's position and pose information and geometric information, predict possible occlusion events. When an occlusion event is predicted, the shared ambient sound field signal is superimposed with a personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array. The acoustic speaker uses beamforming technology to control delay and gain to ensure that the binaural signal forms a sound pressure peak at the target location in the user's ear.

[0007] Furthermore, the steps for identifying the user's current focus sound source based on user location, head posture, sound source location, and intensity include: The Euclidean distance is obtained by calculating the distance between the sound source and the user's head based on the user's location and the sound source location. Potential focal sound sources are identified based on Euclidean distance; Obtain the user's head orientation based on head posture; If a potential focal sound source is located within a preset angle range in front of the user's head, or if the intensity of the potential focal sound source is greater than a preset threshold, the potential focal sound source is identified as the user's current focal sound source.

[0008] As a technological improvement, the method also includes: Real-time monitoring of processor utilization and processor temperature; Adjust the ambient sound field rendering strategy based on processor utilization and processor temperature to ensure sufficient rendering computing resources for the focal sound source.

[0009] Based on this, the steps for adjusting the ambient sound field rendering strategy according to processor utilization and processor temperature to ensure the rendering computing resources of the focal sound source include: Obtain the processor utilization threshold and the processor temperature threshold; When the processor utilization exceeds a utilization threshold or the processor temperature exceeds a temperature threshold, a rendering strategy that progressively reduces the ambient sound field is adopted. This rendering strategy includes at least one of the following: Reduce the accuracy of reflection calculations from environmental sound sources; Reduce the frequency of background sound updates.

[0010] In one implementation, the steps of obtaining the processor utilization threshold and the processor temperature threshold include: The thermal conductivity performance of the processor's cooling system is evaluated, and the thermal conductivity performance evaluation results are obtained. Based on the thermal conductivity evaluation results, the correction parameters for the processor utilization threshold and the processor temperature threshold are determined; The preset utilization rate and preset temperature are adjusted based on the correction parameters to obtain the processor utilization threshold and processor temperature threshold.

[0011] As a further improvement, the method also includes: Monitoring non-focus acoustic events in a virtual environment; When a non-focal acoustic event occurs, the sound field rendering strategy is adjusted to ensure the rendering resources of the focal sound source. This sound field rendering adjustment strategy includes at least one of the following: The shared ambient sound field signal is rendered with progressive degradation. Reduce the update frequency of haptic feedback and non-critical visual effects to ensure that the focus sound source task can preempt processor resources.

[0012] Based on the above, the steps for monitoring non-focal acoustic events in a virtual environment include: Receive real-time information from all sound sources, including the loudness and frequency of the sound sources; When the loudness of a sound source exceeds the loudness threshold and the frequency exceeds the frequency threshold, the sound source is determined to be a non-focal acoustic event.

[0013] Secondly, this application also discloses an immersive sound field rendering system for acoustic speakers, the system comprising: The user information acquisition module is used to acquire the user's location and head posture in the virtual space; The sound source information acquisition module is used to acquire the location and intensity of sound sources in the virtual space; The focus sound source determination module is used to identify the focus sound source that the user is currently paying attention to based on the user's position, head posture, sound source position, and intensity. The ambient sound source identification module is used to identify sound sources other than the focal sound source as ambient sound sources; The shared ambient sound field signal acquisition module is used to perform sound field rendering on the ambient sound source to obtain the shared ambient sound field signal. The computational resources of the ambient sound source are lower than those of the focus sound source. The personalized high-resolution signal acquisition module performs a convolution operation on the original audio signal of the focal sound source and the head acoustic transfer function to obtain a personalized high-resolution signal. The binaural signal acquisition module is used to superimpose the shared ambient sound field signal with the personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array.

[0014] This application discloses an immersive sound field rendering method using acoustic speakers. By acquiring the user's position, head posture, sound source position, and intensity in a virtual space, it intelligently identifies the user's current focus sound source and identifies other sound sources as ambient sound sources. For the focus sound source, this application performs high-resolution convolution operations on its original audio signal and the head acoustic transfer function to generate a personalized high-resolution signal, ensuring the user's accurate perception of the core sound and an immersive experience. Simultaneously, for ambient sound sources, this application employs a low-computational-resource sound field rendering strategy to generate a shared ambient sound field signal, effectively reducing the overall computational load. Finally, the shared ambient sound field signal and the personalized high-resolution signal are superimposed to generate a binaural signal, which is then played through an acoustic speaker array.

[0015] Through the above technical solution, this application effectively solves the problem of balancing spatial resolution and computational efficiency in existing acoustic speaker immersive sound field rendering. Specifically, this application ensures accurate spatial positioning and externalization of key sounds by performing high-precision rendering of the focal sound source, overcoming the defect of ambiguous sound source positioning in existing technologies. At the same time, by performing low-resource rendering of ambient sound sources, the overall computational load of the system is significantly reduced, avoiding system sluggishness and overload of core computing units caused by excessive computation. This hierarchical rendering strategy enables the system to maintain sufficient computational efficiency while ensuring high spatial resolution. Especially in professional-grade virtual reality training scenarios with extremely high requirements for realism and interactivity, it can provide a smooth, synchronized, and highly immersive auditory experience, thereby overcoming the performance bottleneck caused by multiple intertwined factors in existing technologies and significantly improving the immersion and actual effect of virtual reality training.

[0016] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0017] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0018] Figure 1 This is a flowchart illustrating an immersive sound field rendering method for acoustic speakers provided in one embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical methods, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. It should be noted that the meaning of "multiple" (or "more than") in the description of the embodiments of this application refers to two or more, and "greater than," "less than," "exceeding," etc. are understood to exclude the number itself, while "above," "below," "within," etc. are understood to include the number itself. If "first," "second," etc. are used in the description, they are only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated. Based on the above, this application proposes an immersive sound field rendering method and system for acoustic speakers, aiming to solve the problems of imbalance between spatial resolution and computational efficiency when performing sound field rendering with existing acoustic speakers, which makes it difficult for users to accurately determine the source of sound, overloads the system's computing unit, and causes problems such as blurred sound source localization, slow system response, and destruction of audiovisual synchronization in multi-user collaborative training scenarios.

[0020] The acoustic speaker immersive sound field rendering method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms; the software can be an application that implements the acoustic speaker immersive sound field rendering method, but is not limited to the above forms. This application can be applied to numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via communication networks. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices. It should be noted that in various specific embodiments of this invention, when processing is required based on data related to the characteristics of an object (e.g., user attributes or sets of attribute information), permission or consent from the corresponding object is obtained first, and the collection, use, and processing of this data comply with relevant laws and standards. Furthermore, when the embodiments of the present invention need to obtain the attribute information of an object, they will obtain the separate permission or separate consent of the corresponding object through pop-up windows or redirection to a confirmation page. After obtaining the separate permission or separate consent of the corresponding object, they will then obtain the relevant data of the object necessary for the embodiments of the present invention to operate normally.

[0021] See Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of an acoustic speaker immersive sound field rendering method provided in this application. The embodiment of this application provides an acoustic speaker immersive sound field rendering method, including but not limited to steps S110 to S170, which are described in detail below.

[0022] S110. Obtain the user's location and head posture in the virtual space; S120. Obtain the location and intensity of the sound source in the virtual space; S130. Identify the current focus sound source of the user based on the user's location, head posture, sound source location, and intensity. S140. Identify sound sources other than the focal sound source as ambient sound sources; S150. Perform sound field rendering on the ambient sound source to obtain a shared ambient sound field signal. The computational resources of the ambient sound source are lower than those of the focus sound source. S160. Perform convolution operation on the original audio signal of the focal sound source and the head acoustic transfer function to obtain a personalized high-resolution signal; S170: The shared ambient sound field signal is superimposed with the personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array.

[0023] To better understand the acoustic horn immersive sound field rendering method proposed in this application, it is necessary to explain some of the key terms involved.

[0024] Virtual space refers to a digital environment simulated and generated using computer technology, allowing users to interact with it, such as virtual reality (VR) or augmented reality (AR) scenes. User position and head posture are the user's real-time spatial coordinates and head orientation information within the virtual space, crucial for accurately simulating sound source direction. Sound source position and intensity describe the spatial coordinates of various sound sources in the virtual space and their loudness. The head acoustic transfer function (HRTF) is an acoustic characteristic function describing the propagation of sound from a point in space to the eardrum. It incorporates the diffraction, reflection, and absorption effects of sound by the auricle, head, and torso, and is key to achieving three-dimensional sound field rendering.

[0025] In implementing the method of this application, it is first necessary to obtain the user's position and head posture in the virtual space. The user's position can be obtained through various methods, such as an inertial measurement unit (IMU), an optical tracking system, or an electromagnetic tracking system. For example, the user can wear a head-mounted display integrated with an IMU, which can output the user's three-axis acceleration, angular velocity, and magnetic field information in real time. The precise position and head posture of the user in the virtual space can then be calculated using a sensor fusion algorithm. Alternatively, multiple cameras can be set up in the virtual space, and computer vision technology can be used to track specific markers worn by the user to obtain the user's position and head posture.

[0026] Simultaneously, the system needs to acquire the location and intensity of all sound sources in the virtual space. This information is typically preset by the creator of the virtual scene or dynamically generated within the virtual environment based on specific events. For example, in a virtual battlefield scene, gunshots, explosions, footsteps, etc., all have specific virtual locations and intensities, and this information is transmitted to the sound field rendering system in real time.

[0027] Next, based on the acquired user location, head posture, sound source location, and intensity, the system identifies the sound source currently in focus of the user's attention. For example, the system can calculate the Euclidean distance between each sound source and the user's head, and combine this with the user's head orientation to determine which sound sources are located in front of the user's line of sight or within a preset area of ​​focus. Furthermore, the intensity of the sound source can also be used as a basis for determining the focus sound source; for example, sound sources with an intensity exceeding a certain threshold may be identified as focus sound sources.

[0028] Once the focal sound source is identified, the remaining sound sources are identified as ambient sound sources. For these ambient sound sources, the system performs sound field rendering to obtain a shared ambient sound field signal. The rendering strategy for ambient sound sources typically employs lower computational resources; for example, it may use a simplified HRTF model, reduce the precision of reflection calculations, or decrease the update frequency of background noise. For instance, a simplified algorithm based on acoustic ray tracing can be used, calculating only the primary reflection path while ignoring secondary reflections and scattering, thereby reducing computational complexity.

[0029] Meanwhile, for the focal sound source, the system performs a convolution operation on its original audio signal and the head acoustic transfer function (HRTF) to obtain a personalized high-resolution signal. Since the focal sound source is the user's current focus, high-resolution rendering is crucial. The convolution operation combines the original audio signal of the focal sound source with the user's HRTF, simulating the complex acoustic path of sound propagating from the focal sound source to the user's ears, thus providing accurate spatial localization and externalization effects.

[0030] Finally, the shared ambient sound field signal is superimposed with the personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array. The acoustic speaker array can use beamforming technology to precisely control the delay and gain of the sound, ensuring that the binaural signal forms a sound pressure peak at the target location of the user's ears, thereby providing the user with an immersive three-dimensional sound field experience.

[0031] The acoustic horn immersive sound field rendering method of this application effectively solves the problem of balancing spatial resolution and computational efficiency in sound field rendering in the prior art by dividing the sound sources in the virtual space into focal sound sources and ambient sound sources and adopting differentiated rendering strategies for the two.

[0032] Compared to existing technologies that blindly increase the amount of digital filtering data or introduce complex personalized information to improve the detail of sound field rendering, leading to a surge in computational load, the method in this application utilizes computing resources more intelligently and efficiently. By refining the processing of focal sound sources and simplifying the processing of ambient sound sources, this application effectively avoids latency and performance degradation caused by excessive computational load while ensuring high-resolution perception of key information by the user. This hierarchical rendering strategy enables the system to provide users with a high-quality, low-latency immersive sound field experience in complex and dynamic virtual environments, significantly improving the realism and interactivity of applications such as virtual reality training.

[0033] In some embodiments described above, the original audio signal of the focal sound source is convolved with a head acoustic transfer function to obtain a personalized high-resolution signal. However, in practical applications, if the head acoustic transfer function used fails to dynamically adjust according to the user's real-time head posture, it may lead to a decrease in the immersiveness of the sound field rendering and the accuracy of spatial positioning. This is because the human ear's perception of a sound source is highly dependent on head posture; a static head acoustic transfer function cannot accurately simulate the propagation path of sound waves under different head orientations, thus affecting the user's judgment of the sound source's direction and distance.

[0034] In this regard, this application further proposes the following steps for performing convolution operations on the original audio signal of the focal sound source and the head acoustic transfer function to obtain a personalized high-resolution signal: The user's head orientation is obtained based on the head posture. Based on the head orientation, retrieve the corresponding head acoustic transfer function from the pre-stored head transfer function database; A personalized high-resolution signal is obtained by convolving the original audio signal of the focal sound source with the retrieved head acoustic transfer function.

[0035] Specifically, after acquiring the user's position and head posture in the virtual space, obtaining the user's head orientation based on the head posture refers to determining the specific orientation angle of the user's head in the virtual space by analyzing the head posture data, such as three-dimensional rotation information provided by an inertial measurement unit (IMU) or optical tracking system. This head orientation can be represented in the form of Euler angles, quaternions, or rotation matrices, etc., with the aim of accurately capturing the real-time pointing of the user's head relative to the virtual environment.

[0036] The step of retrieving the head acoustic transfer function (HRTF) corresponding to the head orientation from a pre-stored HRTF database can be understood as the system maintaining a dataset containing HRTFs for various head orientations. This database can be pre-generated through measurement or simulation, covering all or most possible head orientations of the user. Once the user's real-time head orientation is obtained, the system searches the database or calculates the HRTF that best matches the current head orientation using an interpolation algorithm. For example, if the database stores HRTFs for every 5 or 10 degrees of orientation, when the user's head orientation falls between two stored angles, an accurate HRTF can be generated using linear interpolation or more complex spatial interpolation methods. The purpose is to provide an acoustic model that highly matches the user's current auditory state for subsequent convolution operations.

[0037] In practical applications, the convolution operation between the original audio signal from the focal sound source and the retrieved head acoustic transfer function specifically utilizes digital signal processing technology to mathematically convolve the original audio signal from the focal sound source with the head acoustic transfer function dynamically retrieved based on the user's head orientation. This convolution operation simulates the frequency and phase changes caused by the physiological structures of the head and auricle during the propagation of sound waves from the sound source to the user's ears. Its purpose is to convert the original audio signal from the focal sound source into a spatially perceptible binaural signal, enabling the user to perceive the precise location and direction of the sound source in virtual space when played through an acoustic speaker array.

[0038] This application's solution dynamically obtains the user's head orientation based on their head posture and retrieves the corresponding head acoustic transfer function from a pre-stored head transfer function database. This solves the problem of mismatch between the head acoustic transfer function and the user's actual head orientation that may exist in traditional solutions. Because the head acoustic transfer function can adapt to the user's head movements in real time, convolution operations can generate more accurate and personalized binaural signals. This dynamic adaptability ensures that when sound waves reach the user's ears, their spatial characteristics accurately reflect the relative position of the sound source in virtual space, maintaining the stability and accuracy of sound source localization even when the user's head rotates frequently.

[0039] Through the above technical solution, this application can significantly improve the realism and spatial positioning accuracy of immersive sound field rendering using acoustic speakers. Since the head acoustic transfer function is dynamically selected based on the user's real-time head orientation, the generated personalized high-resolution signal can more accurately simulate the arrival of sound waves at the user's ears, effectively avoiding sound source positioning drift or blurring caused by changes in head posture. This allows users to obtain a more natural and immersive auditory experience in the virtual environment, enhancing their perception of the direction and distance of virtual sound sources, thereby improving the overall quality of the immersive experience.

[0040] In some preferred embodiments, assuming a user is exploring a virtual environment, their head posture is acquired in real time via an inertial measurement unit (IMU) built into the head-mounted display (HMD). When the user slowly turns their head 30 degrees to the left from directly in front, the system immediately detects this head posture change and calculates the new head orientation. Subsequently, based on this new head orientation, the system retrieves the most closely matched head acoustic transfer function (HRTF) from a pre-established database of head transfer functions (HRTFs) for a 30-degree left turn. If no precisely matching 30-degree HRTF is found in the database, the system can interpolate HRTFs for 25-degree and 35-degree (assuming they exist) to generate an HRTF suitable for a 30-degree orientation. Next, the system convolves the original audio signal of the current focused sound source with this newly retrieved or generated HRTF to generate an updated, personalized, high-resolution signal. In this way, even with continuous head movement, the system ensures that the sound image of the focused sound source remains stably positioned in the correct location within the virtual space, providing the user with coherent and highly realistic auditory feedback.

[0041] Traditional immersive sound field rendering methods using acoustic speakers, when superimposing shared environmental sound field signals with personalized high-resolution signals to generate binaural signals and playing them through an acoustic speaker array, may not fully consider the impact of complex physical obstacles on sound propagation in the virtual environment. For example, suppose there is a wall in the virtual space, the sound source is behind the wall, and the user is in front of the wall. If this problem is not addressed, the system may directly play unadjusted binaural signals, causing the sound perceived by the user to deviate from the laws of sound propagation in the actual physical world, thereby reducing immersion and realism.

[0042] In this regard, this application further proposes a step of superimposing the aforementioned shared environmental sound field signal with the aforementioned personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array, including: Obtain the position and pose information of virtual objects in virtual space, as well as the geometric information of virtual objects; For each focal sound source, calculate the sound propagation path from the aforementioned focal sound source to the user's ears; Based on the preset operation path of the current training task, and according to the above-mentioned virtual object position and pose information and the above-mentioned virtual object geometric information, the possible occlusion events are predicted. When the aforementioned occlusion event is predicted, the shared ambient sound field signal and the personalized high-resolution signal are superimposed to generate a binaural signal, which is then played through an acoustic speaker array. The acoustic speakers use beamforming technology to control delay and gain to ensure that the binaural signal forms a sound pressure peak at the target location of the user's ear.

[0043] Specifically, virtual object position and orientation information refers to the spatial attribute data such as the three-dimensional coordinates and rotation angles of all virtual objects in virtual space that may affect sound propagation. Its purpose is to provide basic data for sound propagation path calculation and occlusion event prediction. Virtual object geometric information can be understood as the geometric feature data of virtual objects such as shape, size, and material. For example, it can be a polygon mesh, voxel data, or a simplified collider model. Its purpose is to accurately simulate the interaction between sound waves and virtual objects, such as reflection, absorption, and diffraction.

[0044] The sound propagation path refers to the physical path of a sound wave from its source, through various media and obstacles in the virtual environment, to its final arrival at the user's ears. This path can be calculated using various acoustic simulation techniques such as ray tracing and wave field synthesis, with the aim of accurately simulating the propagation behavior of sound waves in complex virtual environments.

[0045] In practical applications, occlusion events refer to situations where a virtual object blocks the direct path of sound waves from the sound source to the user's ears. Predicting potential occlusion events can be achieved by combining the preset operation paths of the current training task. For example, in virtual reality games, these paths might include the user's possible movement trajectory and the areas the virtual character might traverse. By analyzing the relationship between these paths and the virtual object's position, posture, and geometric information, it's possible to predict whether the sound source will be occluded by the virtual object in a specific scenario. The purpose is to identify potential obstacles in sound propagation in advance, providing a basis for subsequent adjustments to the sound field rendering.

[0046] Furthermore, beamforming technology refers to controlling the phase and amplitude of each speaker in an acoustic speaker array to create a specific directional sound beam in space, thereby generating a sound pressure level peak at the target location in the user's ear. This involves precise control of the delay and gain of each speaker to compensate for the attenuation and phase difference of the sound wave during propagation. The aim is to ensure that the user can clearly and accurately perceive the binaural signal after the masking process, improving the sound localization accuracy and immersion.

[0047] Through the above technical solution, this application can significantly improve the realism and immersion of immersive sound field rendering using acoustic speakers. By accurately predicting occlusion events in the virtual environment and combining beamforming technology to finely control the acoustic speaker array, sound waves can accurately form the target sound pressure peak at the user's ear. This not only solves the problem of poor sound field rendering effects when traditional methods deal with complex virtual environments (such as those with occlusions), but also ensures that users can obtain an auditory experience in virtual space that is highly consistent with the acoustic experience in the real world, thereby greatly enhancing the user's perception depth and interactive experience of the virtual environment.

[0048] Specifically, the steps described above for identifying the user's current focus sound source based on user location, head posture, sound source location, and intensity include: The Euclidean distance is obtained by calculating the distance between the sound source and the user's head based on the user's location and the sound source location. The potential focal sound source is determined based on the Euclidean distance; The user's head orientation is obtained based on the head posture. If the potential focus sound source is located within a preset angle range in front of the user's head, or if the intensity of the potential focus sound source is greater than a preset threshold, the potential focus sound source is identified as the focus sound source currently being focused on by the user.

[0049] The calculation of the Euclidean distance between the sound source and the user's head refers to the straight-line distance between each sound source and the center point of the user's head in a virtual three-dimensional space, calculated using the standard Euclidean distance formula. This distance can serve as a direct indicator for assessing the spatial proximity between the sound source and the user, with the aim of initially identifying sound sources that are relatively close to the user in physical space.

[0050] Furthermore, potential focal sound sources are determined based on the Euclidean distance. Specifically, a distance threshold can be set, and all sound sources whose distance from the user's head is less than this threshold are identified as potential focal sound sources. For example, when the Euclidean distance between a sound source and the user's head is less than 2 meters, the sound source is considered a potential focal sound source. This step aims to narrow down the scope of subsequent processing and improve the efficiency of focal sound source identification.

[0051] Furthermore, obtaining the user's head orientation based on the head posture refers to using data provided by a user head posture sensor (such as an inertial measurement unit, IMU) to determine the orientation vector or angle of the user's head in virtual space in real time. This head orientation information is crucial for determining the user's visual and auditory attention direction.

[0052] Specifically, if the potential focal sound source is located within a preset angle range in front of the user's head, or if the intensity of the potential focal sound source is greater than a preset threshold, the potential focal sound source is identified as the user's current focal sound source. The preset angle range can be understood as the user's visual and auditory attention area; for example, it can be set as a fan-shaped area 30 degrees to the left and right of the user's head. When a potential focal sound source falls within this area, it indicates that the user may be actively paying attention to that sound source. The preset threshold refers to the minimum standard for sound source intensity; for example, when the sound source intensity exceeds 80 decibels, even if it is not in front of the user's head, it may be identified as a focal sound source due to its prominence. This comprehensive judgment mechanism aims to more accurately capture the user's actual focus, taking into account both the user's active attention (head orientation) and significant events in the environment (sound source intensity).

[0053] The aforementioned technical solution enables more accurate and intelligent identification of the focal sound sources that a user is currently focusing on within a virtual environment. This identification method considers not only the spatial proximity of the sound source to the user but also the user's subjective attention (head orientation) and the objective salience (intensity) of the sound source itself, thus avoiding misjudgments or omissions that may result from relying on a single dimension. Consequently, it ensures that subsequent sound field rendering resources can be prioritized and accurately allocated to the sound sources that the user is truly interested in, significantly enhancing the realism of immersive sound field rendering and the user experience, allowing users to perceive and interact with key acoustic information in the virtual environment more naturally.

[0054] In some embodiments described above, an immersive sound field rendering method using acoustic speakers is proposed. This method aims to provide a high-quality immersive sound field experience by distinguishing between focal sound sources and ambient sound sources. However, in practical applications, sound field rendering, especially high-resolution focal sound source rendering, demands significant computational resources. If the system faces excessive processor load during operation, such as due to increased complexity of the virtual scene or simultaneous execution of other computationally intensive tasks, rendering performance may degrade, thereby affecting the rendering quality and real-time performance of the focal sound source and ultimately impairing the user's immersion.

[0055] In this regard, this application further proposes that the above method also includes: Real-time monitoring of processor utilization and processor temperature; The ambient sound field rendering strategy is adjusted based on the processor utilization and the processor temperature to ensure the rendering computing resources for the focal sound source.

[0056] Specifically, processor utilization refers to the percentage of time the processor is occupied by tasks, reflecting its workload. Processor temperature refers to the core temperature of the processor during operation, and is an important indicator of processor workload and heat dissipation. These parameters can be obtained in real time through the application programming interface (API) provided by the operating system or hardware sensors. For example, processor performance counters or temperature sensor data can be queried periodically. Adjusting the ambient sound rendering strategy refers to dynamically changing the computational complexity or resource consumption of ambient sound source rendering based on the real-time load of the processor. Its purpose is to reduce the resource consumption of non-core tasks (i.e., ambient sound source rendering) when system resources are strained, freeing up necessary computing resources for core tasks (i.e., focus sound source rendering), thereby ensuring the rendering quality and real-time performance of focus sound sources. Ensuring the rendering computing resources of focus sound sources can be understood as ensuring that the rendering task of focus sound sources can obtain sufficient central processing unit (CPU), graphics processing unit (GPU), or memory resources to maintain its preset rendering accuracy, frame rate, and latency requirements, avoiding stuttering, degraded sound quality, or increased latency due to insufficient resources.

[0057] This application's solution dynamically senses the system's computational load by introducing real-time monitoring of processor utilization and temperature. When the system load approaches or reaches a preset threshold—that is, when processor utilization or temperature is too high—it indicates that system resources may be insufficient to maintain high-quality operation of all rendering tasks simultaneously. At this point, by adjusting the ambient sound field rendering strategy, such as reducing its computational precision or update frequency, the computational resource consumption of ambient sound source rendering can be effectively reduced. It is precisely this dynamic resource allocation mechanism that, under conditions of limited system resources, prioritizes the rendering quality of focal sound sources crucial to user immersion, preventing them from being affected by system load fluctuations.

[0058] Through the above technical solution, this application effectively addresses the computational resource bottleneck problem that sound field rendering may face under complex virtual environments or high-load conditions. This solution intelligently monitors system performance and dynamically adjusts strategies for non-core rendering tasks, ensuring that the focal sound source always receives sufficient computational resources, thereby guaranteeing high-quality, low-latency rendering of the focal sound source. This not only enhances the auditory realism and presence of users in immersive experiences but also improves the robustness and stability of the entire sound field rendering system, avoiding performance degradation due to system overload and significantly optimizing the user experience.

[0059] In some of the embodiments described above in this application, a strategy is proposed to acquire processor utilization and processor temperature in real time, and adjust the ambient sound field rendering based on the processor utilization and processor temperature to ensure rendering computing resources for the focal sound source. However, in its implementation, if the specific method and triggering conditions of the adjustment strategy are not clearly defined, the adjustment of the ambient sound field rendering may be too abrupt, affecting the user's immersive experience, or it may fail to perform timely and effective resource scheduling when processor resources are scarce.

[0060] In response, this application further proposes the following steps for adjusting the ambient sound field rendering strategy based on processor utilization and processor temperature to ensure the rendering computing resources of the focal sound source: Obtain the utilization threshold and the temperature threshold of the processor; When the processor utilization exceeds a utilization threshold or the processor temperature exceeds a temperature threshold, a rendering strategy that progressively reduces the ambient sound field is adopted, and the rendering strategy includes at least one of the following: Reduce the accuracy of reflection calculations from environmental sound sources; Reduce the frequency of background sound updates.

[0061] Specifically, the processor utilization threshold and processor temperature threshold refer to the critical values ​​that the system pre-sets or dynamically calculates to determine whether the processor is in an overloaded state. These thresholds can be configured based on factors such as the processor model, heat dissipation performance, and expected system load. For example, the utilization threshold can be set as a percentage of the processor's maximum utilization, while the temperature threshold can be set as the upper limit of the processor's safe operating temperature. These thresholds are the basis for triggering subsequent rendering strategy adjustments. When the processor utilization exceeds the utilization threshold or the processor temperature exceeds the temperature threshold, it indicates that processor resources are becoming strained or overheating, requiring measures to be taken for resource scheduling. At this time, the system will adopt a rendering strategy that progressively reduces the ambient sound field. Progressive reduction means that the decrease in rendering quality or computational load is gradual and non-abrupt, aiming to minimize the impact on the user's immersion. This strategy avoids sudden performance degradation, thereby providing a smoother user experience. In practical applications, the rendering strategy includes at least one of the following: reducing the accuracy of reflection calculations of ambient sound sources; reducing the update frequency of background sound. Reducing the accuracy of ambient sound source reflection calculations can be understood as decreasing the number of simulations of sound wave reflection paths, simplifying the complexity of reflection models, or lowering the sampling rate of reflected sound during sound field rendering. For example, calculating third-order reflections can be simplified to calculating first-order or second-order reflections, or a coarser mesh can be used for reflection calculations. The aim is to reduce the computational load required for ambient sound source rendering, thereby freeing up processor resources. Reducing the update frequency of background sounds refers to lowering the real-time update or resampling frequency of non-critical ambient sounds (such as distant wind sounds, ambient noise, etc.). For example, background sounds that are updated 30 times per second can be reduced to 10 or 5 times per second, or updates can be paused at certain non-critical moments. The aim is to further save processor resources by reducing the computational overhead of infrequently updated background sounds.

[0062] This application's solution introduces processor utilization and temperature thresholds, providing the system with a clear and quantifiable criterion to identify states of processor resource scarcity or overheating. When processor resource constraints are detected, the system no longer simply adjusts the rendering strategy but instead employs a progressive reduction in the ambient sound field rendering strategy. This progressive adjustment mechanism combines specific measures such as reducing the accuracy of reflection calculations for ambient sound sources and decreasing the update frequency of background sounds. This allows the system to smoothly and non-aggressively reduce the rendering quality of non-critical ambient sound fields while ensuring sufficient rendering computational resources for focal sound sources. This avoids a sudden drop in rendering quality due to resource constraints, effectively solving the problem of abrupt rendering adjustments and impact on user immersion that may occur in basic solutions.

[0063] Through the above technical solution, this application can intelligently and precisely adjust the ambient sound field rendering strategy according to the actual operating state of the processor. By setting clear utilization and temperature thresholds, the system can respond promptly to changes in processor load, avoiding the impact of excessive resource consumption on the rendering of the core focus sound source. More importantly, by adopting a progressive degradation strategy and specifically reducing the reflection calculation accuracy of ambient sound sources or reducing the update frequency of background sounds, the degradation process of the ambient sound field becomes smoother and more natural, significantly reducing the user's perception of changes in rendering quality. Thus, while ensuring high-quality rendering of the focus sound source, it maximizes the continuity and comfort of the overall immersive experience.

[0064] In some embodiments described above, this application proposes obtaining processor utilization and temperature thresholds, and adjusting the ambient sound field rendering strategy based on these thresholds to ensure sufficient computational resources for rendering the focal sound source. However, in practical applications, simply obtaining preset or static processor utilization and temperature thresholds may not adequately adapt to performance differences under varying processor hardware configurations, cooling conditions, and long-term operating conditions. Improperly set thresholds may lead to the system failing to fully utilize its performance when cooling capacity allows, or facing overheating risks when cooling capacity is limited, thus affecting system stability and rendering effects. Therefore, this application further proposes a more intelligent and adaptive threshold acquisition method that dynamically determines these thresholds by evaluating the processor's cooling performance, thereby more accurately ensuring sufficient computational resources for rendering the focal sound source and optimizing the balance between system performance and stability.

[0065] The steps for obtaining the processor utilization threshold and the processor temperature threshold mentioned above include: The thermal conductivity performance of the processor's heat dissipation system was evaluated, and the thermal conductivity performance evaluation results were obtained. Based on the thermal conductivity evaluation results, correction parameters for the processor utilization threshold and the processor temperature threshold are determined; The processor utilization threshold and the processor temperature threshold are obtained by correcting the preset utilization rate and preset temperature according to the correction parameters.

[0066] Specifically, evaluating the thermal conductivity performance of a processor's cooling system involves quantitatively analyzing the heat transfer efficiency of the processor and its cooling components (such as heat sinks, fans, and heat pipes) under different loads and ambient temperatures. This can include measuring the processor's core temperature rise rate and stable temperature value at a specific power consumption, or analyzing heat distribution using thermal imaging technology. The thermal conductivity evaluation result can be a numerical value, such as thermal resistance, or a set of parameters describing heat dissipation efficiency. Its purpose is to obtain a true picture of the processor's heat dissipation capabilities during actual operation.

[0067] The correction parameters for determining the processor utilization and temperature thresholds based on the thermal conductivity evaluation results can be understood as adjusting the preset performance limits according to the processor's actual heat dissipation capacity. For example, if the evaluation results show excellent heat dissipation performance, a correction parameter that allows for higher utilization and temperature can be determined; conversely, if the heat dissipation performance is poor, the correction parameter will tend to reduce the allowable utilization and temperature. These correction parameters can be multiplicative factors, additive offsets, or other forms used to adjust the initial preset values.

[0068] In practical applications, adjusting preset utilization and temperature based on correction parameters to obtain processor utilization and temperature thresholds involves combining pre-set general utilization and temperature values ​​with correction parameters determined based on heat dissipation performance to arrive at dynamic thresholds that better suit current hardware and environmental conditions. For example, the preset utilization might be a general value; by multiplying or adding correction parameters, a utilization threshold optimized for a specific system can be obtained. The aim is to enable the system to respond more flexibly and accurately to the processor's actual thermal load capacity, avoiding performance bottlenecks or overheating risks caused by fixed thresholds.

[0069] The solution presented in this application addresses the aforementioned issues by obtaining precise data on the processor's actual thermal management capabilities through an evaluation of the processor's heat dissipation system's thermal conductivity. This precise thermal conductivity evaluation allows the system to determine correction parameters reflecting the current hardware's thermal characteristics. These correction parameters are used to dynamically adjust preset processor utilization and temperature thresholds, ensuring that the final thresholds are no longer static, universal values ​​but adaptable to the specific processor's thermal conditions. This dynamic adjustment mechanism ensures that when the processor's cooling capacity is strong, the performance threshold can be appropriately increased to allow for higher computational loads; conversely, when cooling capacity is limited, the threshold can be promptly reduced to prevent overheating. This maximizes system performance and maintains stability while ensuring sufficient computational resources for focused sound source rendering.

[0070] Through the above technical solution, this application can dynamically adjust the processor utilization threshold and processor temperature threshold according to the actual heat dissipation capacity of the processor. Compared with the solution of using fixed or empirical thresholds, this application can more accurately reflect the performance boundary of the current hardware environment and effectively avoid performance waste or system overheating caused by improper threshold settings. Therefore, while ensuring the computing resources for rendering the focus sound source, the system can make fuller use of processor performance, improve overall rendering efficiency and user experience stability, and extend hardware life.

[0071] In some of the embodiments described above in this application, although it is possible to distinguish and render the focus sound source and the ambient sound source separately, in complex virtual environments, when sudden non-focus acoustic events occur, these events may unexpectedly consume a large amount of computing resources, thereby affecting the rendering quality and real-time performance of the focus sound source that the user is currently focusing on. If the above problems are not resolved, users may experience instability or delay in the rendering of the focus sound source during the immersive experience, thereby reducing the overall immersion and user experience.

[0072] In this regard, this application further proposes an immersive sound field rendering method for acoustic speakers, which also includes: Monitoring non-focus acoustic events in a virtual environment; When the non-focus acoustic event occurs, the sound field rendering strategy is adjusted to ensure the rendering resources of the focus sound source. The sound field rendering adjustment strategy includes at least one of the following: The shared ambient sound field signal is rendered with progressively downgraded rendering. Reduce the update frequency of haptic feedback and non-critical visual effects to ensure that the focus sound source task can preempt processor resources.

[0073] Specifically, monitoring non-focal acoustic events in a virtual environment refers to the system continuously monitoring the activity of sound sources other than the focal sound source in the virtual environment. These non-focal acoustic events may include sudden background noise, collision sounds of distant objects, random sound effects in the environment, etc., and are characterized by the potential to generate high computational load in a short period of time.

[0074] When a non-focus acoustic event is detected, the system adjusts the sound field rendering according to a preset strategy. This progressive degradation rendering of the shared ambient sound field signal can be understood as freeing up computational resources by reducing its rendering accuracy or complexity without completely sacrificing ambient sound field perception. For example, this could involve reducing the number of reflection calculations for ambient sound sources, simplifying reverberation models, and lowering the sampling rate or bit rate of ambient sound sources. This degradation is progressive, meaning the system adjusts gradually based on resource constraints, avoiding sudden quality drops.

[0075] Furthermore, adjustment strategies can also include reducing the update frequency of haptic feedback and non-critical visual effects. Haptic feedback refers to providing tactile sensations to the user through vibration or other physical means, while non-critical visual effects refer to visual effects that have a minor impact on the core immersive experience but still consume rendering resources, such as particle effects and secondary lighting calculations. By reducing the update frequency of these non-core functions, processor resources can be effectively prioritized for rendering tasks of focused sound sources, ensuring that the audio processing of focused sound sources can be completed in a timely and high-quality manner.

[0076] This application's solution introduces a monitoring mechanism for non-focal acoustic events in the virtual environment, enabling timely detection of potential resource contention. Once a non-focal acoustic event that may affect the rendering of the focal sound source is identified, the system proactively implements resource allocation strategies. Specifically, by progressively downgrading the rendering of shared environmental sound field signals and reducing the update frequency of haptic feedback and non-critical visual effects, the system strategically releases computing resources occupied by non-core tasks. It is precisely this dynamic resource management and priority adjustment that ensures the rendering task of the focal sound source is prioritized when resources are limited or unexpected events occur, thereby maintaining its high resolution and low latency characteristics.

[0077] Through the above technical solution, this application can effectively address the resource contention problem caused by sudden non-focus acoustic events in virtual environments. This solution ensures that, in complex and dynamic environments, the focus sound source that the user is concerned with always has sufficient computing resources for high-quality rendering, avoiding a decrease or delay in the rendering quality of the focus sound source due to resource contention by other non-critical events. This significantly improves the stability and reliability of the immersive experience, enabling users to continuously obtain clear and accurate perception of the focus sound source, thereby enhancing the overall immersion and user satisfaction.

[0078] According to the solution proposed in this application, the system will detect the occurrence of this thunderstorm as a non-focal acoustic event. To ensure the rendering quality of the bird song, the focal sound source, the system will immediately adjust the sound field rendering strategy. Specifically, the system may progressively degrade the rendering of the waterfall sound, a shared ambient sound field signal, for example, by reducing its reverberation calculation accuracy from high to medium, or reducing the number of calculations for its reflection paths. At the same time, the system will also reduce the update frequency of non-critical visual effects related to the thunderstorm (such as the particle effects of lightning), and even temporarily reduce the intensity or frequency of haptic feedback (such as the vibration of the controller). Through these measures, the computing resources originally used to render the thunderstorm and ambient sound sources are partially released and prioritized for the rendering task of the bird song. As a result, users can still hear the bird song clearly and in real time, and although the sound and visual effects of the thunderstorm are simplified, the overall immersive experience is not seriously affected, and the experience of the focal sound source is effectively guaranteed.

[0079] In some embodiments described above in this application, a method is proposed to monitor non-focal acoustic events in a virtual environment and adjust the sound field rendering strategy according to the events to ensure rendering resources for focal sound sources. Specifically, the steps for monitoring non-focal acoustic events in a virtual environment can be further refined.

[0080] The steps for monitoring non-focal acoustic events in a virtual environment include: Receive real-time information from all sound sources, including the loudness and frequency of the sound sources; When the loudness of the sound source exceeds a loudness threshold and the frequency exceeds a frequency threshold, the sound source is determined to be a non-focal acoustic event.

[0081] Specifically, receiving real-time information from all sound sources means that the system continuously acquires the current acoustic attribute data of each sound source existing in the virtual environment. This real-time information may include, but is not limited to, the instantaneous loudness, dominant frequency, and spectral distribution of the sound source. Loudness can be understood as the perceived intensity of sound, usually related to sound pressure level, but focusing more on the subjective perception of sound volume by the human ear; frequency represents the speed of sound wave vibration, determining the pitch of the sound. This real-time information can be acquired through acoustic simulation engines or sensor data within the virtual environment.

[0082] Furthermore, when the loudness of the sound source exceeds a preset loudness threshold and the frequency exceeds a preset frequency threshold, the sound source is identified as a non-focus acoustic event. The loudness threshold and frequency threshold are preset parameters used to define what level of acoustic event is considered significant and likely to interfere with the user's perception of the focus sound source. For example, the loudness threshold can be set to a certain decibel value; when the loudness of a sound source exceeds this value, it indicates that it has high energy. The frequency threshold can be set to a certain Hertz range to filter out sounds within a specific frequency range, such as sudden sharp sounds or low rumbling sounds. By combining the loudness and frequency dimensions for judgment, non-focus acoustic events that may have a significant impact on the user's immersion can be identified more accurately.

[0083] This application's solution, by setting loudness and frequency thresholds, effectively filters out potentially disruptive acoustic events from numerous non-focal sound sources in a virtual environment. Traditionally, systems may require complex analysis of all non-focal sound sources or rely on a single dimension (such as distance) for judgment, which can lead to misjudgments or wasted resources. By introducing loudness and frequency as judgment criteria, the system can more accurately identify non-focal acoustic events that are more audibly prominent and more likely to distract the user. For example, an explosion sound that is far from the user but has extremely high loudness and a unique frequency, or a background ambient sound that is close but has low loudness, can be distinguished through this mechanism, ensuring that only non-focal events that truly require system intervention trigger subsequent resource adjustment strategies. This acoustic feature-based identification method enables the system to respond more intelligently to changes in the virtual environment, avoiding unnecessary resource allocation adjustments and thus optimizing overall rendering efficiency.

[0084] Through the above technical solution, this application can achieve accurate identification of non-focal acoustic events in a virtual environment. This judgment mechanism based on both loudness and frequency avoids the inaccuracies that may arise from judging based on a single dimension. For example, it avoids misjudging some nearby sound sources with insignificant loudness or frequency as interference events, and also avoids overlooking important events that are far away but have prominent acoustic characteristics. Therefore, the system can more effectively identify non-focal acoustic events that truly require resource allocation, ensuring that when such events occur, the sound field rendering strategy can be adjusted promptly and accurately, guaranteeing rendering resources for focal sound sources, maintaining the user's immersion in the core content, and optimizing the efficiency of system resource utilization.

[0085] In some embodiments, this application proposes an acoustic speaker immersive sound field rendering system, the system comprising: The user information acquisition module is used to acquire the user's location and head posture in the virtual space; The sound source information acquisition module is used to acquire the location and intensity of sound sources in the virtual space; The focus sound source determination module is used to identify the focus sound source that the user is currently paying attention to based on the user's position, head posture, sound source position, and intensity. An ambient sound source determination module is used to identify sound sources other than the focal sound source as ambient sound sources; A shared ambient sound field signal acquisition module is used to perform sound field rendering on the ambient sound source to obtain a shared ambient sound field signal, wherein the computing resources of the ambient sound source are lower than those of the focal sound source; The personalized high-resolution signal acquisition module performs a convolution operation on the original audio signal of the focal sound source and the head acoustic transfer function to obtain a personalized high-resolution signal. The binaural signal acquisition module is used to superimpose the shared ambient sound field signal and the personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array.

[0086] Specifically, the user information acquisition module can be configured to be integrated into the head-mounted display device, acquiring the user's head motion data in real time through a built-in inertial measurement unit (IMU) and converting it into position and posture information in virtual space. As a preferred implementation, the module can also interact with external optical or electromagnetic tracking systems to receive user position and head posture data provided by these systems.

[0087] The sound source information acquisition module can be designed to interface with the rendering engine or physics engine of the virtual scene, receiving preset or dynamically generated sound source data from the virtual environment in real time. For example, in a virtual training scene, the sound source location and intensity information of events such as explosions and gunshots will be calculated by the virtual engine and transmitted to this module.

[0088] The focus sound source identification module can have a built-in recognition algorithm. For example, by calculating the Euclidean distance between the sound source and the user's head, and combining this with the user's head orientation, it can determine whether the sound source is located within a preset focus area in front of the user. Furthermore, the intensity of the sound source can also be used as a criterion; for example, sound sources with an intensity exceeding a certain threshold will be preferentially identified as potential focus sound sources.

[0089] The ambient sound source identification module has a relatively straightforward function: it is configured to classify all sound sources other than the focus sound source identified by the focus sound source identification module as ambient sound sources. This means that once a sound source is not determined to be the user's current focus, it will be automatically marked as an ambient sound source by this module for subsequent differentiated processing.

[0090] The shared ambient sound field signal acquisition module is designed to render using computational resources lower than those of the focal sound source. For example, a simplified head acoustic transfer function (HRTF) model can be used, or in acoustic ray tracing algorithms, only the primary reflection paths can be calculated while ignoring secondary reflections and scattering, thereby reducing computational complexity. Furthermore, this module can further conserve computational resources by reducing the update frequency of background sound.

[0091] The personalized high-resolution signal acquisition module is responsible for convolving the raw audio signal from the focal sound source with the head acoustic transfer function to obtain a personalized high-resolution signal. This module typically includes a high-performance digital signal processor (DSP) or graphics processing unit (GPU) to perform complex real-time convolution operations. The head acoustic transfer function can be retrieved from a pre-stored database and adjusted in real time according to the user's head posture to ensure accurate spatial localization and externalization of the sound.

[0092] The binaural signal acquisition module outputs the generated binaural signals to the acoustic speaker array for playback. The acoustic speaker array can integrate beamforming technology, which, by precisely controlling the delay and gain of each speaker, ensures that a sound pressure peak is formed at the target location of the user's ear, thereby providing the user with a highly immersive three-dimensional sound field experience.

[0093] The acoustic horn immersive sound field rendering system of this application effectively solves the problem of balancing spatial resolution and computational efficiency in sound field rendering in the prior art through its modular design and differentiated sound source processing strategy.

[0094] 1. Differentiated Rendering of Focused and Ambient Sound Sources: Traditional methods often treat all sound sources equally and render them with high precision, resulting in wasted computational resources. This application intelligently identifies the focused sound sources that the user is interested in and renders them with high resolution and personalization, while using a low-computational-resource rendering strategy for non-focused ambient sound sources. This significantly reduces the overall computational load while ensuring the user's core auditory experience.

[0095] 2. Combination of Personalized High-Resolution and Shared Ambient Sound Field: This application superimposes the personalized high-resolution signal obtained by convolving the focal sound source with the head acoustic transfer function, and the shared ambient sound field signal obtained by rendering the ambient sound source. This combination method not only ensures the user's accurate perception and externalization of key sound sources, but also provides a rich and realistic background sound field, enhancing the sense of immersion.

[0096] 3. Optimized allocation of computing resources: By employing different rendering strategies for different types of sound sources, this application can allocate computing resources more rationally, ensuring the rendering quality of the focus sound source while avoiding performance bottlenecks caused by processing a large number of non-critical sound sources. The computing resources for ambient sound sources are lower than those for focus sound sources, which directly solves the system sluggishness problem caused by excessive computation in existing technologies.

[0097] Compared to existing technologies that blindly increase the amount of digital filtering data or introduce complex personalized information to improve the precision of sound field rendering, leading to a surge in computational load, the system in this application can utilize computing resources more intelligently and efficiently. This hierarchical rendering strategy enables the system to provide users with a high-quality, low-latency immersive sound field experience in complex and dynamic virtual environments, significantly improving the realism and interactivity of applications such as virtual reality training. The foregoing has provided a detailed description of the preferred embodiments of this application. However, this application is not limited to the above-described embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined in this application.

Claims

1. A method for immersive sound field rendering using acoustic speakers, characterized in that, include: Obtain the user's location and head pose in the virtual space; Obtain the location and intensity of a sound source in virtual space; The user's current focus sound source is identified based on the user's location, head posture, sound source location, and intensity. Identify sound sources other than the focal sound source as ambient sound sources; Sound field rendering is performed on the ambient sound source to obtain a shared ambient sound field signal, wherein the computational resources of the ambient sound source are lower than those of the focal sound source; A personalized high-resolution signal is obtained by convolving the original audio signal of the focal sound source with the head acoustic transfer function. The shared ambient sound field signal is superimposed with the personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array.

2. The method according to claim 1, characterized in that, The step of convolving the original audio signal from the focal point sound source with the head acoustic transfer function to obtain a personalized high-resolution signal includes: The user's head orientation is obtained based on the head posture. Based on the head orientation, retrieve the corresponding head acoustic transfer function from the pre-stored head transfer function database; A personalized high-resolution signal is obtained by convolving the original audio signal of the focal sound source with the retrieved head acoustic transfer function.

3. The method according to claim 1, characterized in that, The step of superimposing the shared ambient sound field signal with the personalized high-resolution signal to generate a binaural signal, and playing it through an acoustic speaker array, includes: Obtain the position and pose information of virtual objects in virtual space, as well as the geometric information of virtual objects; For each focal sound source, calculate the sound propagation path from the focal sound source to the user's ears; Based on the preset operation path of the current training task, and according to the position and pose information of the virtual object and the geometric information of the virtual object, possible occlusion events are predicted. When the occlusion event is predicted, the shared ambient sound field signal and the personalized high-resolution signal are superimposed to generate a binaural signal, which is played through an acoustic speaker array. The acoustic speakers use beamforming technology to control the delay and gain to ensure that the binaural signal forms a sound pressure peak at the target position of the user's ear.

4. The method according to claim 1, characterized in that, The step of identifying the user's current focus sound source based on the user's location, head posture, sound source location, and intensity includes: The Euclidean distance is obtained by calculating the distance between the sound source location and the user's head based on the user's location and the sound source location. The potential focal sound source is determined based on the Euclidean distance; The user's head orientation is obtained based on the head posture. If the potential focus sound source is located within a preset angle range in front of the user's head, or if the intensity of the potential focus sound source is greater than a preset threshold, the potential focus sound source is identified as the focus sound source currently being focused on by the user.

5. The method according to claim 1, characterized in that, The method further includes: Real-time monitoring of processor utilization and processor temperature; The ambient sound field rendering strategy is adjusted based on the processor utilization and the processor temperature to ensure the rendering computing resources for the focal sound source.

6. The method according to claim 5, characterized in that, The step of adjusting the ambient sound field rendering strategy based on the processor utilization and the processor temperature to ensure the rendering computing resources of the focal sound source includes: Obtain the processor utilization threshold and the processor temperature threshold; When the processor utilization exceeds a utilization threshold or the processor temperature exceeds a temperature threshold, a rendering strategy that progressively reduces the ambient sound field is adopted, and the rendering strategy includes at least one of the following: Reduce the accuracy of reflection calculations from environmental sound sources; Reduce the frequency of background sound updates.

7. The method according to claim 6, characterized in that, The steps of obtaining the processor utilization threshold and the processor temperature threshold include: The thermal conductivity performance of the processor's heat dissipation system was evaluated, and the thermal conductivity performance evaluation results were obtained. Based on the thermal conductivity evaluation results, the correction parameters for the processor utilization threshold and the processor temperature threshold are determined; The processor utilization threshold and the processor temperature threshold are obtained by correcting the preset utilization rate and preset temperature according to the correction parameters.

8. The method according to claim 1, characterized in that, The method further includes: Monitoring non-focus acoustic events in a virtual environment; When the non-focus acoustic event occurs, the sound field rendering strategy is adjusted to ensure the rendering resources of the focus sound source. The sound field rendering adjustment strategy includes at least one of the following: The shared ambient sound field signal is rendered with progressively downgraded rendering. Reduce the update frequency of haptic feedback and non-critical visual effects to ensure that the focus sound source task can preempt processor resources.

9. The method according to claim 8, characterized in that, The steps for monitoring non-focal acoustic events in a virtual environment include: Receive real-time information from all sound sources, including the loudness and frequency of the sound sources; When the loudness of the sound source exceeds a loudness threshold and the frequency exceeds a frequency threshold, the sound source is determined to be a non-focal acoustic event.

10. An immersive sound field rendering system for acoustic speakers, characterized in that, The system includes: The user information acquisition module is used to acquire the user's location and head posture in the virtual space; The sound source information acquisition module is used to acquire the location and intensity of sound sources in the virtual space; The focus sound source determination module is used to identify the focus sound source that the user is currently paying attention to based on the user's position, head posture, sound source position, and intensity. An ambient sound source determination module is used to identify sound sources other than the focal sound source as ambient sound sources; A shared ambient sound field signal acquisition module is used to perform sound field rendering on the ambient sound source to obtain a shared ambient sound field signal, wherein the computing resources of the ambient sound source are lower than those of the focal sound source; The personalized high-resolution signal acquisition module performs a convolution operation on the original audio signal of the focal sound source and the head acoustic transfer function to obtain a personalized high-resolution signal. The binaural signal acquisition module is used to superimpose the shared ambient sound field signal and the personalized high-resolution signal to generate a binaural signal, which is then played through an acoustic speaker array.