Audio processing method and apparatus

By setting up multiple surround sound units in smart devices, and combining real-time user location information and target filtering functions to calculate output parameters, the problems of poor low-frequency response and flat sound field in smart device audio output are solved, achieving high-quality audio immersion and spatial directivity.

CN122120694APending Publication Date: 2026-05-29SHENZHEN XINYANG CHUANGZHI TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN XINYANG CHUANGZHI TECHNOLOGY CO LTD
Filing Date
2026-04-24
Publication Date
2026-05-29

Smart Images

  • Figure CN122120694A_ABST
    Figure CN122120694A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an audio processing method and device, comprising: obtaining current position information of a user; determining a virtual sound emitting position corresponding to each sound emitting unit according to the current position information and a sound emitting position of each sound emitting unit; determining a target filter coefficient of a target filter function matched with the current position information; determining a target output parameter corresponding to each sound emitting unit according to the target filter coefficient, the current position information, the virtual sound emitting position and the sound emitting position of each sound emitting unit, the target output parameter comprising one or more of a sound wave delay time, a phase offset and a sampling domain delay parameter of the sound emitting unit when outputting a target audio signal; and controlling each sound emitting unit to output the target audio signal based on the target output parameter. In this way, the user can perceive a stereo surround sound, and the immersion and experience of the user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer technology, and to, but is not limited to, an audio processing method and apparatus. Background Technology

[0002] With the development of technology, smart devices are increasingly widely used in various fields, encompassing smart wearable devices, portable smart terminals, smart home control devices, smart vehicle terminals, and many other forms. They have become indispensable tools in people's work, life, and entertainment. However, due to the pursuit of portability and miniaturization in most smart devices, their internal space is extremely limited, making it impossible to accommodate large speaker units and independent acoustic cavities. This results in problems such as poor low-frequency response, thin sound quality, and a flat soundstage in audio output, making it difficult to meet users' needs for a high-quality audio experience.

[0003] Therefore, how to achieve high-quality audio output within the limited space of smart devices, while providing accurate spatial information and dynamic adaptation capabilities to enhance user immersion and recognition ability, is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the audio processing method provided in this application embodiment can significantly improve the audio immersion and spatial directivity, and solve the problems of flat sound field and lack of spatial sense in traditional devices.

[0005] The audio processing method and apparatus provided in this application are implemented as follows: In a first aspect, embodiments of this application provide an audio processing method applied to a smart device, the smart device including multiple sound-emitting units arranged in a surround manner, the method including: Obtain the user's current location information; Based on the current location information and the sound emission position of each of the sound emission units, a virtual sound emission position corresponding to each of the sound emission units is determined, and the virtual sound emission position is located in a specific direction of the user; Determine the target filter coefficients of a target filter function that matches the user's current location information. The target filter function is used to characterize the relationship between the vocal parameters and the vocal frequency when the audio signal propagates from the virtual vocal position to the user. The vocal parameters include amplitude and phase. Based on the target filtering coefficient, the current position information, the virtual sound position, and the sound position of each sound unit, the target output parameters corresponding to each sound unit are determined. The target output parameters include one or more of the following: sound wave delay time, phase offset, and sampling domain delay parameters when the sound unit outputs the target audio signal. Each of the aforementioned sound-producing units is controlled to output the target audio signal based on the target output parameters.

[0006] By implementing this embodiment, the user's location is acquired in real time and a virtual sound position is constructed by combining the surround sound unit arrangement. With the target filtering coefficient, accurate sound field rendering is achieved, allowing the listener to perceive stable stereo surround sound, significantly improving the audio immersion and spatial directivity, and solving the problem of flat sound field and lack of spatial sense in traditional devices.

[0007] In one embodiment, the method further includes: If the user's current location information is updated, the target output parameters corresponding to each sound unit are re-determined based on the updated location information to obtain the updated output parameters; Each of the aforementioned sound-producing units is controlled to output the target audio signal based on the updated output parameters.

[0008] By implementing this embodiment, the sound field parameters are refreshed in real time when the user's position changes dynamically, ensuring that the virtual sound source always follows the listener's movement, forming a dynamic surround sound experience where "sound follows people," and improving the naturalness and coherence of human-computer interaction.

[0009] In one embodiment, determining the target filter coefficients of the target filter function that match the user's current location information includes: Obtain the initial filter coefficients of the target filter function that match the current location information; The initial filter coefficients are subjected to bilinear interpolation to obtain target filter coefficients that match the current position information.

[0010] By implementing this embodiment, high-precision matching of the target filter coefficients is achieved through bilinear interpolation, which accurately matches the user's head position and auditory characteristics, improves spatial audio fidelity, and avoids sound field distortion and positioning ambiguity.

[0011] In one embodiment, determining the target output parameters corresponding to each sound unit based on the target filtering coefficients, the current position information, the virtual sound position, and the sound position of each sound unit includes: Based on the location of the virtual sound source and the sound emission location of each of the sound emission units, the sound wave delay time of each of the sound emission units when outputting the target audio signal is determined; Based on the sound-emitting position of each sound-emitting unit, the target filtering function, and the current position information, the phase offset and sampling domain delay parameters of the sound-emitting unit when outputting the target audio signal are determined.

[0012] By implementing this embodiment, precise calculation of delay, phase shift and sampling domain delay is achieved to realize accurate interference synthesis of sound waves from multiple sound units, construct a stable virtual sound source, and realize a surround sound field comparable to a large multi-channel system in a small device space.

[0013] In one embodiment, controlling each of the sound-producing units to output the target audio signal based on the target output parameters includes: Based on the current location information, the location information of the smart device, and the target filtering coefficient, the driving signal parameters of each sound-emitting unit are determined. The driving signal parameters include an amplitude coefficient, which is used to characterize the output amplitude of the sound-emitting unit. The target audio signal is generated based on the target filter coefficient, the sound wave delay time, the phase shift, the sampling domain delay parameter, and the amplitude coefficient. Each of the aforementioned sound-producing units is controlled to output the target audio signal based on the driving signal parameters.

[0014] By implementing this embodiment, the amplitude coefficient and multi-dimensional output parameters are jointly controlled to achieve independent and precise driving of each sound unit, ensuring balanced and stable audio output at different positions and distances, and improving sound quality clarity and sound field consistency.

[0015] In one embodiment, controlling each of the sound-producing units to output the target audio signal based on the drive signal parameters includes: Obtain the operating noise of the smart device; Generate the antiphase acoustic wave of the operating noise; The antiphase sound wave is superimposed on the target audio signal to obtain a canceled audio signal; Each of the aforementioned sound-generating units is controlled to output the canceled audio signal based on the driving signal parameters.

[0016] By implementing this embodiment, the volume is dynamically optimized when playing specific audio such as key voices and prompts, and noise is further suppressed by combining psychoacoustic masking effects to ensure that important information is clearly audible and improve the reliability of interaction.

[0017] In one embodiment, the method further includes: When specific audio is detected being played, the volume of the canceled audio signal is adjusted to the target volume based on the user's current location information and the location information of the smart device. Each of the aforementioned sound-producing units is controlled to output the canceled audio signal at the target volume.

[0018] By implementing this embodiment, the volume is dynamically optimized when playing specific audio such as key voices and prompts, and noise is further suppressed by combining psychoacoustic masking effects to ensure that important information is clearly audible and improve the reliability of interaction.

[0019] In one embodiment, the sound guiding channel of the smart device is spiral-shaped.

[0020] By implementing this embodiment, the spiral sound guide channel physically extends the propagation path of low-frequency sound waves in a confined space, significantly enhancing the low-frequency response and solving the defects of insufficient bass and thin sound quality in micro devices.

[0021] Secondly, embodiments of this application provide an audio processing apparatus, including: The acquisition module is used to obtain the user's current location information; The determining module is used to determine the virtual sound position corresponding to each of the sound units based on the current location information and the sound position of each sound unit. The virtual sound position is located in a specific direction of the user. The sound units are arranged in a surround manner and the sound units are set in the smart device. Determine the target filter coefficients of a target filter function that matches the user's current location information. The target filter function is used to characterize the relationship between the vocal parameters and the vocal frequency when the audio signal propagates from the virtual vocal position to the user. The vocal parameters include amplitude and phase. Based on the target filtering coefficient, the current position information, the virtual sound position, and the sound position of each sound unit, the target output parameters corresponding to each sound unit are determined. The target output parameters include one or more of the following: sound wave delay time, phase offset, and sampling domain delay parameters when the sound unit outputs the target audio signal. The control module is used to control each of the sound-generating units to output the target audio signal based on the target output parameters.

[0022] Thirdly, embodiments of this application provide a computer device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement the method described in embodiments of this application.

[0023] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in embodiments of this application. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0025] Figure 1 This is an application scenario diagram of the audio processing method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the smart device provided in the embodiments of this application; Figure 3 This is a schematic diagram illustrating the implementation flow of an audio processing method provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the implementation process of a parameter update method provided in an embodiment of this application; Figure 5 This is a schematic diagram illustrating the implementation flow of another audio processing method provided in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the implementation process of outputting the target audio signal according to an embodiment of this application; Figure 7 This is a schematic diagram of the structure of the audio processing device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0028] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0029] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0030] With the development of technology, smart devices are increasingly widely used in various fields, encompassing smart wearable devices, portable smart terminals, smart home control devices, smart vehicle terminals, and many other forms. They have become indispensable tools in people's work, life, and entertainment. However, due to the pursuit of portability and miniaturization in most smart devices, their internal space is extremely limited, making it impossible to accommodate large speaker units and independent acoustic cavities. This results in problems such as poor low-frequency response, thin sound quality, and a flat soundstage in audio output, making it difficult to meet users' needs for a high-quality audio experience.

[0031] Therefore, how to achieve high-quality audio output within the limited space of smart devices, while providing accurate spatial information and dynamic adaptation capabilities to enhance user immersion and recognition ability, is an urgent problem to be solved.

[0032] In view of this, embodiments of this application provide an audio processing method applied to a smart device. Figure 1 This is an application scenario diagram of the audio processing method provided in one embodiment. For example... Figure 1 As shown, smart devices may include, but are not limited to, smart robots, mobile phones, wearable devices (such as smartwatches, smart bracelets, smart glasses, etc.), tablets, laptops, in-vehicle terminals, PCs (Personal Computers), etc.

[0033] The functionality achieved by this method can be implemented by the processor in the smart device calling program code. Of course, the program code can be stored in computer storage media. Therefore, the smart device includes at least a processor and a storage media.

[0034] Figure 2 This is a schematic diagram of the structure of a smart device provided in an embodiment of this application.

[0035] For example, such as Figure 2 As shown, the smart device 20 may include a processor 201, an external memory interface 202, an internal memory 203, and a universal serial bus (USB) interface 204.

[0036] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the smart device 20. In other embodiments of this application, the smart device 20 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0037] Processor 201 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors. For example, processor 201 may be a smart terminal CPU, such as a Snapdragon series processor. In some embodiments, processor 201 may include one or more interfaces. Interfaces may include inter-integrated circuit (I2C) interfaces, inter-integrated circuit sound (I2S) interfaces, pulse code modulation (PCM) interfaces, universal asynchronous receiver / transmitter (UART) interfaces, mobile industry processor interfaces (MIPI) interfaces, general-purpose input / output (GPIO) interfaces, subscriber identity module (SIM) interfaces, and / or universal serial bus (USB) interfaces, etc.

[0038] The external storage interface 202 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the smart device 20. The external storage card communicates with the processor 201 through the external storage interface 202 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0039] The internal memory 203 can be used to store computer executable program code, which includes instructions. The internal memory 203 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of the smart device 20 (such as audio data, phonebook, etc.).

[0040] In addition, the smart devices involved in the embodiments of this application may also be equipped with an operating system, on which various types of applications can be installed and run. The embodiments of this application do not limit this.

[0041] Figure 3 This is a schematic diagram illustrating the implementation flow of the audio processing method provided in an embodiment of this application. Figure 3 As shown, the method may include the following steps: Step 301: Obtain the user's current location information.

[0042] The audio processing method of this application embodiment can be applied to various smart devices. For ease of description, a robot is used as an example of a smart device.

[0043] Understandably, humanoid robots or micro-robots, which can be used in public places or homes, have limited internal space, making it impossible to accommodate large speaker units and independent chambers. This results in poor low-frequency response and thin sound quality. In addition, the sound-emitting units (such as speakers) are usually concentrated in the head or torso, resulting in a concentrated sound output that cannot naturally form a surround sound field.

[0044] Based on this, in this embodiment of the application, the sound guiding channel of the smart device can be configured to be spiral-shaped. Alternatively, the sound guiding channel of the smart device can be configured to be maze-shaped.

[0045] For example, by using 3D printing technology to create compact acoustic cavities with spiral or labyrinthine sound guiding channels within the limited space of a robot's head or torso, the problem of limited internal space in a robot can be effectively solved.

[0046] For example, a spiral sound guide channel in the head (50mm long, 5mm diameter) or a labyrinthine sound guide channel in the torso (80mm long, 8mm diameter) can be used to extend the low-frequency propagation path and enhance the low-frequency response. This structure physically extends the propagation path of low-frequency sound waves, achieving low-frequency enhancement in a very small space, significantly improving sound quality and response frequency range, enabling the robot to achieve high-quality audio output.

[0047] Optionally, the smart device provided in this application embodiment may include multiple sound-emitting units, which are arranged in a surround or spherical pattern. In this way, by employing multiple sound-emitting units in a spherical or hemispherical surround array, replacing traditional stereo speakers or mono output methods, a natural surround sound field is achieved. This design breaks through physical limitations, achieving a surround effect previously only achievable with multi-speaker systems, thus enhancing the user's immersion and experience.

[0048] Optionally, multiple sound-emitting units can be evenly distributed around the robot's head, shoulders, and chest.

[0049] Optionally, the sound-generating unit may include a speaker and / or a microphone.

[0050] Alternatively, smart devices may also incorporate vision devices such as vision sensors, LiDAR, and depth cameras.

[0051] Optionally, a vision device can be used to detect whether a user has entered the interaction area (e.g., 2 to 5 meters away from the robot). If the user has been detected to have entered the interaction area, the robot's head can be controlled to automatically face the user.

[0052] Optionally, after detecting that a user has entered the interactive area, there are no restrictions on how the user's current location information is obtained.

[0053] For example, the Time Difference of Arrival (TDOA) algorithm can be used to calculate the straight-line distance r between the user and the robot, the azimuth angle θ, and the pitch angle φ, thereby solving for the user's current position information.

[0054] Alternatively, a delay-based sound unit method can be used for localization, consisting of M microphones, with microphone coordinates P. m (x m ,y m ,z m The user's sound source coordinates are P(x,y,z), the sound wave propagation speed is c=343m / s, and the time difference between the sound waves arriving at the sound source unit and the reference sound source unit (m=1) is τ. m Then, as shown in Formula 1: τ m =(|P P m | |P P1|) / c;|P P m |= (Formula 1); Therefore, the straight-line distance r between the user and the robot can be determined. .

[0055] In this way, by combining the robot's built-in vision module and sound unit, real-time perception of the user's position, distance, and movement trajectory relative to the robot is achieved.

[0056] Optionally, an initial wavefield synthesis parameter library can also be obtained, including the number of sound-generating units and the sampling frequency of the sound-generating units.

[0057] Step 302: Based on the current location information and the sound position of each sound unit, determine the virtual sound position corresponding to each sound unit. The virtual sound position is located in a specific direction of the user.

[0058] Optionally, based on the user's current location information, such as position P(x,y,z), direction angle θ, φ, and combined with interaction requirements (to allow the user to perceive natural spatial audio), the virtual sound position corresponding to each sound unit can be preset (such as directly in front of the user, to the side, to match the user's hearing habits).

[0059] By combining the distance between the user and the robot, the virtual sound source position and the virtual distance between the user and the robot are determined (it needs to match human auditory perception to ensure the spatial sense of the virtual sound source is realistic), and then the virtual sound source coordinates V(xv,yv,zv) are initially calculated.

[0060] Optionally, the initial sound source volume corresponding to each sound unit can be further determined, such as through the basic logic of the sound pressure level attenuation formula (SPL(r)=SPL(r0)). 20×lg(r / r0), where r0=1m is the reference distance), calculate the initial sound source volume corresponding to each sound unit to ensure that the user can hear a moderate sound source volume at the current position.

[0061] In this way, by adjusting the virtual sound-emitting position of the sound-emitting unit according to the user's position, the sound emitted by the robot is not fixed, so that the user can perceive that the sound is coming from all directions, thus realizing the 3D audio function.

[0062] Step 303: Determine the target filter coefficients of the target filter function that match the user's current location information. The target filter function is used to characterize the relationship between the sound parameters and the sound frequency when the audio signal propagates from the location of the smart device to the user. The sound parameters include amplitude and phase.

[0063] Understandably, sound will sound different to a user when it comes from different directions, not because of the volume, but because the timbre of the audio is altered by the head, ears, and shoulders.

[0064] Therefore, to enable users to clearly perceive the timbre of the audio, a target filtering function can be used to process the audio. The target filtering function characterizes the relationship between the sound parameters and the sound frequency as the audio signal propagates from the location of the smart device to the user. These sound parameters include amplitude and phase. The target filtering function can be used to simulate the filtering parameters resulting from the reflection, scattering, and diffraction of sound waves through the human head, auricle, and torso.

[0065] Optionally, the target filtering function can be the head-related transfer function (HRTF), which is a digital filter that can make the sound directional, so that the sound can present a stereo effect.

[0066] Optionally, when determining the target filter coefficients of the target filter function that matches the user's current location information, a pre-stored HRTF database can be invoked.

[0067] For example, the database can contain HRTF data for 500 different head models, with a sampling frequency of 48kHz and a frequency band of 20Hz-20kHz.

[0068] Subsequently, based on the user's current location information, the corresponding header-related transfer function can be extracted, and the corresponding target filtering coefficients can be determined.

[0069] Optionally, the target filtering coefficients can be different depending on the sampling frequency of the sound-generating unit.

[0070] Optionally, in some embodiments, the target filter coefficients can be determined by performing steps 503 to 504 in the following embodiments.

[0071] Step 304: Based on the target filtering coefficient, current position information, virtual sound position, and the sound position of each sound unit, determine the target output parameters corresponding to each sound unit. The target output parameters include one or more of the following parameters when the sound unit outputs the target audio signal: sound wave delay time, phase offset, and sampling domain delay.

[0072] Optionally, the user's current location information can be used as the center of the sound field, and the preset virtual sound source position can be used as the virtual sound source point. Combined with the physical installation coordinates of each sound unit on the smart device, the path difference and time difference of the audio signal from the virtual sound source to the user and from each sound unit to the user are calculated to obtain the sound wave delay time. This time is used to control the signal triggering time of each sound unit, so that the sound waves of multiple units are accurately superimposed at the user's ear to form a stable and directional virtual sound image.

[0073] Based on the angular frequency of the sound wave and the aforementioned sound wave delay time, the phase shift is further calculated, enabling the output signals of each sound unit to form controllable interference in space, accurately constructing a three-dimensional spatial orientation sense of front, back, left, right, up, and down, and avoiding sound image drift and sound field blurring.

[0074] Meanwhile, by combining the system sampling frequency, the sound wave delay time is converted into a sampling domain delay parameter, and precise time delay control is achieved through digital signal processing. This adapts to the real-time processing and output of digital audio systems, ensuring that wave field synthesis operates stably and with low latency on embedded platforms.

[0075] By jointly calculating the aforementioned sound wave delay time, phase shift, and sampling domain delay parameters, each sound unit outputs an independent and coordinated audio signal. Based on the different time delay and other parameters of multiple sound units, a continuous, natural, and highly accurate dynamic spatial sound field is reconstructed without relying on a large-size acoustic cavity, significantly enhancing the user's auditory immersion.

[0076] Step 305: Control each sound unit to output the target audio signal based on the target output parameters.

[0077] Optionally, the calculated acoustic delay time, phase shift, and sampling domain delay parameters can be fused with the target filter coefficient matching the user position and the amplitude coefficient obtained based on distance attenuation to generate an independent driving signal for the corresponding sound unit.

[0078] Intelligent devices can perform digital-to-analog conversion and power amplification of the driving signal according to a preset sampling rate, so that each sound unit can emit sound synchronously according to a specified timing, phase and amplitude. By utilizing the propagation and interference effect of sound waves in space, a spatial sound image consistent with the virtual sound position can be accurately reconstructed at the user's ear position, so that the user can perceive a three-dimensional surround sound audio with clear direction, distance and layer.

[0079] During signal output, each sound unit works in concert, without relying on large acoustic cavities and multi-track hardware, to achieve stable, continuous, and highly accurate dynamic spatial audio playback in confined spaces such as humanoid robots, thereby enhancing auditory immersion and natural interaction.

[0080] Understandably, a user's location may change while listening to audio signals using a smart device. Therefore, after the user's location changes, the target output parameters of the sound unit may no longer be adapted, failing to achieve the desired "sound follows the user" listening effect.

[0081] Based on this, in the embodiments of this application, the following can also be performed: Figure 4 The parameter update method shown is designed to ensure that users can still have a continuous, seamless, and highly immersive spatial audio experience while moving.

[0082] Step 401: If the user's current location information is updated, the target output parameters corresponding to each sound unit are re-determined based on the updated location information to obtain the updated output parameters.

[0083] The system monitors the user's position changes in real time using a visual sensor and microphone array. When the user's orientation, distance, or movement trajectory relative to the smart device changes, the sound field parameter recalculation process is immediately triggered. Based on the updated user's 3D coordinates, virtual sound-emitting position, and physical coordinates of each sound-emitting unit, the sound wave delay time, phase shift, and sampling domain delay parameters are recalculated. Then, bilinear interpolation is used to match the target filter coefficients corresponding to the new position, and sound pressure level attenuation correction is completed simultaneously. Finally, a complete set of updated output parameters adapted to the user's new position is generated.

[0084] Step 402: Control each sound unit to output the target audio signal based on the updated output parameters.

[0085] Optionally, the calculated updated output parameters can be sent to the driver modules of each sound unit in real time, replacing the original output parameters and driving each sound unit to emit sound synchronously according to the new timing, phase, and amplitude. Through real-time and continuous parameter refresh and signal output, the spatial virtual sound image is always stably pointed to the user's current position, realizing a dynamic surround sound field effect of "sound following the person's movement," ensuring that the user can still obtain a continuous, uninterrupted, and highly immersive spatial audio experience while moving.

[0086] In this embodiment, by acquiring the user's location in real time and constructing a virtual sound position by combining the surround sound unit arrangement, and by using the target filtering coefficient to achieve accurate sound field rendering, the listener can perceive stable stereo surround sound, significantly improving the audio immersion and spatial directivity, and solving the problem of flat sound field and lack of spatial sense in traditional devices.

[0087] Figure 5 This is a schematic diagram illustrating the implementation flow of the audio processing method provided in an embodiment of this application. Figure 5 As shown, the method may include the following steps: Step 501: Obtain the user's current location information.

[0088] The method of executing step 501 is the same as the method of executing step 301 in the above embodiment, and will not be repeated here. For example, the depth camera or vision sensor on the smart device can be used in conjunction with a microphone array to realize the real-time detection of the user's three-dimensional position, and obtain the user's azimuth angle, pitch angle and straight-line distance relative to the device, so as to provide accurate positional basis for subsequent sound field calculation.

[0089] Step 502: Based on the current location information and the sound position of each sound unit, determine the virtual sound position corresponding to each sound unit. The virtual sound position is located in a specific direction of the user.

[0090] Here, the method of executing step 502 is the same as the method of executing step 302 in the above embodiment, and will not be described again here. Taking the user's current position as the listening center, and combining the preset interaction logic and spatial audio rendering requirements, stable virtual sound points are set in front of, to the side of or around the user, so that each sound unit coordinates to point to the virtual point, providing a unified spatial sound source reference for subsequent wave field synthesis.

[0091] Step 503: Obtain the initial filter coefficients of the target filter function that match the current location information.

[0092] Optionally, the user's current location information may include three-dimensional coordinates, azimuth angle θ, and pitch angle φ.

[0093] Optionally, the target filtering function can be the Head-Related Transfer Function (HRTF), which is a digital filter that can make the sound directional, so that the sound can present a stereo effect.

[0094] Optionally, based on the user's azimuth and pitch angles, several sets of standard HRTF initial filter coefficients closest to that azimuth can be retrieved from a pre-stored Head Relational Transfer Function (HRTF) database. HRTF filter coefficients characterize the amplitude attenuation and phase change of sound waves propagating from a certain direction in space to the human ear, and are core parameters for reproducing realistic spatial hearing. By loading initial filter coefficients corresponding to the user's azimuth, the differences in human ear perception of sound from different directions can be simulated, providing a filtering foundation for subsequent accurate spatial sound field construction.

[0095] Step 504: Perform bilinear interpolation on the initial filter coefficients to obtain the target filter coefficients that match the current position information.

[0096] Understandably, since the HRTF database only stores the standard filter coefficients for discrete azimuths, while the user's actual azimuth changes continuously, a bilinear interpolation algorithm is used to interpolate the initial filter coefficients for adjacent standard azimuths to obtain the target filter coefficients H(f,θ,φ) that perfectly match the user's real-time azimuth.

[0097] Here, interpolation is used to eliminate sound field breaks and distortions caused by orientation discretization, making the spatial audio orientation smoother and the positioning more accurate, adapting to the user's auditory characteristics at any position.

[0098] Step 505: Determine the sound wave delay time of each sound unit when outputting the target audio signal based on the virtual sound source location and the sound output location of each sound unit.

[0099] Optionally, the virtual sound source coordinates V(x) are used. v ,y v ,z v The sound-emitting position S of the sound-emitting unit k Using the user's real-time location coordinates P as the calculation reference, and based on the speed of sound in air c, the difference between the propagation distance from each sound unit to the user and the propagation distance from the virtual sound source to the user is calculated, thus obtaining the sound wave delay time τ of each sound unit. k , v The calculation formula is shown in Formula 2 below: τ k , v =(|P S k | |P V|) / c (Formula 2); This delay time is used to control the timing of the sound emission of each sound unit, so that the sound waves emitted by multiple units can be synchronously superimposed at the user's ear position, accurately constructing a sound image consistent with the virtual sound source position, and realizing spatial sound source localization.

[0100] Step 506: Determine the phase offset and sampling domain delay parameters of the sound unit when outputting the target audio signal based on the sound position of each sound unit, the target filtering function, and the current position information.

[0101] Optionally, the sound wave delay time τ corresponding to each sound-generating unit can be used as a basis. k , v Based on the angular frequency of the audio signal ω=2πf, the phase shift φ is calculated. k Specifically, as shown in Formula 3 below: φ k =ω×τ k , v =2πf×(|P S k | |P V|) / c (Formula 3); Here, the phase shift of the sound-generating unit is used to adjust the phase relationship of the output signals of each sound-generating unit, so that the sound waves form controllable coherent interference in space, avoiding sound image shift and sound field chaos.

[0102] At the same time, it can be combined with the system's sampling frequency f s Transform continuous time delay into a digital domain executable sample-domain delay n k Specifically, as shown in Formula 4 below: n k =round(τ k , v ×fs )(Formula 4); In one optional embodiment, if the sampling frequency is 48kHz, then n k =round(τ k , v ×48000); if the sampling frequency is 96kHz, then n k =round(τ k , v ×96000).

[0103] Sampling domain delay enables precise timing control of digital audio systems, ensuring that the sound waves output by each speaker can be accurately superimposed to form a virtual sound source, thus guaranteeing the real-time performance and stability of wave field synthesis.

[0104] Step 507: Based on the current location information, the location information of the smart device, and the target filtering coefficient, determine the driving signal parameters of each sound unit. The driving signal parameters include the amplitude coefficient, which is used to characterize the output amplitude of the sound unit.

[0105] Optionally, the amplitude coefficient A of each sound unit can be calculated by combining the distance between the user and the sound unit, the spatial sound pressure level attenuation law, and the target filtering coefficient. k It is used to adjust the output amplitude of the sound unit and compensate for distance attenuation and frequency response differences.

[0106] In some embodiments, an FMOD audio engine or a dedicated DSP processor can be used to fit the frequency response characteristics of the sound unit through a mathematical model, and combined with the dynamic correction amplitude based on the user's position, to calculate the amplitude and phase difference of the drive signal of each speaker, so that the output amplitude of the sound unit at different positions and distances is balanced, ensuring a uniform overall sound field and stable sound quality.

[0107] Step 508: Generate the target audio signal based on the target filter coefficient, sound wave delay time, phase offset, sampling domain delay parameter, and amplitude coefficient.

[0108] Optionally, the original audio signal s(t) can be sequentially processed by timing delay, phase modulation, amplitude adjustment, and HRTF filtering to generate a target audio signal suitable for each sound unit. This can be achieved through the following formula 5: s k (t)=A k ×s(t τ k , v )×cos(φ k )×H k (f,θ,φ) (Formula 5); The target audio signal integrates spatial orientation, distance attenuation, and human hearing characteristics, enabling each sound unit to output an audio component with a clear spatial direction. Multiple units working together can reconstruct a three-dimensional virtual sound source in space, achieving high-quality spatial audio without the need for multiple audio tracks or large cavities.

[0109] Step 509: Control each sound unit to output the target audio signal based on the drive signal parameters.

[0110] Optionally, the target audio signal can be converted from digital to analog and amplified in power and then sent to the corresponding sound unit in real time, so that each sound unit can emit sound synchronously according to the calculated delay, phase and amplitude.

[0111] By propagating and coherently superimposing sound waves in space, a stable, clear, and directional three-dimensional spatial sound field is formed at the user's location. The entire output process dynamically responds to changes in the user's position, achieving immersive, high-precision spatial audio playback in confined spaces such as humanoid robots without the need for complex hardware structures.

[0112] Understandably, smart devices generate operating noise during operation, which can interfere with the output of audio signals.

[0113] Based on this, in some embodiments, it can be achieved by performing, as follows Figure 6 The method shown is used to output the target audio signal. For example... Figure 6 As shown, the method may include the following steps: Step 601: Obtain the operating noise of the smart device.

[0114] Optionally, the smart device incorporates an array of multi-channel sound-generating units, which are fixedly deployed inside the device body, near the noise acquisition area of ​​the joint motor and power module. The array of sound-generating units collects composite noise signals generated by the device during operation in real time, including mechanical vibration noise, motor electromagnetic noise, and wind noise. These signals are then processed by a preamplifier circuit and an anti-aliasing filter to obtain a clean noise reference signal x(n).

[0115] This ensures that interfering noises affecting the auditory experience can be fully captured, providing accurate input for subsequent active noise cancellation.

[0116] Step 602: Generate the antiphase acoustic wave of the operating noise.

[0117] Optionally, the intelligent device can use an adaptive filtering algorithm to perform waveform inversion calculation on the acquired operating noise signal x(n). The core lies in updating the filtering weight vector w(n) in real time, and by minimizing the energy of the error signal e(n), fitting an inverted sound wave waveform y(n) that is opposite in phase and matches the amplitude of the original noise.

[0118] This process features rapid convergence and can dynamically track changes in the noise spectrum, ensuring the real-time performance and accuracy of the phase inversion effect.

[0119] Optionally, the adaptive filtering algorithm can employ a least mean square adaptive filtering algorithm. Using the collected noise from the robot joint motors as the input signal, the algorithm iteratively updates the filter weights to ensure that the out-of-phase sound wave output by the filter maintains equal amplitude and opposite phase to the original noise signal. In each iteration, the algorithm adjusts the weights based on the current error signal, gradually approaching the optimal noise reduction waveform, ultimately maximizing the cancellation of the superimposed noise components.

[0120] Alternatively, adaptive filtering algorithms can employ recursive least squares algorithms. Aiming to minimize the weighted sum of squared noise errors, these algorithms recursively update the covariance and gain matrices to rapidly correct filter weights, enabling the out-of-phase sound wave to achieve a precise inverse match with the original noise in a very short time. Compared to least mean square adaptive filtering algorithms, recursive least squares algorithms respond faster to noise changes, have smaller steady-state errors, and maintain excellent noise reduction performance even in scenarios involving high-speed robot movement and drastic noise fluctuations.

[0121] Step 603: The antiphase sound wave is superimposed on the target audio signal to obtain the canceled audio signal.

[0122] Optionally, the target audio signal s, after wave field synthesis processing, can be... k (t) is linearly superimposed with the generated antiphase sound wave y(n). Through the principle of superposition cancellation, the energy of the original operating noise is directly eliminated or greatly attenuated in space.

[0123] Optionally, a dynamic gain coefficient γ can be introduced during the superposition process to adaptively adjust the intensity of the antiphase sound wave, ensuring that noise is effectively suppressed while preventing excessive noise reduction from distorting the target audio signal, thus achieving a balanced output of noise and speech.

[0124] Step 604: Control each sound unit to output a canceled audio signal based on the drive signal parameters.

[0125] Optionally, the canceled audio signal, which incorporates the inverted sound waves, can be combined with the sound wave delay time τ calculated in the previous step. k , v Phase offset φ k Sampling domain delay n k and amplitude coefficient A k The system generates the final drive signal and transmits it to each sound unit. Based on precise timing and phase control, the sound unit synchronously outputs the cancelling audio signal, using the principle of spatial sound wave interference to physically cancel noise at the user's ear position, thereby purifying the listening environment.

[0126] In some embodiments, when specific audio is detected being played, the volume of the cancelling audio signal can be adjusted to the target volume based on the user's current location information and the location information of the smart device; and each sound unit can be controlled to output the cancelling audio signal at the target volume.

[0127] Here, specific audio includes, but is not limited to, important information audio such as interactive voice, prompts, and instructions.

[0128] When a smart device detects that such audio is about to play or is currently playing, it can trigger an adaptive volume adjustment mechanism. This mechanism dynamically calculates the target volume level by considering the real-time distance between the user and the device, their azimuth, and ambient noise. For example, when the user moves away from the device, the output volume is appropriately increased to compensate for sound pressure level attenuation; when the user moves closer or ambient noise suddenly increases, the volume is appropriately increased based on the psychoacoustic masking effect to ensure that key information is clearly audible. Ultimately, the sound unit is controlled to output the target volume to cancel the audio signal, significantly improving the clarity and reliability of human-computer interaction voice while ensuring noise reduction, achieving an intelligent experience of "noise reduction without lost sound."

[0129] In this embodiment, by acquiring the user's location in real time and constructing a virtual sound position by combining the surround sound unit arrangement, and by using the target filtering coefficient to achieve accurate sound field rendering, the listener can perceive stable stereo surround sound, significantly improving the audio immersion and spatial directivity, and solving the problem of flat sound field and lack of spatial sense in traditional devices.

[0130] Based on the foregoing embodiments, this application provides an audio processing device, which includes various modules and units included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.

[0131] Figure 7 This is a schematic diagram of the structure of the audio processing device provided in the embodiments of this application, such as... Figure 7 As shown, the device 700 includes an acquisition module 701, a determination module 702, and a control module 703, wherein: Module 701 is used to obtain the user's current location information; The determining module 702 is used to determine the virtual sound position corresponding to each of the sound units based on the current location information and the sound position of each sound unit. The virtual sound position is located in a specific direction of the user. The sound units are arranged in a surround manner and the sound units are set in the smart device. Determine the target filter coefficients of a target filter function that matches the user's current location information. The target filter function is used to characterize the relationship between the vocal parameters and the vocal frequency when the audio signal propagates from the location of the smart device to the user. The vocal parameters include amplitude and phase. Based on the target filtering coefficient, the current position information, the virtual sound position, and the sound position of each sound unit, the target output parameters corresponding to each sound unit are determined. The target output parameters include one or more of the following: sound wave delay time, phase offset, and sampling domain delay parameters when the sound unit outputs the target audio signal. The control module 703 is used to control each of the sound-generating units to output the target audio signal based on the target output parameters.

[0132] In some embodiments, the determining module 702 is further configured to, upon detecting an update to the user's current location information, re-determine the target output parameters corresponding to each of the sound-emitting units based on the updated location information to obtain updated output parameters; Each of the aforementioned sound-producing units is controlled to output the target audio signal based on the updated output parameters.

[0133] In some embodiments, the acquisition module 701 is further configured to acquire the initial filter coefficients of the target filter function that matches the current location information; The determining module 702 is further configured to perform bilinear interpolation on the initial filter coefficients to obtain target filter coefficients that match the current position information.

[0134] In some embodiments, the determining module 702 is further configured to determine the sound wave delay time of each of the sound-generating units when outputting the target audio signal, based on the virtual sound source location and the sound-generating location of each of the sound-generating units. Based on the sound-emitting position of each sound-emitting unit, the target filtering function, and the current position information, the phase offset and sampling domain delay parameters of the sound-emitting unit when outputting the target audio signal are determined.

[0135] In some embodiments, the determining module 702 is further configured to determine the driving signal parameters of each of the sound-emitting units based on the current location information, the location information of the smart device, and the target filtering coefficient. The driving signal parameters include an amplitude coefficient, which is used to characterize the output amplitude of the sound-emitting unit. The target audio signal is generated based on the target filter coefficient, the sound wave delay time, the phase shift, the sampling domain delay parameter, and the amplitude coefficient. The control module 703 is also used to control each of the sound-generating units to output the target audio signal based on the drive signal parameters.

[0136] In some embodiments, the acquisition module 701 is further configured to acquire the operating noise of the smart device; The determining module 702 is further configured to generate an antiphase sound wave of the operating noise; and to superimpose the antiphase sound wave with the target audio signal to obtain a canceled audio signal; The control module 703 is also used to control each of the sound-generating units to output the canceled audio signal based on the drive signal parameters.

[0137] In some embodiments, the determining module 702 is further configured to, upon detecting the playback of specific audio, adjust the volume of the canceled audio signal to a target volume based on the user's current location information and the location information of the smart device; The control module 703 is also configured to control each of the sound-generating units to output the canceled audio signal at the target volume.

[0138] In some embodiments, the sampling frequency of the sound-generating unit is different, and the target filtering coefficient is different.

[0139] In some embodiments, the sound channel of the smart device is spiral-shaped.

[0140] It should be noted that, in the embodiments of this application... Figure 7 The module division of the image processing apparatus shown is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or be integrated into one unit with two or more units. The integrated units described above can be implemented in hardware, as software functional units, or in a combination of software and hardware.

[0141] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a smart device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0142] This application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the methods described above.

[0143] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in the above embodiments.

[0144] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.

[0145] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0146] In one embodiment, the image processing apparatus provided in this application can be implemented as a computer program, and the computer program can be implemented as follows: Figure 8The device operates on the computer device shown. The memory of the computer device can store the various program modules that make up the above-described apparatus. The computer program, composed of the various program modules, causes the processor to execute the steps of the methods in the various embodiments of this application described in this specification.

[0147] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0148] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.

[0149] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0150] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0151] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.

[0152] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0153] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.

[0154] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0155] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a smart device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0156] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0157] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0158] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0159] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An audio processing method, characterized in that, Applied to smart devices, the smart devices including multiple sound-emitting units arranged in a surround manner, the method includes: Obtain the user's current location information; Based on the current location information and the sound emission position of each of the sound emission units, a virtual sound emission position corresponding to each of the sound emission units is determined, and the virtual sound emission position is located in a specific direction of the user; Determine the target filter coefficients of a target filter function that matches the user's current location information. The target filter function is used to characterize the relationship between the vocal parameters and the vocal frequency when the audio signal propagates from the location of the smart device to the user. The vocal parameters include amplitude and phase. Based on the target filtering coefficient, the current position information, the virtual sound position, and the sound position of each sound unit, the target output parameters corresponding to each sound unit are determined. The target output parameters include one or more of the following: sound wave delay time, phase offset, and sampling domain delay parameters when the sound unit outputs the target audio signal. Each of the aforementioned sound-producing units is controlled to output the target audio signal based on the target output parameters.

2. The method according to claim 1, characterized in that, The method further includes: If the user's current location information is updated, the target output parameters corresponding to each sound unit are re-determined based on the updated location information to obtain the updated output parameters; Each of the aforementioned sound-producing units is controlled to output the target audio signal based on the updated output parameters.

3. The method according to claim 1, characterized in that, The determination of the target filter coefficients of the target filter function that matches the user's current location information includes: Obtain the initial filter coefficients of the target filter function that match the current location information; The initial filter coefficients are subjected to bilinear interpolation to obtain target filter coefficients that match the current position information.

4. The method according to claim 3, characterized in that, The step of determining the target output parameters corresponding to each sound unit based on the target filtering coefficients, the current position information, the virtual sound position, and the sound position of each sound unit includes: Based on the location of the virtual sound source and the sound emission location of each of the sound emission units, the sound wave delay time of each of the sound emission units when outputting the target audio signal is determined; Based on the sound-emitting position of each sound-emitting unit, the target filtering function, and the current position information, the phase offset and sampling domain delay parameters of the sound-emitting unit when outputting the target audio signal are determined.

5. The method according to claim 4, characterized in that, The control of each of the sound-producing units to output the target audio signal based on the target output parameters includes: Based on the current location information, the location information of the smart device, and the target filtering coefficient, the driving signal parameters of each sound-emitting unit are determined. The driving signal parameters include an amplitude coefficient, which is used to characterize the output amplitude of the sound-emitting unit. The target audio signal is generated based on the target filter coefficient, the sound wave delay time, the phase shift, the sampling domain delay parameter, and the amplitude coefficient. Each of the aforementioned sound-producing units is controlled to output the target audio signal based on the driving signal parameters.

6. The method according to claim 5, characterized in that, The control of each of the sound-producing units to output the target audio signal based on the drive signal parameters includes: Obtain the operating noise of the smart device; Generate the antiphase acoustic wave of the operating noise; The antiphase sound wave is superimposed on the target audio signal to obtain a canceled audio signal; Each of the aforementioned sound-generating units is controlled to output the canceled audio signal based on the driving signal parameters.

7. The method according to claim 6, characterized in that, The method further includes: When specific audio is detected being played, the volume of the canceled audio signal is adjusted to the target volume based on the user's current location information and the location information of the smart device. Each of the aforementioned sound-producing units is controlled to output the canceled audio signal at the target volume.

8. The method according to claim 3, characterized in that, The sampling frequency of the sound-generating unit is different, and the target filtering coefficient is different.

9. The method according to claim 1, characterized in that, The sound guiding channel of the smart device is spiral-shaped.

10. An audio processing apparatus, characterized in that, The device includes: The acquisition module is used to obtain the user's current location information; The determining module is used to determine the virtual sound position corresponding to each of the sound units based on the current location information and the sound position of each sound unit. The virtual sound position is located in a specific direction of the user. The sound units are arranged in a surround manner and the sound units are set in the smart device. Determine the target filter coefficients of a target filter function that matches the user's current location information. The target filter function is used to characterize the relationship between the vocal parameters and the vocal frequency when the audio signal propagates from the location of the smart device to the user. The vocal parameters include amplitude and phase. Based on the target filtering coefficient, the current position information, the virtual sound position, and the sound position of each sound unit, the target output parameters corresponding to each sound unit are determined. The target output parameters include one or more of the following: sound wave delay time, phase offset, and sampling domain delay parameters when the sound unit outputs the target audio signal. The control module is used to control each of the sound-generating units to output the target audio signal based on the target output parameters.