Audio signal processing method and electronic equipment

By separating and adjusting the sound source signal of the audio signal, and allocating signal paths based on the scene information and the receiver's position, the problem of mismatch between sound and scene in audio playback is solved, and the sense of space and presence is enhanced.

CN120544593APending Publication Date: 2025-08-26LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510724780.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing audio playback methods cannot adjust the sound according to the scene information, resulting in the sound mismatch between the scene information and lack of sense of space and presence.

Method used

By acquiring the audio signal, separating the audio information to form the first sound source signal, adjusting the parameters of the sound source signal based on the scene information, allocating the signal path to adapt to the temporal and spatial information of the receiver, and separating the two-channel sound source using a neural network model to adjust the sound image, phase and amplitude, and selecting a suitable signal path for output.

Benefits of technology

It realizes the matching of audio signals and scene information, enhances the sense of space and presence, and provides the effect of virtual sound surround.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544593A_ABST
    Figure CN120544593A_ABST
Patent Text Reader

Abstract

The invention provides an audio signal processing method which can be applied to the technical field of audio and video. The method comprises the steps that audio signals are acquired, audio information in the audio signals is separated to form first sound source signals, and the audio signals comprise at least one first sound source signal; adjusting parameters of the first sound source signal based on scene information corresponding to the audio signal to form a second sound source signal, wherein the second sound source signal is adapted to an audio signal receiver to receive spatio-temporal information of the first sound source signal; and distributing a signal path for a second sound source signal, and obtaining a first signal path, the first signal path being determined based on the spatio-temporal information and the position of the audio signal receiver.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio and video technology, and more specifically, to an audio signal processing method and electronic device. Background Art

[0002] In existing audio playback scenarios, the limitations on the number of speakers and audio source channels can lead to discrepancies between the sound played through the speakers and the scene it corresponds to. However, existing audio playback methods are unable to adjust the sound based on scene information, making it impossible to align the played sound with the scene information. This results in a lack of spatial perception and a lack of synchronization between the sound and the scene information. For example, if the scene depicts a horse galloping from a distance, the sound of the horse's hooves should be played from far to near. However, with existing audio playback methods, users always hear the sound in the center of the screen, without any directional information, resulting in a poor listening experience. Summary of the Invention

[0003] In view of the above problems, the present application provides an audio signal processing method and electronic device for playing sounds in 3D and creating a sense of presence and space.

[0004] According to a first aspect of the present application, a method for processing an audio signal is provided, comprising: acquiring an audio signal, separating audio information in the audio signal to form a first sound source signal, the audio signal comprising at least one first sound source signal; adjusting parameters of the first sound source signal based on scene information corresponding to the audio signal to form a second sound source signal, the second sound source signal being adapted to the spatiotemporal information of an audio signal receiver receiving the first sound source signal; and allocating a signal path to the second sound source signal to acquire a first signal path, the first signal path being determined based on the spatiotemporal information and the position of the audio signal receiver.

[0005] According to an embodiment of the present application, the audio signal includes a two-channel sound source, and separating the audio information in the audio signal to form a first sound source signal includes: identifying a sound object in the two-channel sound source, the sound object marking the main body that generates the sound source signal in the two-channel sound source; and separating the audio information in the two-channel sound source based on the sound object to form the first sound source signal.

[0006] According to an embodiment of the present application, adjusting the parameters of the first sound source signal based on the scene information corresponding to the audio signal to form a second sound source signal includes: adjusting the parameters of the first sound source signal according to the characteristics of the first sound source signal in the scene information, the parameter adjustment including at least one of sound image adjustment, phase adjustment and amplitude adjustment, and the characteristics including at least one of the sound image, phase and amplitude of the first sound source signal.

[0007] According to an embodiment of the present application, the method also includes: identifying an object in the displayed video corresponding to the audio signal; and obtaining the action of the object and / or the type of the object as the scene information, the action or type of the object matching the first sound source signal; wherein, adjusting the parameters of the first sound source signal based on the scene information corresponding to the audio signal to form a second sound source signal includes: adjusting the parameters of the first sound source signal based on the action of the object and / or the type of the object to form a second sound source signal.

[0008] According to an embodiment of the present application, the spatiotemporal information includes: azimuth information, time information and intensity information, the azimuth information marks the direction in which the audio signal receiver receives the first sound source signal, the time information marks the time when the audio signal receiver receives the first sound source signal, and the intensity information marks the size of the first sound source signal received by the audio signal receiver.

[0009] According to an embodiment of the present application, the method also includes: determining a transfer function based on the spatiotemporal information of the first sound source signal and the position of the audio signal receiver; calculating an output signal corresponding to the second sound source signal based on the transfer function and the second sound source signal; and outputting the output signal through the first signal path.

[0010] According to an embodiment of the present application, determining the transfer function based on the spatiotemporal information of the first sound source signal and the position of the audio signal receiver includes: determining a first transfer function based on the spatiotemporal information and the position of the audio signal receiver; determining a second transfer function based on the first signal path and the position of the audio signal receiver; and combining the first transfer function and the second transfer function to obtain the transfer function.

[0011] According to an embodiment of the present application, the assigning of a signal path to the second sound source signal and the obtaining of the first signal path include: selecting an output channel of the second sound source signal based on the spatiotemporal information and the position of the audio signal receiver, the output channel including at least one of a left channel and a right channel; and using the signal path corresponding to the output channel as the first signal path.

[0012] According to an embodiment of the present application, the number of signal paths included in the left channel is the same as the number of signal paths included in the right channel.

[0013] The second aspect of the present application provides an electronic device, wherein the electronic device includes: one or more processors; one or more memories for storing executable instructions, and when the executable instructions are executed by the processor, the following method is performed: obtaining an audio signal, separating audio information in the audio signal to form a first sound source signal, the audio signal including at least one first sound source signal; adjusting parameters of the first sound source signal based on scene information corresponding to the audio signal to form a second sound source signal, the second sound source signal adapting to the spatiotemporal information of the audio signal receiver receiving the first sound source signal, the scene information including at least one picture layer corresponding to the audio signal; and assigning a signal path to the second sound source signal, obtaining a first signal path, the first signal path being determined based on the spatiotemporal information and the position of the audio signal receiver. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:

[0015] Figure 1 A diagram schematically illustrates an application scenario of a method for processing an audio signal according to an embodiment of the present application;

[0016] Figure 2 The flowchart schematically shows a method for processing an audio signal according to an embodiment of the present application;

[0017] Figure 3 A flowchart of a method for separating a binaural sound source according to an embodiment of the present application is schematically shown;

[0018] Figure 4 Schematically shows a flow chart for allocating signal paths for audio signals according to an embodiment of the present application;

[0019] Figure 5 Schematically shows a flow chart for allocating signal paths for audio signals according to another embodiment of the present application;

[0020] Figure 6 A structural diagram schematically shows the configuration of a speaker according to an embodiment of the present application;

[0021] Figure 7 A schematic diagram schematically illustrates determining an output signal using a spherical coordinate system according to an embodiment of the present application; and

[0022] Figure 8 The block diagram schematically shows an electronic device suitable for implementing the method for processing audio signals according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0024] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0026] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0027] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0028] An embodiment of the present application provides a method for processing an audio signal, comprising: obtaining an audio signal, separating audio information from the audio signal to form a first sound source signal, the audio signal including at least one first sound source signal; adjusting parameters of the first sound source signal based on scene information corresponding to the audio signal to form a second sound source signal, the second sound source signal being adapted to the spatiotemporal information of the first sound source signal received by an audio signal receiver; and assigning a signal path to the second sound source signal, obtaining a first signal path, the first signal path being determined based on the spatiotemporal information and the position of the audio signal receiver. This method enables different sound source signals in the audio signal to adapt to scene information, creating a sense of presence and spatiality, and achieving a virtual surround sound effect.

[0029] Figure 1 The following schematically illustrates an application scenario of the method for processing an audio signal according to an embodiment of the present application.

[0030] like Figure 1 As shown, the application scenario 100 according to this embodiment may include audio playback and video playback. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables.

[0031] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0032] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0033] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0034] It should be noted that the audio signal processing method provided in the embodiment of the present application can generally be performed by the server 105. The audio signal processing method provided in the embodiment of the present application can also be performed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0035] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0036] The following will be based on Figure 1 The scene described by Figures 2 to 7 The audio signal processing method according to the embodiment of the present application is described in detail.

[0037] Figure 2 The flowchart schematically shows a method for processing an audio signal according to an embodiment of the present application.

[0038] like Figure 2 As shown, the audio signal processing method 200 of this embodiment includes operations S210 to S230.

[0039] In operation S210 , an audio signal is acquired, and audio information in the audio signal is separated to form a first sound source signal. The audio signal includes at least one first sound source signal.

[0040] In operation S220, parameters of the first sound source signal are adjusted based on scene information corresponding to the audio signal to form a second sound source signal, and the second sound source signal is adapted to the spatiotemporal information of the first sound source signal received by the audio signal receiver.

[0041] The second sound source signal adapts the spatiotemporal information of the first sound source signal received by the audio signal receiver, for example, so that the sound image and size of the adjusted sound source signal match the sound image and size of the sound source signal marked by the receiver.

[0042] In operation S230, a signal path is allocated to the second sound source signal, and a first signal path is obtained, where the first signal path is determined based on the spatiotemporal information and the position of the audio signal receiver.

[0043] The audio signal processing method provided in the embodiment of the present application adjusts the parameters of each sound source signal in the audio signal according to the scene information corresponding to the audio signal, so that the adjusted sound source signal matches the spatial information of the sound source signal that the receiver wants to hear. The first signal path is determined based on the spatiotemporal information and the position of the audio signal receiver as the signal path allocated to the adjusted sound source signal, which can make the played sound source signal adapt to the scene information and create a real sense of presence and space.

[0044] The audio signal includes audio information in the video, such as audio information corresponding to a 3D display; or separate audio information, such as audio from an audio playback device. The audio signal includes at least one first sound source signal, such as background sound, human voice, sound events, etc.

[0045] In operation S210, the audio information in the audio signal is separated to form a first sound source signal. Specifically, the processing method is determined based on the type of audio signal. The audio signal can be a multi-channel sound source signal, and the multi-channel sound source signal includes at least three high-frequency sound source signals, such as a 5.1 sound source, a 7.1 sound source, a 9.1 sound source, etc. Alternatively, the audio signal can be a dual-channel sound source signal, which only includes two left and right channels, but the two channels contain various sound source signals, including but not limited to sound events, background sounds, and human voices. For example, when watching a movie on a tablet computer, the audio file of the movie is a dual-channel sound source file, and two speakers are used to play a variety of sound source signals during playback.

[0046] Because multi-channel sound sources already separate various sound source signals, no additional processing is required. For example, a 5.1 sound source includes left, right, center, left surround, right surround, and a low-frequency effects channel. However, for two-channel sound sources, pre-processing is required to separate the different sound source signals.

[0047] Figure 3 A flowchart of a method for separating a two-channel sound source according to an embodiment of the present application is schematically shown. Figure 3 Provide a detailed description.

[0048] According to an embodiment of the present application, the audio signal includes a two-channel sound source, and separating the audio information in the audio signal to form a first sound source signal includes: identifying the sound object in the two-channel sound source, and the sound object labeling generates the main body of the sound source signal in the two-channel sound source; separating the audio information in the two-channel sound source based on the sound object to form the first sound source signal.

[0049] For example, a sound object represents the subject that generates a sound source signal. A sound source signal can be the emitter of the sound source signal. For example, a person is the sound object corresponding to a human voice, and a car is the sound object corresponding to traffic noise. Alternatively, a sound object can be the recipient of a sound source signal. The recipient represents the object to which the sound source signal is applied (or acts upon). For example, when playing a video, the background of the screen is the sound object corresponding to background music.

[0050] like Figure 3 As shown, the method 300 for separating two-channel sound sources in this embodiment includes operations S310 to S330.

[0051] In operation S310 , a data set of aliased multiple sound source signals is created.

[0052] In operation S320 , the data set is used as input of a neural network model, and the multiple sound source signals before aliasing are used as output to train the neural network model.

[0053] In operation S330 , the binaural sound source to be separated is input into the trained neural network model, and a first sound source signal is output.

[0054] Exemplarily, the separation of binaural sound sources can be achieved through a multi-classification deep neural network. Specifically, the audio of the sound event marked, the audio of the background sound, and the audio of the human voice are aliased to create a data set. The data in the data set is used as the input of the multi-classification deep neural network, and the audio of the sound event before aliasing, the audio of the background sound, and the audio of the human voice are used as the output to train the network. After the parameters converge, a network model for sound event separation, background sound extraction, and human voice separation can be obtained. Furthermore, a large model can be used for fine-tuning to achieve the separation of binaural sound sources.

[0055] This application does not limit the specific method of separating the two-channel sound source. Any method that can separate the two-channel sound source to form the first sound source signal can be applied to the embodiments of this application.

[0056] An object-based method is used to separate multiple different types of sound source signals from a two-channel sound source, so that the two-channel sound source can also be converted into multiple independent sound source signals, thereby using the audio signal processing method of the embodiment to achieve a surround sound effect when playing audio.

[0057] The first sound source signal is a sound source signal separated from the audio signal. Typically, the audio signal includes multiple first sound source signals. The first sound source signal can be background sound, human voice, or sound event, etc. This application does not specifically limit this. The first sound source signal can be any sound source signal that constitutes the audio signal. In fact, the first sound source signal can also include multiple sound source signals. The multiple sound source signals include at least two sound source signals. The at least two sound source signals can be at least two sound event signals, at least two human voice signals, at least two background sound signals, or a combination of multiple signals among background sound, human voice, and sound event. For example, the first sound source signal can be background sound composed of piano and violin sounds, the first sound source signal can be the sound of multiple people talking, or the first sound source signal can be vehicle sounds, dog barking, and human voices.

[0058] A sound event represents the sound source signal corresponding to an event in a scene. For example, a dog barking or a vehicle humming can be considered a sound event. In some cases, human voices can also be considered sound events. If the predominant sound in a scene is another type of sound, while the human voice plays a supporting role, the human voice can be considered a sound event. For example, in a fight scene, the predominant sound is the fighting, while the surrounding human voices can be considered sound events.

[0059] For example, in some cases, human voices and sound events can also be background sounds, which mainly depends on the main prominent sounds of the scene. For example, in a 3D movie, the protagonist is reading in a noisy vegetable market. At this time, the hawking sounds and the sounds of chopping meat in the vegetable market are all background sounds.

[0060] According to the embodiments of the present application, there are no clear boundaries between background sounds, human voices, and sound events. Background sounds, human voices, and sound events can be converted into each other. The content of the scene information determines which are background sounds, which are sound events, and which are the human voices that need to be highlighted. The content of the scene information can include the main prominent sound in the scene. For example, if the protagonist is reading in a noisy vegetable market, the main highlight is the protagonist's reading sound.

[0061] According to an embodiment of the present application, adjusting the parameters of the first sound source signal based on the scene information corresponding to the audio signal to form the second sound source signal includes: adjusting the parameters of the first sound source signal according to the characteristics of the first sound source signal in the scene information, the parameter adjustment includes at least one of sound image adjustment, phase adjustment and amplitude adjustment, and the characteristics include at least one of the sound image, phase and amplitude of the first sound source signal.

[0062] For example, the scene information corresponding to the audio signal may include the sound information, object information, behavior information, etc. included in the scene rendered by the audio signal or the scene in which the subject generating the audio signal is located. For example, when listening to a game commentary, the scene information may include the movements of the players at the game and the shouts of the audience.

[0063] Parameters of the first sound source signal separated from the audio signal are adjusted based on characteristics of the first sound source signal in the scene information. Specifically, the parameters of the first sound source signal separated from the audio signal are adjusted based on at least one of the image, phase, and amplitude of the first sound source signal in the scene. For example, the volume of the audience cheers separated from the game commentary can be adjusted based on the relative volume of the audience cheers at the game.

[0064] For different sound source signals, parameter adjustments can be made using at least one of image adjustment, phase adjustment, and amplitude adjustment to produce different playback effects, giving the sound source signal to the receiver a sense of space. For example, if a horse in a scene is running from far to near, the sound of horse hooves should be played from far to near. The image of the horse hoof sounds separated from the audio signal can be adjusted to achieve a playback effect from far to near. The amplitude of the separated horse hoof sounds can also be adjusted to make the played hoof sounds gradually louder.

[0065] Sound image adjustment is used to control the position of sound in the stereo field (left and right channels) to create a sense of space and stereo width. Sound image adjustment can be achieved by using the gain difference method, delay difference method, etc., but is not limited to this. Any method that can achieve sound image adjustment can be applied to the embodiments of this application. The gain difference method locates the sound image by adjusting the volume difference between the left and right channels. For example, when the gain of the left channel is increased, the sound image moves to the left; when the gain of the right channel is increased, the sound image moves to the right. The delay difference method inserts a small delay (for example, less than 1ms) between the left and right channels, and uses the difference in the arrival time of the sound to perceive the position of the human ear. Sound image adjustment can achieve: stereo field expansion, for example, distributing different instruments in the left and right channels to avoid the sound piling up in the center; special spatial sense, for example, simulating the movement of sound from left to right; mono to stereo, and enhancing the richness of the listening experience through sound image separation.

[0066] Amplitude adjustment controls the volume and dynamic range of an audio signal. This can be achieved using methods such as gain control and dynamic processing, but is not limited to these methods. Any method that can achieve amplitude adjustment can be applied to the embodiments of this application. Gain control directly increases or decreases the overall volume; dynamic processing uses a compressor to reduce peak volume above a threshold, narrowing the dynamic range, and a limiter to strictly limit the maximum level to prevent clipping. Amplitude adjustment can balance the volume and ensure that the volume of different tracks is coordinated, for example, the lead vocals stand out above the accompaniment.

[0067] Phase adjustment changes the relative timing of audio waveforms to resolve phase issues or create specific effects. Phase adjustment can be achieved using methods such as phase inversion, phase shifting, and all-pass filtering. Phase inversion flips the polarity of a waveform, useful for eliminating phase cancellation in multi-microphone recordings; phase shifting adjusts the time alignment of waveforms through delay or filtering; and all-pass filtering adjusts the phase response without changing the amplitude. Phase adjustment can resolve phase cancellation, create stereo enhancement, and create special sound effects.

[0068] According to an embodiment of the present application, the parameters of the separated first sound source signal are adjusted according to the characteristics of the first sound source signal in the scene. Since different sound source signals ultimately aim to achieve different playback effects, different sound source signals are configured with different adjustment methods to make the various sound sources played match the actual scene, highlight specific sounds, and enhance the sense of space.

[0069] According to an embodiment of the present application, the method for processing an audio signal also includes: identifying an object in a displayed video corresponding to the audio signal; and obtaining the action of the object and / or the type of the object as scene information, the action of the object and / or the type of the object matching the first sound source signal; wherein, adjusting the parameters of the first sound source signal based on the scene information corresponding to the audio signal to form a second sound source signal includes: adjusting the parameters of the first sound source signal based on the action of the object and / or the type of the object to form the second sound source signal.

[0070] The scene information of the audio signal can also include the video screen when the video is played, for example, a stereoscopic screen or screen layer of a 3D display scene. The objects in the displayed video can include objects in the video screen or stereoscopic objects in the 3D display. The action and type of the object in the video screen (or the stereoscopic object in the 3D display) can be used as scene information to adjust the parameters of the first sound source signal to form a second sound source signal. For example, in a 3D movie being played, a car rushes into a vegetable market. The sound source signal corresponding to the car in the movie video can be adjusted for sound and image based on the car's action. The sound source signal corresponding to the vegetable market can also be adjusted for amplitude and phase based on the characteristics of the background sound (here, the vegetable market type is the background).

[0071] Exemplarily, the action of the object and / or the type of the object matches the first sound source signal. The action of the object can generate a first sound source signal. For example, in the video screen, the protagonist is cutting vegetables, and cutting vegetables corresponds to the sound of chopping vegetables in the audio signal. The type of the object can be, for example, background, etc., and the electronic device can also be the background. The electronic device can be a device that can play audio signals. According to the type of the object, the first sound source signal corresponding to the type of the object can be determined. For example, there is an electronic device playing audio in the 3D display. The electronic device is an object in the 3D display, and the audio it plays is the first sound source signal. For example, for a 3D display scene composed of picture layers, the action of the object can be determined by analyzing the differences between the picture layers in the 3D display, and the type of the object can also be determined by analyzing the differences between the picture layers in the 3D display. The area where there is no change between the picture layers is usually the background area.

[0072] For scenes that include video images and audio signals, there is a correspondence between the audio signal and the video display content. Specifically, the various sound source signals in the audio signal are matched with the objects in the displayed video. The movement of the object implies the auditory effect to be achieved when playing the sound source signal, or the type of the object implies the auditory effect to be achieved when playing the sound source signal. For example, if the object in the video is the background, when playing the background music corresponding to the background, the effect of the background sound emanating from the surround direction must be achieved.

[0073] For scenes that include video images and audio signals, the parameters of a first sound source signal are adjusted based on the motion and / or type of objects in the video to form a second sound source signal. For example, because the audio signal receiver needs to hear a sound source signal that is synchronized with the video image, the parameters of the first sound source signal can be adjusted based on the motion and / or type of objects in the video image to form a second sound source signal that is synchronized with the image. For 3D displays, the parameters of the first sound source signal can also be adjusted based on changes between image layers. For example, the amplitude of the sound source signal (such as background music) corresponding to areas where no changes occur between image layers can be reduced. The parameter adjustment can include at least one of image and sound image adjustment, amplitude adjustment, and phase adjustment. The second sound source signal is the sound source signal after parameter adjustment, and the image, amplitude, and phase of the second sound source signal match the motion and / or type of objects in the video image. The specific method for adjusting the parameters of the first sound source signal to form the second sound source signal has been described above and will not be elaborated here.

[0074] For example, in a 3D movie, when a horse gallops from far to near, the sound of the horse's hooves should be heard from far to near and gradually increase in size. To this end, the horse's hoof sound is first separated from the audio signal corresponding to the 3D display, and then the sound image and amplitude of the horse's hoof sound are adjusted according to the horse's movements to form a second sound source signal. The sound image of the second sound source signal presents an auditory effect from far to near, and the amplitude of the second sound source signal presents an auditory effect from small to large. A signal path is allocated to the second sound source signal for output.

[0075] The audio signal processing method of this embodiment adjusts the sound source signal separated from the audio signal according to the object in the video picture, and can be used in 3D display scenes to avoid deviations between the played sound information and the picture information, so that the played sound information is synchronized with the video picture, creating a real sense of presence and space.

[0076] In operation S220, the second sound source signal adaptation audio signal receiver receives the spatiotemporal information of the first sound source signal.

[0077] According to an embodiment of the present application, the spatiotemporal information includes: azimuth information, time information and intensity information. The azimuth information marks the direction in which the audio signal receiver receives the first sound source signal, the time information marks the time when the audio signal receiver receives the first sound source signal, and the intensity information marks the size of the first sound source signal received by the audio signal receiver.

[0078] Exemplarily, the directional information indicates the direction from which the sound is intended to be emitted. Typically, human voices are emitted from the center, background sounds are emitted from surround directions, and sound events are emitted from a specified direction, for example, the sound of horse hooves is emitted from left to right, or from far to near.

[0079] Exemplarily, the second sound source signal obtained after parameter adjustment of the first sound source signal is adapted to the spatiotemporal information of the first sound source signal received by the audio signal receiver, that is, the first sound source signal is adjusted based on the requirements of the direction, time and size of the received sound source signal when the reference audio signal receiver receives the first sound source signal. For example, the sound image of the sound event separated from the audio signal is adjusted according to the azimuth information of the sound event that the receiver wants to hear.

[0080] For example, the spatiotemporal information of the first sound source signal received by the receiver can be the actual location, time, and magnitude of the sound source signal heard by the human ear when the user is in the scene. The spatiotemporal information can be the spatiotemporal information corresponding to an object in a 3D display, where the object's action and / or type annotates the location, time, and magnitude of the sound source signal. For example, if a horse is running from a distance, the direction of the horse's hoofbeats is the location information, and the intensity of the hoofbeats is the intensity information from small to large.

[0081] When the receiver receives different sound source signals in the audio signal, in order to achieve the surround sound playback effect, there are requirements for the direction, time and intensity of the received sound source signal. Therefore, the second sound source signal needs to adapt to the temporal and spatial information of the first sound source signal received by the audio signal receiver.

[0082] The parameters of the first sound source signal are adjusted according to the spatiotemporal information of the first sound source signal received by the receiver, thereby helping to enhance the sense of space when playing audio and avoid deviation between sound and picture when playing video.

[0083] After the second sound source signal is acquired, in operation S230 , a signal path is allocated to the second sound source signal, and a first signal path is acquired, where the first signal path is determined based on the spatiotemporal information and the position of the audio signal receiver.

[0084] According to an embodiment of the present application, a signal path is assigned to the second sound source signal, and obtaining the first signal path includes: selecting an output channel of the second sound source signal based on spatiotemporal information and the position of the audio signal receiver, the output channel including at least one of a left channel and a right channel; and using the signal path corresponding to the output channel as the first signal path.

[0085] According to an embodiment of the present application, the number of signal paths included in the left channel is the same as the number of signal paths included in the right channel.

[0086] Exemplarily, the signal path corresponding to the output channel is the playback device for the sound source signal, such as a speaker. Based on the azimuth information of the first sound source signal received by the receiver and the receiver's location, at least one of the left and right channels is selected. The signal path corresponding to the selected output channel serves as the first signal path. That is, the selected output channel and the signal paths on the output channel (e.g., multiple speakers) serve as the first signal path for the second sound source signal. For example, if the azimuth of horse hoofbeats in an audio signal indicates that the horse is running from the left, the left channel is selected based on this azimuth and the receiver's ear location. In practice, the channel selection can change over time. For example, if the horse approaches, the azimuth of the hoofbeats is noted as emanating from the center, and both the left and right channels are selected as the output channels. If the horse runs to the other side, the azimuth of the hoofbeats is noted as emanating from the right, and the right channel is selected as the output channel.

[0087] Since the output channel of the second sound source signal is determined based on the spatiotemporal information and the position of the audio signal receiver, the signal path corresponding to the output channel is the first signal path. Therefore, the first signal path is also determined based on the spatiotemporal information and the position of the audio signal receiver.

[0088] According to the embodiments of the present application, the output channel of the second sound source signal is selected based on spatiotemporal information and the location of the audio signal receiver, and the signal path corresponding to the output channel is used as the first signal path. This enables automatic matching and automatic division of signal output paths. Different sound source signals are output through appropriate channels, allowing the receiver to clearly perceive the spatial sense of the sound.

[0089] Exemplarily, the same number of sound source signal playback devices are symmetrically placed on the left and right channels. The following will take speakers as an example to illustrate the configuration of the playback devices.

[0090] For example, embodiments may include two speaker configurations: a dual tweeter + one woofer system and a four or more tweeter + one woofer system. The number of tweeters is an even number to facilitate left and right channel distribution and ensure a balanced sound field. The woofer plays the low-frequency portion of the signal. It should be noted that a woofer is also a speaker. It should be noted that a woofer is not required; a system can be composed of two full-band speakers, saving one woofer.

[0091] For the dual high-frequency speaker + 1 woofer system, it is used to match the current integrated speakers and other products. In the following, reference will be made to Figure 4 The system is described for distributing signal paths for audio signals.

[0092] The system of 4 high-frequency speakers or more + 1 woofer is a minimum configuration proposed to meet the requirements of better listening quality. The more speakers there are, the greater the freedom of the audio system, which is more conducive to the realization of 3D audio playback. For example, some high-end theaters contain at least 63 speakers, including not only woofers but also sky channels, so as to achieve a better listening experience. In the following, we will refer to Figure 5 Describes the signal path distribution for audio signals based on a system of 4 high-frequency speakers + 1 low-frequency unit.

[0093] Figure 4 The flowchart for allocating signal paths for audio signals according to an embodiment of the present application is schematically shown. Figure 4 The following description is made by taking as an example the case where the first sound source signal separated from the audio signal includes background sound, human voice and sound event.

[0094] Exemplarily, the background sound, human voice and sound event are respectively adjusted in parameters to obtain the corresponding second sound source signals; based on the temporal and spatial information of the background sound, human voice and sound event and the position of the audio signal receiver, the output channel corresponding to the corresponding second sound source signal is selected, and the output channel and the corresponding speaker are used as the first signal path; the second sound source signal is converted into an output signal, which is output through the first signal path (such as channel + speaker).

[0095] like Figure 4 As shown, the output channel includes a left channel and a right channel, and each channel includes 1 speaker (1 high-frequency speaker). The dotted line indicates selectivity, and at least one of the left channel and the right channel can be selected to output the output signal. For example, the background sound after parameter adjustment selects the right channel and the corresponding speaker as the first signal path. Furthermore, the second sound source signal is converted into an output signal based on the transfer function, and output through the first signal path. For example, the background sound after parameter adjustment is converted into a corresponding output signal based on the transfer function, and output through the right channel and 1 speaker on the right channel. In the following, a method for converting the second sound source signal into an output signal based on the transfer function has been explained, and Figure 7 The specific implementation process is explained in the form of coordinates.

[0096] The dual high-frequency speaker + 1 woofer system can only achieve virtual sound from 0 to 180 degrees. With the receiver's position as the origin, it can achieve virtual sound in front of the receiver.

[0097] Figure 5 The flowchart of allocating signal paths for audio signals according to another embodiment of the present application is schematically shown. Figure 5 The following description is made by taking as an example the case where the first sound source signal separated from the audio signal includes background sound, human voice and sound event.

[0098] Exemplarily, the background sound, human voice and sound event are respectively adjusted in parameters to obtain the corresponding second sound source signals; based on the temporal and spatial information of the background sound, human voice and sound event and the position of the audio signal receiver, the output channel corresponding to the corresponding second sound source signal is selected, and the output channel and the corresponding speaker are used as the first signal path; the second sound source signal is converted into an output signal, which is output through the first signal path (such as channel + speaker).

[0099] like Figure 5 As shown, the output channel includes a left channel and a right channel, and each channel includes 2 speakers (2 high-frequency speakers). The dotted line indicates selectivity, and at least one of the left channel and the right channel can be selected to output the output signal. For example, the background sound after parameter adjustment selects the right channel and the corresponding 2 speakers as the first signal path. Furthermore, the second sound source signal is converted into an output signal based on the transfer function, and output through the first signal path. For example, the background sound after parameter adjustment is converted into a corresponding output signal based on the transfer function, and output through the right channel and the 2 speakers on the right channel. In the following, a method for converting the second sound source signal into an output signal based on the transfer function has been explained, and Figure 7 The specific implementation process is explained in the form of coordinates.

[0100] When the number of speakers increases, such as 8 or 16, they can be divided into groups of 4 or 8, divided into left and right channels, and the channels can be allocated according to the above figure.

[0101] Figure 6 The figure schematically shows a structural diagram of the configuration of a speaker according to an embodiment of the present application.

[0102] Figure 6 Shown Figure 5 The corresponding structure is a 4-high-frequency speaker + 1 woofer system. The 4 high-frequency speakers are symmetrically distributed on both sides of the device in groups of two. The two speakers on the left are connected to the left channel, and the two speakers on the right are connected to the right channel.

[0103] According to an embodiment of the present application, an even number of audio playback devices are symmetrically distributed to the left and right channels, which can achieve a balanced sound field.

[0104] Exemplarily, in order for the receiver to receive the adjusted sound source signal from the surrounding space, it is necessary to perform signal conversion on the second sound source signal and convert it into an output signal output through the first signal path.

[0105] According to an embodiment of the present application, the method for processing an audio signal also includes: determining a transfer function based on the spatiotemporal information of the first sound source signal and the position of the audio signal receiver; calculating an output signal corresponding to the second sound source signal based on the transfer function and the second sound source signal; and outputting the output signal through the first signal path.

[0106] According to an embodiment of the present application, determining the transfer function based on the spatiotemporal information of the first sound source signal and the position of the audio signal receiver includes: determining a first transfer function based on the spatiotemporal information and the position of the audio signal receiver; determining a second transfer function based on the first signal path and the position of the audio signal receiver; and combining the first transfer function and the second transfer function to obtain a transfer function.

[0107] Exemplarily, the first transfer function is determined based on the orientation information of the first sound source signal and the position of the audio signal receiver. The first transfer function may be a transfer function from the position determined based on the orientation information of the first sound source signal to the human ear. For example, the first transfer function is a transfer function from 50 meters southeast to the human ear.

[0108] Exemplarily, the second transfer function is determined based on the position of the signal path of the output sound source signal and the position of the audio signal receiver. The second transfer function can be a transfer function from the audio playback device to the human ear. For example, the second transfer function is a transfer function from the terminal speaker to the human ear.

[0109] The first transfer function and the second transfer function are combined to form a transfer function, which is used to convert the second sound source signal into a corresponding output signal. The specific formula is as follows:

[0110] (1)

[0111] in, is the first transfer function, is the angle of the sound source signal relative to the receiver, denote the transfer function to the left ear and the transfer function to the right ear, respectively. is the second transfer function, n is the number of signal paths (for example, the number of speakers), H nL ,H nR They represent the transfer function to the left ear and the transfer function to the right ear respectively, S is the second sound source signal, and Output is the output signal.

[0112] Figure 7 The figure schematically shows a schematic diagram of determining an output signal using a spherical coordinate system according to an embodiment of the present application.

[0113] In the audio playback scenario, since the original configuration of the speakers is known, the listening position of the receiver can be designed and determined in advance. Therefore, a spherical coordinate system is established with the listening position of the receiver as the origin to calculate the output signal corresponding to the second sound source signal. Figure 7 As shown, take 4 speakers as an example:

[0114] The transfer functions from speakers #1 to #4 to the human ear are H nL (n=1,2,3,4), HnR (n=1,2,3,4), H nL , H nR are the transfer functions to the left ear and the right ear respectively. Assuming that in a spherical coordinate system, a speaker is placed every θ from the South Pole to the North Pole, the transfer functions of each speaker to the human ear are , . , These are the transfer functions to the left ear and the right ear, respectively. Considering that small angle changes cannot be detected by the human ear, it is not recommended to use angle values ​​that are too small. The recommended value is 30°. Therefore, according to different orientation annotation information, select the corresponding transfer function , Therefore, the signal that should be applied to speakers #1 to #4, that is, the output signal corresponding to the second sound source signal is:

[0115] (1)

[0116] Using four speakers The sound from the direction can be simulated based on the fact that the product of the output signals of the four speakers and the transfer functions of the four speakers to the human ear is equal to the second sound source signal and The output signals applied to the four speakers are obtained by multiplying the transfer functions of the directional speakers to the human ear, which is expressed as the above formula (1).

[0117] According to an embodiment of the present application, by determining the output signal applied to the speaker corresponding to the adjusted sound source signal, the receiver can hear the sound from the surrounding space from the speaker, thereby achieving the effect of virtual sound surround playback.

[0118] According to the embodiment of the present application, the first signal path is determined based on the spatiotemporal information and the position of the audio signal receiver. The position of the audio signal receiver is actually the position of the receiver's ear, and the automatic ear tracking function can be used to determine the position of the human ear. Specifically, the camera information is used to detect the key points of the face, and the face key point recognition algorithm is used to identify the nose, ears and head. The direction of the human ear is obtained based on the coordinate information of the key points in the single frame image. The direction is defined as , , , They are respectively in the left ear direction and the right ear direction.

[0119] The control logic of the automatic ear tracking function is as follows: a) If the ear information can be directly identified from the image, the direction of the ear is calculated based on the position of the ear in the single-frame image; b) If the ear is occluded, the direction of the ear is estimated by translating 0.08m left and right based on the position of the nose in the single-frame image; c) If both the ear and nose are occluded, the left and right boundaries of the head in the image are used as the direction of the ear based on the position of the head in the single-frame image.

[0120] According to the direction of the human ear , , the Output in formula (1) can be deflected, and the deflection matrix is ​​defined as W, so the output signal can be expressed as:

[0121] OUPUT = Output · W (2)

[0122] Among them, OUTPUT is the signal output by the final audio playback device.

[0123] The automatic ear tracking function can accurately determine the position of the human ear, improve the accuracy of the receiver's position positioning, and achieve better surround sound playback effects.

[0124] In an embodiment of the present application, the user's consent or authorization may be obtained before obtaining the user's information. For example, before automatically tracking the ear, a request to obtain the user's information may be sent to the user. If the user agrees or authorizes the acquisition of the user's information, the automatic ear tracking function is executed.

[0125] According to the audio signal processing method of the embodiment of the present application, including the configuration of the speaker, preprocessing of the audio signal, parameter adjustment, path allocation of the audio signal, and automatic tracking of the human ear, the control of the sound field is realized under the condition that the hardware remains unchanged. The sound field can be automatically adjusted and adjusted according to the playback content, thereby improving the user's listening experience.

[0126] Based on the above audio signal processing method, the present application also provides an electronic device.

[0127] The electronic device includes: one or more processors; one or more memories for storing executable instructions, and when the executable instructions are executed by the processors, the following method is performed: obtaining an audio signal, separating audio information in the audio signal to form a first sound source signal, where the audio signal includes at least one first sound source signal; adjusting parameters of the first sound source signal based on scene information corresponding to the audio signal to form a second sound source signal, where the second sound source signal adapts to the spatiotemporal information of the audio signal receiver receiving the first sound source signal, where the scene information includes at least one picture layer corresponding to the audio signal; and allocating a signal path for the second sound source signal, obtaining a first signal path, where the first signal path is determined based on the spatiotemporal information and the position of the audio signal receiver.

[0128] Figure 8 The block diagram schematically shows an electronic device suitable for implementing the method for processing audio signals according to an embodiment of the present application.

[0129] like Figure 8 As shown, an electronic device 800 according to an embodiment of the present application includes a processor 801, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 802 or a program loaded from a storage unit 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.

[0130] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in the one or more memories.

[0131] According to an embodiment of the present application, electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to bus 804. Electronic device 800 may also include one or more of the following components connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 808 including a hard disk; and a communication section 809 including a network interface card such as a LAN card or modem. Communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. Removable media 811, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 810 as needed, so that computer programs read from the removable media can be installed into storage section 808 as needed.

[0132] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.

[0133] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.

[0134] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the audio signal processing method provided in the embodiments of the present application.

[0135] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the processor 801 executes the computer program. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0136] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0137] In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0138] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0140] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.

Claims

1. A method for processing an audio signal, wherein: The method comprises: Acquire an audio signal, separate audio information from the audio signal to form a first sound source signal, wherein the audio signal includes at least one first sound source signal; Adjusting parameters of the first sound source signal based on scene information corresponding to the audio signal to form a second sound source signal, wherein the second sound source signal is adapted to the spatiotemporal information of the first sound source signal received by an audio signal receiver; and A signal path is allocated to the second sound source signal, and a first signal path is obtained, where the first signal path is determined based on the spatiotemporal information and the position of the audio signal receiver.

2. The method according to claim 1, wherein The audio signal includes a two-channel sound source, and separating audio information from the audio signal to form a first sound source signal includes: Identifying a sound object in the binaural sound source, wherein the sound object labels a subject that generates a sound source signal in the binaural sound source; and The audio information in the two-channel sound source is separated based on the sound object to form the first sound source signal.

3. The method according to claim 1, wherein The adjusting the parameters of the first sound source signal based on the scene information corresponding to the audio signal to form the second sound source signal includes: Parameter adjustment is performed on the first sound source signal according to characteristics of the first sound source signal in the scene information, where the parameter adjustment includes at least one of sound image adjustment, phase adjustment, and amplitude adjustment, and the characteristics include at least one of the sound image, phase, and amplitude of the first sound source signal.

4. The method according to claim 1, wherein The method further comprises: identifying an object in the displayed video corresponding to the audio signal; and acquiring a motion of the object and / or a type of the object as the scene information, wherein the motion of the object and / or the type of the object matches the first sound source signal; The adjusting the parameters of the first sound source signal based on the scene information corresponding to the audio signal to form the second sound source signal includes: The parameters of the first sound source signal are adjusted based on the motion of the object and / or the type of the object to form a second sound source signal.

5. The method according to claim 1, wherein The spatiotemporal information includes: direction information, time information, and intensity information. The direction information marks the direction in which the audio signal receiver receives the first sound source signal, the time information marks the time when the audio signal receiver receives the first sound source signal, and the intensity information marks the magnitude of the first sound source signal received by the audio signal receiver.

6. The method according to claim 1, wherein The method further comprises: determining a transfer function based on the spatiotemporal information of the first sound source signal and the position of the audio signal receiver; Calculating an output signal corresponding to the second sound source signal based on the transfer function and the second sound source signal; and The output signal is output through the first signal path.

7. The method according to claim 6, wherein: The determining of the transfer function based on the spatiotemporal information of the first sound source signal and the position of the audio signal receiver includes: determining a first transfer function based on the spatiotemporal information and the position of the audio signal receiver; determining a second transfer function based on the first signal path and the location of the audio signal recipient; and The first transfer function and the second transfer function are combined to obtain the transfer function.

8. The method according to claim 1, wherein Allocating a signal path for the second sound source signal to obtain the first signal path includes: selecting an output channel of the second sound source signal based on the spatiotemporal information and the position of the audio signal receiver, the output channel comprising at least one of a left channel and a right channel; and The signal path corresponding to the output channel is used as the first signal path.

9. The method according to claim 8, wherein The number of signal paths included in the left channel is the same as the number of signal paths included in the right channel.

10. An electronic device, wherein: The electronic device comprises: one or more processors; One or more memories for storing executable instructions, wherein when the executable instructions are executed by the processor, the following method is performed: Acquire an audio signal, separate audio information from the audio signal to form a first sound source signal, wherein the audio signal includes at least one first sound source signal; Adjusting parameters of the first sound source signal based on scene information corresponding to the audio signal to form a second sound source signal, wherein the second sound source signal is adapted to the spatiotemporal information of the first sound source signal received by an audio signal receiver, wherein the scene information includes at least one picture layer corresponding to the audio signal; and A signal path is allocated to the second sound source signal, and a first signal path is obtained, where the first signal path is determined based on the spatiotemporal information and the position of the audio signal receiver.