Audio processing method and apparatus, and display device

WO2026174893A1PCT designated stage Publication Date: 2026-08-27SHENZHEN TCL NEW-TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/139741
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-20
Filing Date
2025-12-03
Publication Date
2026-08-27

Smart Images

  • Figure CN2025139741_27082026_PF_FP_ABST
    Figure CN2025139741_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses an audio processing method and apparatus, and a display device. The method comprises: acquiring a camera viewing angle corresponding to a currently played image of a target display device; on the basis of the camera viewing angle, determining a primary audio channel corresponding to the currently played image; and on the basis of the primary audio channel and an original audio signal corresponding to the currently played image, determining a corresponding target audio signal for each of a plurality of audio channels.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing method, device and display device

[0001] The present application claims priority from the Chinese patent application No. 202510189452.7, filed on February 20, 2025, and entitled "Audio processing method, device, display device and computer readable storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the technical field of audio processing, in particular to an audio processing method, device and display device. BACKGROUND

[0003] Currently, the mainstream terminal device for displaying images, such as a television, is 2D, and the camera angle for shooting a video is the same as the human eye angle, so the best experience for the audience should be to follow the camera angle to watch the image and listen to the sound. TECHNICAL PROBLEM

[0004] However, the picture scene of the existing terminal device has the problem of inaccurate sound image position, for example, a person stands at the edge of a cliff to watch a waterfall, at this time the picture angle is a downward view, but the audio signal is output from the left and right or surround sound channels. TECHNICAL SOLUTION

[0005] The embodiments of the present application provide an audio processing method, device and display device, which can determine the target audio signal corresponding to each sound channel in combination with the camera angle corresponding to the current playing picture, thereby improving the accuracy of the sound image position of the current playing picture.

[0006] In a first aspect, the present application provides an audio processing method applied to a target display device, the target display device comprising a plurality of sound channels, and the audio processing method comprising:

[0007] obtaining a camera angle corresponding to a current playing picture of the target display device;

[0008] determining a main sound channel corresponding to the current playing picture based on the camera angle;

[0009] determining a target audio signal corresponding to each sound channel in the plurality of sound channels based on the main sound channel and an original audio signal corresponding to the current playing picture.

[0010] In some embodiments of the present application, the main sound channel corresponding to the current playing picture is determined based on the camera angle, comprising:

[0011] if the camera angle is a downward view angle, determining the main sound channel corresponding to the current playing picture as a sky sound channel;

[0012] if the camera angle is a horizontal view angle, determining the main sound channel corresponding to the current playing picture as a left-right sound channel.

[0013] If the camera view angle is a top view angle, the main channel corresponding to the current playing picture is determined as a ground channel.

[0014] In some embodiments of the present application, based on the main channel and the original audio signal corresponding to the current playing picture, the target audio signal corresponding to each channel in the plurality of channels is determined, including:

[0015] Based on the main channel, the channel gain value corresponding to each channel in the plurality of channels is determined;

[0016] Based on the channel gain value and the original audio signal corresponding to the current playing picture, the target audio signal corresponding to each channel is determined.

[0017] In some embodiments of the present application, based on the main channel, the channel gain value corresponding to each channel in the plurality of channels is determined, including:

[0018] The channel gain value corresponding to the main channel in the plurality of channels is determined as a first gain value;

[0019] The channel gain value corresponding to the other channels in the plurality of channels except the main channel is determined as a second gain value; wherein the first gain value is greater than the second gain value.

[0020] In some embodiments of the present application, based on the channel gain value and the original audio signal corresponding to the current playing picture, the target audio signal corresponding to each channel is determined, including:

[0021] The channel gain value corresponding to each channel and the original audio signal corresponding to the current playing picture are multiplied to obtain the target audio signal corresponding to each channel.

[0022] In some embodiments of the present application, when the plurality of channels of the target display device includes a sky channel, a ground channel and a left-right channel, and the sky channel is the main channel, the determination process of the target audio signal corresponding to each channel can be represented as: HeadLROut = γ1*LRIn, LROut = γ2*LRIn, SurLROut = γ2*LRIn, HeadLROut represents the target audio signal corresponding to the sky channel, LROut represents the target audio signal corresponding to the left-right channel, SurLROut represents the target audio signal corresponding to the ground channel, LRIn represents the original audio signal, γ1 represents the first gain value, and γ2 represents the second gain value.

[0023] In some embodiments of the present application, the plurality of channels includes a sky channel, a ground channel and a left-right channel, and the target display device includes: a central processing unit, a first speaker, a second speaker and a third speaker, the first speaker, the second speaker and the third speaker are connected with the central processing unit respectively;

[0024] The target audio signal corresponding to the sky channel is output through the first loudspeaker, the target audio signal corresponding to the ground channel is output through the second loudspeaker, and the target audio signal corresponding to the left-right channel is output through the third loudspeaker.

[0025] In some embodiments of the present application, the first loudspeaker and the second loudspeaker are external loudspeakers, or the first loudspeaker and the second loudspeaker are built-in loudspeakers.

[0026] When the first loudspeaker and the second loudspeaker are built-in loudspeakers, after determining the target audio signal corresponding to each channel in the plurality of channels based on the original audio signal corresponding to the main channel and the current playing picture, the method comprises:

[0027] The target audio signal corresponding to the sky channel and the target audio signal corresponding to the ground channel are respectively filtered.

[0028] The filtered target audio signal corresponding to the sky channel is output through the first loudspeaker.

[0029] The filtered target audio signal corresponding to the ground channel is output through the second loudspeaker.

[0030] The target audio signal corresponding to the left-right channel is output through the third loudspeaker.

[0031] In some embodiments of the present application, when the first loudspeaker and the second loudspeaker are external loudspeakers, after determining the target audio signal corresponding to each channel in the plurality of channels based on the original audio signal corresponding to the main channel and the current playing picture, the method comprises:

[0032] The target audio signal corresponding to the sky channel is output through the first loudspeaker.

[0033] The target audio signal corresponding to the ground channel is output through the second loudspeaker.

[0034] The target audio signal corresponding to the left-right channel is output through the third loudspeaker.

[0035] In some embodiments of the present application, the camera view angle corresponding to the current playing picture of the target display device is obtained, comprising:

[0036] A first image and a second image corresponding to the current playing picture of the target display device are obtained.

[0037] Based on the first image and the second image, a target camera extrinsic parameter of a target camera corresponding to the current playing picture is determined.

[0038] Based on the target camera extrinsic parameter, the camera view angle corresponding to the current playing picture of the target display device is determined.

[0039] In some embodiments of the present application, based on the first image and the second image, the target camera extrinsic parameter of the target camera corresponding to the current playing picture is determined, comprising:

[0040] The first image and the second image are pixel point matched to obtain a plurality of second target pixel points in the second image corresponding to a plurality of first target pixel points in the first image;

[0041] Based on the initial camera intrinsic parameter of the target camera corresponding to the current playing picture, the normalized spatial point coordinates corresponding to each first target pixel point in the plurality of first target pixel points are determined;

[0042] Based on the initial camera extrinsic parameter corresponding to the target camera, the normalized spatial point coordinates are re-projected onto the second image to obtain a re-projection pixel point coordinate corresponding to each second target pixel point;

[0043] Based on the original pixel point coordinates and the re-projection pixel point coordinates corresponding to each second target pixel point, a re-projection error corresponding to each second target pixel point is determined;

[0044] Based on the re-projection error, the initial camera intrinsic parameter and the initial camera extrinsic parameter are updated, and the step of determining the normalized spatial point coordinates corresponding to each first target pixel point in the plurality of first target pixel points based on the initial camera intrinsic parameter of the target camera corresponding to the current playing picture is continued to be executed until the re-projection error meets the first condition;

[0045] The updated initial camera extrinsic parameter is determined as the target camera extrinsic parameter of the target camera corresponding to the current playing picture.

[0046] In some embodiments of the present application, the determination process of the normalized spatial point coordinates corresponding to each first target pixel point is: X i = Z i *(u i -c x ) / f x , Y i = Z i *(v i -c y ) / f y , Z = depth / s, (X i , Y i , Z i ) represents the normalized spatial point coordinates corresponding to the i-th first target pixel point, (u i , v i ) represents the original pixel coordinates corresponding to the i-th first target pixel point, represents the initial camera intrinsic parameter, depth represents the depth value of the target camera, and s represents the conversion parameter of the depth value and the actual distance.

[0047] In some embodiments of the present application, the determination process of the re-projection pixel point coordinates is as follows: λ ij denotes a scale factor, (X i , Y i , Z i ) denotes the normalized spatial point coordinates corresponding to the i-th first target pixel point, P(R j , t j ) denotes the projection matrix of the target camera j, R j and t j denote the initial camera extrinsic parameters of the target camera j, R j denotes the rotation matrix R j of the target camera j, t j denotes the translation vector of the target camera j, and (u ij , v ij ) denotes the re-projection pixel point coordinates of the second target pixel point corresponding to the i-th first target pixel point.

[0048] In some embodiments of the present application, the determination process of the re-projection error is as follows: wherein (u ij , v ij ) denotes the re-projection pixel point coordinates of the second target pixel point corresponding to the i-th first target pixel point, denotes the original pixel coordinates of the second target pixel point corresponding to the i-th first target pixel point.

[0049] In some embodiments of the present application, the first condition satisfied by the re-projection error is that the re-projection error corresponding to all the second target pixel points is less than a first threshold, or the first condition satisfied by the re-projection error is that the sum of squares of the re-projection errors corresponding to all the second target pixel points is less than a second threshold, or the first condition satisfied by the re-projection error is that the sum of squares of the re-projection errors corresponding to all the second target pixel points obtained by the previous and subsequent times is less than a third threshold.

[0050] In some embodiments of the present application, the camera view angle corresponding to the current playing picture of the target display device is obtained, comprising:

[0051] obtaining the current playing picture of the target display device;

[0052] performing target recognition on the current playing picture to obtain a target recognition result of the current playing picture;

[0053] determining the camera view angle corresponding to the current playing picture based on the target recognition result.

[0054] In some embodiments of the present application, target recognition is performed on the current playing picture to obtain a target recognition result of the current playing picture, including:

[0055] The current playing picture is input into the target recognition model, and the target recognition result of the current playing picture is output by the target recognition model.

[0056] In some embodiments of the present application, based on the target recognition result, the camera view corresponding to the current playing picture is determined, including:

[0057] Based on the target recognition result and the first association information, the camera view corresponding to the current playing picture is determined, wherein the first association information is used to represent the corresponding relationship between the recognition result and the camera view.

[0058] In a second aspect, the embodiments of the present application also provide an audio processing apparatus applied to a target display device, the target display device including a plurality of sound channels, and the audio processing apparatus including:

[0059] An information acquisition module is configured to acquire a camera view corresponding to a current playing picture of the target display device;

[0060] A first determination module is configured to determine a main sound channel corresponding to the current playing picture based on the camera view;

[0061] A second determination module is configured to determine a target audio signal corresponding to each sound channel in the plurality of sound channels based on the main sound channel and an original audio signal corresponding to the current playing picture.

[0062] In a third aspect, the present application also provides a display device, including:

[0063] One or more processors;

[0064] A memory; and

[0065] One or more application programs, wherein the one or more application programs are stored in the memory and are configured to be executed by the processor to implement the audio processing method of any one of the first aspect.

[0066] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to execute the steps in the audio processing method of any one of the first aspect. Advantages

[0067] The target audio signal of each channel can be determined in combination with the camera view corresponding to the current playing picture, thereby improving the accuracy of the sound image position of the current playing picture. BRIEF DESCRIPTION OF DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0069] FIG. 1 is a scene schematic diagram of a target display device according to an embodiment of the present application;

[0070] FIG. 2 is a scene schematic diagram of the installation position of a loudspeaker of a target display device according to an embodiment of the present application;

[0071] FIG. 3 is a flow schematic diagram of an audio processing method according to an embodiment of the present application;

[0072] FIG. 4 is a specific embodiment flow schematic diagram of obtaining a camera view corresponding to a current playing picture according to an embodiment of the present application;

[0073] FIG. 5 is a specific embodiment flow schematic diagram of determining a target audio signal according to an embodiment of the present application;

[0074] FIG. 6 is a principle block diagram of an audio processing device according to an embodiment of the present application;

[0075] FIG. 7 is a structure schematic diagram of a display device according to an embodiment of the present application. DETAILED DESCRIPTION

[0076] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0077] In the description of the present application, the terms "first", "second", "third", etc. are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second", "third", etc. can be explicitly or implicitly included one or more of the features.

[0078] In the present application, the word "exemplary" is used to mean "serving as an example, instance, or illustration". Any implementation described as "exemplary" in the present application is not necessarily to be construed as preferred or advantageous over other implementations. The following description is presented to enable any person skilled in the art to make and use the present application. In the following description, for the purposes of explanation, details are set forth in order to provide a thorough understanding of the present application. It should be appreciated that one of ordinary skill in the art will realize that the present application can be practiced without the use of these specific details. In other instances, well-known structures and processes are not elaborated in order not to obscure the description of the present application with unnecessary details. Thus, the present application is not intended to be limited by the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0079] It should be noted that the method of the present application is executed in a computer device, and the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It can be understood that in subsequent embodiments, if the size, quantity, position, etc. are mentioned, they all exist in the form of corresponding data for processing by the computer device, and specific details are not repeated here.

[0080] The present application provides an audio processing method and device, a display device and a computer readable storage medium, which are described in detail below.

[0081] The audio processing method provided by the present application is applied to a target display device, which can be any device with display function such as a television, a notebook computer, a tablet computer, a projector, etc. Please refer to FIG. 1, which is a schematic diagram of a scene of a target display device provided by the present application. The target display device can include a central processing unit (CPU) 110 and a plurality of sound channels. A sound channel refers to an independent audio signal collected or played back at different spatial positions during recording or playing. The central processing unit 110 is mainly used to obtain a camera viewing angle corresponding to a current playing picture of the target display device; determine a main sound channel corresponding to the current playing picture based on the camera viewing angle; and determine a target audio signal corresponding to each sound channel in the plurality of sound channels based on the main sound channel and an original audio signal corresponding to the current playing picture. The target audio signal corresponding to each sound channel can be determined in combination with the camera viewing angle corresponding to the current playing picture, thereby improving the accuracy of the sound image position of the current playing picture.

[0082] Further, continuing to refer to FIG. 1, the plurality of sound channels include a sky sound channel, a ground sound channel, and a left-right sound channel, and the target display device further includes a first speaker 120, a second speaker 130, and a third speaker 140, which are respectively connected to the central processor 110, the audio signal corresponding to the sky sound channel is output by the first speaker 120, the audio signal corresponding to the ground sound channel is output by the second speaker 130, and the audio signal corresponding to the left-right sound channel is output by the third speaker 140.

[0083] Optionally, continuing to refer to FIG. 1, the plurality of sound channels further include a surround sound channel, and the target display device further includes a fourth speaker 150 connected to the central processor 110, and the audio signal corresponding to the surround sound channel is output by the fourth speaker 150. For example, referring to FIG. 2, the target display device includes a plurality of loudspeakers, which are respectively arranged above, below, in front of, and behind the user, and the plurality of loudspeakers can output the audio signals corresponding to the sky, ground, left-right, and surround sound channels respectively, so as to generate a 360° spatial sound field in the horizontal and vertical directions.

[0084] In some embodiments, the target display device can further include a memory 160 connected to the central processor 110, and the memory 160 is configured to store data, such as camera information, for example, camera view angle, initial camera external parameter, initial camera internal parameter, etc., and audio information, for example, original audio signal, target audio signal, etc.

[0085] In addition, continuing to refer to FIG. 1, the target display device further includes a microphone 170 connected to the central processor 110, and the target display device can obtain a voice control instruction of the user through the microphone 170, and perform a corresponding control operation based on the voice control instruction. For example, the user issues a voice control instruction of “turn off” to the target display device, and the target display device performs a turn-off operation after receiving the control instruction through the microphone 170.

[0086] It should be noted that the scene diagram of the target display device shown in FIG. 1 is only an example, and the target display device and the scene described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. It is known to those skilled in the art that, as the target display device evolves and new business scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0087] First, the embodiment of the present application provides an audio processing method, the execution subject of the audio processing method is an audio processing device, the audio processing device is applied to a target display device, the target display device includes a plurality of sound channels, and the audio processing method includes the following steps: acquiring a camera view angle corresponding to a current playing picture of the target display device; determining a main sound channel corresponding to the current playing picture based on the camera view angle; and determining a target audio signal corresponding to each sound channel in the plurality of sound channels based on the main sound channel and an original audio signal corresponding to the current playing picture.

[0088] As shown in FIG. 3, it is a flowchart of one embodiment of the audio processing method in the embodiment of the present application, which can include the following steps S201-S203, as follows:

[0089] Step S201: acquiring a camera view angle corresponding to a current playing picture of a target display device.

[0090] In the embodiment of the present application, the camera view angle corresponding to the current playing picture refers to the relative position and angle between the target camera used for shooting the current playing picture and the object being shot in the current playing picture, and the camera view angle corresponding to the current playing picture can be any one of a horizontal view angle, a top view angle and a downward view angle. The horizontal view angle refers to the view angle in which the target camera and the object being shot in the current playing picture are at approximately the same height, the top view angle refers to the view angle in which the target camera is higher than the object being shot in the current playing picture, and the downward view angle refers to the view angle in which the target camera is lower than the object being shot in the current playing picture.

[0091] Optionally, the camera view angle corresponding to the current playing picture can be determined by the target display device based on the current playing picture, and the camera view angle corresponding to the current playing picture can also be acquired from other computer devices through a network, Bluetooth, infrared or the like, which is not limited in the embodiment.

[0092] In some embodiments, the step of acquiring the camera view angle corresponding to the current playing picture of the target display device specifically includes: acquiring the current playing picture of the target display device; performing target recognition on the current playing picture to obtain a target recognition result of the current playing picture; and determining the camera view angle corresponding to the current playing picture based on the target recognition result.

[0093] In the embodiment of the present application, the step of performing target recognition on the current playing picture to obtain the target recognition result of the current playing picture specifically includes: inputting the current playing picture into a target recognition model to output the target recognition result of the current playing picture through the target recognition model. The target recognition model is a pre-trained neural network model used for target recognition of the current playing picture.

[0094] Further, when determining the camera view corresponding to the current playing picture based on the target recognition result, the camera view corresponding to the current playing picture can be determined based on the target recognition result and first association information. The first association information is used to represent the corresponding relationship between the recognition result and the camera view. For example, when the target recognition result is an airplane, the camera view corresponding to the current playing picture is determined to be an overhead view. When the target recognition result is a waterfall, the camera view corresponding to the current playing picture is determined to be a downward view.

[0095] In some embodiments, referring to FIG. 4, the above-mentioned obtaining the camera view corresponding to the current playing picture of the target display device can include steps S301-S303, which are as follows:

[0096] In step S301, the first image and the second image corresponding to the current playing picture of the target display device are obtained.

[0097] In the embodiments of the present application, the current playing picture includes multiple continuous images. The first image and the second image can be any two continuous images in the multiple images, or the first image and the second image can be any two discontinuous images in the multiple images, which is not limited in the embodiments. Alternatively, the first image and the second image can be any two continuous images in the multiple images, which can improve the accuracy of the determined camera view.

[0098] In step S302, the target camera extrinsic parameter of the target camera corresponding to the current playing picture is determined based on the first image and the second image.

[0099] In the embodiments of the present application, the target camera extrinsic parameter is the camera extrinsic parameter of the target camera corresponding to the current playing picture, which is obtained by camera calibration based on the first image and the second image. Camera calibration refers to the process of determining the intrinsic parameter and the extrinsic parameter of the camera. The camera intrinsic parameter includes the focal length and the distortion coefficient of the camera, and the camera extrinsic parameter includes the position and the attitude of the camera.

[0100] In some embodiments, the step of determining the target camera extrinsic parameter of the target camera corresponding to the current playing picture based on the first image and the second image specifically comprises: performing pixel point matching on the first image and the second image to obtain a plurality of second target pixel points in the second image corresponding to a plurality of first target pixel points in the first image; determining a normalized spatial point coordinate corresponding to each first target pixel point in the plurality of first target pixel points based on the initial camera intrinsic parameter of the target camera corresponding to the current playing picture; re-projecting the normalized spatial point coordinate onto the second image based on the initial camera extrinsic parameter corresponding to the target camera to obtain a re-projected pixel point coordinate corresponding to each second target pixel point; determining a re-projection error corresponding to each second target pixel point based on an original pixel point coordinate corresponding to the second target pixel point and the re-projected pixel point coordinate; updating the initial camera intrinsic parameter and the initial camera extrinsic parameter based on the re-projection error, and continuing to perform the step of determining the normalized spatial point coordinate corresponding to each first target pixel point in the plurality of first target pixel points based on the initial camera intrinsic parameter of the target camera corresponding to the current playing picture until the re-projection error meets a first condition; and determining the updated initial camera extrinsic parameter as the target camera extrinsic parameter of the target camera corresponding to the current playing picture.

[0101] Optionally, the determination process of the normalized spatial point coordinate corresponding to each first target pixel point can be represented as: X i = Z i *(u i -c x ) / f x , Y i = Z i *(v i -c y ) / f y , Z = depth / s, (X i , Y i , Z i ) represents the normalized spatial point coordinate corresponding to the i-th first target pixel point, (u i , v i ) represents the original pixel coordinate corresponding to the i-th first target pixel point, represents the initial camera intrinsic parameter, depth represents the depth value of the target camera, and s represents the conversion parameter of the depth value and the actual distance. Optionally, when the target camera is not a depth camera, Z = 1.

[0102] In some embodiments, the determination process of the re-projected pixel point coordinate is: λ ij represents a scale factor, (X i , Y i , Z i) represents the normalized spatial point coordinate corresponding to the i-th first target pixel point, P(R j , t j ) represents the projection matrix of the target camera j, R j and t j represent the initial camera extrinsic parameters of the target camera j, R j represents the rotation matrix R j of the target camera j, t j represents the translation vector of the target camera j, and (u ij , v ij ) represents the re-projection pixel point coordinate of the second target pixel point corresponding to the i-th first target pixel point.

[0103] Further, the determination process of the re-projection error can be represented as: wherein (u ij , v ij ) represents the re-projection pixel point coordinate of the second target pixel point corresponding to the i-th first target pixel point, and P(R represents the original pixel coordinate of the second target pixel point corresponding to the i-th first target pixel point.

[0104] In some embodiments, the first condition satisfied by the re-projection error can be that the re-projection error corresponding to all the second target pixel points is less than a first threshold, the first condition satisfied by the re-projection error can also be that the sum of squares of the re-projection error corresponding to all the second target pixel points is less than a second threshold, and the first condition satisfied by the re-projection error can also be that the sum of squares of the re-projection error corresponding to all the second target pixel points obtained twice in succession is less than a third threshold, and the present embodiment is not limited thereto.

[0105] In step S303, the camera view angle corresponding to the current playing picture of the target display device is determined based on the target camera extrinsic parameters.

[0106] In the present embodiment, the target camera extrinsic parameters include the rotation matrix and the translation matrix of the target camera. By analyzing the rotation matrix in the target camera extrinsic parameters, the orientation of the target camera can be determined, so as to determine whether the camera view angle corresponding to the current playing picture of the target display device is a level view, a downward view or an upward view.

[0107] In step S202, the main channel corresponding to the current playing picture is determined based on the camera view angle.

[0108] In the embodiments of the present application, the main channel is a channel in the multiple channels of the target display device that matches the camera perspective corresponding to the current playing picture. The main channel can be any one or more of the multiple channels. For example, when the camera perspective is an upward-looking perspective, the main channel corresponding to the current playing picture is the sky channel and the left-right channel. When the camera perspective is a level-looking perspective, the main channel corresponding to the current playing picture is the left-right channel. When the camera perspective is a downward-looking perspective, the main channel corresponding to the current playing picture is the ground channel and the action channel.

[0109] For another example, when the camera perspective is an upward-looking perspective, the main channel corresponding to the current playing picture is the sky channel. When the camera perspective is a level-looking perspective, the main channel corresponding to the current playing picture is the left-right channel. When the camera perspective is a downward-looking perspective, the main channel corresponding to the current playing picture is the ground channel.

[0110] In some embodiments, the multiple channels include a left-right channel, a sky channel and a ground channel, and the main channel is one channel in the multiple channels that matches the camera perspective corresponding to the current playing picture. Optionally, the step of determining the main channel corresponding to the current playing picture based on the camera perspective specifically includes: if the camera perspective is an upward-looking perspective, determining that the main channel corresponding to the current playing picture is the sky channel; if the camera perspective is a level-looking perspective, determining that the main channel corresponding to the current playing picture is the left-right channel; and if the camera perspective is a downward-looking perspective, determining that the main channel corresponding to the current playing picture is the ground channel. In this embodiment, the main channel corresponding to the current playing picture is determined based on the camera perspective, which can adjust the audio in combination with the camera perspective corresponding to the current playing picture, thereby improving the accuracy of the sound image position of the current playing picture.

[0111] In step S203, target audio signals corresponding to each channel in the multiple channels are determined based on the main channel and the original audio signal corresponding to the current playing picture.

[0112] In the embodiments of the present application, the original audio signal is an original audio signal corresponding to the current playing picture. For example, for a playing picture in which a person stands on the edge of a cliff and looks at a waterfall, the original audio signal is the impact sound of the water flow. For a playing picture in which an airplane flies over the head of a person, the original audio signal is the sound of the airplane flying over the head. In this embodiment, target audio signals corresponding to each channel in the multiple channels are determined based on the main channel and the original audio signal corresponding to the current playing picture, which can make the sound of the current playing picture output by the channel that matches the current playing picture, thereby improving the accuracy of the sound image position of the current playing picture.

[0113] In some embodiments, referring to FIG. 5, the step S203 of determining the target audio signals corresponding to each channel in the multiple channels based on the main channel and the original audio signal corresponding to the current playing picture can include steps S401-S402, which are specifically as follows:

[0114] S401, determine a channel gain value corresponding to each of the plurality of channels based on the main channel.

[0115] In the embodiments of the present application, the step of determining a channel gain value corresponding to each of the plurality of channels based on the main channel specifically includes: determining the channel gain value corresponding to the main channel in the plurality of channels as a first gain value; determining the channel gain value corresponding to the other channels in the plurality of channels except the main channel as a second gain value; wherein the first gain value is greater than the second gain value.

[0116] For example, the first gain value is 1 and the second gain value is 0.7, if the main channel is the ground channel, the channel gain value corresponding to the ground channel is set to 1, and the channel gain value corresponding to the other channels in the plurality of channels except the ground channel is set to 0.7, so that the sound of the current playing picture can be mainly output by its main channel, thereby improving the accuracy of the sound image position of the current playing picture.

[0117] S402, determine a target audio signal corresponding to each of the plurality of channels based on the channel gain value and the original audio signal corresponding to the current playing picture.

[0118] In some embodiments, the step of determining a target audio signal corresponding to each of the plurality of channels based on the channel gain value and the original audio signal corresponding to the current playing picture specifically includes: multiplying the channel gain value corresponding to each of the plurality of channels and the original audio signal corresponding to the current playing picture to obtain the target audio signal corresponding to each of the plurality of channels.

[0119] For example, when the plurality of channels of the target display device includes a sky channel, a ground channel and a left-right channel, and the sky channel is the main channel, the determination process of the target audio signal corresponding to each of the plurality of channels can be represented as: HeadLROut = γ1*LRIn, LROut = γ2*LRIn, SurLROut = γ2*LRIn, HeadLROut represents the target audio signal corresponding to the sky channel, LROut represents the target audio signal corresponding to the left-right channel, SurLROut represents the target audio signal corresponding to the ground channel, LRIn represents the original audio signal, γ1 represents the first gain value, and γ2 represents the second gain value.

[0120] In some embodiments, the plurality of channels includes a sky channel, a ground channel and a left-right channel, and the target display device includes: a central processor, a first speaker, a second speaker and a third speaker, the first speaker, the second speaker and the third speaker are connected with the central processor respectively. Wherein the target audio signal corresponding to the sky channel is output through the first speaker, the target audio signal corresponding to the ground channel is output through the second speaker, and the target audio signal corresponding to the left-right channel is output through the third speaker.

[0121] In some embodiments, the first speaker and the second speaker are external speakers, and after determining the target audio signal corresponding to each sound channel in the plurality of sound channels based on the main sound channel and the original audio signal corresponding to the currently played picture, the method further includes: outputting the target audio signal corresponding to the sky sound channel through the first speaker; outputting the target audio signal corresponding to the ground sound channel through the second speaker; and outputting the target audio signal corresponding to the left-right sound channel through the third speaker.

[0122] In other embodiments, the first speaker and the second speaker are internal speakers, and after determining the target audio signal corresponding to each sound channel in the plurality of sound channels based on the main sound channel and the original audio signal corresponding to the currently played picture, the method further includes: respectively filtering the target audio signal corresponding to the sky sound channel and the target audio signal corresponding to the ground sound channel; outputting the filtered target audio signal corresponding to the sky sound channel through the first speaker; outputting the filtered target audio signal corresponding to the ground sound channel through the second speaker; and outputting the target audio signal corresponding to the left-right sound channel through the third speaker.

[0123] In the embodiments of the present application, the target audio signal corresponding to the sky sound channel and the target audio signal corresponding to the ground sound channel can be respectively filtered by a preset spatial filter. The spatial filter refers to a digital filter generated using a head-related transfer function (HRTF) technique. When the first speaker and the second speaker are internal speakers, and the target display device needs to form a corresponding sound field by reflection, respectively filtering the target audio signal corresponding to the sky sound channel and the target audio signal corresponding to the ground sound channel in the embodiments can improve the quality of the target audio signals corresponding to the sky sound channel and the ground sound channel, and thus can improve the sound field effect corresponding to the currently played picture.

[0124] To better implement the audio processing method in the embodiments of the present application, based on the audio processing method, an audio processing device is further provided in the embodiments of the present application, as shown in FIG. 6, the audio processing device 600 includes:

[0125] An information acquisition module 610 is configured to acquire a camera view angle corresponding to a currently played picture of a target display device.

[0126] A first determination module 620 is configured to determine a main sound channel corresponding to the currently played picture based on the camera view angle.

[0127] A second determination module 630 is configured to determine a target audio signal corresponding to each sound channel in a plurality of sound channels based on the main sound channel and an original audio signal corresponding to the currently played picture.

[0128] In this embodiment, the main channel corresponding to the current playback screen is determined based on the camera view corresponding to the current playback screen of the target display device. The target audio signal corresponding to each channel in the multiple channels is determined based on the main channel and the original audio signal corresponding to the current playback screen. The target audio signal corresponding to each channel can be determined by combining the camera view corresponding to the current playback screen, thereby improving the accuracy of the sound image position of the current playback screen.

[0129] In some embodiments of this application, the first determining module 620 determines the main audio channel corresponding to the currently playing screen based on the camera's viewpoint, including:

[0130] If the camera's view is from below, determine that the main audio channel corresponding to the currently playing image is the sky channel;

[0131] If the camera's view is at eye level, determine that the main audio channel corresponding to the currently playing image is the left or right audio channel;

[0132] If the camera's view is from above, determine that the main audio channel corresponding to the currently playing image is the ground channel.

[0133] In some embodiments of this application, the second determining module 630 determines the target audio signal corresponding to each of the multiple channels based on the original audio signal corresponding to the main channel and the currently playing screen, including:

[0134] Based on the main channel, determine the channel gain value corresponding to each channel in multiple channels;

[0135] Based on the channel gain value and the original audio signal corresponding to the current playback screen, the target audio signal corresponding to each channel is determined.

[0136] In some embodiments of this application, the second determining module 630 determines the channel gain value corresponding to each channel among multiple channels based on the main channel, including:

[0137] The gain value of the channel corresponding to the main channel in multiple channels is determined as the first gain value;

[0138] The channel gain values ​​corresponding to the channels other than the main channel in the multiple channels are determined as the second gain values; wherein, the first gain value is greater than the second gain value.

[0139] In some embodiments of this application, the multiple sound channels include a sky channel, a ground channel, and left and right channels. The target display device includes a central processing unit, a first speaker, a second speaker, and a third speaker, which are respectively connected to the central processing unit.

[0140] The target audio signal corresponding to the sky channel is output through the first speaker, the target audio signal corresponding to the ground channel is output through the second speaker, and the target audio signals corresponding to the left and right channels are output through the third speaker.

[0141] In some embodiments of this application, the first speaker and the second speaker are external speakers, or the first speaker and the second speaker are internal speakers. When the first speaker and the second speaker are internal speakers, after the second determining module 630 determines the target audio signal corresponding to each of the multiple channels based on the original audio signal corresponding to the main channel and the currently playing screen, the second determining module 630 is further configured to:

[0142] The target audio signals corresponding to the sky channel and the target audio signals corresponding to the ground channel are filtered separately.

[0143] The target audio signal corresponding to the sky channel after filtering is output through the first speaker;

[0144] The target audio signal corresponding to the ground channel is output through the second speaker after filtering.

[0145] The target audio signals corresponding to the left and right channels are output through the third speaker.

[0146] In some embodiments of this application, the information acquisition module 610 acquires the camera view corresponding to the current playback screen of the target display device, including:

[0147] Obtain the first and second images corresponding to the current playback screen of the target display device;

[0148] Based on the first image and the second image, determine the target camera extrinsic parameters of the target camera corresponding to the currently playing screen;

[0149] Based on the target camera's extrinsic parameters, determine the camera's viewing angle corresponding to the current playback screen on the target display device.

[0150] In some embodiments of this application, the information acquisition module 610 determines the target camera extrinsic parameters of the target camera corresponding to the currently playing screen based on the first image and the second image, including:

[0151] Pixel matching is performed on the first image and the second image to obtain multiple second target pixels corresponding to multiple first target pixels in the first image in the second image;

[0152] Based on the initial camera intrinsic parameters of the target camera corresponding to the current playback screen, determine the normalized spatial coordinates of each first target pixel among multiple first target pixels;

[0153] Based on the initial camera extrinsic parameters corresponding to the target camera, the normalized spatial point coordinates are reprojected onto the second image to obtain the reprojected pixel coordinates corresponding to each second target pixel.

[0154] Based on the original pixel coordinates and reprojected pixel coordinates corresponding to the second target pixel, the reprojection error corresponding to each second target pixel is determined.

[0155] Based on the reprojection error, the initial camera intrinsic and extrinsic parameters are updated, and the process of determining the normalized spatial coordinates of each first target pixel in the multiple first target pixels is continued based on the initial camera intrinsic parameters of the target camera corresponding to the current playback screen, until the reprojection error meets the first condition.

[0156] The updated initial camera extrinsic parameters are then used as the target camera extrinsic parameters for the target camera corresponding to the currently playing screen.

[0157] This application embodiment also provides a display device that integrates any of the audio processing devices provided in this application embodiment. The display device includes:

[0158] One or more processors;

[0159] Memory; and

[0160] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the audio processing method in any of the embodiments described above.

[0161] This application also provides a display device that integrates any of the audio processing devices provided in this application. Figure 7 shows a schematic diagram of the structure of the display device involved in this application. Specifically:

[0162] The display device may include components such as a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, and an input unit 804. Those skilled in the art will understand that the display device structure shown in FIG7 does not constitute a limitation on the display device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0163] The processor 801 is the control center of the display device. It connects various parts of the display device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 802, and by calling data stored in the memory 802, it performs various functions and processes data of the display device, thereby providing overall monitoring of the display device. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 801.

[0164] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the display device, etc. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.

[0165] The display device also includes a power supply 803 that supplies power to the various components. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 803 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0166] The display device may also include an input unit 804, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0167] Although not shown, the display device may also include display units, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 801 in the display device loads the executable files corresponding to the processes of one or more application programs into the memory 802 according to the following instructions, and the processor 801 runs the application programs stored in the memory 802 to realize various functions, as follows:

[0168] Obtain the camera viewpoint corresponding to the current playback screen of the target display device;

[0169] Based on the camera's perspective, determine the main audio channel corresponding to the currently playing image;

[0170] Based on the original audio signal corresponding to the main channel and the currently playing screen, determine the target audio signal corresponding to each channel in the multiple channels.

[0171] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0172] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the audio processing methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps:

[0173] Obtain the camera viewpoint corresponding to the current playback screen of the target display device;

[0174] Based on the camera's perspective, determine the main audio channel corresponding to the currently playing image;

[0175] Based on the original audio signal corresponding to the main channel and the currently playing screen, determine the target audio signal corresponding to each channel in the multiple channels.

[0176] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0177] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0178] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0179] The above describes in detail the audio processing method and device, the display device and the computer readable storage medium provided by the embodiments of the present application. The principles and implementation manners of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed. In conclusion, the content of the specification should not be understood as a limitation of the present application.

Claims

1. An audio processing method, wherein, Applied to a target display device, the target display device including multiple audio channels, the audio processing method includes: Obtain the camera viewpoint corresponding to the current playback screen of the target display device; Based on the camera's perspective, determine the main audio channel corresponding to the currently playing image; Based on the original audio signal corresponding to the main channel and the currently playing screen, the target audio signal corresponding to each of the multiple channels is determined.

2. The audio processing method according to claim 1, wherein, Determining the main audio channel corresponding to the currently playing image based on the camera's view includes: If the camera's viewing angle is a downward angle, the main audio channel corresponding to the currently playing image is determined to be the sky channel; If the camera's viewing angle is a head-up view, then the main audio channel corresponding to the currently playing image is determined to be the left or right audio channel; If the camera's view is a top-down view, then the main audio channel corresponding to the currently playing image is determined to be the ground audio channel.

3. The audio processing method according to claim 1, wherein, The step of determining the target audio signal corresponding to each of the multiple channels based on the original audio signal corresponding to the main channel and the currently playing screen includes: Based on the main channel, determine the channel gain value corresponding to each of the plurality of channels; Based on the channel gain value and the original audio signal corresponding to the currently playing screen, the target audio signal corresponding to each channel is determined.

4. The audio processing method according to claim 3, wherein, The step of determining the channel gain value corresponding to each of the plurality of channels based on the main channel includes: The gain value of the channel corresponding to the main channel among the plurality of channels is determined as the first gain value; The channel gain values ​​corresponding to the other channels besides the main channel are determined as the second gain values; wherein the first gain value is greater than the second gain value.

5. The audio processing method according to claim 3, wherein, The step of determining the target audio signal corresponding to each channel based on the channel gain value and the original audio signal corresponding to the current playback screen includes: The channel gain value corresponding to each channel is multiplied with the original audio signal corresponding to the currently playing screen to obtain the target audio signal corresponding to each channel.

6. The audio processing method according to claim 3, wherein, When the target display device has multiple channels including a sky channel, a ground channel, and left and right channels, and the sky channel is the main channel, the process of determining the target audio signal corresponding to each channel can be expressed as: HeadLROut = γ1*LRIn, LROut = γ2*LRIn, SurLROut = γ2*LRIn, where HeadLROut represents the target audio signal corresponding to the sky channel, LROut represents the target audio signal corresponding to the left and right channels, SurLROut represents the target audio signal corresponding to the ground channel, LRIn represents the original audio signal, γ1 represents the first gain value, and γ2 represents the second gain value.

7. The audio processing method according to claim 1, wherein, The plurality of sound channels include a sky channel, a ground channel, and left and right channels. The target display device includes a central processing unit, a first speaker, a second speaker, and a third speaker, wherein the first speaker, the second speaker, and the third speaker are respectively connected to the central processing unit. The target audio signal corresponding to the sky channel is output through the first speaker, the target audio signal corresponding to the ground channel is output through the second speaker, and the target audio signals corresponding to the left and right channels are output through the third speaker.

8. The audio processing method according to claim 7, wherein, The first speaker and the second speaker are external speakers, or the first speaker and the second speaker are internal speakers; When the first speaker and the second speaker are built-in speakers, after determining the target audio signal corresponding to each of the plurality of channels based on the original audio signal corresponding to the main channel and the currently playing screen, the process includes: The target audio signals corresponding to the sky channel and the target audio signals corresponding to the ground channel are respectively filtered; The target audio signal corresponding to the sky channel is output through the first speaker after filtering. The target audio signal corresponding to the ground channel is output through the second speaker after filtering. The target audio signals corresponding to the left and right channels are output through the third speaker.

9. The audio processing method according to claim 8, wherein, When the first speaker and the second speaker are external speakers, after determining the target audio signal corresponding to each of the plurality of channels based on the original audio signal corresponding to the main channel and the currently playing screen, the process includes: The target audio signal corresponding to the sky channel is output through the first speaker; The target audio signal corresponding to the ground channel is output through the second speaker; The target audio signals corresponding to the left and right channels are output through the third speaker.

10. The audio processing method according to claim 1, wherein, The step of obtaining the camera viewpoint corresponding to the current playback screen of the target display device includes: Obtain the first image and the second image corresponding to the current playback screen of the target display device; Based on the first image and the second image, determine the target camera extrinsic parameters of the target camera corresponding to the currently playing screen; Based on the target camera's extrinsic parameters, the camera's viewing angle corresponding to the current playback screen of the target display device is determined.

11. The audio processing method according to claim 10, wherein, The step of determining the target camera extrinsic parameters of the target camera corresponding to the currently playing screen based on the first image and the second image includes: Pixel matching is performed on the first image and the second image to obtain multiple second target pixels corresponding to multiple first target pixels in the first image in the second image; Based on the initial camera intrinsic parameters of the target camera corresponding to the current playback screen, the normalized spatial coordinates of each first target pixel among multiple first target pixels are determined. Based on the initial camera extrinsic parameters corresponding to the target camera, the normalized spatial point coordinates are reprojected onto the second image to obtain the reprojected pixel coordinates corresponding to each second target pixel. Based on the original pixel coordinates corresponding to the second target pixel and the reprojected pixel coordinates, the reprojection error corresponding to each second target pixel is determined. Based on the reprojection error, the initial camera intrinsic parameters and the initial camera extrinsic parameters are updated, and the step of determining the normalized spatial coordinates of each of the multiple first target pixels based on the initial camera intrinsic parameters of the target camera corresponding to the current playback screen continues to be executed until the reprojection error satisfies the first condition. The updated initial camera extrinsic parameters are determined as the target camera extrinsic parameters of the target camera corresponding to the currently playing screen.

12. The audio processing method according to claim 11, wherein, The process of determining the normalized spatial coordinates of each of the first target pixels is as follows: X i =Z i *(u i -c x ) / f x Y i =Z i *(v i -c y ) / f y Z = depth / s, (X i Y i Z i (u) represents the normalized spatial coordinates of the i-th first target pixel. i v i 0 represents the original pixel coordinates corresponding to the i-th first target pixel. This represents the initial camera intrinsic parameters, depth represents the depth value of the target camera, and s represents the conversion parameter between the depth value and the actual distance.

13. The audio processing method according to claim 11, wherein, The process of determining the coordinates of the reprojected pixel is as follows: λ ij Represents the scaling factor, (X) i Y i Z i P(R) represents the normalized spatial coordinates of the i-th first target pixel. j , t j R represents the projection matrix of the target camera j. j and t j R represents the initial extrinsic parameters of the target camera j. j R represents the rotation matrix of the target camera j. j , t j Let u represent the translation vector of the target camera j. ij v ij ) represents the coordinates of the reprojected second target pixel corresponding to the i-th first target pixel.

14. The audio processing method according to claim 11, wherein, The process for determining the reprojection error is as follows: Among them, (u ij v ij () represents the coordinates of the reprojected second target pixel corresponding to the i-th first target pixel. This represents the original pixel coordinates of the second target pixel corresponding to the i-th first target pixel.

15. The audio processing method according to claim 11, wherein, The reprojection error satisfies the first condition that the reprojection error corresponding to all second target pixels is less than a first threshold, or the reprojection error satisfies the first condition that the sum of the squares of the reprojection errors corresponding to all second target pixels is less than a second threshold, or the reprojection error satisfies the first condition that the sum of the squares of the reprojection errors corresponding to all second target pixels obtained in the previous two iterations is less than a third threshold.

16. The audio processing method according to claim 1, wherein, The step of obtaining the camera viewpoint corresponding to the current playback screen of the target display device includes: Obtain the current playback screen of the target display device; Perform target recognition on the currently playing screen to obtain the target recognition result of the currently playing screen; Based on the target recognition result, the camera viewpoint corresponding to the currently playing screen is determined.

17. The audio processing method according to claim 16, wherein, The step of performing target recognition on the currently playing screen to obtain the target recognition result of the currently playing screen includes: The currently playing screen is input into the target recognition model, and the target recognition model outputs the target recognition result of the currently playing screen.

18. The audio processing method according to claim 16, wherein, Determining the camera viewpoint corresponding to the currently playing screen based on the target recognition result includes: Based on the target recognition result and the first association information, the camera viewpoint corresponding to the current playback screen is determined; wherein, the first association information is used to characterize the correspondence between the recognition result and the camera viewpoint.

19. An audio processing apparatus, wherein, Applied to a target display device, the target display device including multiple audio channels, the audio processing device includes: The information acquisition module is used to acquire the camera viewpoint corresponding to the current playback screen of the target display device; The first determining module is used to determine the main audio channel corresponding to the currently playing screen based on the camera's viewpoint; The second determining module is used to determine the target audio signal corresponding to each of the multiple channels based on the original audio signal corresponding to the main channel and the current playback screen.

20. A display device, wherein, The display device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the audio processing method of any one of claims 1 to 18.