Audio playback method and related device
By dynamically adjusting sound output based on the visual position of sound-emitting objects within videos, the method addresses the issue of sound-video separation in large terminal devices, enhancing the overall viewing experience through improved audio-video synchronization.
Patent Information
- Application Number
- JP2024563421
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-27
- Filing Date
- 2023-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
As terminal device display sizes increase, the separation between sound and video becomes more pronounced, especially in scenarios like conversations, airplane flights, and passing cars, leading to a disjointed user experience.
An audio playback method that dynamically changes the sound generation position of target sound-emitting objects in videos based on their relative distance to multiple speakers, ensuring that the sound aligns with the visual position of the object on the display.
This method achieves an "audio-video fusion" effect, significantly enhancing the user's video viewing experience by integrating sound with visual elements more seamlessly.
Smart Images

Figure 2025516202000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technologies, and more particularly, to an audio playback method and related devices.
Background Art
[0002] With the development of terminal technologies, in order to meet the various requirements of users, the sizes and functions of terminal devices are becoming increasingly diverse. For example, when the display area of a large display is large, users can play videos on the large display, which brings a better video viewing experience to users.
[0003] However, as the display size of terminal devices increases, the moving area of objects in the video displayed on the display becomes wider. As a result, when watching a movie or a TV program, users often feel that the sound and the video are separated and not well integrated. Specifically, in some scenarios such as conversations with people, flights of airplanes, and passing of cars, the separation of sound and images is more obvious. Therefore, how to integrate the sound in the video and the sound generating object in the video is an urgent problem to be solved.
Summary of the Invention
Means for Solving the Problems
[0004] This application provides an audio playback method and related devices such that the sound generation position of a target sound generating object in a video changes along with the position of the target sound generating object in the video. This realizes the "audio-video fusion" effect and improves the user's video viewing experience.
[0005] According to a first aspect, the present application provides an audio playback method applied to an electronic device including a plurality of speakers, the plurality of speakers including a first speaker and a second speaker, the method including the following. The electronic device starts playing a first video clip, the video of the first video clip including a first sound generation target, the electronic device obtains, from the audio data of the first video clip, a first audio emitted by the first sound generation target, and when the electronic device determines that, at a first moment, the position of the first sound generation target in the video of the first video clip is a first position, the electronic device outputs the first audio via the first speaker, the electronic device obtains, from the audio data of the first video clip, a second audio emitted by the first sound generation target, and when the electronic device determines that, at a second moment, the position of the first sound generation target in the video of the first video clip is a second position, the electronic device outputs the second audio via the second speaker, the first moment being different from the second moment, the first position being different from the second position, and the first speaker being different from the second speaker.
[0006] The electronic device obtains a first audio emitted by the first sound generation target, and the electronic device obtains a second audio emitted by the first sound generation target. The first audio and the second audio may be extracted and obtained in real time by the electronic device, or the electronic device may obtain in advance the complete audio emitted by the first sound generation target, and the first audio and the second audio are audio clips in the complete audio emitted by the first sound generation target at different moments.
[0007] According to the method provided by the first aspect, the electronic device can change the sound generation position of the target sound-emitting organism based on the relative distance between the target sound-emitting organism in the image frame and each loudspeaker of the electronic device, whereby the sound generation position of the target sound-emitting organism in the video changes together with the display position of the target sound-emitting organism on the display. This realizes "audio-video fusion" and improves the user's video viewing experience.
[0008] Regarding the first aspect, in a possible implementation, when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio through the first speaker. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker, the electronic device outputs the first audio through the first speaker. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the second position, the electronic device outputs the second audio through the second speaker. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker, the electronic device outputs the second audio through the second speaker. In this way, after determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device further determines the speaker closest to the first position and outputs the audio emitted by the first sound generation target at the corresponding moment through the closest speaker, whereby the sound generation position of the target sound-emitting organism changes together with the display position of the target sound-emitting organism on the display.
[0009] Regarding the first aspect, in a possible implementation, when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio through the first speaker. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker, the electronic device outputs the first audio through the first speaker at the first volume value and outputs the first audio through the second speaker at the second volume value, where the first volume value is greater than the second volume value. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the second position, the electronic device outputs the second audio through the second speaker. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker, the electronic device outputs the first audio through the second speaker at the third volume value and outputs the first audio through the first speaker at the fourth volume value, where the third volume value is greater than the fourth volume value. In this way, after determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device further determines the distance between each speaker and the first position. When the distances are different, the output volumes of the speakers are different, that is, each speaker may play sound simultaneously. The shorter the distance, the greater the output volume of the speaker. The longer the distance, the smaller the output volume of the speaker.
[0010] Regarding the first aspect, in one possible implementation, the video of the first video clip includes a second sound generation target, and the method further includes the following. When the electronic device extracts the third audio emitted by the second sound generation target from the audio data of the first video clip, and determines at the third moment that the position of the second sound generation target in the video of the first video clip is the third position, the electronic device outputs the third audio via the first speaker. When the electronic device extracts the fourth audio emitted by the second sound generation target from the audio data of the first video clip, and determines at the fourth moment that the position of the second sound generation target in the video of the first video clip is the fourth position, the electronic device outputs the fourth audio via the second speaker. The third moment is different from the fourth moment, and the third position is different from the fourth position. In this way, the electronic device can simultaneously detect the positions of multiple sound generation targets and speakers, and change the sound generation positions of multiple sound generation targets.
[0011] Regarding the first aspect, in one possible implementation, the multiple speakers further include a third speaker. After the electronic device outputs the first audio via the first speaker, the method further includes the following. After the first time has elapsed, or after the number of image frames exceeds the first number, if the electronic device does not detect the position of the first sound generation target in the video of the first video clip, the electronic device outputs the audio of the first sound generation target via the third speaker. In this way, when the electronic device only detects the audio of the first sound generation target and does not detect the image position of the first sound generation target in the image data, the electronic device outputs the audio of the first sound generation target via a predetermined speaker.
[0012] Regarding the first aspect, in a possible implementation, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker, the electronic device outputs the first audio via the first speaker. Specifically, the electronic device obtains the position information of the first speaker and the position information of the second speaker, and based on the first position of the first sound generation target in the video of the first video clip, the position information of the first speaker, and the position information of the second speaker, the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker. In this way, the position of each speaker of the electronic device is unique, and the electronic device can determine the distance between the first sound generation target and each speaker based on the position of each speaker and the position of the first sound generation target in the video.
[0013] Regarding the first aspect, in a possible implementation, when the electronic device obtains the first audio emitted by the first sound generation target from the audio data of the first video clip, specifically, the electronic device obtains multiple types of audio from the audio data of the first video clip based on multiple types of predetermined audio characteristics, and the electronic device determines the first audio emitted by the first sound generation target from the multiple types of audio. In this way, the electronic device can calculate the similarity between multiple types of audio in the audio data and multiple types of predetermined audio characteristics based on multiple types of predetermined audio characteristics to determine the first audio emitted by the first sound generation target.
[0014] Regarding the first aspect, in a possible implementation, for the electronic device to determine that the position of the first sound generation target in the video of the first video clip is the first position, specifically, the electronic device identifies the first target image corresponding to the first sound generation target from the video of the first video clip based on multiple types of predetermined image features, and the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position based on the display area of the first target image in the video of the first video clip. In this way, the electronic device can determine the image features of the first target image corresponding to the first sound generation target based on multiple types of predetermined image features in order to determine the position of the first target image in the video of the first video clip.
[0015] Regarding the fourth aspect, in a possible implementation, the plurality of speakers further includes a fourth speaker, and before the electronic device outputs the first audio, the method further includes the following. The electronic device obtains predetermined audio channel information from the audio data of the first video clip, and the predetermined audio channel information includes outputting the first audio and the first background sound from the fourth speaker. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, for the electronic device to output the first audio through the first speaker, specifically, when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio through the first speaker and outputs the first background sound through the fourth speaker. In this way, when determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device can provide the first audio to the first speaker, output the first audio through the first speaker, and output other audio such as background sound and music sound through the predetermined speaker.
[0016] Regarding the first aspect, in one possible implementation, the position information of a plurality of speakers of an electronic device is different. In this way, the sound generation position of the target sound generation object changes together with the display position of the target sound generation object on the display.
[0017] Regarding the first aspect, in one possible implementation, the type of the first sound generation target is any one of a person, an animal, an object, and a landscape.
[0018] Regarding the first aspect, in one possible implementation, the type of the first audio is any one of a human voice, an animal sound, an environmental sound, a music sound, and an object sound.
[0019] According to the second aspect, an embodiment of the present application provides an audio playback method, and the method includes the following. An electronic device starts playing a first video clip, the video of the first video clip includes a first sound generation target, the electronic device extracts a first audio emitted by the first sound generation target from the audio data of the first video clip, and when the electronic device determines that the position of the first sound generation target in the video of the first video clip is a first position at a first moment, the electronic device outputs the first audio via a first audio output device. The electronic device extracts a second audio emitted by the first sound generation target from the audio data of the first video clip, and when the electronic device determines that the position of the first sound generation target in the video of the first video clip is a second position at a second moment, the electronic device outputs the second audio via a second audio output device. The first moment is different from the second moment, the first position is different from the second position, and the first audio output device is different from the second audio output device.
[0020] The electronic device acquires first audio emitted by a first sound - generating target, and the electronic device acquires second audio emitted by the first sound - generating target. The first audio and the second audio may be extracted and acquired in real - time by the electronic device, or the electronic device may acquire the complete audio emitted by the first sound - generating target in advance, and the first audio and the second audio are audio clips in the complete audio emitted by the first sound - generating target at different instants.
[0021] According to the method provided in the second aspect, when the electronic device is externally connected to an audio output device, the electronic device can change the sound - generating position of the target sound - generating object based on the relative distance between the target sound - generating object in the image frame and the audio output device. Thereby, the sound - generating position of the target sound - generating object in the video changes together with the display position of the target sound - generating object on the display. This realizes "audio - video fusion" and improves the user's video viewing experience.
[0022] Regarding the second aspect, in a possible implementation form, the type of the first audio output device is any one of a sound box, earphones, a power amplifier, a multimedia console, and an audio adapter.
[0023] Regarding the second aspect, in a possible implementation, when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio via the first audio output device. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first audio output device is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second audio output device, the electronic device outputs the first audio via the first audio output device. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the second position, the electronic device outputs the second audio via the second audio output device. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the second audio output device is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the first audio output device, the electronic device outputs the second audio via the second audio output device. In this way, after determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device further determines the audio output device closest to the first position and outputs, via the closest audio output device, the audio emitted by the first sound generation target at the corresponding moment. Thereby, the sound generation position of the target sound generation entity changes together with the display position of the target sound generation entity on the display.
[0024] Regarding the second aspect, in one possible implementation, when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio via the first audio output device. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first audio output device is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second audio output device, the electronic device outputs the first audio via the first audio output device at the first volume value and outputs the first audio via the second audio output device at the second volume value, where the first volume value is greater than the second volume value. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the second position, the electronic device outputs the second audio via the second audio output device. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the second audio output device is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the first audio output device, the electronic device outputs the first audio via the second audio output device at the third volume value and outputs the first audio via the first audio output device at the fourth volume value, where the third volume value is greater than the fourth volume value. In this way, after determining that the position of the first sound generation object in the video of the first video clip is the first position, the electronic device further determines the distance between each audio output device and the first position. When the distances are different, the output volumes of the audio output devices are different, that is, the audio output devices may play sounds simultaneously. The shorter the distance, the greater the output volume of the audio output device. The longer the distance, the smaller the output volume of the audio output device.
[0025] Regarding the second aspect, in one possible implementation, the video of the first video clip includes a second sound generation target, and the method further includes the following. When the electronic device extracts the third audio emitted by the second sound generation target from the audio data of the first video clip, and determines at the third moment that the position of the second sound generation target in the video of the first video clip is the third position, the electronic device outputs the third audio via the first audio output device. When the electronic device extracts the fourth audio emitted by the second sound generation target from the audio data of the first video clip, and determines at the fourth moment that the position of the second sound generation target in the video of the first video clip is the fourth position, the electronic device outputs the fourth audio via the second audio output device. The third moment is different from the fourth moment, and the third position is different from the fourth position. In this way, the electronic device can simultaneously detect the positions of multiple sound generation targets and audio output devices, and change the sound generation positions of multiple sound generation targets.
[0026] Regarding the second aspect, in one possible implementation, the multiple audio output devices further include a third audio output device. After the electronic device outputs the first audio via the first audio output device, the method further includes the following. After the first time has elapsed, or after the number of image frames exceeds the first number, if the electronic device does not detect the position of the first sound generation target in the video of the first video clip, the electronic device outputs the audio of the first sound generation target via the third audio output device. In this way, when the electronic device only detects the audio of the first sound generation target and does not detect the image position of the first sound generation target in the image data, the electronic device outputs the audio of the first sound generation target via a pre-determined audio output device.
[0027] Regarding the second aspect, in one possible implementation, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first audio output device is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second audio output device, the electronic device outputs the first audio via the first audio output device. Specifically, the electronic device obtains the position information of the first audio output device and the position information of the second audio output device, and based on the first position of the first sound generation target in the video of the first video clip, the position information of the first audio output device, and the position information of the second audio output device, the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first audio output device is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second audio output device. In this way, the position of each audio output device of the electronic device is unique, and the electronic device can determine the distance between the first sound generation target and each audio output device based on the position of each audio output device and the position of the first sound generation target in the video.
[0028] Regarding the second aspect, in a possible implementation, for the electronic device to obtain the first audio emitted by the first sound generation target from the audio data of the first video clip, specifically, the electronic device obtains multiple types of audio from the audio data of the first video clip based on multiple types of predetermined audio features, and the electronic device determines the first audio emitted by the first sound generation target from the multiple types of audio. In this way, in order to determine the first audio emitted by the first sound generation target, the electronic device can calculate the similarity between the multiple types of audio in the audio data and the multiple types of predetermined audio features based on the multiple types of predetermined audio features.
[0029] Regarding the second aspect, in a possible implementation, for the electronic device to determine that the position of the first sound generation target in the video of the first video clip is the first position, specifically, the electronic device identifies the first target image corresponding to the first sound generation target from the video of the first video clip based on multiple types of predetermined image features, and the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position based on the display area of the first target image in the video of the first video clip. In this way, in order to determine the position of the first target image in the video of the first video clip, the electronic device can determine the image features of the first target image corresponding to the first sound generation target based on multiple types of predetermined image features.
[0030] Regarding the second aspect, in one possible implementation, the plurality of audio output devices further includes a fourth audio output device, and before the electronic device outputs the first audio, the method further includes the following. When the electronic device obtains predetermined audio channel information from the audio data of the first video clip, and the predetermined audio channel information includes outputting the first audio and the first background sound from the fourth audio output device, and when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio via the first audio output device, specifically, when determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio via the first audio output device and outputs the first background sound via the fourth audio output device. In this way, when determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device provides the first audio to the first audio output device, outputs the first audio via the first audio output device, and can output other audio such as background sound and music sound via a predetermined audio output device.
[0031] Regarding the second aspect, in one possible implementation, the position information of the plurality of audio output devices of the electronic device is different. In this way, the sound generation position of the target sound generation object changes together with the display position of the target sound generation object on the display.
[0032] According to a third aspect, the present application provides an electronic device, the electronic device including one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, the one or more memories being configured to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions, whereby the electronic device executes an audio playback method provided in any possible implementation form of any one of the foregoing aspects.
[0033] According to a fourth aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium storing instructions. When the instructions are executed on an electronic device, the electronic device is enabled to execute an audio playback method provided in any possible implementation form of any one of the foregoing aspects.
[0034] According to a fifth aspect, the present application provides a computer program product. When the computer program product is executed by an electronic device, the electronic device is enabled to execute an audio playback method provided in any possible implementation form of any one of the foregoing aspects.
[0035] According to a sixth aspect, the present application provides a chip or chip system including a processing circuit and an interface circuit. The interface circuit is configured to receive code instructions and transmit the code instructions to the processing circuit, the processing circuit being configured to execute the code instructions to execute an audio playback method provided in any possible implementation form of any one of the foregoing aspects.
Brief Description of the Drawings
[0036]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 8C
Figure 9A
Figure 9B
Figure 9C
Figure 10A
Figure 10B
Figure 10C
Figure 11
Figure 12
DETAILED DESCRIPTION OF THE INVENTION
[0037] The technical solutions according to the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. In the description of the embodiments of the present application, unless otherwise stated, " / " represents "or". For example, A / B may represent A or B. In this specification, "and / or" only describes the association relationship between related objects and represents that three relationships may exist. For example, A and / or B may represent three cases: only A exists, both A and B exist, and only B exists. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more.
[0038] The terms "first" and "second" mentioned below are only intended for the purpose of explanation and should not be understood as an indication of relative importance or a suggestion, nor an implicit indication of the quantity of the technical features shown. Therefore, the features defined by "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, unless otherwise stated, "a plurality of" means two or more.
[0039] The term "user interface" (UI) in the following embodiments of the present application is a media interface for interaction and information exchange between an application or an operating system and a user, and implements conversion between the internal form of information and a form acceptable to the user. The user interface is source code written in a specific computer language such as Java or an extensible markup language (XML). The interface source code is analyzed and rendered on an electronic device and finally presented as content recognizable by the user. The user interface is in the general form of a graphical user interface (GUI), is related to the operation of a computer and is graphically displayed, and can be visual interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, or widgets displayed on the display of an electronic device.
[0040] The fusion of the sound in the video and the sound generating object in the video can be realized in the following manner.
[0041] Method 1: The fusion of the sound in the video and the sound generating object in the video is realized based on the organic light emitting diode (OLED) sound-on-display technology.
[0042] OLED sound-on-display technology is to fix a plurality of vibrating loudspeakers behind an OLED display and drive the display to emit sound through the vibrating loudspeakers, so that to a certain extent, viewers can feel that the sound is coming from the display. However, the OLED display is large, and the objects in the video may move freely within the display area of the OLED display. However, the sound corresponding to the video can only be produced through specific loudspeakers behind the OLED display, that is, the sound generation position is fixed. As a result, the sound generation position corresponding to the video cannot change along with the movement of the objects in the video, that is, the sound of the moving objects cannot be tracked.
[0043] Based on this, an embodiment of the present application provides an audio playback method. The method includes the following steps.
[0044] Step 1: An electronic device acquires video data, and the video data includes audio data and image data.
[0045] The video data may be real-time data acquired by the electronic device, or may be cached data acquired by the electronic device.
[0046] The audio data includes an audio file and predetermined audio channel information of the audio file, and the predetermined audio channel information restricts the audio channel through which the electronic device outputs the audio file. The predetermined audio channel information may be, for example, a left audio channel, a right audio channel, or a center audio channel. The left audio channel may be classified into a left front channel, a left rear channel, etc. The right audio channel may be classified into a right front channel, a right rear channel, etc.
[0047] When the predetermined channel information in the audio data is the left front channel, the right front channel, and the center voice channel, the electronic device plays the audio file through the left front loudspeaker, the right front loudspeaker, and the center loudspeaker of the electronic device.
[0048] Step 2: The electronic device extracts the predetermined type of audio from the audio file. The predetermined type of audio includes human voices, animal sounds, environmental sounds, music sounds, object sounds, etc.
[0049] In other words, to form a complete audio file, a plurality of different types of audio, such as human voices, animal sounds, environmental sounds, music sounds, object sounds, are integrated in the audio file in the video data. The electronic device may separate the plurality of types of audio integrated in the audio file to obtain the plurality of different types of audio.
[0050] Step 3: The electronic device identifies the target sound-emitting organism in the image data and determines the position coordinates of the target sound-emitting organism on the display of the electronic device.
[0051] Next, the electronic device determines the distance between the target sound-emitting organism and each loudspeaker based on the position coordinates of the target sound-emitting organism on the display and the positions of the plurality of loudspeakers of the electronic device, and finally determines the loudspeaker closest to the target sound-emitting organism.
[0052] Step 4: The electronic device outputs the audio corresponding to the target sound-emitting organism through the loudspeaker closest to the target sound-emitting organism.
[0053] The electronic device outputs the audio corresponding to the target sound-emitting organism through the first loudspeaker closest to the target sound-emitting organism, and outputs an audio of a type not determined in advance through the loudspeaker corresponding to the original voice channel information.
[0054] Alternatively, the electronic device outputs the audio corresponding to the target sound-emitting organism through the first audio output device corresponding to the first loudspeaker closest to the target sound-emitting organism, and outputs an audio of a type not determined in advance through the audio output device corresponding to the loudspeaker corresponding to the original voice channel information. The audio output device can be a sound amplification device such as a sound box. In this way, the electronic device can output the audio through the audio output device corresponding to the loudspeaker to improve the sound quality and volume of the output audio.
[0055] In some embodiments, when there are a plurality of loudspeakers closest to the target sound-emitting organism, the electronic device can simultaneously output the audio corresponding to the target sound-emitting organism through the plurality of loudspeakers.
[0056] In some embodiments, the position of the target sound-emitting organism on the display changes in real time. After the position of the target sound-emitting organism changes, the distance between the target sound-emitting organism and each loudspeaker of the electronic device also changes. After the electronic device determines the second loudspeaker closest to the target sound-emitting organism, the electronic device outputs the audio corresponding to the target sound-emitting organism through the second loudspeaker closest to the target sound-emitting organism. The position of the first loudspeaker is different from the position of the second loudspeaker.
[0057] Alternatively, the electronic device outputs the audio corresponding to the target sound-emitting organism through the second audio output device corresponding to the second loudspeaker closest to the target sound-emitting organism.
[0058] Obtaining the audio emitted by the target sound-emitting organism by the electronic device can be that the electronic device extracts the audio output instantaneously in real time, or that the electronic device can obtain in advance the complete audio emitted by the target sound-emitting organism.
[0059] Optionally, in addition to emitting sound through the nearest loudspeaker, the electronic device may further emit sound simultaneously through a plurality of loudspeakers, but if the distances are different, the output volume of each speaker is different. The shorter the distance, the greater the output volume of the speaker. The longer the distance, the smaller the output volume of the speaker.
[0060] For example, if the electronic device determines that the distance between the target sound-emitting organism and the first speaker is shorter than the distance between the target sound-emitting organism and the second speaker, the electronic device can simultaneously output the first audio emitted by the target sound-emitting organism through the first speaker and the second speaker, but the volume of the first audio emitted by the first speaker is greater than the volume of the first audio emitted by the second speaker.
[0061] According to the audio playback method provided in the embodiments of the present application, the electronic device determines the loudspeaker closest to the target sound-emitting organism at that relative distance based on the relative distance between the target sound-emitting organism in the image frame and each loudspeaker of the electronic device, and outputs the audio corresponding to the target sound-emitting organism through the closest loudspeaker or the audio output device corresponding to the closest loudspeaker, whereby the sound generation position of the target sound-emitting organism in the video changes together with the display position of the target sound-emitting organism on the display. This realizes "audio-video fusion" and improves the user's video viewing experience.
[0062] FIG. 1 is a diagram of the structure of the electronic device 100.
[0063] The electronic device 100 is a device configured using loudspeakers with different arrangement directions, and a device configured using loudspeakers having at least two different arrangement directions. The arrangement direction of the loudspeaker is relative to the electronic device 100. For example, the center of the display of the electronic device 100 is used as the center point, and the loudspeakers of the electronic device 100 can be classified into a left front loudspeaker, a right front loudspeaker, a left rear loudspeaker, a right rear loudspeaker, a center loudspeaker (i.e., a loudspeaker located at the center point of the electronic device), etc. In another embodiment, the electronic device 100 may further include other loudspeakers having additional arrangement directions. This is not limited in the embodiments of the present application.
[0064] The type of the electronic device 100 includes, but is not limited to, a large display, a projector, a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, an in-vehicle device, a smart home device, a smart city device, etc. The specific type of the electronic device 100 is not limited in the embodiments of the present application. In the following embodiments of the present application, an example where the electronic device 100 is a large display is used for illustration.
[0065] The electronic device 100 may include a processor 110, a wireless communication module 120, an audio module 130, an internal memory 140, buttons 190, a motor 191, an indicator 192, a camera 193, a display 194, a sensor module 150, etc. The audio module 130 may include a speaker 130A, a receiver 130B, a microphone 130C, and a headset interface 130D. The sensor module 150 may include an acceleration sensor 150A, a distance sensor 150B, a light proximity sensor 150C, a temperature sensor 150D, a touch sensor 150E, an ambient light sensor 150F, etc.
[0066] In some embodiments, the electronic device 100 may not include the microphone 130C and the headset interface 130D.
[0067] In some embodiments, the electronic device 100 may not include any one or more of the sensor module 150.
[0068] It can be understood that the structure shown in this embodiment of the present invention is not a specific limitation on the electronic device 100. In some other embodiments of this application, the electronic device 100 may include more or fewer components than those shown in the figure, or some components may be combined, or some components may be divided, or different component arrangements may be made. The components shown in the figure may be implemented by hardware, software, or a combination of software and hardware.
[0069] Processor 110 may include one or more processing units. For example, processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, a neural-network processing unit (NPU), and the like. Different processing units may be independent components or may be integrated into one or more processors. Processor 110 may be configured to extract a predetermined type of audio from an audio file, and the predetermined type of audio may include human voice, animal sound, environmental sound, music sound, and object sound. Processor 110 may further identify a target sound-emitting object in the image data, determine the position coordinates of the target sound-emitting object on the display of the electronic device, and output the audio corresponding to the target sound-emitting object based on the position coordinates of the target sound-emitting object on the display of the electronic device via the loudspeaker closest to the target sound-emitting object. For details, please refer to the detailed description in the subsequent embodiments. In this embodiment of the present application, the details are not described here.
[0070] The controller may generate an arithmetic control signal based on the instruction operation code and the time series signal to completely control the fetching of instructions and the execution of instructions.
[0071] The memory may be disposed within the processor 110 and is configured to store instructions and data. In some embodiments, the memory within the processor 110 is cache memory. The memory may store instructions or data that are being used or are periodically used by the processor 110. When the processor 110 needs to reuse an instruction or data, the processor may directly call the instruction or data from the memory. This avoids repetitive access, reduces the latency of the processor 110, and improves the efficiency of the system.
[0072] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, a universal serial bus (USB) interface, etc.
[0073] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple groups of the I2C bus.
[0074] The I2S interface can be configured to perform audio communication. In some embodiments, the processor 110 may include a plurality of groups of I2S buses. The processor 110 may be coupled to the audio module 130 through an I2S bus to perform communication between the processor 110 and the audio module 130.
[0075] The PCM interface can also be used to perform audio communication, sample, quantize, and encode analog signals. In some embodiments, the audio module 130 may be coupled to the wireless communication module 120 through a PCM bus interface.
[0076] The UART interface is a general-purpose serial data bus and is configured to perform asynchronous communication. This bus can be a bidirectional communication bus. The UART interface converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically configured to connect the processor 110 to the wireless communication module 120.
[0077] The MIPI interface can be configured to connect the processor 110 to peripheral components such as the display 194 or the camera 193. The MIPI interface includes, for example, a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 communicates with the camera 193 via the CSI to implement the camera function of the electronic device 100. The processor 110 communicates with the display 194 via the DSI interface to implement the display function of the electronic device 100.
[0078] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be configured to connect the processor 110 to the camera 193, the display 194, the wireless communication module 120, the audio module 130, the sensor module 150, etc. The GPIO interface can alternatively be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0079] The USB interface is an interface that complies with the specifications of the USB standard, and specifically can be a mini USB interface, a micro USB interface, a USB type-C interface, etc. The USB interface may be configured to perform data transmission between the electronic device 100 and a peripheral device, or may be configured to connect to a headset for playing audio through the headset. The interface can be configured to connect to another electronic device such as an AR device.
[0080] It can be understood that the interface connection relationship between the modules shown in this embodiment of the present invention is only an example for explanation and does not constitute a constraint on the structure of the electronic device 100. In some other embodiments of the present application, the electronic device 100 can alternatively use an interface connection method different from that in the foregoing embodiments, or use a combination of multiple interface connection methods.
[0081] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the wireless communication module 120, the modem processor, the baseband processor, etc.
[0082] Antenna 1 is configured to transmit and receive electromagnetic wave signals. Each antenna of the electronic device 100 can be configured to cover one or more communication frequency bands. Different antennas can be multiplexed to improve the utilization rate of the antennas. For example, Antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0083] The wireless communication module 120 is applied to the electronic device 100 and can provide a wireless communication solution including a wireless local area network (WLAN) (for example, a wireless fidelity (Wi-Fi) network), Bluetooth (registered trademark) (BT), a global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC) technology, infrared (IR) technology, etc. The wireless communication module 120 can be one or more components integrating at least one communication processor module. The wireless communication module 120 receives electromagnetic waves through Antenna 1, performs demodulation and filtering processing on the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 120 can further receive the signals to be transmitted from the processor 110, perform frequency modulation and amplification on the signals, and convert the signals into electromagnetic waves for radiation through Antenna 1.
[0084] In some embodiments, in the electronic device 100, the antenna 1 is coupled to the wireless communication module 120, whereby the electronic device 100 can communicate with a network and another device by using wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA (registered trademark)), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, IR technology, etc. GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), BeiDou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation system (SBAS).
[0085] The electronic device 100 can implement a display function through a GPU, a display 194, an application processor, etc. The GPU is a microprocessor for image processing and is connected to the display 194 and the application processor. The GPU is configured to execute mathematical and geometric calculations and render images. The processor 110 includes one or more GPUs and can execute program instructions for generating or changing display information.
[0086] The display 194 is configured to display images, videos, etc. The display 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini-LED, a micro-LED, a micro-OLED, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0087] The electronic device 100 can implement a camera function through an ISP, a camera 193, a video codec, a GPU, a display 194, an application processor, etc.
[0088] The ISP is configured to process data fed back by camera 193. For example, during shooting, the shutter is pressed and light is transmitted through the lens to the photosensitive element of the camera. The optical signal is converted into an electrical signal, and in order to convert the electrical signal into a visible image, the photosensitive element of the camera transmits the electrical signal to the ISP for processing. The ISP can further perform algorithm optimization regarding noise and image brightness. The ISP can further optimize parameters such as exposure and color temperature of the shooting scenario. In some embodiments, the ISP can be disposed in camera 193.
[0089] Camera 193 is configured to capture a still image or video. The optical image of an object is generated through the lens and projected onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal and then transmits the electrical signal to the ISP to convert the electrical signal into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1. In some embodiments, electronic device 100 may not include camera 193.
[0090] The digital signal processor is configured to process digital signals and can process another digital signal in addition to the digital image signal. For example, when electronic device 100 selects a frequency, the digital signal processor is configured to perform a Fourier transform on the frequency energy.
[0091] The video codec is configured to compress or decompress digital video. The electronic device 100 may support one or more video codecs. In this way, the electronic device 100 may play or record video in multiple coding formats, such as Moving Picture Experts Group (MPEG)-1, MPEG-2, MPEG-3, and MPEG-4.
[0092] The NPU is a neural-network (NN) computing processor that simulates the biological neural network structure like the transmission mode between neurons in the human brain to perform high-speed processing on input information and can perform continuous self-learning. Applications such as intelligent cognition of the electronic device 100, such as image recognition, face recognition, speech recognition, and text understanding, can be realized through the NPU.
[0093] The internal memory 140 may include one or more random access memories (RAMs) and one or more non-volatile memories (NVMs).
[0094] Random access memory may include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), such as the fifth generation DDR SDRAM generally called DDR5 SDRAM. Non-volatile memory may include magnetic disk storage devices and flash memory.
[0095] Flash memory may be classified into NOR flash, NAND flash, 3D NAND flash, etc. according to the operating principle, or may be classified into single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc. based on the amount of cell potential level, or may be classified into universal flash storage (UFS), embedded multimedia card (eMMC), etc. according to the storage specification.
[0096] Random access memory may be directly read and written by using the processor 110, may be configured to store executable programs (such as machine instructions) for the operating system or another running program, or may be configured to store user data, application data, etc.
[0097] The non-volatile memory may also store executable programs, user data, applications, etc., and may be pre-loaded into the random access memory so as to be directly read and written by the processor 110.
[0098] The electronic device 100 can implement audio functions, such as music playback and recording, through an audio module 130, a speaker 130A, a receiver 130B, a microphone 130C, a headset jack 130D, an application processor, etc.
[0099] The audio module 130 is configured to convert digital audio information into an analog audio signal output, and is also configured to convert an analog audio input into a digital audio signal. The audio module 130 may be configured to encode and decode audio signals. In some embodiments, the audio module 130 may be disposed within the processor 110, or some functional modules within the audio module 130 may be disposed within the processor 110.
[0100] The speaker 130A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into an audio signal. The electronic device 100 can listen to music via the speaker 130A. Optionally, the electronic device 100 may have multiple loudspeakers, and the positions of the multiple loudspeakers on the electronic device 100 are different. For example, the center point of the display of the electronic device 100 is used as the center point, and the loudspeakers of the electronic device 100 can be classified into a left front loudspeaker, a right front loudspeaker, a left rear loudspeaker, a right rear loudspeaker, a center loudspeaker (i.e., the loudspeaker located at the center point of the electronic device), etc. In another embodiment, the electronic device 100 may further include other loudspeakers with additional arrangement directions. This is not limited in the embodiments of the present application.
[0101] In some embodiments, the electronic device 100 may be externally connected to one or more audio output devices (e.g., a sound box) in a wired or wireless manner, whereby the electronic device 100 can output audio through the externally connected one or more audio output devices, improving the sound quality of the output audio.
[0102] The receiver 130B, also referred to as an "earpiece", is configured to convert an electrical audio signal into an audio signal. When a call is answered through the electronic device 100 or when speech information is received, the receiver 130B can be placed near a person's ear to listen to the speech. In some embodiments, the electronic device 100 may not include the receiver 130B.
[0103] The microphone 130C, also referred to as a "mike" or "mic", is configured to convert an audio signal into an electrical signal. When making a call or transmitting speech information, the user can speak near the microphone 130C through the user's mouth to input the audio signal into the microphone 130C. At least one microphone 130C can be disposed within the electronic device 100. In some other embodiments, two microphones 130C can be disposed within the electronic device 100 to collect audio signals and implement a noise reduction function. In some other embodiments, alternatively, three, four, or more microphones 130C can be disposed in the electronic device 100 to collect audio signals, perform noise reduction, identify the sound source, and thereby implement functions such as a directional recording function. In some embodiments, the electronic device 100 may not include the microphone 130C.
[0104] The headset jack 130D is configured to connect to a wired headset. The headset jack 130D may be a USB interface, or may be a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface. In some embodiments, the electronic device 100 may not include the headset jack 130D.
[0105] The acceleration sensor 150A can detect the acceleration of the electronic device 100 in various orientations (usually on three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. The acceleration sensor 150A may be configured to identify the posture of the electronic device and is used for applications such as switching between landscape mode and portrait mode or a pedometer.
[0106] The distance sensor 150B is configured to measure distance. The electronic device 100 can measure distance by an infrared method or a laser method. In some embodiments, in a shooting scenario, the electronic device 100 can measure the distance through the distance sensor 150B to perform fast focusing.
[0107] The proximity light sensor 150C may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device 100 emits infrared light by using the light emitting diode. The electronic device 100 detects the infrared light reflected by a nearby object through the photodiode. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 may determine that there is no object near the electronic device 100. The electronic device 100 may detect that the user is holding the electronic device 100 near the ear for a call by using the proximity light sensor 150C to automatically turn off the display for power saving. The proximity light sensor 150C may also be used in a smart cover mode or a pocket mode to automatically perform screen unlocking or locking.
[0108] The temperature sensor 150D is configured to detect temperature. In some embodiments, the electronic device 100 executes a temperature processing policy through the temperature detected by the temperature sensor 150D. For example, when the temperature reported by the temperature sensor 150D exceeds a threshold, the electronic device 100 reduces the performance of the processor near the temperature sensor 150D to reduce power consumption for heat prevention. In some other embodiments, when the temperature is lower than another threshold, the electronic device 100 heats the battery to prevent the electronic device 100 from abnormally shutting down due to the low temperature. In some other embodiments, when the temperature is lower than yet another threshold, the electronic device 100 increases the output voltage of the battery to avoid abnormal shutdown caused by the low temperature.
[0109] The touch sensor 150E is also referred to as a "touch component". The touch sensor 150E may be disposed on the display 194, and the touch sensor 150E and the display 194 form a touch screen, which is also referred to as a "touch screen". The touch sensor 150E is configured to detect touch operations performed on or near the touch sensor. The touch sensor can communicate the detected touch operation to the application processor to determine the type of touch event. Visual output regarding the touch operation can be provided on the display 194. In some other embodiments, the touch sensor 150E may also be disposed on the surface of the electronic device 100 at a position different from that of the display 194.
[0110] Note that the electronic device 100 may include other sensors, or the device 100 may not include one or more of the above sensors.
[0111] The buttons 190 include a power button, a volume button, etc. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 can receive button inputs and generate button signal inputs regarding user settings and function controls of the electronic device 100.
[0112] The motor 191 can generate vibration prompts. The motor 191 can be configured to provide incoming call vibration prompts and touch vibration feedback. For example, touch operations performed in different applications (such as audio playback) can correspond to different vibration feedback effects. The motor 191 can also correspond to different vibration feedback effects for touch operations performed in different regions of the display 194. Different application scenarios (such as a time reminder, an alarm clock, and a game) can also correspond to different vibration feedback effects. The touch vibration feedback effect can be customized.
[0113] Indicator 192 may be indicator light and may be configured to indicate the charging state and changes in power, or may be configured to indicate messages, notifications, etc.
[0114] The following describes the layout diagram of the loudspeakers of the electronic device 100.
[0115] FIG. 2 is a diagram of the positions of a plurality of loudspeakers of the electronic device 100 on the electronic device 100.
[0116] As shown in FIG. 2, by using the center point of the display of the electronic device 100 as the origin, the horizontal right direction as the x-axis, and the upward direction perpendicular to the horizontal direction as the y-axis direction, a rectangular coordinate system is established. The space between the positive x-axis direction and the positive y-axis direction is the first quadrant, the space between the negative x-axis direction and the positive y-axis direction is the second quadrant, the space between the negative x-axis direction and the negative y-axis direction is the third quadrant, and the space between the positive x-axis direction and the negative y-axis direction is the fourth quadrant.
[0117] Only when there are loudspeakers in at least two different arrangement directions on the electronic device 100, the sound generation position of the target sound generation object in the video can change together with the position of the target sound generation object on the display. One embodiment of the present application is described by using an example in which the electronic device 100 has loudspeakers in five different arrangement directions. Note that different devices have loudspeakers with different quantities and arrangement directions. This is not limited in the embodiments of the present application.
[0118] For example, the five loudspeakers of the electronic device 100 are the loudspeaker 201, the loudspeaker 202, the loudspeaker 203, the loudspeaker 204, and the loudspeaker 205, respectively. The loudspeaker 201 is located in the first quadrant, that is, since the loudspeaker 201 is located at the upper right front of the electronic device 100, the loudspeaker 201 can also be called the upper right front loudspeaker. The loudspeaker 203 is located in the second quadrant, that is, since the loudspeaker 203 is located at the upper left front of the electronic device 100, the loudspeaker 203 can also be called the upper left front loudspeaker. The loudspeaker 204 is located in the third quadrant, that is, since the loudspeaker 204 is located at the lower left of the electronic device 100, the loudspeaker 204 can also be called the lower left loudspeaker. The loudspeaker 202 is located in the fourth quadrant, that is, since the loudspeaker 202 is located at the lower right of the electronic device 100, the loudspeaker 202 can also be called the lower right loudspeaker. The loudspeaker 205 may be located at the origin, that is, the intersection of the x-axis and the y-axis, or the loudspeaker 205 may be located at the center of the display of the electronic device 100, and thus the loudspeaker 205 can also be called the central loudspeaker.
[0119] The loudspeaker 201 may play the audio output by the upper right front audio channel, the loudspeaker 202 may play the audio output by the lower right audio channel, the loudspeaker 203 may play the audio output by the upper left front audio channel, the loudspeaker 204 may play the audio output by the lower left rear audio channel, and the loudspeaker 205 may play the audio output by the central audio channel.
[0120] Loudspeakers 201, 202, 203, 204, and 205 are each positioned in different arrangement directions. After the position of the target sound-generating body on the display of the electronic device 100 changes, the electronic device 100 can output the audio of the target sound-generating body via different loudspeakers. Therefore, the sound generation position of the target sound-generating body (i.e., the sound generation location of the loudspeaker) changes along with the position of the target sound-generating body, thereby achieving the "sound following video" effect.
[0121] It should be noted that loudspeakers 201, 202, 203, 204, and 205 may be positioned behind the display of the electronic device 100, or loudspeakers 201, 202, 203, 204, and 205 may be positioned at the edge of the display of the electronic device 100. This is not limited in the embodiments of the present application.
[0122] In addition to the above five arrangement directions, the electronic device 100 can further divide the display of the electronic device 100 into further different arrangement directions. For example, the electronic device 100 can subsequently divide the second quadrant and divide the second quadrant into a first region and a second region. The sum of the display areas of both the first region and the second region is the display area of the second quadrant. Based on this, the electronic device 100 can divide the positions of the multiple loudspeakers of the electronic device 100 into loudspeakers with further arrangement directions.
[0123] Figure 3 is another view of the positions of the multiple loudspeakers of the electronic device 100 on the electronic device 100.
[0124] As shown in FIG. 3, the center point of the display of the electronic device 100 is used as the origin, and the upward direction orthogonal to the horizontal direction is the y-axis direction. In the horizontal direction, the display of the electronic device 100 is divided into three regions, and the x1-axis and the x2-axis orthogonal to the y-axis are separately drawn. The x1-axis and the x2-axis divide the display area of the electronic device 100 into three equal regions in the horizontal direction. In FIG. 3, the y-axis, the x1-axis, and the x2-axis divide the display area of the display of the electronic device 100 into six equally divided regions.
[0125] For example, the seven loudspeakers of the electronic device 100 are the loudspeaker 206, the loudspeaker 207, the loudspeaker 208, the loudspeaker 209, the loudspeaker 210, the loudspeaker 211, and the loudspeaker 212, respectively. The loudspeaker 206 is located in the display area on the right side of the y-axis and above the x1-axis. Therefore, the loudspeaker 206 can also be called the right front loudspeaker. The loudspeaker 207 is located in the display area on the right side of the y-axis, below the x1-axis, and above the x2-axis. Therefore, the loudspeaker 206 can also be called the right center loudspeaker. The loudspeaker 208 is located in the display area on the right side of the y-axis and below the x2-axis. Therefore, the loudspeaker 208 can also be called the right rear loudspeaker. The loudspeaker 209 is located in the display area on the left side of the y-axis and above the x1-axis. Therefore, the loudspeaker 209 can also be called the left front loudspeaker. The loudspeaker 210 is located in the display area on the left side of the y-axis, below the x1-axis, and above the x2-axis. Therefore, the loudspeaker 210 can also be called the left center loudspeaker. The loudspeaker 211 is located in the display area on the left side of the y-axis and below the x2-axis. Therefore, the loudspeaker 211 can also be called the left rear loudspeaker. The loudspeaker 212 is located at the center of the display of the electronic device 100, or the loudspeaker 212 is located between the loudspeaker 210 and the loudspeaker 211. Therefore, the loudspeaker 212 can also be called the center loudspeaker.
[0126] Loudspeaker 206 may play the audio output by the right front audio channel, loudspeaker 207 may play the audio output by the right center audio channel, loudspeaker 208 may play the audio output by the right rear audio channel, loudspeaker 209 may play the audio output by the left front audio channel, loudspeaker 210 may play the audio output by the left center audio channel, loudspeaker 211 may play the audio output by the left rear audio channel, and loudspeaker 212 may play the audio output by the center audio channel.
[0127] It should be noted that loudspeakers 206, 207, 208, 209, 210, 211, and 212 may be located behind the display of the electronic device 100, or loudspeakers 206, 207, 208, 209, 210, 211, and 212 may be located at the edge of the display of the electronic device 100. This is not limited in the embodiments of the present application.
[0128] In addition to the method of classifying the arrangement directions of the loudspeakers shown in FIGS. 2 and 3, the electronic device 100 may further perform classification of the arrangement directions regarding the positions of the loudspeakers on the electronic device 100 in any other manner. This is not limited in the embodiments of the present application.
[0129] In addition to playing sound through the loudspeaker of the electronic device 100, the electronic device 100 can also be externally connected to an audio output device (e.g., a sound box). The electronic device 100 outputs audio through the loudspeaker via the corresponding sound box, thereby improving the volume and sound quality of the output audio. For example, in a home theater scenario or a movie-watching scenario in a theater, the video audio played on a large display is output via a sound box. This can obtain higher sound quality and improve the user's video-watching experience.
[0130] The following embodiments of this application are based on the specific implementation principle of "sound following video" in a home theater scenario where the electronic device 100 outputs audio via a sound box.
[0131] FIG. 4 is a diagram of an example where the electronic device 100 outputs audio via an externally connected sound box.
[0132] The electronic device 100 can be connected to the sound box in a wired or wireless manner. The electronic device 100 may be externally connected to a plurality of sound boxes, and each of the plurality of sound boxes corresponds to a different audio channel of the electronic device 100. In this way, the electronic device 100 can output the audio of the audio channel via a plurality of sound boxes.
[0133] For example, as shown in FIG. 4, the electronic device 100 may be connected to five sound boxes. The five sound boxes are the sound box 213, the sound box 214, the sound box 215, the sound box 216, and the sound box 217. The sound box 213 may be configured to output audio issued by the right front audio channel, the sound box 214 may be configured to output audio issued by the right rear audio channel, the sound box 215 may be configured to output audio issued by the center audio channel, the sound box 216 may be configured to output audio issued by the left front audio channel, and the sound box 217 may be configured to output audio issued by the left rear audio channel.
[0134] In this way, after the electronic device 100 is connected to the sound box 213, the sound box 214, the sound box 215, the sound box 216, and the sound box 217, the electronic device 100 may not output audio through the loudspeaker 201, the loudspeaker 202, the loudspeaker 203, the loudspeaker 204, and the loudspeaker 205. When the electronic device 100 outputs audio through the right front audio channel, the electronic device 100 first detects whether the sound box is connected to the right front audio channel. If the sound box is connected, the electronic device 100 outputs audio through the sound box 213 instead of the loudspeaker 201. When the electronic device 100 outputs audio through the right rear audio channel, the electronic device 100 first detects whether the sound box is connected to the right rear audio channel. If the sound box is connected, the electronic device 100 outputs audio through the sound box 214 instead of the loudspeaker 202. When the electronic device 100 outputs audio through the left front audio channel, the electronic device 100 first detects whether the sound box is connected to the left front audio channel. If the sound box is connected, the electronic device 100 outputs audio through the sound box 216 instead of the loudspeaker 203. When the electronic device 100 outputs audio through the left rear audio channel, the electronic device 100 first detects whether the sound box is connected to the left rear audio channel. If the sound box is connected, the electronic device 100 outputs audio through the sound box 217 instead of the loudspeaker 204. When the electronic device 100 outputs audio through the center audio channel, the electronic device 100 first detects whether the sound box is connected to the center audio channel. If the sound box is connected, the electronic device 100 outputs audio through the sound box 215 instead of the loudspeaker 205.
[0135] It should be noted that the electronic device 100 may be further connected to another sound box and output audio of further different audio channels via the other sound box. This is not limited in the embodiments of the present application.
[0136] The following describes the functional modules of the electronic device 100 provided in this embodiment of the present application to realize that the sound generation position of the target sound generating object in the video changes together with the position of the target sound generating object on the display.
[0137] FIG. 5 is a diagram of an example of the functional modules of the electronic device 100.
[0138] The functional modules include, but are not limited to, a sound extraction module 501, a position determination module 502, an audiovisual rendering module 503, and an audio control module 504.
[0139] Optionally, before the audio data is input to the sound extraction module 501 and before the image data is input to the position determination module 502, the video data can be preprocessed. That is, the video data is processed into video data in a predetermined format, so that the sound extraction module 501 and the position determination module 502 can execute processing on the data in the predetermined format. For example, the preprocessing includes multi-channel downmixing, data frame splicing, and short-time Fourier transform operations. The downmixing is performed for both stereo and dual-channel multi-channel inputs to ensure object-based left and right sound images. In the data frame splicing, the input time-domain series of the AI audio model is constructed based on the past buffer, the current buffer, and the future buffer.
[0140] The audio extraction module 501 is configured to obtain audio data from video data. After obtaining the audio data, the audio extraction module 501 is further configured to identify predetermined audio channel information in the audio data based on the audio data. For example, if it is predetermined that the audio data may be output by two audio channels (for example, a left audio channel and a right audio channel), the predetermined audio channel information in the audio data includes the left audio channel, the right audio channel, the audio data output by the left audio channel, and the audio output by the right audio channel.
[0141] The audio extraction module 501 is further configured to extract one or more types of audio from the audio data based on predetermined audio characteristics. The one or more types of audio include, but are not limited to, human voices, animal sounds, environmental sounds, music sounds, object sounds, etc. That is, the audio data obtained by the audio extraction module 501 is integrated with multiple different types of audio. The audio extraction module 501 separates the obtained audio to obtain multiple types of audio. How the audio extraction module 501 extracts one or more types of audio will be described in detail in subsequent embodiments. In this embodiment of the present application, the details are not described here.
[0142] The location identification module 502 is configured to obtain image data in the video data. After the location identification module 502 obtains the image data in the video data, the location identification module 502 is further configured to identify the first location information of the target sound-emitting organism on the display in the image data. The target sound-emitting organism can be a person, an animal, an object (such as an airplane or a car), etc. The location identification module 502 can identify the location of the target sound-emitting organism on the display in the image data based on an image identification algorithm. How the location identification module 502 identifies the location of the target sound-emitting organism on the display in the image data will be described in detail in subsequent embodiments. In this embodiment of the present application, the details are not described here.
[0143] In some embodiments, the location identification module 502 may identify a plurality of target sound-emitting organisms from the image data, and the location identification module 502 may obtain the locations of the plurality of target sound-emitting organisms on the display. The location identification module 502 may send the locations of the plurality of target sound-emitting organisms on the display to the audio-visual rendering module 503.
[0144] After the sound extraction module 501 obtains one or more types of audio, the sound extraction module 501 is further configured to send the one or more types of audio to the audio-visual rendering module 503.
[0145] After the location identification module 502 identifies the first location information of the target sound-emitting organism on the display in the image data, the location identification module 502 is further configured to send the first location information of the target sound-emitting organism on the display to the audio-visual rendering module 503.
[0146] The audio-visual rendering module 503 is configured to receive one or more types of audio transmitted by the audio extraction module 501, and the audio-visual rendering module 503 is further configured to receive first position information of the target sound-emitting object on the display transmitted by the positioning module 502.
[0147] The audio-visual rendering module 503 is further configured to perform post-processing on the audio data, for example, perform an inverse short-time Fourier transform, smooth the data, and integrate the object audio channel into the source audio channel. The inverse short-time Fourier transform returns the time spectrum to the time-domain signal. The smoothing of the data can fade in and fade out the signal between frames to remove the popping sounds caused by the model strategy. The integration of the object audio channel into the source audio channel is used to alternately arrange the processed object audio channel with the source audio channel and then output it.
[0148] Next, the audio-visual rendering module 503 is configured to determine the first loudspeaker closest to the target sound-emitting object based on the first position information and the position information of a plurality of loudspeakers on the electronic device 100.
[0149] The audio-visual rendering module 503 is further configured to determine the first audio of the target sound-emitting object from one or more types of audio. If there are multiple target sound-emitting objects, the audio-visual rendering module 503 can determine the target sound-emitting object corresponding to each type of audio from the one or more types of audio.
[0150] The audio-visual rendering module 503 is further configured to transmit to the audio control module 504 a first audio channel identifier corresponding to the first loudspeaker and the first audio.
[0151] The audio control module 504 is configured to receive a first channel identifier and first audio corresponding to a first loudspeaker and transmitted by the audio-visual rendering module 503. First, the audio control module 504 determines whether the first audio channel is connected to the sound box. If the first audio channel is not connected to the sound box, the audio control module 504 outputs the first audio via the first loudspeaker corresponding to the first audio channel. Or, if the first audio channel is connected to the sound box, the audio control module 504 outputs the first audio via the first sound box corresponding to the first audio channel.
[0152] In this way, the audio of the target sound-emitting object can be output from the loudspeaker or sound box closest to the target sound-emitting object based on the position of the target sound-emitting object on the display in the image data. Therefore, the audio output position of the target sound-emitting object changes together with the position of the target sound-emitting object on the display.
[0153] It should be noted that the sound extraction module 501, the position identification module 502, the audio-visual rendering module 503, and the audio control module 504 may be randomly combined to implement the above functions. Or, the sound extraction module 501, the position identification module 502, the audio-visual rendering module 503, and the audio control module 504 may be separately used as one module to implement the above functions. This is not limited in the embodiments of this application.
[0154] The following will explain in detail how the electronic device 100 extracts one or more types of audio from the audio data. The one or more types of audio include, but are not limited to, human voices, animal sounds, environmental sounds, music sounds, object sounds (such as the sound of an airplane or a car), etc.
[0155] After the electronic device 100 obtains the audio data in the video data, the electronic device 100 extracts one or more types of audio from the audio data.
[0156] In some embodiments, the electronic device 100 may obtain an audio extraction model in advance through training, and extract one or more types of audio in the audio data in the video data by using the trained audio extraction model.
[0157] Before the electronic device 100 extracts one or more types of audio from the audio data, the electronic device 100 needs to obtain one or more types of predetermined audio features.
[0158] [Table 1]
[0159] Table 1 shows examples of several types of audio features. The audio features may include, but are not limited to, tone, loudness, timbre, melody, etc. After the audio features of the audio are extracted, the audio features of the audio may be represented by a feature vector.
[0160] Table 1 shows examples of vector representations of several types of audio features. The audio features of a human voice may be represented by the feature vector V1, where V1 = [a1, a2, a3, a4, a5, a6..., an]. The audio features of a barking dog may be represented by the feature vector V2, where V2 = [b1, b2, b3, b4, b5, b6..., bn]. The audio features of the sound of wind may be represented by the feature vector V3, where V3 = [c1, c2, c3, c4, c5, c6..., cn]. The audio features of the sound of an airplane may be represented by the feature vector V4, where V4 = [d1, d2, d3, d4, d5, d6..., dn]. The audio features of the sound of a car may be represented by the feature vector V5, where V5 = [e1, e2, e3, e4, e5, e6..., en]. The audio features of the sound of a train may be represented by the feature vector V6, where V6 = [f1, f2, f3, f4, f5, f6..., fn]. Other audio features of more types may be included. In this embodiment of the present application, the details are not described here.
[0161] Note that audio of the same type but with different attributes has different audio features.
[0162] For example, a human voice can be a female voice or a male voice according to gender. According to age, a human voice can be the voice of a person from 0 to 5 years old, the voice of a person from 5 to 10 years old, the voice of a person from 10 to 15 years old, the voice of a person from 15 to 20 years old, the voice of a person from 20 to 25 years old, the voice of a person from 25 to 30 years old, and so on. The same type of human voice at different stages further has different audio features. In this embodiment of the present application, the details are not described here.
[0163] To form a complete audio file, multiple different types of audio, such as human voices, animal sounds, environmental sounds, music sounds, and object sounds, are integrated in an audio file within video data. An electronic device can separate the multiple types of audio integrated in the audio file to obtain the multiple audio.
[0164] Next, the electronic device 100 separately extracts the audio features of the multiple audio and compares the audio features of the multiple audio with predetermined audio features. When the similarity is higher than a predetermined value (for example, 90%), the electronic device 100 can determine the audio type of the audio. For example, if the similarity between the audio features of one of the multiple audio and the audio features of the sound of a car is 95%, and the similarity between the audio features of that audio and the audio features of a human voice, the sound of a barking dog, the sound of the wind, and the audio features of the sound of a car is lower than 90%, for example, the similarity is 20%, the electronic device 100 can determine that the audio type of the audio is the sound of an airplane.
[0165] It should be noted that the audio data in the video data may include one or more types of audio, and the electronic device 100 can extract different types of audio based on the above method, that is, separate different types of audio.
[0166] After the electronic device 100 extracts multiple different types of audio from the audio data in the video data, the electronic device 100 further determines the target sound-emitting organisms of the multiple different types of audio and the positions of the target sound-emitting organisms on the display of the electronic device 100, that is, it is necessary to determine the positions of the target sound-emitting organisms of each type of audio on the display of the electronic device 100.
[0167] The electronic device 100 can determine an image of a target sound-emitting organism in a video frame by using an image identification model.
[0168] First, the image identification model needs to be trained by using a large number of sample photos. The sample photos include a specific quantity of various types of photos, such as portrait photos, photos of animals (e.g., photos of dogs), and photos of objects (e.g., photos of airplanes or cars). The specific training process is as follows. The sample photos are used as the input of the image identification model, and the image identification model outputs the image features of the input sample photos. Then, the image features of the sample photos output by the image identification model are compared with the predetermined image features of the sample photos, and it is observed whether the similarity between the two image features is greater than a predetermined value. If the similarity is greater than the predetermined value, it indicates that the image identification model can accurately identify the features of the photo, and the training of the image identification model is completed. If the similarity is less than the predetermined value, the above steps are repeated until the similarity between the image features of the sample photos output by the image identification model and the predetermined image features of the sample photos is greater than the predetermined value.
[0169] Note that the image identification model may be trained by the electronic device 100, or the image identification model may be trained on the server, and the trained image identification model is transmitted to the electronic device 100. This is not limited in the embodiments of the present application.
[0170]
Table 2
[0171] Table 2 shows examples of several types of image features. The image features may include, but are not limited to, color features, texture features, contour features, etc. After the image features of the image are extracted, the image features of the image can be represented by a feature vector.
[0172] Table 2 shows examples of vector representations of several types of image features. The image features of a face image may be represented by feature vector F1, where F1 = [A1, A2, A3, A4, A5, A6..., An]. The image features of a dog image may be represented by feature vector F2, where F2 = [B1, B2, B3, B4, B5, B6..., Bn]. The image features of an airplane image may be represented by feature vector F3, where F3 = [C1, C2, C3, C4, C5, C6..., Cn]. The image features of a car image may be represented by feature vector F4, where F4 = [D1, D2, D3, D4, D5, D6..., Dn]. The image features of a train image may be represented by feature vector F5, where F5 = [E1, E2, E3, E4, E5, E6..., En]. Other image features of more types may be included. In this embodiment of the present application, the details are not described here.
[0173] It should be noted that images of the same type but with different attributes have different image features.
[0174] For example, a face image can be an image of a female or a male according to gender. According to age, a face image can be an image of a person's face from 0 to 5 years old, an image of a person's face from 5 to 10 years old, an image of a person's face from 10 to 15 years old, an image of a person's face from 15 to 20 years old, an image of a person's face from 20 to 25 years old, an image of a person's face from 25 to 30 years old, etc. Face images of the same type at different stages further have different image features. In this embodiment of the present application, the details are not described here.
[0175] The image data in the video data may include multiple types of images, such as face images, animal images, car images, airplane images, and train images. The electronic device 100 needs to identify multiple types of images from the image data and determine the target sound-emitting organisms from the multiple types of images.
[0176] For example, the electronic device 100 can identify a predetermined type of image in the image data, that is, determine the target sound-emitting organism by using an image identification model.
[0177] Specifically, the electronic device 100 separately extracts image features from the image data and compares the image features with predetermined image features. When the similarity is higher than a predetermined value (for example, 90%), the electronic device 100 can determine the image type of the image data. For example, if the similarity between the image features of the image in the image data and the image features of an automobile image is 95%, and the similarity between the image features of that image and the image features of a face image, the image features of a (complete) animal image, and the image features of a train image is less than 90%, for example, the similarity is 20%, the electronic device 100 can determine that the image type of the image is a train.
[0178] In some embodiments, after the electronic device 100 separates a plurality of types of audio, such as a first type of first audio (e.g., audio of a person's voice), from video data, the electronic device 100 needs to track the sound-generating body corresponding to the first audio from the image data of the video data. Specifically, the electronic device 100 determines an image of the first type based on the first type of the first audio, and the electronic device 100 finds the image features of the first type from a plurality of predetermined image features. Then, the electronic device 100 extracts the features of a plurality of images in the image data of the video data, compares the extracted image features of the plurality of image data of the video data with the predetermined image features of the first type, and finds the position of the image of the first type in the image. In this way, the electronic device 100 only needs to compare the extracted image features of the plurality of images in the image data of the video data with the image features of the first type (i.e., determine whether the similarity is greater than a predetermined value), and does not need to compare the image features with all the plurality of predetermined image features. This can reduce the time for the electronic device 100 to determine the position of the image of the first type in the image. For example, when the type of the first audio is a person's voice, the electronic device 100 needs to determine the position of the person in the image from the image data of the video data. The electronic device 100 finds the image features of a person from a plurality of types of predetermined image features, the electronic device 100 extracts the features of a plurality of images in the image data of the video data, compares the extracted image features of the plurality of image data of the video data with the image features of a person, and determines the position of the image of the person in the image. In this way, the electronic device 100 only needs to compare the image features of the plurality of image data of the video data with the image features of a person, and does not need to compare the image features of the plurality of image data of the video data with the plurality of predetermined image features of the image (e.g., image features of an animal or an object). This can reduce the time for the electronic device 100 to determine the position of the image of the person in the image.
[0179] In another embodiment, after the electronic device 100 separates a plurality of types of audio, such as a first type of first audio (e.g., audio of a person's voice), from the video data, the electronic device 100 needs to track the sound-generating object corresponding to the first audio from the image data of the video data. Specifically, the electronic device 100 may extract features of a plurality of images in the image data of the video data and obtain a plurality of types of predetermined image features. The electronic device 100 compares the image features of the plurality of images in the image data of the video data with the plurality of types of predetermined image features to determine each type of the image and the position of each image. Then, the electronic device 100 determines the image type of the sound-generating object corresponding to the first audio from each type of the image and further determines the position of the image of the sound-generating object corresponding to the first audio in the image.
[0180] Alternatively, the position of the image of the sound-generating object in the image may be determined in another manner. This is not limited in the embodiments of the present application.
[0181] In another embodiment, the electronic device 100 may first identify the image features of the target sound-generating object in the image data, and then determine the type of the target sound-generating object and the position of the target sound-generating object in the image. Then, the electronic device 100 separates the corresponding type of audio from the audio data of the video data based on the type of the target sound-generating object. For example, the electronic device 100 first identifies the image of a person in the image data and the position of the person in the image and determines that the person is making a sound. Then, based on the case where the type of the target sound-generating object is a person, the electronic device 100 extracts the audio emitted by the person from the audio data of the video data and outputs the audio of the person from the corresponding audio channel based on the position of the person in the image.
[0182] After the electronic device 100 obtains the image type in the image data, the electronic device 100 may determine the target sound-generating object in the image data. The target sound-generating object includes, but is not limited to, a person, an animal, an object, etc.
[0183] Video data image data usually has images of a plurality of different target sound-emitting organisms. Each target sound-emitting organism corresponds to the output audio. The electronic device 100 further needs to associate the image features of each target sound-emitting organism with the audio features of the output audio of the target sound-emitting organism in a one-to-one correspondence, whereby the electronic device 100 can track the position of the image of the target sound-emitting organism on the electronic device 100 and the output position of the audio emitted by the target sound-emitting organism.
[0184] Specifically, the electronic device 100 extracts the image features and audio features of a target sound-emitting organism (for example, the first target object). The electronic device 100 determines that the similarity is greater than a predetermined value based on the predetermined image features of the image and the image features of the first target object, and the electronic device 100 determines the image type of the first target object. The electronic device 100 further needs to determine that the similarity is greater than a predetermined value based on the predetermined audio features of the audio type and the audio features of the first target object, and the electronic device 100 determines the audio type of the first target object. When the image type of the first target object and the audio type of the first target object are the same type, for example, when they are the image of a face and the sound of a person respectively, the electronic device 100 establishes a one-to-one association relationship between the audio features and the image features of the first target object. Next, the electronic device 100 can track the display position of the first target object on the display in the image data and audio data output later, and output the audio of the first target object at the corresponding position. If the electronic device 100 identifies a new target sound-emitting organism in the image data and audio data output later, and both the audio features and the image features of the new target sound-emitting organism are different from the audio features of the previous target sound-emitting organism, the electronic device 100 may determine that a new target sound-emitting organism has appeared. The electronic device 100 establishes a one-to-one association relationship between the audio features and the image features of the new target sound-emitting organism, tracks the display position of the new target sound-emitting organism on the display, and outputs the audio of the corresponding target sound-emitting organism at the corresponding position. Similarly, the electronic device 100 can track the position of each target sound-emitting organism in the video data on the display and output the audio of the target sound-emitting organism at the corresponding position.
[0185]
Table 3
[0186] Table 3 shows the relationship between the image features and audio features of the target sound-emitting organisms, and examples of the identified target sound-emitting organisms when the electronic device 100 plays video data. When the target sound-emitting organism is a human, the human image features are F1 = [A1, A2, A3, A4, A5, A6..., An], the human audio features are V1 = [a1, a2, a3, a4, a5, a6..., an], and the electronic device 100 establishes a one-to-one relationship between F1 and V1. When the target sound-emitting organism is an animal (dog), the animal (dog) image features are F2 = [B1, B2, B3, B4, B5, B6..., Bn], the animal (dog) audio features are V2 = [b1, b2, b3, b4, b5, b6..., bn], and the electronic device 100 establishes a one-to-one relationship between F2 and V2. When the target sound-emitting organism is a car, the car image features are F4 = [D1, D2, D3, D4, D5, D6..., Dn], the car audio features are V5 = [e1, e2, e3, e4, e5, e6..., en], and the electronic device 100 establishes a one-to-one relationship between F4 and V5. When the target sound-emitting organism is an airplane, the airplane image features are F3 = [C1, C2, C3, C4, C5, C6, ..., Cn], the car audio features are V4 = [d1, d2, d3, d4, d5, d6, ..., dn], and the electronic device 100 establishes a one-to-one relationship between F3 and V4. When the target sound-emitting organism is a train, the train image features are F5 = [E1, E2, E3, E4, E5, E6, ..., En], the train audio features are V6 = [f1, f2, f3, f4, f5, f6, ..., fn], and the electronic device 100 establishes a one-to-one relationship between F5 and V6.
[0187] Regarding people, it should be noted that image features and audio features are different for people with different attributes. For example, the image features of men are different from those of women, and the audio features of men are also different from those of women. In this way, the electronic device 100 can also distinguish people with different attributes.
[0188] The electronic device 100 tracks the position of the image of the target sound-emitting object on the display, that is, determines the loudspeaker closest to the position of the target sound-emitting object on the display, and determines the change in the position of the target sound-emitting object on the display in order to output the audio corresponding to the target sound-emitting object through the closest loudspeaker.
[0189] The following describes how the electronic device 100 determines the position of the image of the target sound-emitting object (for example, Person 1) on the display.
[0190] The electronic device 100 identifies the contour of the target sound-emitting object and determines the position of the target sound-emitting object on the display based on the contour of the target sound-emitting object.
[0191] For example, regarding people and animals, people and animals make sounds based on their lips, and the electronic device 100 can identify the contour of the head image of a person or an animal based on an algorithm. Furthermore, the electronic device 100 can determine the lip position from the contour of the head image of a person or an animal. After determining the lip position, the electronic device 100 can determine the position of the lip position on the display of the electronic device 100.
[0192] For example, for another object such as a train, a car, or an airplane, the electronic device 100 can identify the outline of the object based on an algorithm. Further, the electronic device 100 determines the center point of the outline of the object based on the outline of the object. The electronic device 100 can use the center point as the sound generation part of the object. After determining the sound generation part, the electronic device 100 can determine the position of the sound generation part of the object on the display of the electronic device 100. In addition to using the center point of the outline of the object as the sound generation part of the object, the electronic device 100 may also determine the sound generation part of the object in another manner. This is not limited in the embodiments of the present application.
[0193] FIG. 6 and FIG. 7 are each diagrams of an example of how to determine the position of person 1 on the display.
[0194] Optionally, in order to determine the position of person 1 on the display of the electronic device 100, with the center point of the electronic device 100 as the origin, the horizontal right direction as the positive direction of the x-axis, the upward direction perpendicular to the x-axis as the positive direction of the y-axis, and the direction perpendicular to the direction towards the inside of the display as the positive direction of the z-axis, a three-dimensional coordinate system is established by using.
[0195] It should be noted that the position of the loudspeaker of the electronic device 100 is unique to the electronic device 100. Therefore, the electronic device 100 can obtain the position information of a plurality of loudspeakers of the electronic device 100. For example, in an embodiment of the present application, the position of the loudspeaker described in FIG. 2 is used as an example for explaining how to determine the position of person 1 on the display, but this should not be a constraint.
[0196] As shown in FIG. 6, it is assumed that the width of the display of the electronic device 100 is M, and the height of the display is N. The loudspeaker 201 is in the first quadrant, and the position coordinates of the loudspeaker 201 may be expressed as A(a, b, 0), where a is greater than 0 and less than M, and b is greater than 0 and less than N. The loudspeaker 202 is in the fourth quadrant, and the position coordinates of the loudspeaker 202 may be expressed as B(c, d, 0), where c is greater than 0 and less than M, and d is greater than -N and less than 0. The loudspeaker 203 is in the second quadrant, and the position coordinates of the loudspeaker 203 may be expressed as C(e, f, 0), where e is greater than -M and less than 0, and f is greater than 0 and less than N. The loudspeaker 204 is in the third quadrant, and the position coordinates of the loudspeaker 204 may be expressed as D(g, h, 0), where g is greater than -M and less than 0, and h is greater than -N and less than 0. The loudspeaker 205 is located at the origin, and the position coordinates of the loudspeaker 205 may be expressed as E(i, j, 0), where i is equal to 0 and j is equal to 0.
[0197] As shown in FIG. 6, at the first moment, the electronic device 100 can identify the head contour of person 1 and determine the coordinates of the sound generation part (for example, the coordinates of the lips) based on the head contour of person 1. It is assumed that the coordinates of the sound generation part of person 1 are F1(o, p, q). For example, o is greater than -M and less than 0, p is greater than -N and less than 0, and q is greater than 0.
[0198] The electronic device 100 can determine the distance between the sound generation part of person 1 and each loudspeaker based on the coordinates F1 of the sound generation part of person 1 and the position coordinates of each loudspeaker.
[0199] For example, if the electronic device 100 determines that the distance between the sound generation part of person 1 and the loudspeaker 201 is r1, the distance between the sound generation part of person 1 and the loudspeaker 202 is r2, the distance between the sound generation part of person 1 and the loudspeaker 203 is r3, the distance between the sound generation part of person 1 and the loudspeaker 204 is r4, and the distance between the sound generation part of person 1 and the loudspeaker 205 is r5, and if r5 < r4 < r3 < r2 < r1, the electronic device 100 determines that the distance between the sound generation part of person 1 and the loudspeaker 205 is the closest, and the electronic device 100 determines that the audio corresponding to person 1 is output via the loudspeaker 205 at the first moment.
[0200] As shown in FIG. 7, person 1 changes in real time, for example, at the second moment, and person 1 moves to the position shown in FIG. 7. Therefore, the electronic device 100 can identify the head contour of person 1 and determine the coordinates of the sound generation part (for example, the coordinates of the lips) based on the head contour of person 1. Assume that the coordinates of the sound generation part of person 1 are F2(u, v, w). For example, u is greater than 0 and less than M, v is greater than 0 and less than N, and q is greater than 0.
[0201] The electronic device 100 can determine the distance between the sound generation part of person 1 and each loudspeaker based on the coordinates F2 of the sound generation part of person 1 and the position coordinates of each loudspeaker.
[0202] For example, if the electronic device 100 determines that the distance between the sound generation part of person 1 and the loudspeaker 201 is r6, the distance between the sound generation part of person 1 and the loudspeaker 202 is r7, the distance between the sound generation part of person 1 and the loudspeaker 203 is r8, the distance between the sound generation part of person 1 and the loudspeaker 204 is r9, and the distance between the sound generation part of person 1 and the loudspeaker 205 is r10, and when r6 < r10 < r7 < r8 < r9, the electronic device 100 determines that the distance between the sound generation part of person 1 and the loudspeaker 201 is the closest, and the electronic device 100 determines that the audio corresponding to person 1 is output via the loudspeaker 201 at the second moment.
[0203] The following describes how the electronic device 100 outputs the audio of the target sound generation object via the closest loudspeaker.
[0204] It can be seen from the above embodiments that after the electronic device 100 acquires the audio data in the video data, it acquires the predetermined audio channel information in the audio data.
[0205]
Table 4
[0206] Table 4 shows examples of pre-determined voice channel information and the types of audio output by each voice channel in the audio data acquired by the electronic device. For example, the pre-determined voice channel information includes a left center voice channel, a right center voice channel, and a center voice channel. The left center voice channel outputs environmental sounds and music, the right center voice channel also outputs environmental sounds and music, and the center voice channel outputs the audio of the first target sound-emitting organism. The audio of the first target sound-emitting organism is the audio extracted from the audio data by the electronic device 100. For example, the audio of the first target sound-emitting organism can be the audio emitted by a person. In other words, the electronic device 100 outputs the audio emitted by a person from the default center voice channel.
[0207] The pre-determined voice channel information shown in Table 4 can be called 2D audio information. The 2D audio information indicates the voice channels through which each type of audio is output and is pre-determined. The electronic device 100 cannot change the voice channel for outputting the audio of the target sound-emitting organism based on the change in the position of the target sound-emitting organism on the display. The user feels that the audio of the target sound-emitting organism is output from a specific position and does not feel the change in sound in space. For example, the audio of the first target sound-emitting organism is pre-determined to be output from the center voice channel. When the position of the first target sound-emitting organism on the display changes in real time in the video image output by the electronic device 100, the electronic device 100 cannot output the audio of the first target sound-emitting organism from the left center voice channel or the right center voice channel. The user feels that the audio of the first target sound-emitting organism is output from a specific position and does not feel the change in sound in space.
[0208] The electronic device 100 can extract other types of audio from the audio data, such as audio emitted by animals or objects. The audio emitted by humans is used here as an example for explanation and should not be restrictive.
[0209] After the electronic device 100 extracts one or more types of audio from the audio data, such as the audio of the first target sound-emitting organism, the audio that would be emitted by the first target sound-emitting organism at the first moment is the first audio. After the electronic device 100 determines the position of the first target sound-emitting organism on the display, the electronic device 100 determines the first loudspeaker closest to the first target sound-emitting organism, and the electronic device 100 may determine the first voice channel identifier corresponding to the first loudspeaker based on the first loudspeaker.
[0210] Next, the electronic device 100 loads the first audio into the corresponding first voice channel. First, the electronic device 100 determines whether the first voice channel is connected to the sound box. If the first voice channel is not connected to the sound box, the electronic device 100 outputs the first audio through the first loudspeaker corresponding to the first voice channel, or if the first voice channel is connected to the sound box, the electronic device 100 outputs the first audio through the first sound box corresponding to the first voice channel. For example, the first voice channel may be the left rear voice channel.
[0211] When the electronic device 100 outputs the first audio, the electronic device 100 outputs other audio (such as background sound or music sound) from a predetermined voice channel.
[0212]
Table 5
[0213] Table 5 shows an example in which at a first moment, based on the position of a first target sound-emitting organism on the display, the electronic device 100 outputs first audio from a left rear audio channel corresponding to the nearest first loudspeaker.
[0214] The electronic device 100 still outputs ambient sound and music from a predetermined left center audio channel and a predetermined right center audio channel. At the first moment, if the first target sound-emitting organism is closest to the first loudspeaker, the electronic device 100 may output the first audio from the left rear audio channel corresponding to the first loudspeaker. In this case, the electronic device 100 no longer outputs the first audio from the predetermined center audio channel.
[0215] In some embodiments, at the first moment, the electronic device 100 outputs the audio of the target sound-emitting organism via the first loudspeaker. The electronic device 100 detects the audio of the target sound-emitting organism that is to be output at a third moment, but the electronic device 100 does not detect an image of the target sound-emitting organism in the image data. In a possible implementation, the electronic device 100 may still output the audio of the target sound-emitting organism via the first loudspeaker. If the electronic device 100 does not detect an image of the target sound-emitting organism in the image data after a specific period or specific image frame, the electronic device 100 may output the audio of the target sound-emitting organism from a predetermined audio channel. For example, the predetermined audio channel may be a center audio channel. In another possible implementation, the electronic device 100 may directly output the audio of the target sound-emitting organism from a predetermined audio channel. For example, the predetermined audio channel may be a center audio channel.
[0216] In some embodiments, at the first moment, if the number of target sound-emitting organisms identified by the electronic device 100 is greater than a threshold (e.g., 5), the electronic device 100 outputs the audio of the plurality of target sound-emitting organisms from a predetermined audio channel, and does not output the audio of the plurality of target sound-emitting organisms through the nearest loudspeaker or sound box based on the positions of the plurality of target sound-emitting organisms on the display.
[0217] Next, at the second moment, the audio that will be emitted by the first target sound-emitting organism is the second audio. After the electronic device 100 determines the position of the first target sound-emitting organism on the display, the electronic device 100 determines the second loudspeaker closest to the first target sound-emitting organism, and the electronic device 100 may determine a second audio channel identifier corresponding to the second loudspeaker based on the second loudspeaker.
[0218] Next, the electronic device 100 loads the second audio into the corresponding second channel. First, the electronic device 100 determines whether the second audio channel is connected to a sound box. If the second audio channel is not connected to a sound box, the electronic device 100 outputs the second audio through the second loudspeaker corresponding to the second audio channel, or if the second audio channel is connected to a sound box, the electronic device 100 outputs the second audio through the second sound box corresponding to the second audio channel. For example, the second audio channel may be the right front audio channel.
[0219] When the electronic device 100 outputs the second audio, the electronic device 100 outputs another audio (e.g., background sound or music sound) from a predetermined audio channel.
[0220]
Table 6
[0221] Table 6 shows an example in which at a second moment, based on the position of the first target sound-emitting organism on the display, the electronic device 100 outputs second audio from the left rear audio channel corresponding to the nearest second loudspeaker.
[0222] The electronic device 100 still outputs ambient sound and music from the predetermined left center audio channel and the predetermined right center audio channel. At the second moment, when the first target sound-emitting organism is closest to the second loudspeaker, the electronic device 100 may output the second audio from the right front audio channel corresponding to the second loudspeaker. In this case, the electronic device 100 no longer outputs the second audio from the predetermined center audio channel or the left rear audio channel.
[0223] The audio output audio channel information shown in Table 5 and Table 6 may also be referred to as 3D audio information. The 3D audio information indicates the audio channels through which several types of audio are output and can change in real time. The electronic device 100 may change the audio channel for outputting the audio of the target sound-emitting organism based on the change in the position of the target sound-emitting organism on the display. The user feels that the audio of the target sound-emitting organism is output from different positions and feels the change of sound in space. For example, the first audio emitted by the first target sound-emitting organism shown in Table 5 is output from the left rear audio channel, and the second audio emitted by the first target sound-emitting organism shown in Table 6 is output from the right front audio channel. In other words, as the position of the first target sound-emitting organism on the display changes in real time, the electronic device 100 can change the audio output position of the first target sound-emitting organism. The user feels that the audio of the first target sound-emitting organism is output from different positions and feels the change of sound in space.
[0224] The following describes how the electronic device 100 changes the audio output position of the target sound-emitting object based on the change in the position of the target sound-emitting object for a specific scenario.
[0225] Figures 8A to 8C are diagrams of examples in which the electronic device 100 outputs the audio of the target sound-emitting object (e.g., a train) through different loudspeakers as the position of the target sound-emitting object changes.
[0226] As shown in Figure 8A, at the first moment, the electronic device 100 identifies the position of the target sound-emitting object on the display, determines that the distance between the target sound-emitting object and the loudspeaker 201 is the shortest, and the electronic device 100 is not externally connected to an audio output device (e.g., a sound box). In this case, the electronic device 100 outputs the first audio emitted by the target sound-emitting object at the first moment through the loudspeaker 201.
[0227] Optionally, at the first moment, the electronic device 100 may simultaneously output the first audio through the loudspeaker 201, the loudspeaker 202, the loudspeaker 203, the loudspeaker 204, and the loudspeaker 205. However, the output volumes of the loudspeaker 201, the loudspeaker 202, the loudspeaker 203, the loudspeaker 204, and the loudspeaker 205 are different, and the output volume of the loudspeaker can be gradually decreased from near to far based on the distance between the target sound-emitting object and the loudspeaker. For example, the volume of the loudspeaker 201 is greater than the output volumes of the loudspeaker 202, the loudspeaker 203, the loudspeaker 204, and the loudspeaker 205.
[0228] As shown in FIG. 8B, at the second instant, the electronic device 100 identifies the position of the target sound-emitting object on the display, determines that the distance between the target sound-emitting object and the loudspeaker 205 is the shortest, and the electronic device 100 is not externally connected to an audio output device (e.g., a sound box). In this case, the electronic device 100 outputs the second audio emitted by the target sound-emitting object at the second instant via the loudspeaker 205, and the second instant is after the first instant.
[0229] Optionally, at the second instant, the electronic device 100 may simultaneously output the first audio via the loudspeakers 201, 202, 203, 204, and 205. However, the output volumes of the loudspeakers 201, 202, 203, 204, and 205 are different, and the output volume of the loudspeaker may be gradually decreased from near to far based on the distance between the target sound-emitting object and the loudspeaker. For example, the volume of the loudspeaker 205 is greater than the output volumes of the loudspeakers 202, 203, 204, and 201.
[0230] As shown in FIG. 8C, at the third instant, the electronic device 100 identifies the position of the target sound-emitting object on the display, determines that the distance between the target sound-emitting object and the loudspeaker 204 is the shortest, and the electronic device 100 is not externally connected to an audio output device (e.g., a sound box). In this case, the electronic device 100 outputs the third audio emitted by the target sound-emitting object at the third instant via the loudspeaker 204, and the third instant is after the second instant.
[0231] Optionally, at a third moment, the electronic device 100 can simultaneously output a first audio via the loudspeaker 201, the loudspeaker 202, the loudspeaker 203, the loudspeaker 204, and the loudspeaker 205. However, the output volumes of the loudspeaker 201, the loudspeaker 202, the loudspeaker 203, the loudspeaker 204, and the loudspeaker 205 are different, and the output volume of the loudspeaker can be gradually decreased from near to far based on the distance between the target sound generating body and the loudspeaker. For example, the volume of the loudspeaker 204 is greater than the output volumes of the loudspeaker 202, the loudspeaker 203, the loudspeaker 201, and the loudspeaker 205.
[0232] Therefore, it can be seen from FIGS. 8A to 8C that as the position of the target sound generating body on the display changes, the electronic device 100 can output the audio emitted by the target sound generating body via different loudspeakers. This realizes "audio-video fusion" and improves the user's video viewing experience.
[0233] FIGS. 9A to 9C are diagrams of examples in which the electronic device 100 outputs the audio of the target sound generating body (for example, a train) via different sound boxes as the position of the target sound generating body changes.
[0234] As shown in FIG. 9A, at a first moment, the electronic device 100 identifies the position of the target sound generating body on the display, determines that the distance between the target sound generating body and the loudspeaker 201 is the shortest, and the loudspeaker 201 corresponds to a first audio channel (for example, a left front audio channel). However, when the electronic device 100 detects that the left front audio channel is connected to the sound box 213, the electronic device 100 outputs the first audio emitted by the target sound generating body at the first moment via the sound box 213.
[0235] Optionally, at a first instant, the electronic device 100 may simultaneously output a first audio via the sound boxes 213, 214, 215, 216, and 217. However, the output volumes of the sound boxes 213, 214, 215, 216, and 217 are different, and the output volume of the sound box may be gradually decreased from near to far based on the distance between the target sound generating body and the sound box. For example, the volume of the sound box 213 is greater than the output volumes of the sound boxes 214, 215, 216, and 217.
[0236] As shown in FIG. 9B, at a second instant, the electronic device 100 identifies the position of the target sound generating body on the display, determines that the distance between the target sound generating body and the loudspeaker 205 is the shortest, and the loudspeaker 205 corresponds to a second audio channel (for example, a central audio channel). However, when the electronic device 100 detects that the central audio channel is connected to the sound box 215, the electronic device 100 outputs a second audio emitted by the target sound generating body at the second instant via the sound box 215, and the second instant is after the first instant.
[0237] Optionally, at a second instant, the electronic device 100 may simultaneously output a first audio via the sound boxes 213, 214, 215, 216, and 217. However, the output volumes of the sound boxes 213, 214, 215, 216, and 217 are different, and the output volume of the sound box may be gradually decreased from near to far based on the distance between the target sound generating body and the sound box. For example, the volume of the sound box 215 is greater than the output volumes of the sound boxes 214, 213, 216, and 217.
[0238] As shown in FIG. 9C, at a third instant, the electronic device 100 identifies the position of the target sound generating body on the display, determines that the distance between the target sound generating body and the loudspeaker 204 is the shortest, and the loudspeaker 204 corresponds to a third audio channel (for example, the left rear audio channel). However, when the electronic device 100 detects that the left rear audio channel is connected to the sound box 217, the electronic device 100 outputs a third audio emitted by the target sound generating body at the third instant via the sound box 217, and the third instant is after the second instant.
[0239] Optionally, at a first moment, the electronic device 100 may simultaneously output first audio via the sound boxes 213, 214, 215, 216, and 217. However, the output volumes of the sound boxes 213, 214, 215, 216, and 217 are different, and the output volume of the sound box may be gradually decreased from near to far based on the distance between the target sound generating body and the sound box. For example, the volume of the sound box 217 is greater than the output volumes of the sound boxes 214, 215, 216, and 213.
[0240] Therefore, it can be seen from FIGS. 9A to 9C that as the position of the target sound generating body on the display changes, the electronic device 100 can output the audio emitted by the target sound generating body via different sound boxes. Outputting audio via a sound box can, in one aspect, improve the sound quality of the output audio, and in another aspect, realize "audio-video fusion" to improve the user's video viewing experience.
[0241] According to the audio output provided in the embodiments of the present application, in the video data played by the electronic device 100, after the position of the target sound generating body displayed on the display changes, the audio output position of the target sound generating body also changes. This realizes the "audio-video fusion" effect. The embodiments of the present application are applicable to scenarios where video is output, such as digital television (DTV) live broadcast scenarios, Huawei video-on-demand scenarios, and local video playback scenarios.
[0242] In the following embodiments of the present application, when playing a video, the electronic device 100 receives a user operation so that the "audio-video fusion" effect is specifically realized.
[0243] Figures 10A to 10C are diagrams of examples in which the electronic device 100 receives a user operation so that the electronic device 100 realizes an "audio-video fusion" effect.
[0244] As shown in Figure 10A, the electronic device 100 receives a user operation and displays the user interface 1001. The user interface 1001 shows examples of one or more channel options. For example, there are options such as My Home, Homepage, VIP, TV programs, movies, animations, teenagers, and games. The user interface 1001 displays one or more recommended videos, such as the recommended video 1002, under the option of TV programs.
[0245] As shown in Figure 10A, the electronic device 100 receives a user input operation on the recommended video 1002 by using a remote control, and in response to the user input operation, the electronic device 100 displays the user interface 1003 shown in Figure 10B.
[0246] The user interface 1003 shows an example of the content of the recommended video 1002, which includes, but is not limited to, the video content of the recommended video 1002, the name of the recommended video 1002 (for example, Jump Over with Courage), and a plurality of function controls. The plurality of function controls may be a fast-forward control, a rewind control, a next episode control, a progress bar, etc., and the plurality of function controls may also be a speed selection control, a high-definition control, a spatial audio control 1004, etc.
[0247] As shown in FIG. 10B, the electronic device 100 can receive a user's input operation for the spatial audio control 1004 by using a remote control. In response to the user's input operation, the electronic device 100 may display the prompt information 1005 shown in FIG. 10C, and the content of the prompt information 1005 includes "Switching to spatial audio... Please wait". The prompt information 1005 notifies the user that the spatial audio has been switched.
[0248] After the electronic device 100 is switched to spatial audio, the electronic device 100 tracks the position of the target sound generating object in "Jump over bravely" on the display based on the method described in the above embodiment. In addition, after the position of the target sound generating object on the display changes, in order to achieve "audio-video fusion", the audio output position of the target sound generating object is changed. In this way, the user can feel that the position of the audio emitted by the electronic device changes in real time, and the user can feel the spatial audio, thereby improving the user's auditory experience.
[0249] FIG. 11 is a schematic flowchart of an audio playback method according to an embodiment of the present application.
[0250] S1101: The electronic device starts playing the first video clip, and the video of the first video clip includes the first sound generation target.
[0251] S1102: When the electronic device obtains the first audio emitted by the first sound generation target from the audio data of the first video clip and determines that the position of the first sound generation target in the video of the first video clip is the first position at the first moment, the electronic device outputs the first audio through the first speaker.
[0252] S1103: When the electronic device acquires the second audio emitted by the first sound generation target from the audio data of the first video clip and determines that the position of the first sound generation target in the video of the first video clip is the second position at the second moment, the electronic device outputs the second audio through the second speaker.
[0253] The first sound generation target may be the target sound generation organism shown in FIG. 8A.
[0254] The first speaker may be the loudspeaker 201 shown in FIG. 8A, and the second speaker may be the loudspeaker 205 shown in FIG. 8B.
[0255] The first audio may be the audio emitted by the loudspeaker 201 shown in FIG. 8A, and the second audio may be the audio emitted by the loudspeaker 205 shown in FIG. 8B.
[0256] The first position may be the position of the target sound generation organism in the image shown in FIG. 8A.
[0257] The second position may be the position of the target sound generation organism in the image shown in FIG. 8B.
[0258] The first moment is different from the second moment, the first position is different from the second position, and the first speaker is different from the second speaker.
[0259] This embodiment of the present application provides an audio playback method applicable to an electronic device including a plurality of speakers, and the plurality of speakers include a first speaker and a second speaker.
[0260] The electronic device acquires first audio emitted by a first sound generation target, and the electronic device acquires second audio emitted by the first sound generation target. The first audio and the second audio may be extracted and acquired in real time by the electronic device, or the electronic device may acquire in advance the complete audio emitted by the first sound generation target, and the first audio and the second audio are audio clips in the complete audio emitted by the first sound generation target at different instants.
[0261] According to the audio playback method provided in this embodiment of the present application, the electronic device can change the sound generation position of the target sound generation object based on the relative distance between the target sound generation object in the image frame and each loudspeaker of the electronic device, whereby the sound generation position of the target sound generation object in the video changes together with the display position of the target sound generation object on the display. This realizes "audio-video fusion" and improves the user's video viewing experience.
[0262] In a possible implementation, when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio through the first speaker. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker, the electronic device outputs the first audio through the first speaker. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the second position, the electronic device outputs the second audio through the second speaker. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker, the electronic device outputs the second audio through the second speaker. In this way, after determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device further determines the speaker closest to the first position and outputs, through the closest speaker, the audio emitted by the first sound generation target at the corresponding moment. Thereby, the sound generation position of the target sound generation organism changes together with the display position of the target sound generation organism on the display.
[0263] For details, please refer to the related descriptions in FIGS. 8A to 8C. Details will not be described again in this embodiment of the present application.
[0264] In a possible implementation, when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio through the first speaker. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker, the electronic device outputs the first audio through the first speaker at the first volume value and outputs the first audio through the second speaker at the second volume value, where the first volume value is greater than the second volume value. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the second position, the electronic device outputs the second audio through the second speaker. Specifically, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker, the electronic device outputs the first audio through the second speaker at the third volume value and outputs the first audio through the first speaker at the fourth volume value, where the third volume value is greater than the fourth volume value. In this way, after determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device further determines the distance between each speaker and the first position. When the distances are different, the output volumes of the speakers are different, that is, the speakers may play sounds simultaneously. The shorter the distance, the greater the output volume of the speaker. The longer the distance, the smaller the output volume of the speaker.
[0265] For details, please refer to the related descriptions in FIGS. 8A to 8C. Details will not be described again in this embodiment of the present application.
[0266] In a possible implementation, the video of the first video clip includes a second sound generation target, and the method further includes the following. When the electronic device extracts the third audio emitted by the second sound generation target from the audio data of the first video clip, and determines that the position of the second sound generation target in the video of the first video clip is the third position at the third moment, the electronic device outputs the third audio via the first speaker. When the electronic device extracts the fourth audio emitted by the second sound generation target from the audio data of the first video clip, and determines that the position of the second sound generation target in the video of the first video clip is the fourth position at the fourth moment, the electronic device outputs the fourth audio via the second speaker. The third moment is different from the fourth moment, and the third position is different from the fourth position. In this way, the electronic device can simultaneously detect the positions of multiple sound generation targets and speakers, and change the sound generation positions of multiple sound generation targets.
[0267] In a possible implementation, the multiple speakers further include a third speaker. After the electronic device outputs the first audio via the first speaker, the method further includes the following. After the first period of time has elapsed, or after the number of image frames exceeds the first number, if the electronic device does not detect the position of the first sound generation target in the video of the first video clip, the electronic device outputs the audio of the first sound generation target via the third speaker. In this way, when the electronic device only detects the audio of the first sound generation target and does not detect the image position of the first sound generation target in the image data, the electronic device outputs the audio of the first sound generation target via a predetermined speaker.
[0268] In one possible implementation, when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker, the electronic device outputs the first audio through the first speaker. Specifically, the electronic device includes obtaining the position information of the first speaker and the position information of the second speaker, and based on the first position of the first sound generation target in the video of the first video clip, the position information of the first speaker, and the position information of the second speaker, determining that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker. In this way, the position of each speaker of the electronic device is unique, and the electronic device can determine the distance between the first sound generation target and each speaker based on the position of each speaker and the position of the first sound generation target in the video.
[0269] Specifically, for how to determine the distance between the first sound generation target and each speaker, please refer to the relevant descriptions in FIGS. 6 and 7. Details will not be described again in this embodiment of the present application.
[0270] In a possible implementation, for the electronic device to obtain the first audio emitted by the first sound generation target from the audio data of the first video clip, specifically, the electronic device obtains multiple types of audio from the audio data of the first video clip based on multiple types of predetermined audio features, and the electronic device determines the first audio emitted by the first sound generation target from the multiple types of audio. In this way, in order to determine the first audio emitted by the first sound generation target, the electronic device can calculate the similarity between the multiple types of audio in the audio data and the multiple types of predetermined audio features based on the multiple types of predetermined audio features.
[0271] For details, please refer to the relevant description in Table 1. Details will not be described again in this embodiment of this application.
[0272] In a possible implementation, for the electronic device to determine that the position of the first sound generation target in the video of the first video clip is the first position, specifically, the electronic device identifies the first target image corresponding to the first sound generation target from the video of the first video clip based on multiple types of predetermined image features, and the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position based on the display area of the first target image in the video of the first video clip. In this way, in order to determine the position of the first target image in the video of the first video clip, the electronic device can determine the image features of the first target image corresponding to the first sound generation target based on multiple types of predetermined image features.
[0273] For details, please refer to the relevant description in Table 2. Details will not be described again in this embodiment of this application.
[0274] In one possible implementation, the plurality of speakers further includes a fourth speaker, and before the electronic device outputs the first audio, the method further includes the following. When the electronic device acquires predetermined audio channel information from the audio data of the first video clip, and the predetermined audio channel information includes outputting the first audio and the first background sound from the fourth speaker, and when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio through the first speaker, specifically, when the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device outputs the first audio through the first speaker and outputs the first background sound through the fourth speaker. In this way, when determining that the position of the first sound generation target in the video of the first video clip is the first position, the electronic device can provide the first audio to the first speaker, output the first audio through the first speaker, and output other audio such as background sound and music sound through the predetermined speaker.
[0275] For how the electronic device provides the first audio to the first speaker, refer to the relevant descriptions in Table 4, Table 5, and Table 6. Details will not be described again in this embodiment of the present application.
[0276] In one possible implementation, the position information of the plurality of speakers of the electronic device is different. In this way, the sound generation position of the target sound generation object changes together with the display position of the target sound generation object on the display.
[0277] In one possible implementation, the type of the first sound generation target is any one of a person, an animal, an object, and a landscape.
[0278] In one possible implementation, the first type of audio is any one of human voice, animal sound, environmental sound, music sound, and object sound.
[0279] FIG. 12 is a schematic flowchart of another audio playback method according to an embodiment of the present application.
[0280] S1201: The electronic device starts playing the first video clip, and the video of the first video clip includes a first sound generation target.
[0281] S1202: When the electronic device obtains the first audio emitted by the first sound generation target from the audio data of the first video clip, and determines that the position of the first sound generation target in the video of the first video clip at the first moment is the first position, the electronic device outputs the first audio through the first audio output device.
[0282] S1203: When the electronic device obtains the second audio emitted by the first sound generation target from the audio data of the first video clip, and determines that the position of the first sound generation target in the video of the first video clip at the second moment is the second position, the electronic device outputs the second audio through the second audio output device.
[0283] The first sound generation target may be the target sound generating body shown in FIG. 9A.
[0284] The first audio output device may be the sound box 213 shown in FIG. 9A, and the second audio output device may be the sound box 215 shown in FIG. 9B.
[0285] The first audio may be the audio emitted by the sound box 213 shown in FIG. 9A, and the second audio may be the audio emitted by the sound box 215 shown in FIG. 9B.
[0286] The first position can be the position of the target sound-emitting object in the image shown in FIG. 9A.
[0287] The second position can be the position of the target sound-emitting object in the image shown in FIG. 9B.
[0288] The first moment is different from the second moment, the first position is different from the second position, and the first audio output device is different from the second audio output device.
[0289] The electronic device acquires the first audio emitted by the first sound-emitting target, and the electronic device acquires the second audio emitted by the first sound-emitting target. The first audio and the second audio may be extracted and acquired in real time by the electronic device, or the electronic device may acquire the complete audio emitted by the first sound-emitting target in advance, and the first audio and the second audio are audio clips in the complete audio emitted by the first sound-emitting target at different moments.
[0290] According to the method provided by the second aspect, when the electronic device is externally connected to the audio output device, the electronic device can change the sound generation position of the target sound-emitting object based on the relative distance between the target sound-emitting object in the image frame and the audio output device, whereby the sound generation position of the target sound-emitting object in the video changes together with the display position of the target sound-emitting object on the display. This realizes "audio-video fusion" and improves the user's video viewing experience.
[0291] When a speaker is mounted on the electronic device 100, the electronic device 100 can obtain the position information of each speaker of the electronic device 100, and based on the position information of each speaker of the electronic device 100 and the position of the first sound generation target in the video of the first video clip, determine the distance between the first sound generation target and each speaker. Next, the electronic device 100 obtains the audio channel information corresponding to each speaker. When an audio output device (for example, a sound box) is connected to the audio channel, the electronic device 100 outputs the audio issued by that audio channel via the audio output device, or when the audio output device (for example, a sound box) is not connected to the audio channel, the electronic device 100 outputs the audio issued by that audio channel via the corresponding speaker.
[0292] When there is no speaker on the electronic device 100 (for example, a projector), the electronic device 100 cannot obtain the position information of each speaker of the electronic device 100, but the electronic device 100 is connected to an audio output device, and the electronic device 100 can determine the position information of the connected audio output device, for example, the position in space. Next, the electronic device 100 determines the distance between the first sound generation target and each audio output device based on the position information of the audio output device and the position of the first sound generation target in the video of the first video clip. Then, the audio of the first sound generation target is output via the corresponding audio output device based on the distance between the first sound generation target and each audio output device.
[0293] In a possible implementation, the type of the first audio output device is any one of a sound box, earphones, a power amplifier, a multimedia console, and an audio adapter.
[0294] Note that the multiple possible implementations provided in the procedure of the method shown in FIG. 11 may also be applicable to the procedure of the method shown in FIG. 12. Details will not be described again in the embodiments of this application.
[0295] The embodiments of this application may be randomly combined to achieve different technical effects.
[0296] All or part of the above embodiments may be implemented by using software, hardware, firmware, or any combination thereof. When software is used to implement the embodiments, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the procedures or functions according to this application are all or partially generated. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or may be transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired (such as coaxial cable, optical fiber, or digital subscriber line) or wireless (such as infrared, wireless, or microwave) manner. The computer-readable storage medium may be any usable medium accessible by a computer, or a data storage device, such as a server or data center including one or more usable media. The usable media may be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), a semiconductor medium (such as a solid state disk (SSD)), etc.
[0297] Those skilled in the art can understand that all or part of the steps of the method of the embodiment can be realized by a computer program that instructs the relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, the steps of the method of the embodiment are implemented. The above storage medium includes any medium that can store program code, such as ROM, random access memory RAM, magnetic disk, or optical disk.
[0298] In conclusion, the above description is only an embodiment of the technical solution of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, and improvements made in accordance with the disclosure of the present invention shall fall within the protection scope of the present invention.
Description of Reference Numerals
[0299] 1 Antenna 100 Electronic Device 110 Processor 120 Wireless Communication Module 130 Audio Module 130A Speaker 130B Receiver 130C Microphone 130D Headset Jack 140 Internal Memory 150 Sensor Module 150A Acceleration Sensor 150B Distance Sensor 150C Light Proximity Sensor 150D Temperature Sensor 150E Touch Sensor 150F Ambient Light Sensor 190 Button 191 Motor 192 Indicator 193 Camera 194 Display 201 Loudspeaker 202 Loudspeaker 203 Loudspeaker 204 Loudspeaker 205 Loudspeaker 206 Loudspeaker 207 Loudspeaker 208 Loudspeaker 209 Loudspeaker 210 Loudspeaker 211 Loudspeaker 212 Loudspeaker 213 Sound box 214 Sound box 215 Sound box 216 Sound box 217 Sound box 501 Sound extraction module 502 Position determination module 503 Audio-visual rendering module 504 Audio control module 1001 User interface 1002 Recommended video 1003 User interface 1004 Spatial audio control
Claims
1. An audio playback method applied to an electronic device having a plurality of speakers, wherein the plurality of speakers include a first speaker and a second speaker, and the method includes: starting, by the electronic device, playback of a first video clip, wherein an image of the first video clip includes a first sound generation target; acquiring, by the electronic device, first audio emitted by the first sound generation target from audio data of the first video clip; when the electronic device determines that a position of the first sound generation target in the image of the first video clip is a first position at a first moment, outputting, by the electronic device, the first audio through the first speaker; acquiring, by the electronic device, second audio emitted by the first sound generation target from the audio data of the first video clip; when the electronic device determines that the position of the first sound generation target in the image of the first video clip is a second position at a second moment, outputting, by the electronic device, the second audio through the second speaker, wherein the first moment is different from the second moment, the first position is different from the second position, and the first speaker is different from the second speaker.
2. When the electronic device determines that the position of the first sound generation target in the image of the first video clip is the first position, the step of outputting, by the electronic device, the first audio through the first speaker specifically includes: when the electronic device determines that a distance between the first position of the first sound generation target in the image of the first video clip and the first speaker is shorter than a distance between the first position of the first sound generation target in the image of the first video clip and the second speaker, outputting, by the electronic device, the first audio through the first speaker. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the second position, the step of outputting the second audio through the second speaker by the electronic device specifically is The method according to claim 1, comprising the step of outputting the second audio through the second speaker by the electronic device when the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker. **Claim 3** When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the step of outputting the first audio through the first speaker by the electronic device specifically is When the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker, the method includes outputting the first audio through the first speaker at a first volume value and outputting the first audio through the second speaker at a second volume value, where the first volume value is greater than the second volume value. When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the second position, the step of outputting the second audio through the second speaker by the electronic device specifically is When the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker, the electronic device outputs the first audio through the second speaker at a third volume value and outputs the first audio through the first speaker at a fourth volume value, and the third volume value is greater than the fourth volume value. The method according to claim 1.
4. The video of the first video clip includes a second sound generation target, and the method includes extracting, by the electronic device, third audio emitted by the second sound generation target from the audio data of the first video clip; when the electronic device determines that the position of the second sound generation target in the video of the first video clip is a third position at a third moment, outputting, by the electronic device, the third audio through the first speaker; extracting, by the electronic device, fourth audio emitted by the second sound generation target from the audio data of the first video clip; when the electronic device determines that the position of the second sound generation target in the video of the first video clip is a fourth position at a fourth moment, outputting, by the electronic device, the fourth audio through the second speaker; and the third moment is different from the fourth moment, and the third position is different from the fourth position. The method according to claim 1.
5. The plurality of speakers further includes a third speaker. After the step of outputting the first audio through the first speaker by the electronic device, the method includes After the first period of time has elapsed or after the number of image frames exceeds the first quantity, if the electronic device does not detect the position of the first sound generation target in the video of the first video clip, the method further comprises the step of the electronic device outputting the audio of the first sound generation target via the third speaker, according to any one of claims 1 to 4.
6. The method according to claim 5, wherein the third speaker is different from the first speaker and the second speaker.
7. When the electronic device determines that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker, the step of the electronic device outputting the first audio via the first speaker specifically comprises: The step of the electronic device obtaining the position information of the first speaker and the position information of the second speaker; The method according to claim 2, further comprising: the step of the electronic device determining, based on the first position of the first sound generation target in the video of the first video clip, the position information of the first speaker, and the position information of the second speaker, that the distance between the first position of the first sound generation target in the video of the first video clip and the first speaker is shorter than the distance between the first position of the first sound generation target in the video of the first video clip and the second speaker.
8. The step of the electronic device obtaining the first audio emitted by the first sound generation target from the audio data of the first video clip specifically comprises: The step of the electronic device obtaining multiple types of audio from the audio data of the first video clip based on multiple types of predetermined audio features; The method according to any one of claims 1 to 7, further comprising: the step of the electronic device determining the first audio emitted by the first sound generation target from the multiple types of audio.
9. The electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, specifically, a step of identifying, by the electronic device, a first target image corresponding to the first sound generation target from the video of the first video clip based on a plurality of types of predetermined image features; a step of determining, by the electronic device, that the position of the first sound generation target in the video of the first video clip is the first position based on a display area of the first target image in the video of the first video clip, the method according to any one of claims 1 to 8.
10. The plurality of speakers includes a fourth speaker, and before the step of outputting the first audio by the electronic device, the method further includes: a step of obtaining, by the electronic device, predetermined audio channel information from the audio data of the first video clip, the predetermined audio channel information including outputting the first audio and the first background sound from the fourth speaker; When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the step of outputting the first audio through the first speaker by the electronic device specifically includes: When the electronic device determines that the position of the first sound generation target in the video of the first video clip is the first position, the step of outputting the first audio through the first speaker and outputting the first background sound through the fourth speaker by the electronic device, the method according to any one of claims 1 to 9.
11. The method according to any one of claims 1 to 10, wherein position information of the plurality of speakers in the electronic device is different.
12. The method according to any one of claims 1 to 11, wherein the type of the first sound generation target is any one of a person, an animal, an object, and a landscape.
13. The method according to any one of claims 1 to 12, wherein the type of the first audio is any one of a human voice, an animal sound, an environmental sound, a music sound, and an object sound.
14. An audio playback method, comprising: starting, by an electronic device, playback of a first video clip, wherein an image of the first video clip includes a first sound generation target; acquiring, by the electronic device, first audio emitted by the first sound generation target from audio data of the first video clip; when the electronic device determines that a position of the first sound generation target in the image of the first video clip is a first position at a first moment, outputting, by the electronic device, the first audio via a first audio output device; acquiring, by the electronic device, second audio emitted by the first sound generation target from the audio data of the first video clip; when the electronic device determines that the position of the first sound generation target in the image of the first video clip is a second position at a second moment, outputting, by the electronic device, the second audio via a second audio output device, wherein the first moment is different from the second moment, the first position is different from the second position, and the first audio output device is different from the second audio output device.
15. The method according to claim 14, wherein the type of the first audio output device is any one of a sound box, earphones, a power amplifier, a multimedia console, and an audio adapter.
16. An electronic device, comprising one or more processors and one or more memories, the one or more memories being coupled to the one or more processors, the one or more memories being configured to store computer program code, the computer program code comprising computer instructions, the one or more processors being capable of calling the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 13, 14 and 15.
17. A computer-readable storage medium storing instructions which, when executed on an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 13, 14 and 15.
18. A computer program product which, when executed by an electronic device, enables the electronic device to execute the method according to any one of claims 1 to 13, 14 and 15.
19. A chip or chip system comprising a processing circuit and an interface circuit, the interface circuit being configured to receive code instructions and transmit the code instructions to the processing circuit, the processing circuit being configured to execute the code instructions to execute the method according to any one of claims 1 to 13, 14 and 15.
Citation Information
Patent Citations
Method, apparatus and device for achieving co-location of voices and images and medium
CN109194999A
Method and apparatus for sound object following
CN111666802A
Control method and device for terminal loudspeaker, and computer readable storage medium
EP3737087A1
Device and method for controlling sound, data structure of stream, and stream generator
JP2010206265A
Sound image play method and apparatus
US20160065791A1