A video playing method and an electronic device
Patent Information
- Application Number
- CN202411686426.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-11-21
AI Technical Summary
但是,存在电子设备所构建的声场与用户的观看位置不匹配的问题
[0022] In understanding, the beneficial effects that can be achieved by the electronic device of any possible design of the second aspect, the chip system of the third aspect, the computer-readable storage medium of the fourth aspect, and the computer program product of the fifth aspect can be referred to as the beneficial effects of the first aspect and any possible design, which will not be repeated here.
Smart Images

Figure CN120434452B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a video playback method and an electronic device. Background Technology
[0002] With the upgrading of electronic devices, 3D audio data technology has become a popular audio data technology. 3D audio data technology can bring users an immersive audio experience.
[0003] Electronic devices can play 3D audio data through speakers while displaying videos. However, there is a problem that the sound field created by the electronic device does not match the user's viewing position. This results in a poor listening experience, and the electronic device cannot provide a truly immersive audio experience. Summary of the Invention
[0004] This application provides a video playback method and an electronic device that can provide users with an immersive sound experience.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0006] Firstly, a video playback method is provided, applied to an electronic device. The method includes: playing a first video stream, the first video stream including a first audio stream, the first audio stream including audio data corresponding to M sound sources; wherein M is an integer greater than or equal to 1; the sound field of the first audio stream is matched with a first relative position between the M sound sources and the user; playing a second video stream, the second video stream including a corresponding second audio stream, the second audio stream including audio data corresponding to P sound sources; wherein P is an integer greater than or equal to 1; the sound field of the second audio stream is matched with a second relative position between the P sound sources and the user; wherein the first relative position and the second relative position are different; the sound field of the first audio stream is different from the sound field of the second audio stream.
[0007] This application can construct a sound field that matches the relative position of the sound source and the user in a video stream. Matching can be understood as the electronic device constructing a user-centered sound field based on the relative position between the sound source and the user, such as a first relative position or a second relative position. When the first and second relative positions are different, the sound field of the first audio stream constructed by the electronic device will differ from that of the second audio stream. Therefore, regardless of the user's position, the user will receive a user-centered sound field, resulting in optimal listening experience.
[0008] Optionally, the sound field of the first audio stream differs from that of the second audio stream in the following ways: the sound field of the first audio stream and the sound field of the second audio stream have different directions and / or the sound pressure of the sound field of the first audio stream and the sound field of the second audio stream are different at the same location.
[0009] In one possible implementation of the first aspect, before playing the first video stream, the electronic device, for each of the M sound sources, determines a first relative position between the sound source and the user based on a first position of the sound source in a first space and a second position of the user in the first space; and renders the audio data corresponding to the sound source based on the first relative position of the sound source and the user. The first audio stream is a mixed audio stream composed of multiple audio data. The electronic device renders the audio data corresponding to the sound source based on the first relative position of the sound source in the first space and the user, and after rendering, the sound field of the audio data matches the first relative position. After executing this process one by one, the sound fields of the audio data of the M sound sources match the first relative positions, therefore, the sound field of the first audio stream matches the first relative positions.
[0010] In one possible implementation of the first aspect, before playing the second video stream, the electronic device, for each of the P sound sources, determines a second relative position between the sound source and the user based on the first position of the sound source in the first space and the second position of the user in the first space; and renders the audio data corresponding to the sound source based on the second relative position of the sound source and the user. The second audio stream is a mixed audio composed of multiple audio data. The electronic device renders the audio data corresponding to the sound source based on the second relative position of the sound source in the first space and the user, and after rendering, the sound field of the audio data matches the second relative position. After executing this process one by one, the sound fields of the audio data of the P sound sources match the second relative positions, therefore, the sound field of the second audio stream matches the second relative positions.
[0011] In one possible implementation of the first aspect, the first space is a space constructed according to the camera coordinate system, or a space constructed according to the world coordinate system. The method further includes: for each of the M sound sources, obtaining the first position of the sound source in the first space; and obtaining the second position of the user in the first space. The first space can be a three-dimensional space, and based on the relative positions of the sound sources and the user in the three-dimensional space, a user-centered three-dimensional stereo sound effect is constructed to improve the user's listening experience.
[0012] In one possible implementation of the first aspect, the electronic device may determine the first position of the sound source in the first space based on the following method: the electronic device acquires the position of the sound source in the image space; and determines the first position of the sound source based on the position of the sound source in the image space.
[0013] In one possible implementation of the first aspect, the electronic device may determine the user's second position in the first space by: acquiring a first image, the first image including an image of the user; acquiring the user's position in the image space based on the first image; and determining the user's second position based on the user's position in the image space.
[0014] In this application, the user and the sound source are placed in a consistent image space, which allows for the establishment of a precise relative positional relationship between the sound source and the user in the image space. Then, based on this relative positional relationship, the sound source and the user are mapped to a first space. In this first space, the relative positional relationship between the sound source and the user remains unchanged and is correct, thus enabling more accurate audio rendering.
[0015] In one possible implementation of the first aspect, before playing the first video stream, the electronic device may first separate a first image group and a first audio stream from the first video stream. Then, the electronic device determines the correspondence between M audio data points and M sound sources based on the first image group and the first audio stream.
[0016] In one possible implementation of the first aspect, the electronic device can determine the correspondence between M audio data and M sound sources based on the following method: identifying M sound sources according to a first image group; identifying M audio data according to a first audio stream; and establishing a one-to-one correspondence between M audio data and M sound sources based on the matching degree between the M sound sources and the M audio data.
[0017] In one possible implementation of the first aspect, the first relative position includes distance, azimuth, and pitch. The electronic device can render audio data corresponding to the sound source based on the first relative position between the sound source and the user by: determining a head transfer function based on the first relative position between the sound source and the user; and rendering the audio data based on the head transfer function.
[0018] In a second aspect, an electronic device is provided, comprising: a memory, a camera, and one or more processors; the camera, the memory, and the processors are coupled; wherein the memory is used to store computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any of the first aspects.
[0019] Thirdly, a chip system is provided that can be applied to an electronic device including memory. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via lines. The interface circuits are used to receive signals from the aforementioned memory and send the signals to the processors, the signals including computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device performs a method as described in the first aspect and any of its possible design embodiments.
[0020] Fourthly, a computer-readable storage medium is provided, including computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any of the first aspects.
[0021] Fifthly, a computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of any of the methods in the first aspect.
[0022] In understanding, the beneficial effects that can be achieved by the electronic device of any possible design of the second aspect, the chip system of the third aspect, the computer-readable storage medium of the fourth aspect, and the computer program product of the fifth aspect can be referred to as the beneficial effects of the first aspect and any possible design, which will not be repeated here. Attached Figure Description
[0023] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;
[0024] Figure 2 A schematic diagram of the software architecture of an electronic device provided in an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the speaker distribution of an electronic device provided in an embodiment of this application;
[0026] Figure 4 A flowchart illustrating a video playback method provided in an embodiment of this application;
[0027] Figure 5 A schematic diagram of a coordinate system provided for an embodiment of this application;
[0028] Figure 6 A schematic diagram illustrating a method for establishing a correspondence between a sound source and audio data, provided in an embodiment of this application;
[0029] Figure 7 A schematic diagram of coordinate mapping provided for an embodiment of this application;
[0030] Figure 8 This is another schematic diagram of coordinate mapping provided in an embodiment of this application;
[0031] Figure 9 This is another schematic diagram of coordinate mapping provided in an embodiment of this application;
[0032] Figure 10 A schematic diagram illustrating a video playback scenario provided in an embodiment of this application;
[0033] Figure 11 A schematic diagram illustrating another method for establishing a correspondence between a sound source and audio data, provided in an embodiment of this application;
[0034] Figure 12 This is an illustration of a method for rendering audio data provided in an embodiment of this application. Detailed Implementation
[0035] With the upgrading of electronic devices, 3D audio data technology has become a popular audio data technology. 3D audio data technology can bring users an immersive audio experience. It can combine the user's three degrees of freedom (3DoF) or six degrees of freedom (6DoF) to construct a user-centric sound field. The sound field can be understood as the distribution and propagation of sound in space.
[0036] Currently, during audio data playback, the relative position between the electronic device and the user can change or remain unchanged. For example, during audio data playback, the relative position between the user and the electronic device remains constant. In this embodiment, the electronic device can be a wearable electronic device, such as headphones, a virtual reality (VR) headset, extended reality (XR) glasses, or augmented reality (AR) glasses. Alternatively, during audio data playback, the relative position between the user and the electronic device may change. In this embodiment, the electronic device can be a non-wearable electronic device, such as a tablet computer, a personal computer (PC), a mobile phone, or a large screen. In this embodiment, the electronic device can play audio data aloud through a speaker.
[0037] In some embodiments, wearable electronic devices can construct a user-centric sound field based on inertial measurement unit (IMU) head tracking technology combined with the user's 3DoF or 6DoF.
[0038] While non-wearable electronic devices can construct three-dimensional stereo sound effects based on 3D audio data technology, a problem exists where the sound field constructed by the electronic device does not match the user's viewing position, resulting in a poor listening experience. For example, when watching videos on an electronic device, a user may not always be in the same position and can adjust their viewing position at any time. In conventional technologies, the sound field constructed by electronic devices is fixed. For example, the electronic device has a preset optimal viewing position, and the sound field is constructed based on this optimal viewing position. However, when a user watches videos on an electronic device, their viewing position may not be in the optimal viewing position, and their viewing position may change. This leads to a mismatch between the sound field constructed by the electronic device and the user's viewing position, resulting in a poor listening experience.
[0039] Therefore, this application provides a video playback method and an electronic device. The electronic device includes a display screen and is capable of constructing a sound field matching the relative position of a sound source in a video stream and the user. Matching can be understood as the electronic device constructing a user-centered sound field based on the relative position between the sound source and the user, such as a first relative position or a second relative position. When the first and second relative positions are different, the sound field of the first audio stream constructed by the electronic device differs from that of the second audio stream. Therefore, regardless of the user's location, the user receives a user-centered sound field, resulting in optimal listening experience. In other words, in this application embodiment, the electronic device can construct a sound field matching the user's position, enhancing the user's immersive audio experience.
[0040] The method provided in this application can be applied to electronic devices with audio and video playback capabilities. These electronic devices may include mobile phones, tablets, laptops, netbooks, personal digital assistants (PDAs), in-vehicle devices, etc., and this application does not impose any limitations on this. In this application, the aforementioned electronic device is an electronic device capable of running an operating system and installing applications. Optionally, the operating system running on the electronic device may be Android. system, system, Systems, etc.
[0041] Take a mobile phone as an example. Figure 1 As shown, the electronic device 100 may include: a processor 110, a memory 120, a universal serial bus (USB) interface 130, a power management module 140, antennas such as antenna 1 and antenna 2, a communication module 150, a display screen 160, an audio module 170, a camera 180, a sensor module 190, etc.
[0042] Processor 110 may include one or more processing units, such as: application processor (AP), central processing unit, modem processor, graphics processing unit (GPU), ISP, controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors. The controller may be the nerve center and command center of electronic device 100. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0043] In this embodiment, the processor 110 can control the camera 180 to periodically acquire a first image, which may include an image of the user. The processor 110 can also perform audio-video separation on the video to be played, obtaining an image group (e.g., a target image group) and an audio stream (e.g., a target audio stream). The processor 110 can also obtain M sound sources and M audio data corresponding to each of the M sound sources based on the image group and the audio stream. The processor 110 can also map the M sound sources into a first space. The processor 110 can also map the user into the first space based on the first image, so that the M sound sources and the user are in the same space (e.g., all in the first space). The processor 110 can render the audio data corresponding to each sound source based on the relative position of each of the M sound sources in the first space to the user. After rendering, the processor 110 can control the audio module 170 to play the rendered audio data. Here, M is an integer greater than or equal to 1.
[0044] The memory 120 can be used to store computer executable program code, which includes instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the memory 120. The memory 120 may include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function (such as sound playback, interface display, etc.). The data storage area may store data created during the use of the electronic device (such as notification messages). Furthermore, the memory 120 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0045] In this embodiment, the memory 120 stores computer-executable program code, which includes instructions. The processor 110 executes a video playback method provided in this embodiment by running the instructions stored in the memory 120.
[0046] The power management module 140 is used to connect the battery to the processor 110. The power management module 140 receives battery and / or power input to power the processor 110, memory 120, communication module 150, display screen 160, and camera 180, etc. The power management module 140 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 140 may also be located within the processor 110.
[0047] The communication module 150 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR). The communication module 150 can be one or more devices integrating at least one communication processing module. The communication module 150 receives electromagnetic waves via an antenna, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to the processor 110. The communication module 150 can also receive signals to be transmitted from the processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via the antenna.
[0048] In some embodiments, the antenna of the electronic device 100 is coupled to the communication module 150, enabling the electronic device 100 to communicate with networks and other devices via wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BitTorrent, Global Navigation Satellite System (GNSS), WLAN, NFC, FM, and / or IR technologies. The GNSS may include Global Positioning System (GPS), BeiDou Navigation Satellite System (BDS), GLONASS, and / or Galileo.
[0049] Electronic device 100 implements display functions through a GPU, a display screen 160, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 160 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0050] The display screen 160 is used to display images, videos, etc. The display screen 160 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.
[0051] Electronic device 100 can achieve shooting and recording functions through ISP, camera 180, video codec, GPU, display screen 160 and application processor.
[0052] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0053] The audio module 170 may include a speaker 170A. The speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or receive hands-free calls through the speaker 170A. In this embodiment, the electronic device 100 can play rendered audio data through the speaker 170A.
[0054] The camera 180 is used to capture still images or videos. An object passes through the lens, generating an optical image that is projected onto a photosensitive element. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP (Image Signal Processor) for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP (Digital Signal Processor) for further processing. The DSP converts the digital image signal into standard image signals in formats such as RGB and YUV.
[0055] The sensor module 190 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, and bone conduction sensors, etc.
[0056] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0057] Taking the aforementioned electronic device 100 as an example, which is a tablet computer, the software system of the electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses a layered architecture... Taking the system as an example, the software structure of electronic device 100 is illustrated.
[0058] Figure 2 This is a software structure block diagram of the electronic device 100 according to an embodiment of this application.
[0059] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the system libraries, and the kernel layer.
[0060] The application layer can include a series of application packages.
[0061] like Figure 2 As shown, the application package may include a first application, a gallery, a caller ID app, a map app, a navigation app, a WLAN app, a Bluetooth app, a music app, a video app, a text messaging app, a social networking app, and other applications. The first application may be a video application. The video application may be a system-installed application or a third-party application downloaded and installed by the user on the electronic device 100. The video application is used to play videos, which may be videos stored on the electronic device 100 or streaming videos.
[0062] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0063] like Figure 2As shown, the application framework layer may include a view system, resource manager, input system, audio / video separation module, matching module, coordinate transformation module, and sound effect rendering module, etc.
[0064] The input system is used to monitor the phone's input modules (such as touchscreen drivers) and convert the parameters input by the input modules into usable events, which are then passed to the relevant upper-layer modules. For example, the input system is used to monitor the phone's touchscreen through the touchscreen driver and convert the touch parameters generated by the touchscreen input into usable events, which are then passed to the upper-layer APP.
[0065] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build the display interface of an application.
[0066] The audio-video separation module is used to separate the audio and video of the video to be played, such as the target video stream, to obtain a target image group including multiple image frames and a target audio stream corresponding to the target video stream.
[0067] The matching module is used to acquire multiple sound sources based on the target image group, acquire multiple audio data based on the target audio stream, and establish a matching relationship between the sound sources and the audio data.
[0068] The coordinate transformation module is used to map the location of the sound source to the first space and the location of the user to the first space.
[0069] The sound rendering module is used to render the audio data corresponding to the sound source based on the relative position of the sound source and the user in the first space.
[0070] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0071] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0072] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0073] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0074] A 2D graphics engine is a graphics engine for 2D drawing.
[0075] The kernel layer is the layer between hardware and software. The kernel layer can contain touchscreen drivers, display drivers, camera drivers, sensor drivers, etc.
[0076] The following section uses a tablet computer as an example to illustrate this solution, along with accompanying diagrams.
[0077] A tablet computer may include at least one speaker. For example, a tablet computer may include one speaker, which can be located on the left, right, top, or bottom of the tablet. As another example, a tablet computer may include four speakers. These four speakers can be located on the left and right sides, top, and bottom of the tablet. Or, as... Figure 3 As shown, each set of two speakers is positioned as a group, with one group located on the left side of the tablet and the other on the right. Alternatively, one group of speakers can be located at the top of the tablet, and the other at the bottom. It should be understood that the tablet may include more than […]. Figure 3 The examples shown may have more or fewer speakers, and the speakers can be placed anywhere on the tablet computer as needed. This application does not impose any specific limitations on this.
[0078] The tablet computer may include a first application for playing videos. The video may be stored on the tablet computer, such as a video taken by the user or a video downloaded by the user. Optionally, the video may also be a streaming video. Streaming video is online video content that is transmitted in real time over a computer network and allows users to watch it instantly without first downloading the entire video file to a local device. The first application may be a system application that comes pre-installed on the tablet computer, or a third-party application downloaded and installed on the tablet computer by the user. This application does not specifically limit this aspect.
[0079] The tablet computer can respond to the first operation by periodically executing the video playback method provided in this application embodiment during video playback. In other words, the first operation is used to trigger the tablet computer to periodically execute the video playback method provided in this application embodiment.
[0080] For example, the first operation could be an operation to trigger video playback. This operation could be, for example, an action on the video within a first application, such as a click. Alternatively, it could be an action on a video playback control within the first application, such as a click. Another example is speaking a voice command near the tablet's microphone, such as "Play xx".
[0081] For example, the first operation could be turning on a preset switch. For instance, a tablet displays a video playback interface that includes a preset switch and a currently playing video. The preset switch can be in a closed state. In response to an operation on the preset switch, such as a click, the preset switch switches from a closed state to an open state. In response to the preset switch being turned on, the tablet can periodically execute the video playback method provided in the application.
[0082] Specifically, the tablet computer can respond to the first operation by acquiring the video stream 1 to be played. The tablet computer can execute the video playback method provided in this application embodiment on the video stream 1 and play the processed video stream 1. During the playback of the processed video stream 1, the tablet computer acquires the video stream 2 to be played, executes the video playback method provided in this application embodiment on the video stream 2, and plays the processed video stream 2 after ending the playback of the video stream 1. This process continues until the video playback method provided in this application embodiment is executed on the video stream i and the processed video stream i is played. Here, video stream i is the last segment of the currently playing or to-be-played video (hereinafter collectively referred to as the first video). In other words, the tablet computer can segment the first video, dividing it into multiple video streams, and process and play the processed video streams sequentially according to the time order.
[0083] In this context, video stream 1, video stream 2, ..., and video stream i are video streams of a preset duration. Alternatively, video stream 1, video stream 2, ..., and video stream i-1 are video streams of a preset duration, where the length of video stream i is less than the preset duration. For example, if the preset duration is 10 seconds and the first video is 1000 seconds, then video stream 1, video stream 2, ..., and the 100th video stream are all 10-second video streams. Another example: if the preset duration is 10 seconds and the first video is 1005 seconds, then video stream 1, video stream 2, ..., and the 100th video stream are all 10-second video streams, and the 101st video stream is a 5-second video stream. It should be understood that the preset duration can also be 20 milliseconds, 1 second, 5 seconds, 15 seconds, etc., and this embodiment does not specifically limit this.
[0084] Video stream 1, video stream 2, ..., and video stream i are consecutive video streams in the first video. That is, video stream 1, video stream 2, ..., and video stream i are continuous. For example, video stream 1 is the video stream from second 0 to second 10 in the first video, video stream 2 is the video stream from second 10 to second 20 in the first video, and the third video stream is the video stream from second 20 to third 30 in the first video.
[0085] In the method provided in this application embodiment, a tablet computer plays a first video stream, which includes a first audio stream. The first audio stream includes audio data corresponding to M sound sources; where M is an integer greater than or equal to 1; the sound field of the first audio stream matches the first relative position of the M sound sources and the user. The tablet computer plays a second video stream, which includes a corresponding second audio stream. The second audio stream includes audio data corresponding to P sound sources; where P is an integer greater than or equal to 1; the sound field of the second audio stream matches the second relative position of the P sound sources and the user. M and P can be equal or unequal, and this application embodiment does not specifically limit this. The first video stream is played before the second video stream. Optionally, the first video stream and the second video stream can be continuous. For example, the first video stream is video stream 1, and the second video stream is video stream 2. Another example is that the first video stream is video stream 2, and the second video stream can be video stream 3. Alternatively, the first video stream and the second video stream can be discontinuous. For example, the first video stream is video stream 1, and the second video stream is video stream 4. For example, the first video stream is video stream 2, and the second video stream could be video stream 7.
[0086] The sound field of the first audio stream corresponding to the first video stream matches the first relative positions of M sound sources and the user. For example, for each of the M sound sources, the sound field of the audio data corresponding to that sound source matches the first relative position of that sound source and the user. The first audio stream is mixed data composed of multiple audio data, and the sound field of the first audio stream composed of the M audio data matches the first relative positions of the M sound sources and the user. The sound field of the second audio stream corresponding to the second video stream matches the second relative positions of P sound sources and the user. For example, for each of the P sound sources, the sound field of the audio data corresponding to that sound source matches the second relative position of that sound source and the user. The second audio stream is mixed data composed of multiple audio data, and the sound field of the second audio stream composed of the P audio data matches the second relative positions of the P sound sources and the user.
[0087] The first relative position and the second relative position are different, and the sound field of the first audio stream, such as sound field 1, is different from the sound field of the second audio stream, such as sound field 2. The differences between sound field 1 and sound field 2 include: different directions, different sound pressure and phase at the same location, and inconsistent sound pressure distribution in the same space. For example, acoustic instruments can be used to measure the direction, sound pressure and phase at the same location, and sound pressure distribution in the same space of these two sound fields. Such acoustic instruments can be, for example, sound level meters, spectrum analyzers, sound field scanners, etc.
[0088] The above matching can be understood as the tablet computer constructing a sound field centered on the user based on the relative position between the sound source and the user, such as a first relative position or a second relative position. Therefore, regardless of the user's location, the sound field constructed by the tablet computer is always centered on the user, resulting in the optimal listening experience for the user.
[0089] The following section uses either the first video stream or the second video stream as an example to illustrate this solution. This arbitrary video stream can be referred to as the target video stream.
[0090] If the target video stream is the first video stream, the tablet computer can respond and execute the first operation as follows: Figure 4 The method shown.
[0091] If the target video stream is not the first video stream, the tablet can perform actions such as... while playing the previous video stream of the target video stream. Figure 4 The method shown. For example, a tablet computer can execute the following at a time (T+A) after a preset duration A, from the time T when the previous video stream started playing. Figure 4 The method shown above. The preset duration A is less than the duration of the previous video stream. This allows for processing time, ensuring the continuity of video playback. For example, the tablet computer can execute the following at time T2, before the time T1 when the first video stream is played: Figure 4 The method shown. The time difference between T2 and T1 is the reserved processing time.
[0092] Figure 4 This is a flowchart illustrating a video playback method provided in an embodiment of this application. Figure 4 The methods shown may include:
[0093] S1, the tablet computer captures the first image.
[0094] As an example, the tablet computer includes a front-facing camera, through which the tablet computer captures the first image.
[0095] Since the tablet uses its front-facing camera to capture the first image, this first image can include the user's image. The user is the person using the tablet to watch the video.
[0096] It should be understood that the collection, storage, use, processing, transmission, provision, and disclosure of user personal images involved in the technical solutions disclosed in this application all comply with relevant laws and regulations and do not violate public order and good morals. It should be noted that the personal images used in the technical solutions of this application are subject to individual consent, including but not limited to notifying and reminding users to read the relevant user agreement (notification) and sign the agreement (authorization) which includes the authorization of relevant user information before the user uses the video processing function.
[0097] S2, the tablet computer acquires information about the sound source in the target video stream.
[0098] The target video stream can be a segment of a video to be played. An introduction to target video streams can be found in the previous text and will not be repeated here.
[0099] A sound source is an object emitting sound in the target video stream. A sound source can be a person such as a singer, a musical instrument such as a drum kit or guitar, or an animal such as a bird. The target video stream can include at least one sound source.
[0100] The information about the sound source includes its location data. This location data can be the sound source's position in image space. Image space can be a two-dimensional space constructed using pixel coordinates, or it can be a two-dimensional space constructed using image coordinates. The sound source's position in image space can be either pixel coordinates or image coordinates. It should be understood that images are composed of pixels, and image coordinates are the positions of pixels within an image, specifically the coordinates of the pixels in an image coordinate system. First, let's introduce image coordinate systems and pixel coordinate systems. For example... Figure 5 As shown, the origin of the pixel coordinate system is located at the top left corner of the image. The horizontal axis to the right of the origin is the u-axis, and the vertical axis downwards is the v-axis. Pixel coordinates can be (u, v). Here, u represents the horizontal position of the pixel, and v represents the vertical position of the pixel. The origin of the image coordinate system is located at the center of the image, as shown below. Figure 5 Point O' in the image. The x-axis extends horizontally to the right from the origin, and the y-axis extends vertically downwards from the origin. Image coordinates can be, for example, (x, y), where x represents the horizontal position of the pixel, and y represents the vertical position of the pixel. From... Figure 5 As can be seen, when a tablet displays an image in full screen, the image coordinate system and the screen coordinate system can be the same. The origin of the screen coordinate system is located at the center of the screen, with the x-axis pointing horizontally to the right and the y-axis pointing vertically downwards. That is, when a tablet displays an image in full screen, the image space can still be a two-dimensional space constructed using the screen coordinate system.
[0101] The information about the sound source also includes the audio data of the sound source. The target video stream includes the corresponding audio stream, as described below as the target audio stream. This target audio stream can be a mixed audio stream composed of multiple audio data. For example, the target audio stream consists of the sound of a drum kit, a singer's voice, and a guitar playing.
[0102] The target audio stream may include at least one audio data point, and each audio data point corresponds to a sound source. For example, the target video stream includes Q sound sources and N audio data points. Here, Q is an integer greater than or equal to 1, and N is an integer greater than or equal to 1. Q and N may be equal or unequal. Q being unequal to N includes Q being less than N. When Q and N are equal, all Q sound sources have corresponding audio data points. When Q is less than N, the remaining NQ audio data points (excluding the Q audio data points corresponding to Q) can be considered background noise. The tablet computer may not process this background noise. When the target video stream is a first video stream, Q = M; when the target video stream is a second video stream, Q = P.
[0103] The tablet computer can identify sound sources and audio data in a target video stream and establish a one-to-one correspondence between sound sources and audio data. For example, the tablet computer can perform... Figure 6 The method shown achieves the above objective. Figure 6 The methods shown include:
[0104] The first step is for the tablet computer to separate the target video stream into audio and video, resulting in a target image group and a target audio stream.
[0105] The tablet computer can separate the video and audio of a target video stream, resulting in a set of images, such as a target image group, and an audio stream, such as a target audio stream. The target image group includes at least one frame. The length of the target audio stream is the same as the length of the target video stream, for example, both being 10 seconds. The tablet computer can separate the target video stream using existing audio-video separation techniques; known techniques can be referenced here, and will not be elaborated further in this application. Optionally, the target image group can be a video stream that does not include audio data; that is, the target image group is obtained by removing the audio data from the target video stream.
[0106] When the target video stream is a first video stream, the target image group is a first image group, and the target audio stream is a first audio stream. When the target video stream is a second video stream, the target image group is a second image group, and the target audio stream is a second audio stream. That is, the tablet computer can separate the first image group and the first audio stream from the first video stream, and determine the correspondence between M audio data points and M sound sources based on the first image group and the first audio stream. The tablet computer can separate the second image group and the second audio stream from the second video stream, and determine the correspondence between P audio data points and P sound sources based on the second image group and the second audio stream. Specifically, the tablet computer can perform steps two through four to achieve the above objectives.
[0107] The second step is for the tablet computer to identify the sound sources in the target video stream.
[0108] The tablet computer can identify sound sources in a target video stream and pinpoint their image coordinates. Two possible methods are described below.
[0109] The first method involves a tablet computer identifying sound sources in a target video stream and obtaining their image coordinates. When the target video stream contains multiple sound sources, the tablet computer can identify the image coordinates of each sound source within the target image. The target image is the image that includes the sound source. The target video stream includes the target image. For example, the tablet computer can identify sound sources in the target video stream and obtain their image coordinates based on a target detection algorithm. For instance, the tablet computer can represent the image coordinates of the sound source using the image coordinates of the smallest bounding box surrounding the sound source within the target image. The bounding box can be the smallest rectangle surrounding the sound source, which is located within that rectangle. Of course, the bounding box can also be various shapes such as circles or ellipses. The image coordinates of the bounding box can be the image coordinates of the center point of the bounding box within the target image, or they can be the image coordinates of the top-left and bottom-right vertices of the bounding box within the target image.
[0110] The number of target images can be one or more. When there are multiple target images, for a target image containing the same sound source, the tablet computer can acquire multiple image coordinates of that sound source. If multiple image coordinates are consistent, it indicates that the sound source has not moved. In this case, the tablet computer can use any one of the image coordinates as the image coordinates of the sound source. If the distance between any two image coordinates is less than a threshold, it indicates that the sound source has moved, but the range of movement is small. In this case, the tablet computer can use any one of the image coordinates as the image coordinates of the sound source, or it can use the center position of the region composed of the multiple image coordinates as the position of the sound source. To avoid the sound source moving and causing the constructed sound field to be inconsistent with reality, in this embodiment, the duration of the target video stream can be small; for example, the target video stream can include 3 frames, i.e., the length of the target video stream can be 0.05 seconds.
[0111] The second method involves the tablet computer using a first neural network model to obtain the image coordinates of the sound sources in the target video stream. After inputting the first image set into the first neural network model, the model can output the image coordinates of the sound sources in the target video stream. For example, after inputting the first image set into the first neural network model, it can output detection values for one or more sound sources. For instance, the detection values could be [x,y,w,h]. Here, [x,y,w,h] represents the image coordinates of the sound source, where (x,y) are the coordinates of the center point of the bounding box surrounding the sound source, and (w,h) are the width and height of the bounding box. Another example is that the detection value could be [x,y]. Here, (x,y) are the coordinates of the center point of the bounding box surrounding the sound source.
[0112] Optionally, the first neural network model can also output the type of sound source. For example, the detection value could be [x, y, w, h, l]. Here, l indicates the type of sound source. This type could be a person, a musical instrument such as a drum kit, guitar, violin, or an animal.
[0113] Optionally, the first neural network model can also output image features of the sound source. These image features include color features, shape features, spatial relationship features, and texture features.
[0114] Optionally, the first neural network model may include convolutional layers, residual blocks, and fully connected layers. The convolutional layers are used to perform convolutional processing on the image to extract low-dimensional image features. The residual blocks are used to further process the low-dimensional image features to obtain higher-dimensional image features. For example, after convolutional and residual processing, the first neural network model can obtain image features corresponding to multiple sound sources. For example, the first neural network model can obtain Q image features, which correspond one-to-one with Q sound sources. The fully connected layers may include a first fully connected layer, or a first fully connected layer and a second fully connected layer. The first fully connected layer is used to output the image coordinates of the Q sound sources based on the Q image features. The second fully connected layer is used to output the types of the Q sound sources based on the Q image features. For example, the first neural network model may be a residual network (ResNet) model.
[0115] The third step involves the tablet extracting multiple audio data points from the target audio stream.
[0116] The tablet computer can perform audio separation on a target audio stream, resulting in multiple audio data sets. For example, the target audio stream might be mixed audio data, consisting of multiple audio data sets. For instance, the target audio stream could be composed of drum sounds, vocals, and guitar sounds. After separating the target audio stream, the tablet computer can obtain three types of audio data, such as audio data a, audio data b, and audio data c. Audio data a is the drum sound, audio data b is the vocals, and audio data c is the guitar sound.
[0117] Tablet computers can perform audio separation on target audio streams based on audio separation technology.
[0118] In some embodiments, the tablet computer can separate the target audio stream based on the amplitude and / or frequency of the corresponding sound signal. Since the target audio stream is a time-domain signal, to facilitate audio separation, the tablet computer can first convert the target audio stream into a frequency-domain signal. For example, the tablet computer performs a Fourier transform on the target audio stream, converting it from a time-domain signal to a frequency-domain signal. After converting the target audio stream from a time-domain signal to a frequency-domain signal, the tablet computer can obtain the spectral information of the target audio stream. This spectral information is used to describe the correspondence between frequency and amplitude in the target audio stream. Then, the tablet computer inputs the spectral information of the target audio stream into a second neural network model. The second neural network model can separate the target audio stream into multiple audio data based on the spectral information of the target audio stream. This second neural network model can be, for example, an audio U-NET model, a Tensorflow audio separation model, etc.
[0119] In other embodiments, the tablet computer can directly input the target audio stream into a second neural network model to obtain multiple audio data. The second neural network model has the ability to separate multiple audio tracks, or audio data, from mixed audio data such as the target audio stream.
[0120] Optionally, the tablet computer can also identify the sound source type corresponding to multiple audio data. For example, the second neural network model can also output the sound source type corresponding to the audio data, that is, the type of sound source emitting the audio data. For example, the second neural network model can output that the sound source type of audio data a is a drum kit, the sound source type of audio data b is a person, and the sound source type of audio data c is a guitar. Among them, the person can be further divided into female and male. It can also be further divided into children, middle-aged and elderly, etc.
[0121] Optionally, the second neural network model may include an encoder, a decoder, and an output layer. The encoder includes multiple convolutional layers and downsampling layers, and the decoder includes multiple transposed convolutional layers and upsampling layers. The encoder and decoder are used to perform convolutional processing on the target audio stream to obtain N audio features. The output layer is used to output N corresponding audio data based on the N audio features. Optionally, the output layer may also output the sound source type corresponding to the N audio data based on the N audio features.
[0122] Where N is greater than or equal to Q
[0123] Optionally, the second neural network model can also output audio features that correspond one-to-one with the audio data. Audio features may include, for example, time-domain features and frequency-domain features.
[0124] When the target video stream is the first video stream, the tablet computer identifies M sound sources based on the first image group. Based on the first audio stream, the tablet computer identifies N audio data points, where N is greater than or equal to M. Then, the tablet computer can perform a fourth step, establishing a one-to-one correspondence between the M audio data points and the M sound sources based on the matching degree between the M sound sources and the N audio data points. When the target video stream is the second video stream, the tablet computer identifies P sound sources based on the second image group. Based on the second audio stream, the tablet computer identifies N audio data points, where N is greater than or equal to P. Then, the tablet computer can perform a fourth step, establishing a one-to-one correspondence between the P audio data points and the P sound sources based on the matching degree between the P sound sources and the N audio data points.
[0125] The fourth step is for the tablet computer to establish a correspondence between the sound source and the audio data.
[0126] This correspondence can be a one-to-one one. The tablet computer establishes a correspondence between the sound source and the audio data; specifically, it can establish a correspondence between the image coordinates of the sound source and the audio data.
[0127] Once the tablet computer acquires the sound source type and the sound source type of the audio data, it can establish a correspondence between the sound source and the audio data based on this correspondence. For example, it can establish a correspondence between sound sources and audio data of the same type. For instance, if the sound source type of audio data 'a' is a drum kit, the tablet computer will establish a correspondence between audio data 'a' and a sound source of the type 'drum kit'. If the sound source type of audio data 'b' is a person, the tablet computer will establish a correspondence between audio data 'b' and a sound source of the type 'person'.
[0128] If the tablet computer does not obtain the type of the sound source or the type of the audio data, it can determine the matching degree between the sound source and the audio data, and establish a correspondence between the sound source and the audio data based on the matching degree. For example, for a given audio data, the tablet computer can calculate the matching degree between the audio data and each sound source. Then, the tablet computer establishes a correspondence between the sound sources with matching degrees greater than a threshold and the audio data. Alternatively, it can establish a correspondence between the sound source with the highest matching degree among multiple matching degrees and the audio data.
[0129] Specifically, the tablet computer can calculate the matching degree between the sound source and the audio data based on a third neural network model, and obtain the correspondence between the sound source and the audio data. The third neural network model can be, for example, an inner product network model. The input to the third neural network model can be the Q image features output by the first neural network model and the N audio features output by the second neural network model. The output of the third neural network model can be the inner product of each of the Q image features and each of the N audio features. The inner product indicates the matching degree between the image features and the audio features. Therefore, for an audio feature, the inner product of that audio feature and the Q image features can be obtained. The tablet computer can identify the one-to-one correspondence between the Q image features and the Q audio data based on this inner product. For example, the tablet computer can establish a correspondence between the image features corresponding to inner products greater than a threshold and the audio feature. Alternatively, it can establish a correspondence between the image features corresponding to the largest inner product among multiple inner products and the audio feature. Finally, the tablet computer can establish a correspondence between the sound source and the audio data based on the correspondence between the image features and the audio features. Optionally, the input to the third neural network model can be the Q image features output by the first neural network model and the N audio features output by the second neural network model, and the output of the third neural network model can be a one-to-one correspondence between the Q image features and the Q audio data.
[0130] For example, the third neural network model can calculate the inner product between the first audio feature and each of the Q image features, thus obtaining the inner product of the first audio feature and the Q image features. Then, the third neural network model can calculate the inner product between the second audio feature and each of the Q image features, thus obtaining the inner product of the second audio feature and the Q image features. And so on, the third neural network model can calculate the inner product between the Nth audio feature and each of the Q image features, thus obtaining the inner product of the Nth audio feature and the Q image features.
[0131] The following example uses any one of N audio features (referred to as the target audio feature) and one of Q image features (referred to as the target image feature) to illustrate how the third neural network model determines the inner product of an audio feature and an image feature.
[0132] For example, a tablet computer can calculate the inner product of the target audio features and the target image features based on the following formula.
[0133] The inner product is calculated as: ai × s + b0. Here, i represents the target image features, and s represents the target audio features. a and b0 are coefficients. a and b0 can be known and pre-stored in the tablet.
[0134] Optionally, before calculating the inner product, the third neural network model can perform dimensionality enhancement on the target audio features and target image features. After dimensionality enhancement, the target audio features and target image features have higher dimensions. For example, dimensionality enhancement includes feature multiplication, multinomial regression, etc.
[0135] Optionally, during dimensionality upscaling, K multidimensional features can be obtained based on the target audio features and target image features. Here, K is an integer greater than or equal to 2. For example, the features before dimensionality upscaling can be called one-dimensional features. Multiplying the target audio features yields two-dimensional target audio features. Multiplying the two-dimensional target audio features yields three-dimensional target audio features. Multiplying the target image features yields two-dimensional target image features. Multiplying the two-dimensional target image features yields three-dimensional target image features, and so on. The tablet computer can calculate the inner product of the i-dimensional target audio features and the i-dimensional target image features, where i takes values sequentially from [1, 2, ..., K]. Then, the tablet computer sums the K inner products to obtain the inner product between the target audio features and the target image features. For example, the tablet computer can calculate the inner product between the target audio features and the target image features based on the following formula.
[0136]
[0137] Among them, i KIt is a k-dimensional target image feature, s K It is a k-dimensional target audio feature. K These are the calculated coefficients corresponding to the dimensions, a K It is known that it can be pre-stored on a tablet.
[0138] In execution Figure 6 Following the method shown, the tablet computer can establish a correspondence between the sound source and the audio data. Then, the tablet computer can map the location of the sound source to three-dimensional space. That is, the tablet computer converts the image coordinates of the sound source into coordinates in three-dimensional space; for example, the tablet computer can execute S3.
[0139] S3, the tablet computer maps the location of the sound source into the first space.
[0140] The first space can be a three-dimensional space, and the coordinate system of this three-dimensional space can be, for example, a camera coordinate system. That is, the tablet computer transforms the image coordinates of the sound source into camera coordinates in the camera coordinate system. The camera coordinate system is a three-dimensional Cartesian coordinate system established with the camera's focus center as the origin and the optical axis as the Z-axis.
[0141] Tablet computers can convert the image coordinates of various sound sources into camera coordinates. For example, a tablet computer can implement the conversion from image coordinates to camera coordinates based on a camera model. The camera model provides the theoretical basis for coordinate transformations between image coordinate systems, normalized planar coordinate systems, and spatial coordinate systems. Below, we introduce a feasible method for converting image coordinates to camera coordinates.
[0142] like Figure 7 As shown, the camera coordinate system O-XYZ can be a Cartesian coordinate system. The origin O is established at the optical center of the camera, and the Z-axis points outward from the camera's optical axis. The XOY plane is perpendicular to the camera's optical axis, the X-axis is parallel to the x-axis in the image coordinate system, and the Y-axis is parallel to the y-axis in the image coordinate system. A point (a, b) in image space corresponds to points (A, B, Z) in three-dimensional space. Based on the principle of similar triangles, the image coordinates (a, b) and the camera coordinates (A, B, Z) have the following relationship:
[0143]
[0144] Where f is the focal length of the front-facing camera on the tablet. f is a fixed value. f can be pre-stored in the tablet, and the tablet can read this value from the preset storage location before use. Therefore, the tablet can convert image coordinates to camera coordinates based on the following formula.
[0145]
[0146] For example, if the image coordinates of the sound source are (a1, b1), and the corresponding camera coordinates are (A1, B1, Z1), then the image coordinates and camera coordinates satisfy the following relationship:
[0147]
[0148] Z1 is the distance between the sound source and the camera. The value of Z1 for the sound source can be determined using the distance between the user and the camera. For example, if the distance between the user and the camera is Z0, and the length of the user image in the first image is b0 in the y-direction and a0 in the x-direction, and the distance between the sound source and the camera is Z1, and the length of the sound source in the image frame is b1 in the y-direction and a1 in the x-direction, then a1 / a0 = b1 / b0 = Z1 / Z0. Let a1 / a0 = b1 / b0 = L, Z1 / Z0 = L, and Z1 = Z0L. Here, L can be called the mapping ratio.
[0149] In some embodiments, the value of L can be a preset value, which is pre-stored in the tablet computer.
[0150] In other embodiments, the tablet computer can determine the value of L based on the size of the user image in the first image and the size of the sound source in the image frame. For example, a1 / a0 = b1 / b0 = L. For example, when performing coordinate transformation for each sound source, the tablet computer can adaptively determine the value of L based on the size of each sound source and the size of the user image. Furthermore, to reduce computational load and improve processing efficiency, the tablet computer can determine the value of L using the size of the target sound source and the size of the user image, and directly use this value of L when performing coordinate transformation for other sound sources. This avoids multiple calculations of L and improves processing efficiency. The target sound source can be any one of the multiple sound sources in the first image group. Optionally, the type of the target sound source can be a person. Optionally, the target sound source can be a sound source located in the center of the display screen and of the type of a person. Figure 8 As shown, the length of the singer's sound source in the image frame can be b1, and the length of the user image in the first image can be b0. L = b1 / b0.
[0151] Z0 represents the distance between the user and the front-facing camera of the tablet. In some embodiments, this value may be a preset value, pre-stored in the tablet. For example, Z0 may be preset to 40cm or 30cm. In other embodiments, during video playback, the tablet can measure the distance between the user and the tablet in real time and use this distance as the distance between the user and the front-facing camera. For example, the tablet may include a ranging device that can be used to measure the distance between the user and the tablet in real time. The ranging device may be one or more of a distance sensor, an ultrasonic detector, and a depth camera.
[0152] because, Z1 = Z0L. Substituting the image coordinates (a1, b1), Z0, L, and f of the sound source into this formula, we can obtain the coordinates (A1, B1, Z1) of the sound source in the camera coordinate system. In this way, the tablet computer can obtain the coordinates (A1, B1, Z1) of the sound source in the camera coordinate system.
[0153] Optionally, the coordinate system in this three-dimensional space can be, for example, the world coordinate system. The world coordinate system is a coordinate system introduced to describe the position of a target in the real world. The coordinates of an object in the world coordinate system can be, for example, (Xw, Yw, Zw), where Xw, Yw, and Zw represent the distances of the object along the X, Y, and Z axes, respectively. The world coordinate system can also be called the geodetic coordinate system. For example, after a tablet computer converts the image coordinates of a sound source into camera coordinates in the camera coordinate system, the tablet computer can convert the camera coordinates of the sound source into world coordinates in the world coordinate system. The transformation between world coordinates and camera coordinates can be performed using rotation and translation matrices. Existing techniques for converting between camera coordinates and world coordinates can be referenced, and will not be elaborated upon here.
[0154] In this context, the three-dimensional space using the camera coordinate system as its coordinate system can be called virtual space, while the three-dimensional space using the world coordinate system as its coordinate system can be called real space. The position of a sound source in the first space can be called its first position. For each of the Q sound sources, the tablet computer can execute the method described above to determine the first position of each sound source in the first space based on its position in the image space. Specifically, for each of the Q sound sources, the tablet computer can execute the method described above to convert the image coordinates of each sound source into its first position in the first space. Optionally, if the position of the sound source in the image space is in pixel coordinates, the tablet computer can first convert the pixel coordinates into image coordinates. Then, the tablet computer can execute the method described above to convert the image coordinates of each of the Q sound sources into its first position in the first space.
[0155] S4, the tablet computer maps the user's position in the first image to the first space.
[0156] The tablet computer can determine the user's position in image space based on the first image. Image space can be a two-dimensional space constructed using pixel coordinates, or it can be a two-dimensional space constructed using image coordinates. The user's position in image space can be either pixel coordinates or image coordinates. (See reference...) Figure 4 The description of image space in S2 will not be repeated here.
[0157] Next, the tablet computer determines the user's second position in the first space based on the user's position in the image space. Taking the user's position in the image space as image coordinates as an example, in the case of a virtual three-dimensional space, the tablet computer can convert the user's image coordinates in the first image into camera coordinates. Specifically, the tablet computer can determine the user's image coordinates in the first image. For example, the tablet computer can represent the user's image coordinates using the image coordinates of the smallest bounding box surrounding the user's image. Another example is that the tablet computer can represent the user's image coordinates using the image coordinates of the smallest bounding box surrounding the user's head image. The smallest bounding box can be the smallest rectangle surrounding the user. Of course, the bounding box can also be various shapes such as circles and ellipses. The image coordinates of the aforementioned smallest bounding box can be the image coordinates of the center point of the bounding box, or it can also be the image coordinates of the top-left and bottom-right vertices of the bounding box.
[0158] After determining the user's image coordinates in the first image, the tablet can convert these coordinates into camera coordinates. For example, the user's image coordinates could be (x... 0, y0), where the user's camera coordinates could be, for example, (X0, Y0, Z0). Because, Substituting Z0, x0, and y0 into the above formula, we can obtain... So, In this way, the tablet computer can obtain the user's position in the camera coordinate system. This also allows the tablet computer to obtain the user's position in virtual space.
[0159] In a real-world 3D space, after acquiring the user's camera coordinates, the tablet can convert them into world coordinates. This allows the tablet to determine the user's location in real space.
[0160] After mapping, the sound source and the user are in the same space, such as the first space. Then, the tablet can render the target audio stream based on the relative positions of the sound source and the user. The first space is a space constructed according to the camera coordinate system, or, alternatively, a space constructed according to the world coordinate system.
[0161] The tablet computer first places the user and the sound source in a consistent image space, establishing a precise relative positional relationship between them. This image space can be the size of the tablet's screen. For example, the tablet displays the target video stream in full screen. The origin of the image coordinate system of the first image is aligned with the origin of the screen coordinate system, and the origins of the image coordinate systems of the frames in the target video stream are also aligned with the origin of the screen coordinate system; for example, the origins are both the midpoint of the screen. Therefore, the image space corresponding to the sound source is consistent with the image space corresponding to the user, and both are the size of the tablet's screen. Then, based on this relative positional relationship, the sound source and user are mapped into the first space. In this way, the relative positional relationship between the sound source and the user remains unchanged and correct within the first space, enabling more accurate audio rendering. Figure 9 As shown, the tablet computer first acquires a first image of the user, placing the user, drum kit, singer, and guitar in a consistent image space. The coordinate system of this image space has its origin at the center of the display screen. The tablet computer can map the drum kit, singer, and guitar from the target video stream into this first space. The tablet computer can also map the user into this first space based on the first image. Then, it renders the corresponding audio data based on the relative positions of the drum kit, singer, guitar, and user in this first space.
[0162] For each of the Q sound sources in the target audio stream: the tablet computer determines the relative position between the sound source and the user based on the first position of the sound source in the first space and the second position of the user in the first space. Based on the relative position between the sound source and the user, the tablet computer renders the audio data corresponding to the sound source. If the target video stream is a first video stream, the aforementioned relative position is the first relative position. If the target video stream is a second video stream, the aforementioned relative position is the second relative position. For example, the tablet computer can execute S5.
[0163] The S5 tablet renders the target audio stream and plays the target video stream based on the relative positions of the sound source and the user in the first space.
[0164] The rendering of the target audio stream can be understood as comprising Q audio data points. Specifically, for each of the Q sound sources, the following method is performed: The relative position between the sound source and the user is calculated based on the first position of the sound source in the first space and the second position of the user in the first space. Based on the relative position, the audio data of the sound source is rendered. Specifically, the tablet computer constructs a user-centric sound field based on the relative position between the sound source and the user, such as the first or second relative position. Since the target audio stream consists of Q audio data points, after rendering, the sound field of each audio data point is user-centric; therefore, the sound field of the target audio stream is also user-centric. Thus, regardless of the user's position, the user receives a user-centric sound field, resulting in optimal listening experience.
[0165] Specifically, the tablet computer performs the following method for each of the Q sound sources:
[0166] Based on the first location of the sound source and the second location of the user, a relative position, such as a first relative position or a second relative position, is determined. The relative position includes the distance from the sound source to the user, the azimuth angle θ, and the pitch angle φ. After determining the relative position, the head-related transfer functions (HRTFs) corresponding to the sound source are determined based on the relative position. During the playback of the target video stream, the tablet computer renders the audio data corresponding to each sound source based on the HRTFs, thus presenting user-centric 3D sound effects.
[0167] HRTF is an audio signal processing technology used to create 3D sound effects. HRTF uses spherical coordinates (r, θ, φ) to represent the position of the sound source. Here, r is the distance between the sound source and the user; θ represents the horizontal angle between the sound source and the user (azimuth); and φ represents the vertical angle between the sound source and the user (pitch). In the vertical plane, φ = -90° represents directly below, φ = 0° represents the horizontal plane, and φ = +90° represents directly above. In the horizontal plane, θ = 0° represents directly in front, θ = 90° represents directly to the right, θ = 180° represents directly behind, and θ = 270° represents directly to the left. In a free field, HRTF is defined as:
[0168]
[0169] H L H R The HRTFs are for the left ear and the right ear, respectively. P L P R These are the sound pressure levels generated by the sound source in the left and right ears, respectively; P0 is the sound pressure level at the user's location when the head is not present.
[0170] In a virtual space where the first space is a camera coordinate system, the tablet computer can determine the relative position of the sound source based on the camera coordinates of the sound source and the user's camera coordinates. In a real space where the first space is a world coordinate system, for a given sound source, the tablet computer can determine the relative position of the sound source based on the world coordinates of the sound source and the user's world coordinates.
[0171] After determining the relative position between the sound source and the user, the tablet computer can determine the HRTF corresponding to the sound source based on the relative position.
[0172] The tablet computer can retrieve the corresponding HRTF from the HRTF library based on its relative location. The HRTF library includes multiple HRTF groups corresponding to preset HRTF parameters. The HRTF library can be located on the tablet computer or on a cloud-based server; this application does not specifically limit this. An HRTF group includes H... L and H R For example, preset pitch angles may include [-40°, -20°, -10°, 0°, 10°, 25°, 40°, 60°], preset azimuth angles may include [-60°, -40°, -20°, 0°, 20°, 25°, 40°, 60°], and preset distances may include [10cm, 20cm, 30cm, 40cm, 50cm, 60cm, 70cm, 80cm, 90cm, 100cm]. For example, an HRTF library includes HRTF groups corresponding to a distance of 100cm, an azimuth angle of 0°, and a pitch angle of 10°. These HRTF groups include H... L and H R. .
[0173] The tablet computer can retrieve a group of HRTFs whose HRTF parameters match the relative position from the HRTF library as the corresponding HRTF for the sound source. Optionally, if some or all parameters in the relative position are inconsistent with the HRTF parameters of each HRTF in the HRTF library, the tablet computer can match the sound source with an HRTF whose HRTF parameters are close. Close means that the difference between the parameters is less than a threshold; for example, the difference between the pitch angle in the relative position and the pitch angle in the parameters of the matched HRTF is less than 5°.
[0174] After determining the HRTF corresponding to the sound source, the tablet computer can render the audio data corresponding to that sound source based on the HRTF. For example, the tablet computer will render the HRTF... L Convolve the audio data of the sound source to obtain the left channel data of the audio data, and then convert H... RThe right channel data of the audio data is obtained by convolving it with the audio data of the sound source. In the case of a tablet computer that includes speakers located on both the left and right sides, such as... Figure 3 As shown, when playing the audio data, the tablet can use a set of speakers located on the left side of the tablet to play the left channel data of the audio data, and simultaneously use a set of speakers located on the right side of the tablet to play the right channel data of the audio data. If the tablet includes a single speaker, the tablet uses that speaker to play both the left and right channel data of the audio data simultaneously. This single speaker can be located anywhere on the tablet.
[0175] In this way, during video playback, the tablet can construct user-centric 3D sound effects in real time based on the user's current viewing position, bringing the user an immersive 3D sound experience.
[0176] For example, such as Figure 10 As shown, the tablet computer plays a first video stream, with the user's viewing position at spatial orientation 1. The first video stream includes a first audio stream. The first audio stream is rendered by the tablet computer based on the relative positions of the drum kit, singer, and guitar in three-dimensional space with spatial orientation 1. Specifically, the tablet computer renders the drum kit sound based on the relative position of the drum kit in three-dimensional space with spatial orientation 1. The tablet computer renders the vocals based on the relative position of the singer in three-dimensional space with spatial orientation 1. The tablet computer renders the guitar sound based on the relative position of the guitar in three-dimensional space with spatial orientation 1. During the viewing of the first video stream, the user's viewing position moves from spatial orientation 1 to spatial orientation 2. The tablet computer then plays a second video stream, which includes a second audio stream. The second audio stream is rendered by the tablet computer based on the relative positions of the drum kit, singer, and guitar in three-dimensional space with spatial orientation 2. This rendering method is consistent with the above. If, during the viewing of the first video stream, the user's viewing position moves from spatial orientation 1 to spatial orientation 3, the tablet computer plays a third video stream, which includes a third audio stream. The third audio stream is rendered by the tablet based on the relative positions of the drum kit, singer, and guitar in three-dimensional space and spatial orientation 3. This rendering method is consistent with the one described above. Therefore, even if the user moves their viewing position, the tablet can always construct a user-centric sound field, enhancing the user's listening experience.
[0177] The following example uses a target video stream comprising three frames, and combines a first neural network model, a second neural network model, and a third neural network model to introduce the video playback method provided in this application. These three frames include three sound sources: a drum kit, a singer, and a guitar.
[0178] like Figure 11As shown, after audio-video separation, the target video stream yields a target image group and a target audio stream. The target image group consists of three frames. Inputting the target image group into a first neural network model, the model outputs three image coordinates and corresponding image features. These image features are the image features of the sound source. For example, the first neural network model can output image coordinates (x...) a ,y a ),(x b ,y b ),(x c ,y c ) and image features a, b and c that correspond one-to-one with these three image coordinates.
[0179] The target audio stream is subjected to a short-time Fourier transform (STFT) to obtain its spectrogram, also known as a speech spectrogram. This spectrogram is then input into a second neural network model. The second neural network model outputs three audio data points and corresponding audio features. For example, it can output audio data points a, b, and c, along with their corresponding audio features a, b, and c.
[0180] Image feature a-image feature c and audio feature a-audio feature c are input into the third neural network model. The third neural network model can output three sets of data, where each set of data includes an image feature and an audio feature, indicating that the image feature and the audio feature have a one-to-one correspondence.
[0181] The tablet computer can derive the correspondence between image coordinates and audio data based on the output of the third neural network model, the correspondence between image coordinates and image features, and the correspondence between audio data and audio features. For example, image coordinates (x... a ,y a The corresponding audio data (a) and image coordinates (x) are given. b ,y b ) corresponds to audio data b, image coordinates (x) c ,y c The image coordinates (x) correspond to the audio data c. a ,y a The image coordinates (x, y) can be the image coordinates of the drum kit, and the audio data 'a' can be the sound of the drum kit being played. b ,y b The image coordinates (x, y) can be the singer's image coordinates, and the audio data (b) can be the singing voice. c ,y c) can be the image coordinates of the guitar, and the audio data c can be the sound of the guitar being played.
[0182] Next, as Figure 12 As shown, the tablet computer obtains the user's image coordinates in the first image, such as image coordinates (x...). d ,y d The tablet will display the image coordinates (x...). a ,y a ),(x b ,y b ),(x c ,y c ) and (x d ,y d Transforming these four image coordinates into three-dimensional space. Taking the conversion of these four image coordinates into camera coordinates by a tablet computer as an example, the tablet computer can perform actions such as... Figure 4 The methods shown in S3 and S4 convert the image coordinates (x) a ,y a ),(x b ,y b ),(x c ,y c ) and (x d ,y d Convert (X) to camera coordinates, such as (X) a ,Y a Z a ),(X b ,Y b Z b ,)(X c ,Y c Z c ,) and (X d ,Y d Z d ).
[0183] Finally, the tablet renders audio data corresponding to the camera coordinates of the sound source and the user's camera coordinates. For example, there is a one-to-one correspondence between the camera coordinates of the sound source and the image coordinates of the sound source, and a one-to-one correspondence between the image coordinates of the sound source and the audio data. The tablet can find the audio data that corresponds one-to-one with the camera coordinates. For example, the tablet renders the audio data based on the camera coordinates of the sound source (X... a ,Y a Z a ) and the user's camera coordinates (X) d ,Y d Z d Rendering audio data a. The tablet computer uses the camera coordinates (X, Y) of the sound source. b ,Y b Z b ) and the user's camera coordinates (X)d ,Y d Z d ) Render audio data b. The tablet computer uses the camera coordinates (X, Y) of the sound source. c ,Y c Z c ) and the user's camera coordinates (X) d ,Y d Z d Rendering audio data (c). For specific rendering methods, please refer to [reference needed]. Figure 4 The S5 component will not be elaborated upon here. This provides the user with a user-centric sound field, ensuring optimal listening quality.
[0184] This application provides an electronic device including a memory, a display screen, and one or more processors. The display screen is coupled to the processors. The memory stores computer program code. The computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the mobile phone in the above method embodiments. The structure of the electronic device can be referred to... Figure 1 The structure of the electronic device 100 shown.
[0185] This application embodiment also provides a computer storage medium, which includes computer instructions, when the computer instructions are executed in the aforementioned electronic device (such as...). Figure 1 When the electronic device 100 shown is run, it causes the electronic device to perform the various functions or steps in the above method embodiments.
[0186] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps described in the above method embodiments.
[0187] This application also provides a chip system including at least one processor and at least one interface circuit. The processor and the interface circuit are interconnected via lines. For example, the interface circuit can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit can be used to send signals to other devices (e.g., the processor). Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, and this application does not specifically limit this.
[0188] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0189] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0190] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0191] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0193] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video playback method, characterized in that, Applied to electronic devices, the method includes: Play a first video stream, which includes a first audio stream, which includes audio data corresponding to M sound sources; where M is an integer greater than or equal to 1; when M is greater than or equal to 2, the positions of each of the M sound sources in the first video stream are different; the sound field of the first audio stream is matched with a first relative position; the first relative position is the relative position of the M sound sources and the user in the same space; Play a second video stream, which includes a corresponding second audio stream. The second audio stream includes audio data corresponding to P sound sources; where P is an integer greater than or equal to 1; when P is greater than or equal to 2, the positions of each of the P sound sources in the second video stream are different; the sound field of the second audio stream matches a second relative position; the second relative position is the relative position of the P sound sources and the user in the same space; the first relative position and the second relative position are different; the sound field of the first audio stream is different from the sound field of the second audio stream. Before playing the video stream, for each of the L sound sources: The relative positions of the sound source and the user are determined based on the first position of the sound source in the first space and the second position of the user in the first space. Based on the relative position of the sound source and the user, render the audio data corresponding to the sound source; Wherein, when the video stream is a first video stream, L is M; when the video stream is a second video stream, L is P.
2. The method according to claim 1, characterized in that, The first space is a space constructed according to the camera coordinate system, or a space constructed according to the world coordinate system. The method further includes: For each of the M sound sources, obtain the first position of the sound source in the first space; Obtain the user's second location in the first space.
3. The method according to claim 2, characterized in that, Obtaining the first position of the sound source in the first space includes: Obtain the position of the sound source in the image space; The first position of the sound source is determined based on its position in the image space.
4. The method according to claim 2 or 3, characterized in that, The step of obtaining the user's second location in the first space includes: Acquire a first image, which includes an image of the user; Based on the first image, the user's position in the image space is obtained; The user's second position is determined based on the user's position in the image space.
5. The method according to any one of claims 1-3, characterized in that, Before playing the first video stream, the method further includes: Separate the first image group and the first audio stream from the first video stream; Based on the first image group and the first audio stream, determine the correspondence between M audio data and the M sound sources.
6. The method according to claim 5, characterized in that, The step of determining the correspondence between the M audio data and the M sound sources based on the first image group and the first audio stream includes: Based on the first image group, the M sound sources are identified; Based on the first audio stream, the M audio data are identified; Based on the matching degree between the M sound sources and the M audio data, a one-to-one correspondence between the M audio data and the M sound sources is established.
7. The method according to claim 2 or 3, characterized in that, The first relative position includes distance, azimuth, and pitch angle; the step of rendering audio data corresponding to the sound source based on the first relative position between the sound source and the user includes: The head transfer function is determined based on the first relative position between the sound source and the user; The audio data is rendered based on the head transfer function.
8. An electronic device, characterized in that, The electronic device includes: a memory, a camera, and one or more processors; the camera, the memory, and the processors are coupled; wherein the memory is used to store computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Sound field calibration method and electronic equipment
CN118233821A