Video playing method, terminal device and storage medium
By acquiring the three-dimensional spatial location information of the target area and audio source during video playback and adjusting the volume accordingly, the problem of insufficient flexibility in video playback volume control is solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2022-08-25
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies have limited flexibility and intelligence in volume control during video playback, which negatively impacts user experience.
When a scaling operation is detected in a target area of a video image, the position information of the target area in three-dimensional space and the position information of the audio source in three-dimensional space are obtained, and the audio volume is adjusted according to this position information.
It improves the flexibility and intelligence of volume control during video playback, enhancing the user experience.
Smart Images

Figure CN117676253B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video technology, and in particular to a video playback method, terminal device, and storage medium. Background Technology
[0002] Currently, when playing videos on a terminal device, scaling the video frame does not change the audio volume; or, scaling the video frame only increases or decreases the overall audio volume. Therefore, existing technology suffers from low flexibility and a low level of intelligence in video playback volume control, negatively impacting the user experience. Summary of the Invention
[0003] The main purpose of this application is to provide a video playback method, terminal device, and storage medium, which aims to improve the flexibility of volume control during video playback and the intelligence of video playback control, thereby enhancing the user experience.
[0004] In a first aspect, embodiments of the present invention provide a method for playing a video, the video including a video image and audio, wherein the audio is acquired from a sound source in a three-dimensional space corresponding to the video image, the method comprising:
[0005] Play the video;
[0006] Upon detecting a scaling operation on a target region in the video image, the system obtains first position information of the spatial region corresponding to the target region in the three-dimensional space, and second position information of the audio source corresponding to the audio in the three-dimensional space.
[0007] The volume of the audio is adjusted based on the first position information of the spatial region corresponding to the target region in the three-dimensional space and the second position information of the audio source corresponding to the audio in the three-dimensional space.
[0008] Secondly, embodiments of the present invention also provide a terminal device, the terminal device including a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for implementing communication between the processor and the memory, wherein the computer program, when executed by the processor, implements the steps of any of the video playback methods provided in this specification.
[0009] Thirdly, embodiments of the present invention also provide a storage medium for computer-readable storage, the storage medium storing one or more programs, the one or more programs being executable by one or more processors to implement the steps of any video playback method provided in this specification.
[0010] This invention provides a video playback method, terminal device, and storage medium. The method involves playing the video; upon detecting a scaling operation on a target region in the video image, acquiring first position information of the spatial region corresponding to the target region in a three-dimensional space, and acquiring second position information of the audio source corresponding to the audio in the three-dimensional space; adjusting the audio volume based on the first position information of the spatial region corresponding to the target region in the three-dimensional space and the second position information of the audio source corresponding to the audio in the three-dimensional space, thereby improving the flexibility and intelligence of volume control during video playback and enhancing the user experience. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating a video playback method provided in an embodiment of the present invention;
[0013] Figure 2 This is a schematic diagram of a terminal device according to an embodiment of the present invention;
[0014] Figure 3 This is a schematic diagram illustrating the second position information of the sound source in three-dimensional space according to an embodiment of the present invention;
[0015] Figure 4 This is a schematic diagram of the video image region corresponding to the sound source involved in an embodiment of the present invention;
[0016] Figure 5 This is a schematic block diagram of a terminal device provided in an embodiment of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0019] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0020] This invention provides a video playback method, a terminal device, and a storage medium. The video playback method can be applied to a terminal device, which may be an electronic device such as a mobile phone, tablet computer, laptop computer, desktop computer, personal digital assistant, or wearable device. It can also be applied to a server, which may be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0021] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0022] Please see Figure 1 , Figure 1 This is a flowchart illustrating a video playback method provided in an embodiment of the present invention.
[0023] For example, the video includes video images and audio, wherein the audio is captured from a sound source in the three-dimensional space corresponding to the video images.
[0024] In one exemplary embodiment, when recording video, a video image is acquired by a shooting device, and audio is acquired by an audio acquisition device. For example, the shooting device and the audio acquisition device can be located on the same terminal device or on different terminal devices; there is no limitation thereto.
[0025] In some embodiments, audio is acquired from sound sources in the three-dimensional space corresponding to video images. For example, the video includes multiple video images, wherein the number of sound sources in the three-dimensional space corresponding to different video images may include at least one, and the number of sound sources in the three-dimensional space corresponding to the same video image may also include at least one. For example, during different time periods of video recording, the three-dimensional space corresponding to the video images may include different sound sources. For instance, during the first time period of video recording, the audio acquisition device collects audio from sound source A1 in the three-dimensional space corresponding to video image A; during the second time period of video recording, the audio acquisition device collects audio from sound source B1 in the three-dimensional space corresponding to video image B. Similarly, during the same time period of video recording, the three-dimensional space corresponding to the video images may include different sound sources. For instance, during the third time period of video recording, the audio acquisition device collects audio from sound sources C1 and C2 in the three-dimensional space corresponding to video image C. The audio acquisition device may distinguish the audio from sound source C1 and the audio from sound source C2 in the third time period based on the different positional information of sound sources C1 and C2. Of course, the number of sound sources, the number of audio files, or the number of audio tracks are not limited to this and are not restricted here.
[0026] like Figure 1 As shown, the video playback method includes steps S101 to S103.
[0027] Step S101: Play the video.
[0028] For example, video playback and video recording can occur on the same terminal device, such as shooting and playing a video through a terminal device like a mobile phone; or they can occur on different terminal devices, such as shooting a video through one terminal device like a mobile phone and playing the video through another terminal device like a laptop. There are no restrictions on this.
[0029] Step S102: When a scaling operation on the target area in the video image is detected, the first position information of the spatial region corresponding to the target area in three-dimensional space is obtained, and the second position information of the sound source corresponding to the audio is obtained in three-dimensional space.
[0030] In one exemplary embodiment, when a scaling operation on a target region in a video image is detected, third position information of the target region in the video image is determined based on the scaling operation; and first position information of the spatial region corresponding to the target region in three-dimensional space is determined based on the third position information of the target region in the video image.
[0031] For example, when shooting a video, it's equivalent to converting the three-dimensional space corresponding to the video image into a video image. For instance, based on the video image, the corresponding three-dimensional space can be determined. Similarly, based on the third position information of the target area in the video image, the first position information of the spatial region corresponding to the target area in three-dimensional space can be determined. Determining the first position information of the spatial region corresponding to the target area in three-dimensional space improves the ease of subsequent judgment on whether the audio source is within the spatial region corresponding to the target area. This, in turn, improves the flexibility of volume control during video playback, enhances the intelligence of video playback control, and ultimately improves the user experience.
[0032] For example, the detected scaling operation on a target region in a video image may occur when a user is editing a video or watching a video, and the user is scaling the target region in the video image. For instance, if a scaling operation is performed on a part or the whole of the area where at least one object is located in the video image, then the part or the whole of the area where the object is located in the video image is determined to be the target region, and in response to the detected scaling operation on the target region in the video image, the first position information of the spatial region corresponding to the target region in three-dimensional space is obtained.
[0033] In some embodiments, the target region in a video image can change after scaling. For example, when a user is editing a video, they may apply a large-head effect to a person in a certain frame of the video image, and that person appears in several consecutive frames after that frame. If the position of the person changes in those frames, i.e., the person moves, the target region can change accordingly. This improves the comprehensiveness and flexibility of obtaining the first position information of the spatial region corresponding to the target region in three-dimensional space.
[0034] In some embodiments, the target region in a video image can remain unchanged after scaling. For example, when a user is editing a video or watching a video, and the user scales a speaker in a certain frame of the video image, and the speaker appears in multiple consecutive frames of the video image after that frame, if the position of the speaker in the multiple frames of the image does not change, the target region can remain unchanged. This improves the comprehensiveness and flexibility of obtaining the first position information of the spatial region corresponding to the target region in three-dimensional space.
[0035] In one exemplary embodiment, the detected scaling operation of the target region in the video image can be used to determine the change in the physical distance between the shooting device and the corresponding spatial region of the target region in three-dimensional space during video recording. For example, zooming in on the target region reduces the physical distance between the shooting device and the corresponding spatial region, and similarly, zooming out increases the physical distance. For instance, based on a decrease in the physical distance between the shooting device and the corresponding spatial region of the target region, it can be determined that the physical distance between the audio acquisition device and the corresponding spatial region of the target region is smaller; conversely, based on an increase in the physical distance, it can be determined that the physical distance between the audio acquisition device and the corresponding spatial region of the target region is larger. This allows for adjustment of the audio volume based on the physical distance between the sound source and the corresponding spatial region of the target region. Therefore, improving the comprehensiveness and flexibility of determining the first position information of the spatial region corresponding to the target region in three-dimensional space is beneficial for subsequently improving the flexibility of volume control during video playback, enhancing the intelligence of video playback control, and improving the user experience.
[0036] In some embodiments, video images are captured by a shooting device and audio is captured by an audio acquisition device. For example, the audio acquisition device is arranged around the shooting device according to a preset layout rule.
[0037] For example, the audio acquisition device includes a first audio acquisition device and a second audio acquisition device, wherein the second audio acquisition device is an audio acquisition device other than the first audio acquisition device, and the preset layout rule is, for example, that any first audio acquisition device is arranged at an angle to at least two second audio acquisition devices. Please refer to [link / reference]. Figure 2 The terminal device is, for example, a mobile phone; the shooting device is, for example, a camera; and the audio acquisition device is, for example, a microphone (MIC). The camera and microphone are located on the same mobile phone. The microphones include, for example, MIC1, MIC2, MIC3, and MIC4, respectively located above, to the right, to the left, and below the camera. For instance, due to the different positions of the microphones, if a sound source exists in the three-dimensional space corresponding to the video image during video recording, the time taken for each microphone to acquire the audio corresponding to that sound source will also differ.
[0038] In one exemplary embodiment, it can be determined whether the audio collected by each audio acquisition device is the same. If they are the same, the second position information of the sound source corresponding to the audio in three-dimensional space can be determined based on the time difference between the audio collected by each audio acquisition device and the correlation between the position of the shooting device and the position of the audio acquisition device. In some embodiments, it can be determined whether the audio collected by each audio acquisition device is the same based on whether the sound waves are the same. For example, based on the difference in the second position information of the sound source corresponding to the audio in three-dimensional space, the audio corresponding to each sound source can be imported into different audio tracks. This allows the volume of the audio in different audio tracks to be adjusted based on the first position information of the spatial region corresponding to the target region in three-dimensional space and the second position information of the sound source corresponding to the audio in three-dimensional space. This improves the flexibility of volume control during video playback, enhances the intelligence of video playback control, and thus improves the user experience.
[0039] Please combine Figure 2 See Figure 3 Audio 1 corresponding to audio source 1 is captured by MIC1, MIC2, MIC3, and MIC4. The distance between audio source 1 and MIC1 is L. 11 If the speed of sound propagation is v, and the time difference between MIC1 and MIC2 in acquiring audio 1 is, for example, t1, then the distance between sound source 1 and MIC2 is L. 12 =L 11 +vt1, if the time difference between MIC1 and MIC3 in capturing audio 1 is, for example, t2, then the distance between audio source 1 and MIC3 is L. 13 =L 11 +vt2, if the time difference between MIC1 and MIC4 in capturing audio 1 is, for example, t3, then the distance between audio source 1 and MIC4 is L. 14 =L 11 +vt3, thus allowing the three-point positioning method, based on L... 11 L 12 L 13 and L 14 The second position information of sound source 1 in three-dimensional space is determined. Similarly, the audio 2 corresponding to sound source 2 in the three-dimensional space corresponding to the video image is captured by MIC1, MIC2, MIC3, and MIC4, and the distance between sound source 2 and MIC1 is L. 21 If the speed of sound propagation is v, then the distance L between sound source 1 and MIC2, MIC3, and MIC4 can be determined based on the time difference between the audio 2 collected by MIC1 and MIC2, MIC3, and MIC4 respectively. 22 L 23 and L 24 Therefore, the three-point positioning method can be used to determine the location based on L.21 L 22 L 23 and L 24 Determine the second position information of sound source 2 in three-dimensional space.
[0040] For example, audio is captured from a sound source in the three-dimensional space corresponding to the video image, which is equivalent to a correlation between the three-dimensional space corresponding to the video image and the three-dimensional space where the sound source is located. For instance, when a shooting device converts an object within the shooting range of a video image into a video image, the audio of the video is captured from at least one of sound sources 1 and 2 in the three-dimensional space, and at least one of sound sources 1 and 2 is within the shooting range of the shooting device. Therefore, a correlation can be established between the position information of at least one of sound sources 1 and 2 in the three-dimensional space and the three-dimensional space corresponding to the video image, determining the video image area corresponding to at least one of sound sources 1 and 2. For example, taking the example where the audio of the video is captured from sound sources 1 and 2 in the three-dimensional space, and the sound sources 1 and 2 are within the shooting range of the shooting device, a correlation can be established between the position information of sound sources 1 and 2 in the three-dimensional space and the three-dimensional space corresponding to the video image, determining the video image area corresponding to each of sound sources 1 and 2. Figure 4 As shown; if the audio source is not within the shooting range of the shooting device, a correlation can be established between the position information of the audio source in three-dimensional space and the three-dimensional space corresponding to the video image. This allows for determining the deviation between the position information of the audio source in three-dimensional space and the three-dimensional space corresponding to the video image, thus more comprehensively determining the position information of the audio source in three-dimensional space. This facilitates more comprehensive control over the volume of the audio included in the video. In some embodiments, the position information of the audio source relative to the shooting device can be determined based on a preset correlation between the position information of the shooting device and the position information of the audio acquisition device, and the position information of the audio source relative to each audio acquisition device. For example, the second position information of the audio source in three-dimensional space is, for instance, the position information of the audio source relative to the shooting device, to establish a correlation between the three-dimensional space corresponding to the video image and the three-dimensional space where the audio source is located.
[0041] For example, based on the correlation between the three-dimensional space corresponding to the video image and the three-dimensional space where the audio source is located, by obtaining the first position information of the spatial region corresponding to the target region in the three-dimensional space and the second position information of the audio source in the three-dimensional space, it is easier to determine whether the audio source is in the spatial region corresponding to the target region. This allows for the adjustment of the audio volume, which in turn improves the flexibility of volume control during video playback, enhances the intelligence of video playback control, and improves the user experience.
[0042] Step S103: Adjust the volume of the audio based on the first position information of the spatial region corresponding to the target region in three-dimensional space and the second position information of the audio source in three-dimensional space.
[0043] For example, if it is determined that the sound source is within the spatial area corresponding to the target area based on the first location information and the second location information, the volume of the audio is adjusted using a first adjustment strategy; and / or, if it is determined that the sound source is not within the spatial area corresponding to the target area based on the first location information and the second location information, the volume of the audio is adjusted using a second adjustment strategy; the second adjustment strategy is different from the first adjustment strategy.
[0044] In some embodiments, when a scaling operation on a target region in a video image is detected, if the audio is collected from at least one sound source in the three-dimensional space corresponding to the video image frame, based on the different second position information of the different sound sources in the three-dimensional space, it is determined whether to adjust the volume of the audio corresponding to the different sound sources using a first adjustment strategy or a second adjustment strategy, thereby distinguishing the volume of the audio corresponding to different sound sources in the three-dimensional space corresponding to the same video image frame.
[0045] In some embodiments, if, before and after a scaling operation on a target region in a video image is detected, the audio is acquired from at least one sound source in a three-dimensional space corresponding to different frame video images, then, based on the second position information of the sound source in the three-dimensional space when a scaling operation on a target region in a video image is detected, it is determined whether the sound source is within the spatial region corresponding to the target region or not, in order to determine whether to apply different volume adjustments to the audio corresponding to the sound source using a first adjustment strategy or a second adjustment strategy, thereby distinguishing the volume of audio corresponding to the same sound source in the three-dimensional space corresponding to different frame video images, or distinguishing the volume of audio corresponding to different sound sources in the three-dimensional space corresponding to different video images.
[0046] For example, by using a second adjustment strategy that differs from the first adjustment strategy, different volume adjustments are made to sound sources located within the spatial area corresponding to the target area and sound sources not located within the spatial area corresponding to the target area, thereby increasing the flexibility of volume control during video playback and improving the user experience.
[0047] In some embodiments, when it is determined that the sound source is in the spatial region corresponding to the target region based on the first location information and the second location information, the volume of the audio is adjusted according to the first adjustment strategy, including: obtaining the scaling ratio corresponding to the target region; increasing the audio volume according to the scaling ratio when the scaling ratio is greater than or equal to a first preset scaling ratio threshold; or decreasing the audio volume according to the scaling ratio when the scaling ratio is less than or equal to a second preset scaling ratio threshold.
[0048] For example, the first preset scaling threshold and the second preset scaling threshold can be the same or different. For instance, the first preset scaling threshold may be greater than or equal to 1, and the second preset scaling threshold may be less than or equal to 1. They can be preset or set by the user, and there is no restriction here.
[0049] For example, the original size can be the original size of the video image; the preset volume can be the audio volume determined based on the user's volume adjustment operation, such as determining the preset volume based on detected operation of the volume up and down keys, or determining the preset volume based on detected operation of the volume adjustment slider.
[0050] In some embodiments, if the detected scaling operation on the target area is to enlarge the target area to 1.1 times the original size, then the scaling ratio corresponding to the target area is 1.1 times; if the detected scaling operation on the target area is to shrink the target area to the original size... The scaling factor corresponding to the target area is the size reduced to the original image size.
[0051] For example, the first preset scaling threshold is, for instance, 1.1 times. If, based on the first and second location information, the sound source is determined to be within the spatial region corresponding to the target area, and the scaling ratio corresponding to the target area is greater than 1.1 times (i.e., the target area is enlarged to more than 1.1 times the original size), then since this scaling ratio is greater than the first preset scaling threshold, the volume of the audio corresponding to the sound source can be increased, for example, increased to more than 1.1 times the preset volume. The preset volume is the volume when the terminal device displays the original size. The second preset scaling threshold is, for instance, reduced to the original size. Based on the second and third position information, the sound source is located within the spatial region corresponding to the target area, and the scaling ratio of the target area is smaller than the original image size. The target area is reduced to a size smaller than the original image. If the scaling ratio is less than the second preset scaling ratio threshold, the volume of the audio source can be reduced, for example, to below the preset volume. The preset volume is the volume when the terminal device displays the original picture size, thereby establishing a correlation between the scaling operation of the target area and the volume of the audio source in the corresponding spatial area of the target area. This improves the flexibility of volume control during video playback, enhances the intelligence of video playback control, and thus improves the user experience.
[0052] In one exemplary embodiment, when it is determined, based on the first location information and the second location information, that the sound source is not within the spatial area corresponding to the target area, the volume of the audio is adjusted using a second adjustment strategy, including: when it is determined, based on the first location information and the second location information, that the sound source is not within the spatial area corresponding to the target area, and the scaling ratio corresponding to the target area is greater than or equal to a first preset scaling ratio threshold, keeping the audio volume unchanged or decreasing the audio volume; when it is determined, based on the first location information and the second location information, that the sound source is not within the spatial area corresponding to the target area, and the scaling ratio corresponding to the target area is less than or equal to a second preset scaling ratio threshold, keeping the audio volume unchanged or increasing the audio volume.
[0053] For example, if it is determined from the first location information and the second location information that the sound source is not in the spatial area corresponding to the target area, in order to distinguish between the sound source in the spatial area corresponding to the target area and the sound source not in the spatial area corresponding to the target area, the volume of the sound source not in the spatial area corresponding to the target area can be adjusted in combination with the scaling ratio corresponding to the target area.
[0054] For example, if the first preset scaling threshold is 1.1x, and it is determined based on the first and second location information that the sound source is not within the spatial region corresponding to the target region, and the scaling ratio corresponding to the target region is 2x, then this scaling ratio is greater than the first preset scaling threshold. Therefore, the audio volume can be kept constant or reduced, for example, reduced to a preset volume. The preset volume is the volume when the terminal device displays the original image size; the second preset scaling threshold is, for example, scaling down to the original image size. Based on the first and second position information, it is determined that the sound source is not within the spatial region corresponding to the target region, and the scaling ratio corresponding to the target region is reduced to the original image size. In certain situations, the volume of the audio can be kept constant or increased, for example, to twice the preset volume (the volume when the terminal device displays the original image size). This allows for the differentiation of whether the corresponding audio source is within the spatial region corresponding to the target area by using different audio volume levels. It also establishes a correlation between scaling operations on the target area and the volume of audio sources not within the spatial region corresponding to the target area. This improves the flexibility and intelligence of volume control during video playback, thereby enhancing the user experience.
[0055] In some embodiments, if the response to the detected scaling operation of a target region in a video image is to first enlarge the target region to a size larger than the original size of the target region, then restore the target region to a size equal to the original size of the target region, and then shrink the target region to a size smaller than the original size of the target region, then the volume of the audio corresponding to the sound source in the spatial region corresponding to the target region is first increased to a volume greater than a preset volume, then restored to a preset volume, and then decreased to a volume less than a preset volume. The volume of the audio corresponding to the sound source not in the spatial region corresponding to the target region is, for example, first decreased to a volume less than a preset volume, then restored to a preset volume, and then increased to a volume greater than a preset volume. Alternatively, the volume of the audio corresponding to the sound source not in the spatial region corresponding to the target region may be, for example, kept at a preset volume.
[0056] In one exemplary embodiment, when it is determined, based on the first location information and the second location information, that the sound source is not within the spatial area corresponding to the target area, the volume of the audio is adjusted using a second adjustment strategy, including: determining the deviation range between the first location information and the second location information based on the first location information and the second location information; and adjusting the volume of the audio based on the scaling ratio corresponding to the target area and the deviation range.
[0057] For example, the deviation between the first position information and the second position information is used to determine the deviation between the sound source and the corresponding spatial region of the target area in three-dimensional space. The larger the deviation between the first position information and the second position information, the larger the deviation between the sound source and the corresponding spatial region of the target area. For example, if the deviation between the first position information and the second position information is less than or equal to a preset deviation threshold, it can be determined that the sound source is located within the spatial region corresponding to the target area. Conversely, if the deviation between the first position information and the second position information is greater than the preset deviation threshold, it can be determined that the sound source is not located within the spatial region corresponding to the target area.
[0058] In some embodiments, when a user zooms in on a target area, it can be determined that the user is more interested in objects and sounds within the corresponding spatial area and wants to see and hear them more clearly. Therefore, while zooming in on the target area, the volume of the audio corresponding to the sound source within the target area can be increased. Conversely, it can be determined that the user is less interested in or uninterested in objects and sounds outside the corresponding spatial area. Therefore, while zooming in on the target area, the volume of the audio corresponding to the sound source outside the target area can be kept constant or decreased. Similarly, when a user zooms out on a target area, it can be determined that the user is less interested in or uninterested in objects and sounds within the corresponding spatial area. Therefore, while zooming out on the target area, the volume of the audio corresponding to the sound source within the target area can be decreased. Conversely, it can be determined that the user is more interested in objects and sounds outside the corresponding spatial area. Therefore, while zooming out on the target area, the volume of the audio corresponding to the sound source outside the target area can be kept constant or increased.
[0059] For example, adjusting the audio volume based on the scaling ratio corresponding to the target region and the deviation magnitude includes: when adjusting the audio volume based on the scaling ratio corresponding to the target region and the deviation magnitude, the adjustment magnitude of the audio volume is positively correlated with the deviation magnitude and negatively correlated with the scaling ratio corresponding to the target region.
[0060] For example, if it is determined that the sound source is not within the spatial region corresponding to the target region, and the scaling ratio of the target region is greater than or equal to a first preset scaling ratio threshold, it can be determined that the user has little or no interest in the sound source. Keeping the audio volume unchanged or decreasing the audio volume is equivalent to the adjustment range of the audio volume being negatively correlated with the scaling ratio of the target region. Correspondingly, the greater the deviation between the sound source and the spatial region corresponding to the target region, the less interested the user may be in the sound source. Therefore, when decreasing the audio volume, the smaller the deviation between the sound source and the spatial region corresponding to the target region, the smaller the adjustment range of the audio volume; that is, the adjustment range of the audio volume is positively correlated with the deviation range. If it is determined that the sound source is not within the spatial region corresponding to the target region... Within the target area, if the scaling ratio corresponding to the target area is less than or equal to the second preset scaling ratio threshold, it can be determined that the user is interested in the audio source. Keeping the audio volume unchanged or increasing the audio volume is equivalent to the adjustment range of the audio volume being negatively correlated with the scaling ratio corresponding to the target area. Correspondingly, the greater the deviation between the audio source and the target area, the greater the user's interest in the audio source may be. Therefore, when increasing the audio volume, the adjustment range of the audio volume can be greater based on the deviation range of the audio source and the target area. That is, the adjustment range of the audio volume is positively correlated with the deviation range, thereby improving the flexibility of volume control during video playback, enhancing the intelligence of video playback control, and improving the user experience.
[0061] In one exemplary embodiment, the video playback method further includes: outputting a prompt message based on the adjustment of the audio volume, the prompt message being used to indicate the change in audio volume.
[0062] For example, the prompt information may include at least one of a volume bar, a volume value, or other form used to display volume. For instance, the position of the audio source in the two-dimensional plane corresponding to the video image is obtained, and the prompt information is displayed on one side of that position. Alternatively, the prompt information may be displayed on one side of a target area, for example, displaying a volume bar on one side of the target area and on one side of the position of the audio source in the two-dimensional plane corresponding to the video image, to display the volume of the audio corresponding to each audio source. In some embodiments, the volume of the audio corresponding to each audio source can also be adjusted according to the displayed volume bars to improve the flexibility of volume control during video playback.
[0063] The video playback method provided in the above embodiments involves playing a video; upon detecting a scaling operation on a target area in the video image, acquiring first position information of the spatial region corresponding to the target area in three-dimensional space, and acquiring second position information of the audio source corresponding to the audio source in three-dimensional space; and adjusting the audio volume based on the first position information of the spatial region corresponding to the target area in three-dimensional space and the second position information of the audio source corresponding to the audio source in three-dimensional space. Adjusting the audio volume based on the second position information of the audio source corresponding to the audio source in three-dimensional space improves the flexibility of volume control during video playback, enhances the intelligence of video playback control, and thus improves the user experience.
[0064] Please see Figure 5 , Figure 5 This is a schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention.
[0065] like Figure 5 As shown, the terminal device 300 includes a processor 301 and a memory 302, which are connected by a bus 303, such as an I2C (Inter-integrated Circuit) bus.
[0066] Specifically, processor 301 provides computing and control capabilities to support the operation of the entire terminal device. Processor 301 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0067] Specifically, the memory 302 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.
[0068] Those skilled in the art will understand that Figure 5The structure shown is merely a block diagram of a portion of the structure related to the embodiments of the present invention, and does not constitute a limitation on the terminal device to which the embodiments of the present invention are applied. A specific server may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0069] The processor is configured to run a computer program stored in a memory, and to implement any of the video playback methods provided in the embodiments of the present invention when executing the computer program.
[0070] In one embodiment, the processor is configured to run a computer program stored in memory, and when executing the computer program, perform the following steps:
[0071] Play the video;
[0072] Upon detecting a scaling operation on a target region in the video image, the system obtains first position information of the spatial region corresponding to the target region in the three-dimensional space, and second position information of the audio source corresponding to the audio in the three-dimensional space.
[0073] The volume of the audio is adjusted based on the first position information of the spatial region corresponding to the target region in the three-dimensional space and the second position information of the audio source corresponding to the audio in the three-dimensional space.
[0074] In one embodiment, when the processor acquires the first position information of the spatial region corresponding to the target region in the three-dimensional space upon detecting a scaling operation on the target region in the video image, it is configured to:
[0075] Upon detecting a scaling operation on a target region in the video image, a third location information of the target region in the video image is determined based on the scaling operation;
[0076] Based on the third position information of the target region in the video image, the first position information of the spatial region corresponding to the target region in the three-dimensional space is determined.
[0077] In one embodiment, when the processor adjusts the volume of the audio based on the first position information of the spatial region corresponding to the target region in the three-dimensional space and the second position information of the sound source corresponding to the audio in the three-dimensional space, it is configured to:
[0078] If, based on the first location information and the second location information, the sound source is determined to be within the spatial region corresponding to the target region, the volume of the audio is adjusted using a first adjustment strategy; and / or,
[0079] If it is determined, based on the first location information and the second location information, that the sound source is not within the spatial area corresponding to the target area, the volume of the audio is adjusted using a second adjustment strategy; the second adjustment strategy is different from the first adjustment strategy.
[0080] In one embodiment, the processor, when determining that the sound source is located within the spatial region corresponding to the target region based on the first location information and the second location information, adjusts the volume of the audio using a first adjustment strategy, including:
[0081] Obtain the scaling ratio corresponding to the target area;
[0082] If the scaling ratio is greater than or equal to a first preset scaling ratio threshold, the volume of the audio is increased according to the scaling ratio; or,
[0083] If the scaling ratio is less than or equal to a second preset scaling ratio threshold, the volume of the audio is reduced according to the scaling ratio.
[0084] In one embodiment, when the processor adjusts the volume of the audio using a second adjustment strategy based on the first location information and the second location information, it is configured to:
[0085] If, based on the first location information and the second location information, it is determined that the sound source is not in the spatial area corresponding to the target area, and the scaling ratio corresponding to the target area is greater than or equal to the first preset scaling ratio threshold, the volume of the audio remains unchanged or the volume of the audio is reduced.
[0086] If, based on the first location information and the second location information, it is determined that the sound source is not within the spatial area corresponding to the target area, and the scaling ratio corresponding to the target area is less than or equal to the second preset scaling ratio threshold, the volume of the audio remains unchanged or the volume of the audio is increased.
[0087] In one embodiment, when the processor adjusts the volume of the audio using a second adjustment strategy based on the first location information and the second location information, it is configured to:
[0088] Based on the first location information and the second location information, determine the deviation range between the first location information and the second location information;
[0089] The volume of the audio is adjusted according to the scaling ratio corresponding to the target area and the deviation magnitude.
[0090] In some embodiments, when the processor adjusts the volume of the audio based on the scaling ratio corresponding to the target region and the deviation magnitude, it is configured to:
[0091] When adjusting the volume of the audio based on the scaling ratio corresponding to the target region and the deviation amplitude, the adjustment amplitude of the audio volume is positively correlated with the deviation amplitude and negatively correlated with the scaling ratio corresponding to the target region.
[0092] In one embodiment, the processor, when implementing the video playback method, is configured to:
[0093] Based on the adjustment of the audio volume, a prompt message is output, which is used to indicate the change in the audio volume.
[0094] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the terminal device described above can be referred to the corresponding process in the aforementioned video playback method embodiment, and will not be repeated here.
[0095] This invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement the steps of any of the video playback methods provided in the specification of this invention.
[0096] The storage medium can be an internal storage unit of the terminal device described in the foregoing embodiments, such as the hard drive or memory of the terminal device. Alternatively, the storage medium can be an external storage device of the terminal device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device.
[0097] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0098] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0099] The sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The above descriptions are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for playing a video, characterized in that, The video includes video images and audio, and the audio is captured from a sound source in the three-dimensional space corresponding to the video images; The method includes: Play the video; Upon detecting a scaling operation on a target region in the video image, the system obtains first position information of the spatial region corresponding to the target region in the three-dimensional space, and second position information of the audio source corresponding to the audio in the three-dimensional space. Based on the first position information of the spatial region corresponding to the target region in the three-dimensional space, and the second position information of the sound source corresponding to the audio in the three-dimensional space, it is determined whether the sound source is located within the spatial region corresponding to the target region; If it is determined that the sound source is located within the spatial region corresponding to the target region, the volume of the audio is adjusted based on a first correlation between the scaling ratio of the target region and the volume of the audio corresponding to the sound source within the spatial region corresponding to the target region; and / or, If it is determined that the sound source is not located in the spatial region corresponding to the target region, the volume of the audio is adjusted based on the second correlation between the scaling ratio of the target region and the volume of the audio corresponding to the sound source not located in the spatial region corresponding to the target region, or the volume of the audio is adjusted based on the scaling ratio of the target region and the deviation between the first position information and the second position information.
2. The playback method according to claim 1, characterized in that, The step of obtaining the first position information of the spatial region corresponding to the target region in the three-dimensional space when a scaling operation on the target region in the video image is detected includes: Upon detecting a scaling operation on a target region in the video image, a third location information of the target region in the video image is determined based on the scaling operation; Based on the third position information of the target region in the video image, the first position information of the spatial region corresponding to the target region in the three-dimensional space is determined.
3. The playback method according to claim 1, characterized in that, When it is determined that the sound source is located within the spatial region corresponding to the target region, the volume of the audio is adjusted based on a first correlation between the scaling ratio of the target region and the volume of the audio corresponding to the sound source within the spatial region corresponding to the target region, including: Obtain the scaling ratio corresponding to the target area; If the scaling ratio is greater than or equal to a first preset scaling ratio threshold, the volume of the audio is increased according to the scaling ratio; or, If the scaling ratio is less than or equal to a second preset scaling ratio threshold, the volume of the audio is reduced according to the scaling ratio.
4. The playback method according to claim 1, characterized in that, When it is determined that the sound source is not located within the spatial region corresponding to the target region, the volume of the audio is adjusted based on a second correlation between the scaling ratio of the target region and the volume of the audio corresponding to a sound source not located within the spatial region corresponding to the target region, including: If it is determined that the sound source is not located in the spatial region corresponding to the target region, and the scaling ratio corresponding to the target region is greater than or equal to the first preset scaling ratio threshold, the volume of the audio is kept unchanged or the volume of the audio is reduced. If it is determined that the sound source is not located in the spatial region corresponding to the target region, and the scaling ratio corresponding to the target region is less than or equal to the second preset scaling ratio threshold, the volume of the audio is kept unchanged or the volume of the audio is increased.
5. The playback method according to claim 1, characterized in that, Adjusting the audio volume based on the scaling ratio corresponding to the target region and the deviation magnitude includes: When adjusting the volume of the audio based on the scaling ratio corresponding to the target region and the deviation amplitude, the adjustment amplitude of the audio volume is positively correlated with the deviation amplitude and negatively correlated with the scaling ratio corresponding to the target region.
6. The playback method according to any one of claims 1 to 5, characterized in that, Also includes: Based on the adjustment of the audio volume, a prompt message is output, which is used to indicate the change in the audio volume.
7. A terminal device, characterized in that, The terminal device includes a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for establishing communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the steps of the video playback method as described in any one of claims 1 to 6.
8. A storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the video playback method according to any one of claims 1 to 6.