Video processing methods, video processing devices, electronic devices and media
By determining the output delay of images and audio in the video processing method at the concert venue and adjusting the output time of images or audio in the video, the problem of asynchronous picture and sound was solved, and the synchronization of picture and sound was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-26
Smart Images

Figure CN119255030B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and specifically relates to a video processing method, video, video processing device, electronic device and medium. Background Technology
[0002] With the increasing popularity of electronic devices with shooting functions such as mobile phones and digital cameras, it has become more and more common to use electronic devices to record videos at concerts.
[0003] Currently, concert venues typically feature large screens that magnify the live performance, allowing audience members, even those some distance away, to clearly see the performers' actions. When recording the concert using electronic devices, both the screen display and the live audio can be captured. Summary of the Invention
[0004] The purpose of this application is to provide a video processing method, processing device, electronic device and medium capable of synchronously recording images and sounds.
[0005] In a first aspect, embodiments of this application provide a video processing method, the method comprising:
[0006] Determine N frames of images and N audio audios in the first video. The N frames of images are images in the first video that have motion features, and the N audio audios are audios in the first video that have preset audio features. N is a positive integer.
[0007] Determine the first output delay based on N frames of images and N audio files;
[0008] Based on the first output delay, the output time of the image or the output time of the audio in the second video are adjusted; wherein, the image and audio in the adjusted second video are synchronized.
[0009] Secondly, embodiments of this application provide a video processing apparatus, which includes: a determining module and an adjusting module; wherein...
[0010] The determination module is used to determine N frames of images and N audios in the first video. The N frames of images are images in the first video that have motion features, and the N audios are audios in the first video that have preset audio features. N is a positive integer.
[0011] A determination module is used to determine the first output delay based on N frames of images and N audio recordings;
[0012] An adjustment module is used to adjust the output time of the image or the output time of the audio in the second video based on the first output delay; wherein the image and audio in the adjusted second video are synchronized.
[0013] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the video processing method as described in the first aspect.
[0014] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the video processing method as described in the first aspect.
[0015] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the video processing method as described in the first aspect.
[0016] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the video processing method as described in the first aspect.
[0017] In this embodiment, N frames of images and N audio tracks are determined in a first video. The N frames are images with motion features in the first video, and the N audio tracks are audio tracks with preset audio features in the first video, where N is a positive integer. Based on the N frames of images and the N audio tracks, a first output delay is determined. Based on the first output delay, the output time of the images or the output time of the audio tracks in the second video are adjusted. The adjusted images and audio tracks in the second video are synchronized. Thus, by determining the first output delay based on the N frames of images and the N audio tracks, the output delay between images with motion features and audio tracks with preset audio features in the video can be determined. Therefore, when adjusting the output time of the images or the output time of the audio tracks in the second video based on the first output delay, the images with motion features and the audio tracks with preset audio features in the second video can be matched, thereby synchronizing the images and sound in the second video. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of a concert scene provided in an embodiment of this application;
[0019] Figure 2 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 1 ;
[0020] Figure 3 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 2 ;
[0021] Figure 4 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 3 ;
[0022] Figure 5 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 4 ;
[0023] Figure 6 This is a schematic diagram of a preview image provided in an embodiment of this application;
[0024] Figure 7 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 5 ;
[0025] Figure 8 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 6 ;
[0026] Figure 9 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application. Figure 1 ;
[0027] Figure 10 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application. Figure 2 ;
[0028] Figure 11 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0029] Figure 12 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0031] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0032] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."
[0033] The video processing method, video processing device, electronic device, and medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0034] The video processing method, video processing device, electronic device, and medium provided in this application can be applied to video recording scenarios, such as recording live videos of large-scale stage events.
[0035] One application scenario is that the video processing method provided in this application embodiment can be applied to the scenario of recording live concert videos.
[0036] At concert venues, large screens typically display magnified images of the performance, allowing audiences to clearly see the performers' actions and performances. For example... Figure 1 As shown, during the concert, while the performers are performing on stage 101, the large screen 102 above the stage will magnify the display of the performers, allowing the audience to clearly see their performance on stage 101. For easy review later, audience members can use mobile phones, digital cameras, or other electronic devices with recording capabilities to record the concert. When recording the concert, electronic devices can simultaneously record the visuals displayed on the large screen 102 and the audio from the concert.
[0037] However, due to the display delay on the large screen at the concert venue, the image displayed on the large screen recorded by the electronic device may be out of sync with the recorded sound from the venue. In other words, the image and sound recorded by the electronic device are out of sync.
[0038] To address this, embodiments of this application provide a video processing method, a video processing apparatus, an electronic device, and a medium. The method involves determining N frames of images and N audio tracks in a first video, where the N frames are images with motion characteristics and the N audio tracks are audio tracks with preset audio characteristics. Based on the N frames and N audio tracks, a first output delay is determined. Based on the first output delay, the output time of the images or audio tracks in a second video is adjusted. The adjusted images and audio tracks in the second video are synchronized. Thus, by determining the first output delay based on the N frames and N audio tracks, the output delay between images with motion characteristics and audio tracks with preset audio characteristics in the video can be determined. Therefore, when adjusting the output time of the images or audio tracks in the second video based on the first output delay, the images with motion characteristics and the audio tracks with preset audio characteristics in the second video can be matched, thereby synchronizing the images and sound in the second video.
[0039] The execution subject of the video processing method provided in this application embodiment can be a video processing device. Exemplarily, the video processing device can be an electronic device, or a functional component or entity within that electronic device. The following will executor an electronic device as an example to illustrate the video processing device provided in this application embodiment.
[0040] Figure 2 This is a flowchart illustrating the video processing method provided in an embodiment of this application, as shown below. Figure 2 As shown, the video processing method provided in this application embodiment may include the following steps 201 to 203.
[0041] Step 201: The electronic device determines N frames of images and N audio clips in the first video.
[0042] In some embodiments of this application, the aforementioned N frames are images with motion features in the aforementioned first video, and the aforementioned N audios are audios with preset audio features in the first video, where N is a positive integer.
[0043] In some embodiments of this application, the first video is a video captured by an electronic device of a target environment, the target environment including a first display screen and a first object, the first display screen being used to display the first object captured by a camera, the first video including M frames of images, each frame including a first image area and a second image area, the first image area being an image area corresponding to the first display screen, and the second image area being an image area corresponding to the first object.
[0044] Wherein, the above N frames are images in the first image region of the above M frames that have motion features, M is a positive integer, and M is greater than or equal to N.
[0045] In some embodiments of this application, the target environment can be the venue environment of a large-scale stage event, such as the venue environment of a concert, and the first display screen is used to magnify and display the first object.
[0046] In some embodiments of this application, the first object mentioned above can be a person, an animal, or an object that can actively or passively perform actions.
[0047] In some embodiments of this application, the aforementioned persons include, but are not limited to, men, women, and children.
[0048] In some embodiments of this application, the animals mentioned above include, but are not limited to, livestock, wild animals, pets, birds, fish, insects, etc.
[0049] In some embodiments of this application, the aforementioned objects include, but are not limited to, inanimate objects such as clothing and accessories, daily necessities, office supplies, toys, tools, and recreational tools.
[0050] In some embodiments of this application, the electronic device can perform a recognition operation on each of the above M-frame images to determine the first image region and the second image region of each frame image.
[0051] In some embodiments of this application, the electronic device can perform motion recognition on a first object displayed in a magnified first image region of each of the M frames in the above-mentioned M frames in order to determine the above-mentioned N frames from the above-mentioned M frames.
[0052] In some embodiments of this application, the first video further includes M audio files. The audio files with the preset audio features refer to the audio files whose frequencies are greater than or equal to a preset frequency. That is, the N audio files are audio files in the first video whose frequencies are greater than or equal to a preset frequency.
[0053] Step 202: The electronic device determines the first output delay based on N frames of images and N audio signals.
[0054] In some embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, step 202 above can be achieved through steps 202a to 202d as follows:
[0055] Step 202a: The electronic device determines the output time difference between the i-th frame and the (i+1)-th frame in the N frames, and obtains K first output time differences.
[0056] In some embodiments of this application, i∈[1,3,...,N), K is N / 2 and K is a positive integer.
[0057] In some embodiments of this application, after the electronic device determines the N frames of images from the first video, it can sort the N frames of images according to the order of their output time.
[0058] In some embodiments of this application, the i-th frame image can be the image whose output time is ranked i-th among the N frames, and the (i+1)-th frame image can be the image whose output time is ranked i+1 among the N frames.
[0059] In some embodiments of this application, the output time difference between the i-th frame image and the (i+1)-th frame image can be the difference between the output time of the (i+1)-th frame image and the output time of the i-th frame image.
[0060] For example, when N is 6, K is 3. If the output time of the first frame in the 6 frames is 10:24:23.05 on September 24, 2024, the output time of the second frame is 10:24:23.26 on September 24, 2024, the output time of the third frame is 10:24:23.47 on September 24, 2024, the output time of the fourth frame is 10:24:23.90 on September 24, 2024, the output time of the fifth frame is 10:25:24.32 on September 24, 2024, and the output time of the sixth frame is 10:25:24.88 on September 24, 2024, then the difference between the output time of the second frame and the output time of the first frame is 0.21 seconds, which is the first of three first output time differences. The difference between the output time of the 4th frame and the output time of the 3rd frame is 0.43 seconds, which is the second of the three first output time differences. The difference between the output time of the 6th frame and the output time of the 5th frame is 0.56 seconds, which is the third of the three first output time differences.
[0061] Step 202b: The electronic device determines the output time difference between the i-th audio and the (i+1)-th audio among the N audios, and obtains K second output time differences.
[0062] In some embodiments of this application, after the electronic device determines the N audio files from the first video, it can sort the N audio files according to the order of their output times.
[0063] In some embodiments of this application, the i-th audio can be the audio with the i-th output time among the N audios, and the (i+1)-th audio can be the audio with the (i+1)-th output time among the N audios.
[0064] In some embodiments of this application, the output time difference between the i-th frame image and the (i+1)-th frame image can be the difference between the output time of the (i+1)-th frame image and the output time of the i-th frame image.
[0065] For example, when N is 6, K is 3. If the output time of the first audio track is 10:24:22.82 on September 24, 2024, the output time of the second audio track is 10:24:23.06 on September 24, 2024, the output time of the third audio track is 10:24:23.23 on September 24, 2024, the output time of the fourth audio track is 10:24:23.70 on September 24, 2024, the output time of the fifth audio track is 10:24:24.11 on September 24, 2024, and the output time of the sixth audio track is 10:25:24.68 on September 24, 2024, then the difference between the output time of the second audio track and the output time of the first audio track is 0.24 seconds, which is the first output time difference among the three second output time differences. The difference between the output time of the 4th audio and the output time of the 3rd audio is 0.47 seconds, which is the second output time difference out of three. The difference between the output time of the 6th audio and the output time of the 5th audio is 0.57 seconds, which is the third output time difference out of three.
[0066] Step 202c: The electronic device determines the output delay between the image corresponding to the third output time difference and the audio corresponding to the fourth output time difference, and obtains K sets of output delays.
[0067] In some embodiments of this application, the third output time difference is any one of the K first output time differences, and the fourth output time difference is the output time difference with the smallest difference between the K second output time differences and the third output time difference.
[0068] For example, when K is 3, and the time differences of the three first outputs are 0.21 seconds, 0.43 seconds and 0.56 seconds respectively, and the time differences of the three second outputs are 0.24 seconds, 0.47 seconds and 0.57 seconds respectively, if the time difference of the third output is 0.21 seconds, the difference between 0.24 seconds and 0.21 seconds among the three second output time differences is the smallest, and 0.24 seconds is the fourth output time difference mentioned above.
[0069] In some embodiments of this application, combined with Figure 3 ,like Figure 4 As shown, step 202c above can be achieved through the following steps 202c1 and 202c2:
[0070] Step 202c1: The electronic device determines the output delay between the i-th frame image corresponding to the third output time difference and the i-th audio corresponding to the fourth output time difference as one of the output delays in a set of output delays.
[0071] In some embodiments of this application, the output delay between the i-th frame image and the i-th audio can be the difference between the output time of the i-th frame image and the output time of the i-th audio.
[0072] For example, when the third output time difference is 0.21 seconds and the fourth output time difference is 0.24 seconds, the output time of the first frame corresponding to the third output time difference is 10:24:23.05 on September 24, 2024, and the output time of the first audio corresponding to the fourth output time difference is 10:24:22.82 on September 24, 2024. Then, the difference of 0.23 seconds between the output time of the first frame and the output time of the first audio is one of the output delays in a set of output delays.
[0073] It is understandable that the motion features of the first image region in the i-th frame image match the audio features of the i-th audio.
[0074] Step 202c2: The electronic device determines the output delay between the (i+1)th frame image corresponding to the third output time difference and the (i+1)th audio corresponding to the fourth output time difference as another output delay in a set of output delays.
[0075] In some embodiments of this application, the output delay between the (i+1)th frame image corresponding to the third output time difference and the (i+1)th audio corresponding to the fourth output time difference can be the difference between the output time of the (i+1)th frame image corresponding to the third output time difference and the output time of the (i+1)th audio corresponding to the fourth output time difference.
[0076] For example, when the third output time difference is 0.21 seconds and the fourth output time difference is 0.24 seconds, the output time of the second frame corresponding to the third output time difference is 10:24:23.26 on September 24, 2024, and the output time of the second audio corresponding to the fourth output time difference is 10:24:23.06 on September 24, 2024. Then, the difference of 0.20 seconds between the output time of the second frame and the output time of the second audio is another output delay in a set of output delays.
[0077] It is understandable that the motion features of the first image region in the (i+1)th frame image match the audio features of the (i+1)th audio.
[0078] It should be noted that the electronic device can perform steps 202c1 and 202c2 in any order. The electronic device can perform step 202c1 first, then step 202c2. Alternatively, it can perform step 202c2 first, then step 202c1. Or, the electronic device can perform both steps 202c1 and 202c2 simultaneously.
[0079] Thus, the electronic device determines the output delay between the i-th frame image corresponding to the third output time difference and the i-th audio corresponding to the fourth output time difference as one output delay in a set of output delays, and determines the output delay between the (i+1)-th frame image corresponding to the third output time difference and the (i+1)-th audio corresponding to the fourth output time difference as another output delay in a set of output delays. Based on K sets of output delays, when determining the first output delay, the electronic device can adjust the output time of the image or audio in the second video according to the output delay between the image and audio that match the motion features and audio features, thereby matching the motion features of the image and the audio features of the audio in the second video, and thus synchronizing the display screen and sound of the second video.
[0080] Step 202d: The electronic device determines the first output delay based on the K group output delays.
[0081] In some embodiments of this application, combined with Figure 3 ,like Figure 5 As shown, step 202d above can be achieved through the following steps 202d1 and 202d2:
[0082] Step 202d1: The electronic device determines the maximum output delay and the minimum output delay from the K groups of output delays.
[0083] In some embodiments of this application, a set of output delays includes two output delays. The electronic device can have K sets of output delays, for a total of N output delays, sorted in ascending or descending order, and the maximum and minimum output delays are determined from the K sets of output delays.
[0084] For example, when N is 6 and K is 3, the three output delays are 0.23 seconds and 0.20 seconds, 0.24 seconds and 0.20 seconds, and 0.21 seconds and 0.20 seconds, respectively. Among the six output delays in total, the maximum output delay is 0.24 seconds and the minimum output delay is 0.20 seconds.
[0085] Step 202d2: The electronic device determines the output time difference between the first image and the second image in the M-frame images that meet the first condition as the first output delay.
[0086] In some embodiments of this application, the first condition includes: the lip-shape features of a first image region in one frame of the M-frame images match those of a second image region in another frame of the M-frame images; and the output time difference between the first frame and the other frame is greater than or equal to the minimum output delay and less than or equal to the maximum output delay.
[0087] In some embodiments of this application, the electronic device can extract facial features of a first object magnified in the first image region of each of the M frames to obtain lip features of the first image region of each of the M frames, and extract facial features of a first object in the second image region of each of the M frames to obtain lip features of the second image region of each of the M frames.
[0088] In some embodiments of this application, the electronic device can perform lip-shape feature matching on the lip-shape features of the first image region of each frame in the M-frame images and the lip-shape features of the second image region of each frame in the M-frame images, and determine the first image and the second image that satisfy the first condition from the M-frame images.
[0089] In some embodiments of this application, if the lip-sync feature of the first image region of the first image matches the lip-sync feature of the second image region of the second image, then the output time difference between the two images can be the difference between the output time of the second image and the output time of the first image; if the lip-sync feature of the second image region of the first image matches the lip-sync feature of the first image region of the second image, then the output time difference between the first image and the second image can be the difference between the output time of the first image and the output time of the second image.
[0090] It is understandable that, since the second image region of each frame in the first video corresponds to the region of the first object, there is no display delay in the second image region, and the second image region is synchronized with the audio. Therefore, when the electronic device adjusts the output time of the image or audio in the second video according to the output time difference between the first and second images that meet the first condition, it can synchronize the first image region of each frame in the second video with the audio, thereby synchronizing the display screen and audio in the second video.
[0091] Thus, the electronic device determines the first output delay by setting the output time difference between the first image and the second image that meet the first condition in the M-frame images as the first output delay, and then adjusts the output time of the image or audio in the second video according to the output time difference between the first image and the second image, thereby synchronizing the display screen and audio in the second video.
[0092] In this way, the electronic device determines the output delay between the image corresponding to the third output time difference and the audio corresponding to the fourth output time difference, and obtains K sets of output delays. Based on the K sets of output delays, the first output delay is determined. When the electronic device adjusts the output time of the image or video in the second video based on the first output delay, it can adjust the output time of the image or audio in the second video according to the output delay between the image and audio that match the motion features and audio features, thereby matching the motion features of the image and the audio features of the audio in the second video, and thus synchronizing the display of the second video with the sound.
[0093] In some embodiments of this application, step 202d above can also be implemented by the following steps 202d3 and 202d4:
[0094] Step 202d3: The electronic device calculates the average or median value of all output delays in the K groups of output delays;
[0095] Step 202d4: The electronic device determines the average or median value as the first output delay.
[0096] In some embodiments of this application, a set of output delays includes two output delays, and K sets of output delays contain a total of 2K output delays. The electronic device can calculate the average or median of the 2K output delays and use the average or median of the 2K output delays as the first output delay.
[0097] In this way, the electronic device calculates the average or median of all output delays in the K groups of output delays, and determines the average or median as the first output delay. When the electronic device adjusts the output time of the image or video in the second video based on the first output delay, it can adjust the output time of the image or audio in the second video according to the output delay between the image and audio that match the motion features and audio features, thereby matching the motion features of the image and the audio features of the audio in the second video, and thus synchronizing the display of the second video with the sound.
[0098] Step 203: The electronic device adjusts the output time of the image or the output time of the audio in the second video based on the first output delay.
[0099] In some embodiments of this application, the output time of the image in the adjusted second video is matched with the output time of the audio.
[0100] In some embodiments of this application, the second video may be the first video and a video of the target environment captured by the electronic device after capturing the first video, or the second video may be a video of the target environment captured by the electronic device after capturing the first video. This embodiment does not impose specific limitations here.
[0101] It is understood that matching the output time of the image and the output time of the audio in the adjusted second video means that the output time of the third image and the third audio in the second video are the same. Specifically, the third image and the third audio are the image and audio that match the motion features of the first image region in the second video with the audio features of the audio in the second video.
[0102] In some embodiments of this application, step 203 above can be implemented by either step 203a or step 203b:
[0103] Step 203a: The electronic device reduces the output time of the image in the second video by the first output delay.
[0104] In some embodiments of this application, the electronic device can reduce the output time of each frame in the second video by the first output delay.
[0105] Step 203b: The electronic device adds the first output delay to the output time of the audio in the second video.
[0106] In some embodiments of this application, the electronic device may add the aforementioned first output delay to the output time of each audio audio in the second video.
[0107] In some embodiments of this application, after the electronic device performs steps 203a or 203b as described above, it can encode the second video using a video encoder. For example, when encoding the second video using a video encoder, the electronic device can encode images and audio with the same output time as a pair of images and audio. This eliminates the display delay on the large screen in the first image area of the second video image, thereby synchronizing the picture and sound in the second video.
[0108] Thus, by reducing the output time of the image in the second video by the first output delay, or by increasing the output time of the audio in the second video by the first output delay, the electronic device can eliminate the display delay of the large screen in the first image area of the second video image, thereby synchronizing the picture and sound in the second video.
[0109] In the video processing method provided in this application embodiment, N frames of images and N audio audios are determined in a first video. The N frames of images are images with motion features in the first video, and the N audio audios are audios with preset audio features in the first video. Based on the N frames of images and N audio audios, a first output delay is determined. Based on the first output delay, the output time of the images or the output time of the audios in the second video are adjusted. The adjusted output time of the images in the second video matches the output time of the audios. Thus, by determining the first output delay based on the N frames of images and N audio audios, the output delay between images with motion features and audios with preset audio features in the video can be determined. Therefore, when adjusting the output time of the images or the output time of the audios in the second video based on the first output delay, the images with motion features in the second video can be matched with the audios with preset audio features, thereby synchronizing the images and sounds in the second video.
[0110] In some embodiments of this application, after step 204d2 above, the video processing method provided in the application embodiments may further include the following step 204 or the following step 205:
[0111] Step 204: The electronic device updates the second image region in the first image to the second image region in the second image.
[0112] In some embodiments of this application, the electronic device can segment and extract the second image to determine a second image region from the second image, and segment and extract the first image to determine a second image region from the first image. Further, the electronic device can update the second image region in the first image to the second image region in the second image, and then merge the image regions in the first image excluding the second image region with the second image region in the second image to form a new first image frame.
[0113] Step 205: The electronic device updates the first image region in the second image to the first image region in the first image.
[0114] In some embodiments of this application, the electronic device can segment and extract the second image to determine a first image region of the second image, and segment and extract the first image to determine a first image region of the first image. Further, the electronic device can update the first image region in the second image to the first image region in the second image, and then merge the image regions in the second image excluding the first image region with the first image region in the first image to form a new second image frame.
[0115] Thus, by updating the second image region in the first image to the second image region in the second image, or updating the first image region in the second image to the first image region in the first image, the electronic device can match the lip-sync features of the two image regions in the image, thereby synchronizing the display of the two image regions in the image.
[0116] In some embodiments of this application, prior to step 201 above, the video processing method provided in this application may further include steps 206 and 207 below, and step 201 above can be implemented through steps 201a and 201b below:
[0117] Step 206: The electronic device displays a preview image.
[0118] In some embodiments of this application, the preview image includes at least one object identifier and at least one audio identifier.
[0119] In some embodiments of this application, the object identifier may be an image of the object indicated by the object identifier.
[0120] In some embodiments of this application, the aforementioned audio identifier may be the audio of the object indicated by the aforementioned audio identifier.
[0121] For example, when the electronic device is a mobile phone, and a user records a video of a live concert, a preview image of the concert can be displayed on the phone. Figure 6 As shown, the preview image 60 on the mobile phone includes sub-regions 601 and 602. Sub-region 601 displays an image of person 1 at the concert, and sub-region 602 displays an image of person 2 at the concert. When person 1 speaks, the mobile phone can capture person 1's voice and display person 1's audio identifier 603 on the preview image 60, for example... Figure 6The sound 1 in the image. When person 2 is speaking, the phone can capture person 2's sound and display person 2's audio identifier 604 on the preview image 60, for example. Figure 6 The sound in the middle 2.
[0122] Step 207: The electronic device responds to the user's first input by associating the first object identifier and the first audio identifier.
[0123] In some embodiments of this application, the first object identifier is an object identifier among the at least one object identifier, and the first audio identifier is an audio identifier among the at least one audio identifier.
[0124] In some embodiments of this application, the first input mentioned above may include any of the following: user click input, swipe input, press input, voice input, gesture input, or other feasible inputs, which are not limited in this application embodiment.
[0125] In some embodiments of this application, the above-mentioned gesture input may include, but is not limited to, at least one of the following: click gesture, swipe gesture, drag gesture, pressure recognition gesture, long press gesture, area change gesture, double press gesture, double tap gesture, specific gesture input or other possible gesture inputs. The specific gesture input form can be determined according to actual needs, and is not limited in some embodiments.
[0126] In some embodiments of this application, the above-mentioned click input can be single-click input, double-click input, or any number of clicks, or it can be long-press input or short-press input. In some embodiments, this is not limited.
[0127] In some embodiments of this application, the above-mentioned sliding input can be a sliding input in any direction, such as sliding up, sliding down, sliding left, or sliding right, etc., and in some embodiments, this is not limited.
[0128] In some embodiments of this application, the first input may include two consecutive first sub-inputs and second sub-inputs. The first sub-input is the user's input on the first object identifier, and the second sub-input is the user's input on the first audio identifier. The first and second sub-inputs are used to indicate the association between the first object identifier and the first audio identifier.
[0129] In some embodiments of this application, the electronic device may associate the first object identifier and the first audio identifier in response to the first sub-input and the second sub-input.
[0130] For example, when the first sub-input and the first sub-input mentioned above are both press inputs, such as Figure 6As shown, the user can continuously press the image of person 1 and the audio identifier 603 of person 1 within the sub-region 601 of the preview image to associate the image of person 1 and the audio identifier of person 1.
[0131] Step 201a: The electronic device determines N frames of images corresponding to the first object from the first video based on the first object identifier.
[0132] In some embodiments of this application, the first object is identified as an image of the first object, and the electronic device can extract features from the image of the first object to obtain the features of the first object.
[0133] In some embodiments of this application, the electronic device can extract features from each frame of the first video and determine N frames of images that include features of the first object from the first video. The N frames of images in the first video that include features of the first object are the N frames of images corresponding to the first object.
[0134] Step 201b: The electronic device determines N audios corresponding to the first object from the first video based on the first audio identifier.
[0135] In some embodiments of this application, the first object is the object indicated by the first object identifier and the first audio identifier.
[0136] In some embodiments of this application, the first audio identifier is the audio of the first object, and the electronic device can extract the voiceprint of the first object from the audio of the first object to obtain the voiceprint of the first object.
[0137] In some embodiments of this application, the electronic device can extract the voiceprint of each audio file in the first video and determine N audio files that include the voiceprint of the first object from the first video. The N audio files in the first video that include the voiceprint of the first object are the N audio files corresponding to the first object.
[0138] Thus, the electronic device determines N frames of images corresponding to the first object from the first video based on the first object identifier, and determines N audios corresponding to the first object from the first video based on the first audio identifier. This enables the electronic device to determine the output delay between the image containing the motion features of the first object and the audio containing the preset audio features of the first object when determining the first output delay between the N frames of images and the N audios. Furthermore, when adjusting the output time of the images or the output time of the audios in the second video based on the first output delay, the electronic device can match the image containing the motion features of the first object in the second video with the audio containing the preset audio features of the first object, thereby synchronizing the image and sound of the first object in the second video.
[0139] The following is combined with Figure 7 Taking a mobile phone as an electronic device and a concert live video as the first and second videos as examples, this application provides a detailed explanation of the video processing method. Figure 7 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 7 As shown, the method may include steps 701 to 705 as follows.
[0140] Step 701: The mobile phone receives the user's video input.
[0141] For example, the above-mentioned video input can be the user's click input on the video recording mode button in the mobile phone camera.
[0142] Step 702: The mobile phone responds to the recording input and records the first video.
[0143] For example, the recording of the first video by the mobile phone includes the mobile phone's image sensor capturing M frames of images from the concert venue, and the mobile phone's speaker capturing M audio signals from the concert venue. Each of the M frames includes a first image region and a second image region. The first image region corresponds to the area on the large screen of the concert, and the second image region corresponds to the area of the performers at the concert venue.
[0144] For example, a mobile phone can use an image sensor to capture images in real time and send them to an ISP, which can then process the images captured in real time by the image sensor.
[0145] Step 703: The mobile phone calculates the audio-visual delay range of the first video based on the character's actions in the first image region of the first video and the audio features of the first video. Within the audio-visual delay range, the mobile phone calculates the audio-visual delay of the first video (i.e., the aforementioned first output delay) based on the character's lip movements in the second image region and the first image region of the first video.
[0146] For example, the mobile phone acquires N frames from the M frames of the first video and N audio clips from the M audio clips of the first video. Here, the N frames are images in the first image region of the M frames that show human motion features, and the N audio clips are audio clips from the M audio clips whose frequencies are greater than or equal to a preset frequency.
[0147] For example, a mobile phone can determine the audio-visual delay range based on N frames of images and N audio files, that is, the range in which the minimum output delay and the maximum output delay mentioned above are located.
[0148] For example, the mobile phone can determine that the lip-sync features of a first image region in one frame of an M-frame image match those of a second image region in another frame of an M-frame image, and that the output time difference between one frame and another frame is within the audio-visual delay range for two frames (i.e., the first image and the second image mentioned above). Then, it can calculate the output timestamps of these two frames as the audio-visual delay of the first video.
[0149] Step 704: The mobile phone uses the audio-visual delay of the first video to modify the output time of the image in the second video or the output time of the image in the second video.
[0150] For example, the second video can be the first video mentioned above, or a video captured by the electronic device at the concert venue after capturing the first video. Alternatively, the second video can be a video captured by the electronic device at the concert venue after capturing the first video. This embodiment does not impose any specific limitations here.
[0151] For example, the mobile phone can reduce the above-mentioned audio-visual delay by reducing the output delay of each frame of the second video, or increase the above-mentioned audio-visual delay by increasing the output delay of each audio in the second video.
[0152] Step 705: The mobile phone encodes the images and audios with the same output time in the second video as a pair of images and audios.
[0153] For example, after the mobile phone modifies the output time of an image in the second video, an encoder can be used to encode the second video. During encoding, the encoder can encode audio and images with output times aligned, resulting in a finished video with complete audio-visual alignment. In this way, the display delay on the large screen in the first image area of the second video image can be eliminated, thereby synchronizing the picture and sound in the second video.
[0154] The video processing method provided in this application calculates the output delay of the video's image and audio output to align the output time of the image and audio in the user's video, thereby reducing or eliminating the problem of audio-visual asynchrony in recorded concert videos and improving the user experience.
[0155] The following is combined with Figure 8 Using a mobile phone as an electronic device, and the first and second videos as videos from a live concert with multiple performers, this paper provides a detailed explanation of the video processing method provided in this application. Figure 8 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 8 As shown, the method may include steps 801 to 808 as follows.
[0156] Step 801: The mobile phone displays a preview interface (i.e., the preview image above). The preview interface includes at least one person icon and at least one audio icon.
[0157] For example, a person identifier can be an image of a person, and an audio identifier can be the audio of a person.
[0158] For example, during the opening stage of a stage performance, when the performers introduce themselves, the preview screen on the phone will display images and audio of each performer at the concert.
[0159] Step 802: The mobile phone responds to the user's first input and associates the first person identifier and the first audio identifier.
[0160] For example, the first object identifier is a person identifier among at least one person identifier, and the first audio identifier is an audio identifier among at least one audio identifier.
[0161] For example, when an actor is speaking at a concert, the phone can capture the actor's voice and display the actor's audio identifier on a preview screen. Simultaneously, the actor's speaking status will be displayed on their image in the preview screen. At this point, the user can input the actor's image and audio identifier as initial input, linking the actor's image and audio.
[0162] Step 803: The mobile phone receives the user's video input.
[0163] Step 804: The mobile phone responds to the recording input and records the first video.
[0164] It should be noted that steps 803 and 804 are similar to steps 701 and 702, and will not be repeated here to avoid repetition.
[0165] Step 805: The mobile phone calculates the audio-visual delay range of the first video based on the character's actions in the first image region of the third image in the first video and the audio features of the third audio in the first video. Within the audio-visual delay range, the mobile phone calculates the audio-visual delay of the first video based on the lip movements of the first character in the second image region of the first video and the lip movements of the first character in the first image region.
[0166] For example, the third image is N frames of images in the first image region of M frames, in which the first person has the characteristics of human movement, and the N videos are audios in M audios, in which the audio frequency of the first person is greater than or equal to a preset frequency.
[0167] For example, a mobile phone can determine the audio-visual delay range based on N frames of images and N audio files, that is, the range in which the minimum output delay and the maximum output delay mentioned above are located.
[0168] For example, the mobile phone can determine that the lip-sync feature of the first person in the first image region of one frame of the M-frame images matches the lip-sync feature of the first person in the second image region of another frame of the M-frame images, and that the output time difference between one frame and another frame is within the audio-visual delay range for two frames (i.e. the first image and the second image mentioned above), and then calculate the output timestamps of these two frames as the audio-visual delay of the first video.
[0169] Step 806: The mobile phone updates the first image region of one frame to the first image region of the other frame in two frames that match the lip-shape features of the first person in the second image region and the lip-shape features of the first person in the first image region, or updates the second image region of one frame to the second image region of the other frame.
[0170] For example, the electronic device can segment and extract each frame of the first video to determine a first image region and a second image region for each frame. Further, the mobile phone can update the first image region of one frame to the first image region of another frame, and then merge the image regions (excluding the first image region) of the first frame with the first image region of the other frame into a new image. Similarly, the mobile phone can update the second image region of another frame to the second image region of the first frame, and then merge the image regions (excluding the second image region) of the other frame with the second image region of the first frame into a new image.
[0171] In this way, by matching the lip-sync features of the first person in the second image region with the lip-sync features of the first person in the first image region, the mobile phone can update the first image region of one frame to the first image region of the other, or update the second image region of one frame to the second image region of the other, thereby achieving full-scene alignment of the real person, screen, and sound in the images of the first video and improving the overall video viewing experience.
[0172] Step 807: The mobile phone uses the audio-visual delay of the first video to modify the output time of the image in the second video or the output time of the image in the second video.
[0173] Step 808: The mobile phone encodes the images and audios with the same output time in the second video as a pair of images and audios.
[0174] It should be noted that the implementation process of steps 807 and 808 is similar to that of steps 704 and 705. To avoid repetition, this embodiment will not describe them again here.
[0175] In the video processing method provided in this application embodiment, on the one hand, by displaying character and audio identifiers of multiple actors on the stage in the preview interface, the character and audio identifiers of different actors can be associated, thereby more accurately identifying the lip movements and audio of different actors, and thus aligning the lip movements and audio of different actors, increasing the audio-visual synchronization accuracy of the video. On the other hand, by segmenting and extracting the image, the different image regions of the live actors and the large screen in the image can be determined, thereby segmenting and aligning the live actors and the large screen. This not only aligns the large screen and sound in the video, but also aligns the output delay difference between the live actors and the large screen, thereby achieving full synchronization of the entire scene.
[0176] It should be noted that the specific implementation process of the video processing method in the above steps can be found in the relevant description of the above embodiments. To avoid repetition, this embodiment will not repeat the details here.
[0177] It should be noted that each of the above method embodiments, or various possible implementations of each method embodiment, can be executed individually or in combination of any two or more. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.
[0178] The video processing method provided in this application can be executed by a video processing device. This application uses a video processing device executing the video processing method as an example to illustrate the video processing device provided in this application.
[0179] Figure 9 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application. The video processing device 900 includes: a determining module 901 and an adjusting module 902.
[0180] The determining module 901 is used to determine N frames of images and N audios in the first video. The N frames of images are images in the first video that have motion features, and the N audios are audios in the first video that have preset audio features. N is a positive integer.
[0181] The determining module 901 is further configured to determine the first output delay based on the N frames of images and the N audio files;
[0182] The adjustment module 902 is used to adjust the output time of the image or the output time of the audio in the second video based on the first output delay; wherein the image and audio in the adjusted second video are synchronized.
[0183] In some embodiments of this application, the determining module 901 is specifically used for:
[0184] Determine the output time difference between the i-th frame and the (i+1)-th frame in N frames, and obtain K first output time differences, i∈[1,3,...,N), where K is N / 2 and K is a positive integer;
[0185] Determine the output time difference between the i-th audio and the (i+1)-th audio among N audios, and obtain K second output time differences;
[0186] Determine the output delay between the image corresponding to the third output time difference and the audio corresponding to the fourth output time difference to obtain K sets of output delays. The third output time difference is any one of the K first output time differences, and the fourth output time difference is the output time difference with the smallest difference from the third output time difference among the K second output time differences.
[0187] The first output delay is determined based on the output delay of K groups.
[0188] In some embodiments of this application, the determining module 901 is specifically used for:
[0189] The output delay between the i-th frame image corresponding to the third output time difference and the i-th audio corresponding to the fourth output time difference is determined as one of the output delays in a set of output delays;
[0190] The output delay between the (i+1)th frame image corresponding to the third output time difference and the (i+1)th audio corresponding to the fourth output time difference is determined as another output delay in a set of output delays.
[0191] In some embodiments of this application, the first video is a video captured from a target environment, the target environment including a first display screen and a first object, the first display screen being used to display the first object captured by the camera, the first video including M frames of images, each frame including a first image area and a second image area, the first image area being an image area corresponding to the first display screen, and the second image area being an image area corresponding to the first object;
[0192] Here, N frames are images in the first image region of M frames that have motion features, M is a positive integer, and M is greater than or equal to N.
[0193] In some embodiments of this application, the determining module 901 is specifically used for:
[0194] Determine the maximum and minimum output delays from the K groups of output delays;
[0195] The output time difference between the first image and the second image in the M-frame images that satisfy the first condition is determined as the first output delay;
[0196] The first condition includes: the lip-sync features of a first image region in one frame match those of a second image region in another frame, and the output time difference between one frame and the other frame is greater than or equal to the minimum output delay and less than or equal to the maximum output delay.
[0197] In some embodiments of this application, combined with Figure 9 ,like Figure 10 As shown, the device 900 also includes:
[0198] The update module 903 is used to update the second image region in the first image to the second image region in the second image; or, update the first image region in the second image to the first image region in the first image.
[0199] In some embodiments of this application, the determining module 901 is specifically used for:
[0200] Calculate the average or median of all output delays in the K groups of output delays;
[0201] The average or median value is determined as the first output delay.
[0202] In some embodiments of this application, the device 900 further includes:
[0203] Display module 904 is used to display a preview image, the preview image including at least one object identifier and at least one audio identifier; association module is used to associate the first object identifier and the first audio identifier in response to a first input from the user, the first object identifier being an object identifier among at least one object identifier, and the first audio identifier being an audio identifier among at least one audio identifier;
[0204] The determining module 901 is specifically used to: determine N frames of images corresponding to the first object from the first video based on the first object identifier;
[0205] Based on the first audio identifier, determine N audio files corresponding to the first object from the first video;
[0206] The first object is the object indicated by the first object identifier and the first audio identifier.
[0207] In some embodiments of this application, the adjustment module 902 is specifically used for:
[0208] Reduce the output time of the image in the second video by the delay of the first output; or...
[0209] Increase the output time of the audio in the second video by the delay of the first output.
[0210] In the video processing apparatus provided in this application, N frames of images and N audio audios are determined in a first video, where the N frames are images with motion features and the N audio audios are audios with preset audio features. Based on the N frames of images and the N audio audios, a first output delay is determined. Based on the first output delay, the output time of the images or the output time of the audios in the second video are adjusted. The adjusted second video synchronizes the images and audios. Thus, by determining the first output delay based on the N frames of images and the N audio audios, the output delay between images with motion features and audios with preset audio features in the video can be determined. Therefore, when adjusting the output time of the images or the output time of the audios in the second video based on the first output delay, the images with motion features and the audios with preset audio features in the second video can be matched, thereby synchronizing the images and sound in the second video.
[0211] The video processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device, augmented reality / virtual reality device, robot, wearable device, super mobile personal computer, netbook, or personal digital assistant, etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific devices.
[0212] The video processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0213] The video processing apparatus provided in this application can implement the various processes implemented in the various embodiments of the above-described video processing method. To avoid repetition, it will not be described again here.
[0214] Optionally, such as Figure 11 As shown, this application embodiment also provides an electronic device 1100, including a processor 1101 and a memory 1102. The memory 1102 stores a program or instructions that can run on the processor 1101. When the program or instructions are executed by the processor 1101, they implement the various steps of the above-described video processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0215] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0216] Figure 12 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0217] The electronic device 1200 includes, but is not limited to, components such as: radio frequency unit 1201, network module 1202, audio output unit 1203, input unit 1204, sensor 1205, display unit 1206, user input unit 1207, interface unit 1208, memory 1209, and processor 1210.
[0218] Those skilled in the art will understand that the electronic device 1200 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1210 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 12 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0219] The processor 1210 is configured to determine N frames of images and N audio audios in a first video, wherein the N frames of images are images with motion features in the first video, and the N audio audios are audios with preset audio features in the first video, and N is a positive integer; determine a first output delay based on the N frames of images and N audio audios; and adjust the output time of the images or the output time of the audio in a second video based on the first output delay; wherein the adjusted images and audio in the second video are synchronized.
[0220] In some embodiments of this application, the processor 1210 is specifically configured to: determine the output time difference between the i-th frame and the (i+1)-th frame in N frames of images, obtaining K first output time differences, i∈[1, 3, ..., N), where K is N / 2 and K is a positive integer; determine the output time difference between the i-th audio and the (i+1)-th audio in N audio, obtaining K second output time differences; determine the output delay between the image corresponding to the third output time difference and the audio corresponding to the fourth output time difference, obtaining K sets of output delays, where the third output time difference is any one of the K first output time differences, and the fourth output time difference is the output time difference with the smallest difference from the third output time difference among the K second output time differences; and determine the first output delay based on the K sets of output delays.
[0221] In some embodiments of this application, the processor 1210 is specifically configured to determine the output delay between the i-th frame image corresponding to the third output time difference and the i-th audio corresponding to the fourth output time difference as one of a set of output delays; and to determine the output delay between the (i+1)-th frame image corresponding to the third output time difference and the (i+1)-th audio corresponding to the fourth output time difference as another of a set of output delays.
[0222] In some embodiments of this application, the first video is a video captured from a target environment, which includes a first display screen and a first object. The first display screen is used to display the first object captured by the camera. The first video includes M frames of images, each frame of images including a first image region and a second image region. The first image region is the image region corresponding to the first display screen, and the second image region is the image region corresponding to the first object. Among them, N frames of images are images in the first image region of the M frames that have motion features, M is a positive integer, and M is greater than or equal to N.
[0223] In some embodiments of this application, the processor 1210 is specifically configured to determine the maximum output delay and the minimum output delay from K groups of output delays; and to determine the output time difference between the first image and the second image in M frames that satisfy a first condition as the first output delay; the first condition includes: the lip features of the first image region in one frame match the lip features of the second image region in another frame, and the output time difference between one frame and another frame is greater than or equal to the minimum output delay and less than or equal to the maximum output delay.
[0224] In some embodiments of this application, the processor 1210 is further configured to update a second image region in a first image to a second image region in a second image; or, update a first image region in a second image to a first image region in a first image.
[0225] In some embodiments of this application, the processor 1210 is specifically used to calculate the average or median of all output delays in the K groups of output delays; and to determine the average or median as the first output delay.
[0226] In some embodiments of this application, a display unit 1206 is used to display a preview image, the preview image including at least one object identifier and at least one audio identifier; a processor 1210 is used to associate a first object identifier and a first audio identifier in response to a first input from a user, wherein the first object identifier is an object identifier among at least one object identifier, and the first audio identifier is an audio identifier among at least one audio identifier; the processor 1210 is specifically used to determine N frames of images corresponding to a first object from a first video based on the first object identifier; and to determine N audios corresponding to the first object from the first video based on the first audio identifier; wherein the first object is the object indicated by the first object identifier and the first audio identifier.
[0227] In some embodiments of this application, the processor 1210 is specifically configured to reduce the output time of the image in the second video by a first output delay; or to increase the output time of the audio in the second video by a first output delay.
[0228] In the electronic device provided in this application, N frames of images and N audio recordings in a first video are determined. The N frames are images with motion features in the first video, and the N audio recordings are audio recordings with preset audio features in the first video. Based on the N frames of images and the N audio recordings, a first output delay is determined. Based on the first output delay, the output time of the images or the output time of the audio recordings in the second video are adjusted. The adjusted images and audio recordings in the second video are synchronized. Thus, by determining the first output delay based on the N frames of images and the N audio recordings, the output delay between images with motion features and audio recordings with preset audio features in the video can be determined. Therefore, when adjusting the output time of the images or the output time of the audio recordings in the second video based on the first output delay, the images with motion features and the audio recordings with preset audio features in the second video can be matched, thereby synchronizing the images and sound in the second video.
[0229] It should be understood that, in this embodiment, the input unit 1204 may include a graphics processing unit (GPU) 12041 and a microphone 12042. The GPU 12041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1206 may include a display panel 12061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1207 includes a touch panel 12071 and at least one of other input devices 12072. The touch panel 12071 is also called a touch screen. The touch panel 12071 may include a touch detection device and a touch controller. Other input devices 12072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0230] The memory 1209 can be used to store software programs and various data. The memory 1209 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1209 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1209 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0231] Processor 1210 may include one or more processing units; optionally, processor 1210 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1210.
[0232] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0233] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0234] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above video processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0235] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0236] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the video processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0237] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0238] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0239] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A video processing method, characterized in that, The method includes: N frames of images and N audio audios are determined in the first video. The N frames of images are images in the first video that have motion features, and the N audio audios are audios in the first video that have preset audio features. N is a positive integer. Based on the N frames of images and the N audio files, the first output delay is determined; Based on the first output delay, the output time of the image or the output time of the audio in the second video are adjusted; wherein, the image and audio in the adjusted second video are synchronized. The determination of the first output delay based on the N frames of images and the N audio audios includes: Determine the output time difference between the i-th frame and the (i+1)-th frame in the N frames, and obtain K first output time differences, where N is an even number greater than or equal to 2, i is an odd number (i∈[1, 3, ..., N)), and K is N / 2 and K is a positive integer; Determine the output time difference between the i-th audio and the (i+1)-th audio among the N audios to obtain K second output time differences; Determine the output delay between the image corresponding to the third output time difference and the audio corresponding to the fourth output time difference to obtain K sets of output delays. The third output time difference is any one of the K first output time differences, and the fourth output time difference is the output time difference with the smallest difference between the third output time difference and the K second output time differences. The first output delay is determined based on the K sets of output delays.
2. The method according to claim 1, characterized in that, The process of determining the output delay between the image corresponding to the third output time difference and the audio corresponding to the fourth output time difference yields K sets of output delays, including: The output delay between the i-th frame image corresponding to the third output time difference and the i-th audio corresponding to the fourth output time difference is determined as one of the output delays in a set of output delays; The output delay between the (i+1)th frame image corresponding to the third output time difference and the (i+1)th audio file corresponding to the fourth output time difference is determined as another output delay in the set of output delays.
3. The method according to claim 1, characterized in that, The first video is a video captured from a target environment, which includes a first display screen and a first object. The first display screen is used to display the first object captured by the camera. The first video includes M frames of images. Each frame of images includes a first image area and a second image area. The first image area is the image area corresponding to the first display screen, and the second image area is the image area corresponding to the first object in the target environment. Wherein, the N frames are images in the first image region of the M frames that have motion features, M is a positive integer, and M is greater than or equal to N.
4. The method according to claim 3, characterized in that, Determining the first output delay based on the K groups of output delays includes: The maximum and minimum output delays are determined from the K sets of output delays; The output time difference between the first image and the second image in the M-frame images that satisfy the first condition is determined as the first output delay; The first condition includes: the lip-sync features of a first image region in one frame match those of a second image region in another frame, and the output time difference between the one frame and the other frame is greater than or equal to the minimum output delay and less than or equal to the maximum output delay.
5. The method according to claim 4, characterized in that, The method further includes: Update the second image region in the first image to the second image region in the second image; or... Update the first image region in the second image to the first image region in the first image.
6. The method according to claim 3, characterized in that, Determining the first output delay based on the K groups of output delays includes: Calculate the average or median of all output delays in the K groups of output delays; The average or median value is determined as the first output delay.
7. The method according to claim 1, characterized in that, The method further includes: Display a preview image, which includes at least one object identifier and at least one audio identifier; In response to a user's first input, a first object identifier and a first audio identifier are associated, wherein the first object identifier is an object identifier among the at least one object identifier, and the first audio identifier is an audio identifier among the at least one audio identifier; The determination of N frames of images and N audio clips in the first video includes: Based on the first object identifier, determine the N frames of images corresponding to the first object from the first video; Based on the first audio identifier, determine the N audios corresponding to the first object from the first video; The first object is the object indicated by the first object identifier and the first audio identifier.
8. The method according to claim 1, characterized in that, The step of adjusting the output time of the image or the output time of the audio in the second video based on the first output delay includes: Reduce the output time of the image in the second video by the first output delay; or... The output time of the audio in the second video is increased by the first output delay.
9. A video processing apparatus, characterized in that, The device includes: a determining module and an adjusting module; wherein... The determining module is used to determine N frames of images and N audios in the first video, wherein the N frames of images are images in the first video that have motion features, and the N audios are audios in the first video that have preset audio features, and N is a positive integer; The determining module is used to determine the first output delay based on the N frames of images and the N audio files; The adjustment module is used to adjust the output time of the image or the output time of the audio in the second video based on the first output delay; wherein the image and audio in the adjusted second video are synchronized. The determining module is specifically used for: Determine the output time difference between the i-th frame and the (i+1)-th frame in the N frames, and obtain K first output time differences, where N is an even number greater than or equal to 2, i is an odd number (i∈[1, 3, ..., N)), and K is N / 2 and K is a positive integer; Determine the output time difference between the i-th audio and the (i+1)-th audio among the N audios to obtain K second output time differences; Determine the output delay between the image corresponding to the third output time difference and the audio corresponding to the fourth output time difference to obtain K sets of output delays. The third output time difference is any one of the K first output time differences, and the fourth output time difference is the output time difference with the smallest difference between the third output time difference and the K second output time differences. The first output delay is determined based on the K sets of output delays.
10. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the video processing method as described in any one of claims 1 to 8.
11. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the video processing method as described in any one of claims 1 to 8.