Video processing method, electronic device and computer readable storage medium

By automatically detecting the image frames of specified actions in the video and processing it, high-quality slow motion or fast motion videos can be generated without manual operation by users, solving the problem that ordinary users find it difficult to shoot these videos.

CN117014686BActive Publication Date: 2025-05-06HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210469119.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-05-06
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

It is difficult for ordinary users to shoot ideal slow motion or fast motion videos, and they need to operate manually and have an understanding of the narrative scene and adjust professional parameters.

Method used

Provide a video processing method, which automatically detects image frames containing specified actions in the video, and collects image frames at the same shooting frame rate as the playback frame rate when recording video, generates processed video files, and realizes automatic slow motion or fast motion playback.

Benefits of technology

Slow-motion or fast-motion videos can be generated without manual operation of the user, which improves the user experience, and using the same shooting frame rate as the playback frame rate can support the use of advanced capabilities such as DCG and PDAF to improve video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117014686B_ABST
    Figure CN117014686B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a video processing method and an electronic device, which relate to the field of terminal technology. When recording a video, the electronic device captures image frames at a first frame rate to generate an original video file. The image frames containing the specified action in the original video file are automatically inserted or extracted. The processed video file is played at the first frame rate. The image frames containing the specified action in a video can be automatically played in slow motion or fast motion without the user manually capturing the specified action; and the use of advanced capabilities when recording videos is supported to ensure video quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a video processing method, an electronic device, and a computer-readable storage medium. Background Art

[0002] Generally speaking, electronic devices use a fixed frame rate to play videos. For example, the frame rate of movies is 24fps (frames per second). 24fps brings natural motion blur, giving moving images a smooth look and feel, which means 24fps represents a movie feel. The video playback frame rate of mobile phones is usually 24fps.

[0003] Usually, the video shooting frame rate is the same as the playback frame rate, which can give users a natural and smooth viewing experience. In some scenarios, the video shooting frame rate can also be different from the playback frame rate. For example, if the video shooting frame rate is higher than the playback frame rate, a slow motion effect will be produced; if the video shooting frame rate is lower than the playback frame rate, a fast motion effect will be produced. Slow motion or fast motion is conducive to expressing special emotions and improving user experience. Slow motion and fast motion shooting are widely welcomed by users.

[0004] However, the slow-motion and fast-motion shooting techniques involve the understanding of narrative scenes and the adjustment of professional parameters; ordinary users find it difficult to master them and often fail to capture ideal slow-motion or fast-motion videos. Summary of the invention

[0005] The embodiments of the present application provide a video processing method and an electronic device, which can automatically generate a slow-motion video or a fast-motion video with good effect without manual operation by the user.

[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0007] In a first aspect, a video processing method is provided, which is applied to an electronic device, wherein the electronic device includes a camera device, and the method includes: in response to an operation of starting video recording, the electronic device collects image frames at a first frame rate through the camera device; upon receiving an operation of ending the recording of the video, the electronic device stops collecting image frames and generates a first video file; wherein the first video file includes a first video portion composed of a first image frame and a second video portion composed of a second image frame, and the first image frame includes a specified action. The electronic device processes the first video file to generate a second video file; the second video file includes a third video portion and a second video portion, wherein the third video portion is obtained by processing the first video portion; the third video portion has a different number of image frames from the first video portion; and the electronic device plays the second video file at the first frame rate.

[0008] In this method, when recording a video, the electronic device captures image frames at a first frame rate to generate an original video file. After the recording is finished, the image frames containing the specified action in the original video file are automatically processed (such as frame insertion or frame extraction). And the processed video file is played at the first frame rate. Automatic slow-motion playback or fast-motion playback of image frames containing specified actions in a video is achieved without the user having to manually capture the specified action. And recording the video at the same shooting frame rate as the playback frame rate can avoid recording at a high frame rate, so that advanced capabilities such as DCG and PDAF can be supported when recording videos to improve video quality.

[0009] In one example, upon receiving the operation of sharing a video, the electronic device forwards the second video file. The processed video file is forwarded. In this way, another electronic device receives the second video file and plays the second video file, thereby realizing slow motion playback or fast motion playback of the motion interval.

[0010] In conjunction with the first aspect, in one implementation, before the electronic device processes the first video file, the method further includes: the electronic device marks the image frames containing the specified action in the first video file to generate marking information, wherein the marking information includes the specified action start information and the specified action end information.

[0011] In this method, after successfully recording a video to obtain an original video stream, the image frames containing the specified action in the original video stream are marked. In this way, the image frames containing the specified action can be directly determined based on the marking information and the original video stream.

[0012] In combination with the first aspect, in one embodiment, before the electronic device marks the image frames containing the specified action in the first video file, the method also includes: reducing the resolution of the image frames captured by the camera device to obtain corresponding low-resolution image frames; and detecting the specified action on the low-resolution image frames.

[0013] In this method, the low-resolution preview stream is analyzed instead of the full-resolution preview stream, which can increase the processing speed of the video pre-processing algorithm unit and improve performance.

[0014] In combination with the first aspect, in one implementation, the electronic device obtains the first video portion according to the tag information and the first video file.

[0015] In combination with the first aspect, in one embodiment, the method also includes: receiving an operation to edit a video, the electronic device displays a first interface; the first interface includes part or all of the image frames of the first video file; receiving an operation by the user to modify the range of the image frame interval containing a specified action on the first interface, the electronic device updates the marking information according to the modified range of the image frame interval containing the specified action.

[0016] In this method, the user is supported to manually adjust the image interval containing the specified action.

[0017] In combination with the first aspect, in one implementation, after receiving an operation to play a video, the electronic device processes the first video file.

[0018] In combination with the first aspect, in one implementation, the electronic device performs frame insertion processing on the first sub-video file; the playback duration of the second video file is greater than the shooting duration of the first video file, so that automatic slow motion playback is achieved.

[0019] In combination with the first aspect, in one implementation, the electronic device performs frame extraction processing on the first video file; and the playback time of the second video file is shorter than the shooting time of the first video file, that is, automatic fast motion playback is achieved.

[0020] In combination with the first aspect, in one embodiment, the electronic device includes a recording device, and the method further includes: in response to an operation of starting video recording, the electronic device collects audio frames through the recording device; upon receiving an operation of ending video recording, the electronic device stops collecting audio frames and generates a first audio frame in a first video file; the first audio frame includes a first audio part corresponding to the timeline of the first video part and a second audio part corresponding to the timeline of the second video part; speech recognition is performed on the first audio part to generate text corresponding to the first audio sub-part containing speech in the first audio part; when the electronic device plays the second video file, the text is displayed in the first video sub-part in the third video part in the form of subtitles; wherein the first video sub-part in the third video part is obtained by interpolating the second video sub-part in the first video part, and the second video sub-part is an image frame corresponding to the timeline of the first audio sub-part.

[0021] In this method, if the specified action interval in the video contains speech, the speech is recognized as text, and the text is displayed in the form of subtitles in the image frame after slow motion processing.

[0022] In one implementation, the duration of the first audio sub-portion is a first duration, and the display duration of the text is N times the first duration; N is an interpolation multiple of the interpolation processing.

[0023] In one implementation, the duration of the audio frame corresponding to the first character in the text is the first duration, and the display duration of the first character is N times the first duration; N is the interpolation multiple of the interpolation processing.

[0024] That is, the subtitle display duration matches the slow-motion processed image.

[0025] In combination with the first aspect, in one implementation, the first frame rate is 24 frames per second.

[0026] In a second aspect, a video processing method is provided, which is applied to an electronic device, the method comprising: the electronic device obtains a first video file, the first video file includes a first image frame and a first audio frame; the shooting frame rate of the first image frame is a first frame rate; the first audio frame includes a first audio part composed of a second audio frame and a second audio part composed of a third audio frame; the third audio frame contains voice; the first image frame includes a first video part composed of a second image frame and a second video part composed of a third image frame; the second image frame corresponds to the second audio frame on a timeline, and the third image frame corresponds to the third audio frame on a timeline; the electronic device processes the first video file to generate a second video file; the second video file includes a third video part and a second video part, the third video part is obtained by processing the first video part; the number of image frames of the third video part and the first video part are different; the electronic device plays the second video file at the first frame rate.

[0027] In this method, an electronic device obtains a video, which can be shot by the electronic device or received from other electronic devices. The image frames that do not contain voice in the video file are automatically processed (such as frame insertion or frame extraction). The processed video file is played at a playback frame rate equal to the shooting frame rate. The part of a video that does not contain voice is automatically played in slow motion or fast motion without manual processing by the user; the part containing voice is not processed to retain the original sound.

[0028] In conjunction with the second aspect, in one implementation, upon receiving the video sharing operation, the electronic device forwards the second video file. In this way, another electronic device receives the second video file and plays the second video file, thereby achieving slow motion playback or fast motion playback of the motion interval.

[0029] In conjunction with the second aspect, in one implementation, the method further includes: when the electronic device plays the third video portion in the second video file, the audio frame stops playing; when the electronic device plays the second video portion in the second video file, the third audio frame plays. That is, when playing a video section containing voice, the original audio of the video is played.

[0030] In combination with the second aspect, in one embodiment, the method also includes: when the electronic device plays the third video part in the second video file, the soundtrack is played at a first volume; when the electronic device plays the second video part in the second video file, the soundtrack is played at a second volume; the second volume is smaller than the first volume, and the second volume is smaller than the playback volume of the original sound of the video.

[0031] In this method, music is automatically added to the processed video, and when playing fast motion or slow motion, the sound of the music is higher; when playing a normal speed video, the original sound of the video is played, and the sound of the music is lower.

[0032] In one implementation, the corresponding soundtrack may be matched according to the overall atmosphere of the video.

[0033] In conjunction with the second aspect, in one implementation, the electronic device processes the first sub-video file including: the electronic device performs frame insertion processing on the first video file; and the playback duration of the second video file is greater than the shooting duration of the first video file, that is, automatic slow motion playback is achieved.

[0034] In conjunction with the second aspect, in one implementation, the electronic device processes the first sub-video file including: the electronic device performs frame extraction processing on the first video file; and the playback duration of the second video file is less than the shooting duration of the first video file, that is, automatic fast motion playback is achieved.

[0035] In combination with the second aspect, in one embodiment, the electronic device includes a camera and a recording device, and the electronic device obtains a first video file including: in response to an operation of starting video recording, the electronic device captures image frames at a first frame rate through the camera and captures audio frames through the recording device; when an operation of ending video recording is received, the electronic device stops capturing image frames and stops capturing audio frames to generate the first video file.

[0036] In conjunction with the second aspect, in one implementation, the electronic device receives an operation of playing the video before processing the first video file, that is, after receiving the operation of playing the video, the electronic device automatically processes the image frames that do not contain voice.

[0037] In combination with the second aspect, in one embodiment, before the electronic device processes the first video file, the method also includes: before receiving the operation of playing the video, the electronic device marks the image frame corresponding to the third audio frame timeline and generates marking information; the electronic device obtains the first video part based on the marking information and the first video file.

[0038] In combination with the second aspect, in one embodiment, when an operation to edit a video is received, the electronic device displays a first interface; the first interface includes part or all of the image frames of the first video file; when an operation to modify the image frame interval range corresponding to the audio frame containing speech is received on the first interface by the user, the electronic device updates the marking information according to the modified image frame interval range corresponding to the audio frame containing speech.

[0039] In this method, the user is supported to manually adjust the image interval for fast motion or slow motion processing.

[0040] In conjunction with the second aspect, in one implementation, the first frame rate is 24 frames per second.

[0041] In a third aspect, an electronic device is provided, which has the function of implementing the method described in the first aspect or the second aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0042] In a fourth aspect, an electronic device is provided, comprising: a processor and a memory; the memory is used to store computer execution instructions, and when the electronic device is running, the processor executes the computer execution instructions stored in the memory to enable the electronic device to perform a method as described in any one of the first or second aspects above.

[0043] In a fifth aspect, an electronic device is provided, comprising: a processor; the processor is used to couple with a memory, and after reading instructions in the memory, execute a method as described in any one of the first aspect or the second aspect according to the instructions.

[0044] In a sixth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is run on a computer, the computer can execute any of the methods in the first aspect or the second aspect.

[0045] In a seventh aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the methods in the first or second aspect.

[0046] In an eighth aspect, a device (for example, the device may be a chip system) is provided, the device including a processor for supporting an electronic device to implement the functions involved in the first aspect or the second aspect. In one possible design, the device also includes a memory for storing program instructions and data necessary for the electronic device. When the device is a chip system, it may be composed of a chip, or may include a chip and other discrete devices.

[0047] Among them, the technical effects brought about by any design method in the third aspect to the eighth aspect can refer to the technical effects brought about by different design methods in the first aspect or the second aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1A A schematic diagram of an example of video acquisition and playback;

[0049] Figure 1B A schematic diagram of an example of video acquisition and playback;

[0050] Figure 1C A schematic diagram of an example of video acquisition and playback;

[0051] Figure 2A A schematic diagram of a scenario example applicable to the video processing method provided in an embodiment of the present application;

[0052] Figure 2B A schematic diagram of a scenario example applicable to the video processing method provided in an embodiment of the present application;

[0053] Figure 2C A schematic diagram of a scenario example applicable to the video processing method provided in an embodiment of the present application;

[0054] Figure 2D A schematic diagram of a scenario example applicable to the video processing method provided in an embodiment of the present application;

[0055] Figure 3 A schematic diagram of the hardware structure of an electronic device applicable to the video processing method provided in an embodiment of the present application;

[0056] Figure 4 A schematic diagram of the software architecture of an electronic device applicable to the video processing method provided in an embodiment of the present application;

[0057] Figure 5 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0058] Figure 6 A schematic diagram of a video processing method provided in an embodiment of the present application;

[0059] Figure 7 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0060] Figure 8 A schematic diagram of a video processing method provided in an embodiment of the present application;

[0061] Fig. 9 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0062] Fig.10 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0063] Fig.11 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0064] Fig. 12A A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0065] Fig. 12BA schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0066] Fig.13A A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0067] Fig. 13B A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0068] Fig.14 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0069] Fig.15 A schematic diagram of a flow chart of a video processing method provided in an embodiment of the present application;

[0070] Fig.16 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0071] Fig.17 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0072] Fig.18 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0073] Fig.19 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0074] Fig. 20 A schematic diagram of a scenario example of a video processing method provided in an embodiment of the present application;

[0075] Fig.21 A schematic diagram of a flow chart of a video processing method provided in an embodiment of the present application;

[0076] Fig. 22 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0077] The technical solutions in the embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to be used as limitations on the present application. As used in the specification and the appended claims of the present application, the singular expressions "a", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear indication to the contrary in the context. It should also be understood that in the following embodiments of the present application, "at least one", "one or more" refer to one or more (including two). The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in a "or" relationship.

[0078] References to "one embodiment" or "some embodiments" etc. described in this specification mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Thus, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways. The term "connection" includes direct connection and indirect connection, unless otherwise specified. "First" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated.

[0079] In the embodiments of the present application, the words "exemplarily" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific way.

[0080] Users can use electronic devices to record videos and generate video files; they can also play video files, that is, play videos. Video files include video streams and audio streams. The video stream is a collection of image frames, and the audio stream is a collection of audio frames. It should be noted that in the embodiments of the present application, processing the video stream (for example, detecting the video stream) means processing the video composed of image frames, and processing the audio stream (for example, detecting the audio stream) means processing the audio composed of audio frames. In the following embodiments, the video stream can be replaced by a video composed of image frames, and the audio stream can be replaced by an audio composed of audio frames.

[0081] The number of image frames captured (recorded) by an electronic device in a unit time is the shooting frame rate of the video, also known as the recording frame rate or video frame rate. The number of image frames played by an electronic device in a unit time is the playback frame rate of the video. In one example, the unit of the shooting frame rate or playback frame rate is frames per second (fps), which indicates the number of image frames captured or played per second.

[0082] It should be noted that the shooting frame rate and playback frame rate in the embodiments of the present application refer to the frame rate of the image frame. The frame rate of the audio frame is related to the encoding format of the audio. The frame rate of the audio frame is not necessarily equal to the frame rate of the image frame, that is, the audio frame and the image frame are not one-to-one corresponding. The audio frame and the image frame are synchronized by consistency on the timeline.

[0083] Generally speaking, when an electronic device normally records a video, the shooting frame rate and the playback frame rate are the same. For example, Figure 1A As shown, the mobile phone records a video of 2 seconds (s) at a shooting frame rate of 24fps; plays it at a playback frame rate of 24fps; and the playback duration is also 2s.

[0084] In some scenarios, the shooting frame rate of a video file is greater than the playback frame rate. For example, Figure 1B As shown in the figure, the mobile phone records a video with a duration of 1s at a shooting frame rate of 48fps and plays it at a playback frame rate of 24fps. The playback duration of the video file is 2s. It can be seen that when the shooting frame rate is twice the playback frame rate, the playback duration is twice the shooting duration, thus producing a 2x slow motion effect.

[0085] In other scenarios, the shooting frame rate of the video file is lower than the playback frame rate. Figure 1C As shown in the figure, the mobile phone records a 2s video file at a shooting frame rate of 12fps and plays it at a playback frame rate of 24fps. The playback duration of the video file is 1s. It can be seen that when the playback frame rate is twice the shooting frame rate, the playback duration is half the shooting duration, thus producing a 2x fast motion effect.

[0086] In the prior art, the slow motion effect or fast motion effect is generally achieved by adjusting the shooting frame rate or the playback frame rate so that the shooting frame rate is different from the playback frame rate. However, in some scenarios, certain shooting frame rates or playback frame rates will affect the video quality. For example, the playback frame rate of mobile phone videos is usually 24fps; to shoot slow motion videos, it is necessary to use a high frame rate greater than 24fps, such as 48fps (2x slow motion) or 96fps (4x slow motion). However, due to the limitations of mobile phone sensors, some advanced capabilities such as DCG (dual conversion gain) and global phase detection auto-focus (PDAF) cannot be supported at high frame rates greater than 24fps, which will cause the frame effect of the slow motion video to be worse than the frame effect of the normal speed video.

[0087] The video processing method provided in the embodiment of the present application records video at the same shooting frame rate as the playback frame rate; for example, the playback frame rate and the shooting frame rate are both 24fps; in this way, advanced capabilities such as DCG and PDAF can be used to improve video quality.

[0088] In some embodiments, the video stream may be detected to determine an image frame containing a specified action.

[0089] In one example, the specified action is a sports action such as shooting, long jump, shooting, swinging a racket in badminton, etc. In one implementation, an interpolation operation is performed on the image frame containing the specified action, and the video file after the interpolation process is played, so that the slow motion playback of the specified action is realized. In another example, the specified action is a martial arts action such as punching and kicking. In one implementation, a frame extraction operation is performed on the image frame containing the specified action, and the video file after the frame extraction process is played, so that the fast motion playback of the specified action is realized. The video processing method provided in the embodiment of the present application, if it is detected that the video includes the specified action, automatically plays the video interval containing the specified action in slow motion or fast motion, and the user does not need to manually select the area for slow motion playback or fast motion playback. Compared with the user manually shooting a slow motion video or a fast motion video, the operation is more convenient and it is easier to capture the specified action. It should be noted that the image frame containing the specified action includes: an image frame containing the start action of the specified action in the image (start image frame), an image frame containing the end action of the specified action in the image (end image frame), and an image frame between the start image frame and the end image frame. For example, if the specified action is long jump, the specified action start action is the action of the foot leaving the ground, and the specified action end action is the action of the foot touching the ground from the air. The start image frame is the image frame containing the action of the foot leaving the ground, and the end image frame is the image frame containing the action of the foot touching the ground from the air.

[0090] For example, Figure 2A As shown in the figure, a video file with a duration of 2s is recorded at a shooting frame rate of 24fps, and the image frames are detected for the specified action, and 6 frames are detected as image frames corresponding to the specified action. The 6 frames are subjected to a 4x interpolation operation, and 3 frames are inserted between every two frames to obtain 24 frames of images; that is, after processing, the video interval containing the specified action becomes 24 frames of images. In this way, when playing at a playback frame rate of 24fps, the playback duration of the video interval containing the specified action becomes longer, achieving a 4x slow motion effect.

[0091] For example, Figure 2B As shown in the figure, a video file with a duration of 2s is recorded at a shooting frame rate of 24fps, and the image frames are detected for the specified action, and 24 frames are detected as image frames corresponding to the specified action. The 24 frames are processed by frame extraction, and 1 frame is extracted from every 4 frames; that is, after processing, the video interval containing the specified action becomes 6 frames. In this way, when playing at a playback frame rate of 24fps, the playback duration of the video interval containing the specified action becomes shorter, achieving a 0.25 times fast motion effect.

[0092] In some embodiments, the audio stream may be detected to determine the audio frames including speech, and the image frames corresponding to the audio frames including speech (referred to as image frames including speech in the embodiments of the present application). It is also possible to determine the audio frames not including speech, and the image frames corresponding to the audio frames not including speech (referred to as image frames not including speech in the embodiments of the present application).

[0093] In one implementation, a fast-motion playback effect can be achieved by performing frame extraction processing on a video stream and playing the video file after the frame extraction processing. In the video processing method provided in the embodiment of the present application, if the video includes voice, the image frames that do not include voice (i.e., the image frames corresponding to the audio frames that do not include voice) are subjected to frame extraction processing, and the image frames that include voice are not subjected to frame extraction processing, and the processed video file is played; in this way, the video interval that does not include voice can achieve a fast-motion playback effect, and the video interval that includes voice can be played normally, so that the original sound of the voice can be retained, and the damage to the voice caused by fast-motion playback can be avoided.

[0094] For example, Figure 2CAs shown in the figure, a video file with a duration of 2s is recorded at a shooting frame rate of 24fps, and voice detection is performed on the audio stream to determine that 12 frames in the video stream are image frames containing voice. Frame extraction is performed on image frames that do not contain voice, and 1 frame is extracted from every 4 frames. Frame extraction is not performed on the 12 frames that contain voice. In this way, when playing at a playback frame rate of 24fps, the playback duration of the video interval that does not contain voice becomes shorter, achieving a 0.25-fold fast motion effect; the image frames that contain voice are played normally, retaining the original sound of the voice.

[0095] In one implementation, a slow motion playback effect can be achieved by performing frame interpolation processing on a video stream and playing the video file after the frame interpolation processing. In the video processing method provided in the embodiment of the present application, if the video includes voice, the image frames that do not include voice are interpolated, the image frames that include voice are not interpolated, and the processed video file is played; in this way, the video interval that does not include voice can achieve a slow motion playback effect, and the video interval that includes voice can be played normally, which can retain the original sound of the voice and avoid the damage of the voice to the slow motion playback.

[0096] For example, Figure 2D As shown, a video file with a duration of 1s is recorded at a shooting frame rate of 24fps, and voice detection is performed on the audio stream to determine the audio stream containing voice, and 12 frames in the video stream are determined to be image frames containing voice. A 4x interpolation operation is performed on the image frames that do not contain voice, and no interpolation is performed on the 12 frames that contain voice. In this way, when playing at a playback frame rate of 24fps, the playback duration of the video interval that does not contain voice becomes longer, achieving a 4x slow motion effect; the image frames that contain voice are played normally, retaining the original sound of the voice.

[0097] The video processing method provided in the embodiment of the present application can be applied to an electronic device with a shooting function. For example, the electronic device can be a mobile phone, a sports camera (GoPro), a digital camera, a tablet computer, a desktop, a laptop, a handheld computer, a notebook computer, a vehicle-mounted device, a smart home device (such as a smart TV, a smart screen, a large screen, a smart speaker, etc.), an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook, and a cellular phone, a personal digital assistant (personal digital assistant, PDA), an augmented reality (augmented reality, AR)\virtual reality (virtual reality, VR) device, etc. The embodiment of the present application does not impose any special restrictions on the specific form of the electronic device.

[0098] like Figure 3As shown, it is a structural schematic diagram of an electronic device 100. The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a temperature sensor, an ambient light sensor, etc.

[0099] It is to be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0100] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0101] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0102] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or cyclically used. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0103] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0104] It is understandable that the interface connection relationship between the modules illustrated in this embodiment is only a schematic illustration and does not constitute a structural limitation of the electronic device. In other embodiments, the electronic device may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0105] The charging management module 140 is used to receive charging input from a charger. The charger may be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 may receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 may receive wireless charging input through a wireless charging coil of an electronic device. While the charging management module 140 is charging the battery 142, it may also power the electronic device through the power management module 141.

[0106] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), etc. In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0107] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0108] The electronic device 100 implements the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.

[0109] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Mini-LED, Micro-OLED, Micro-OLED, quantum dot light-emitting diodes (QLED), etc.

[0110] The electronic device 100 can realize the shooting function through ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0111] ISP is used to process the data fed back by camera 193. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 193.

[0112] Camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device may include 1 or N cameras 193, where N is a positive integer greater than 1. In the embodiment of the present application, the camera 193 can be used to capture video images.

[0113] Digital signal processors are used to process digital signals. In addition to processing digital image signals, they can also process other digital signals. For example, when an electronic device selects a frequency point, a digital signal processor is used to perform Fourier transform on the frequency point energy.

[0114] Video codecs are used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. In this way, the electronic device can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0115] NPU is a neural network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and continuously self-learn. NPU can realize applications such as intelligent cognition of electronic devices, such as image recognition, face recognition, voice recognition, text understanding, etc.

[0116] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, music playing, recording, etc. In the embodiment of the present application, the audio module 170 can be used to collect audio in recording video.

[0117] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or some functional modules of the audio module 170 are arranged in the processor 110. The speaker 170A, also known as the "speaker", is used to convert audio electrical signals into sound signals. The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals. The microphone 170C, also known as the "microphone" and "microphone", is used to convert sound signals into electrical signals. The headphone interface 170D is used to connect wired headphones. The headphone interface 170D can be a USB interface 130, or it can be a 3.5mm open mobile electronic equipment platform (open mobile terminal platform, OMTP) standard interface, and the cellular telecommunications industry association of the USA (cellular telecommunications industry association of the USA, CTIA) standard interface.

[0118] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, audio, video and other files are stored in the external memory card.

[0119] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121. For example, in an embodiment of the present application, the processor 110 can execute instructions stored in the internal memory 121, and the internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area may store data (such as video files) created during the use of the electronic device, etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0120] The button 190 includes a power button, a volume button, etc. The button 190 can be a mechanical button. It can also be a touch button. The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. The indicator 192 can be an indicator light, which can be used to indicate the charging status, power changes, messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into the SIM card interface 195, or pulled out from the SIM card interface 195 to achieve contact and separation with the electronic device. The electronic device can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.

[0121] In some embodiments, the software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, or a cloud architecture. Taking the system as an example, the software structure of the electronic device 100 is exemplarily described.

[0122] Figure 4 A software structure diagram of an electronic device provided in an embodiment of the present application.

[0123] It is understandable that the layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, The system can include the application (App) layer, the framework (FWK) layer, the hardware abstraction layer (HAL) and the kernel layer. Figure 4As shown, The system may also include the Android runtime and system libraries.

[0124] The application layer can include a series of application packages. Figure 4 As shown, the application package may include a camera, a video player, a video editor, etc. The camera application is used to take photos, shoot videos, etc. The video player is used to play video files. The video editor is used to edit video files. In some embodiments, the application package may also include applications such as a gallery, a calendar, a call, music, and short messages.

[0125] The application framework layer provides application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. The application framework layer provides programming services to the application layer through the API interface. Figure 4 As shown in the figure, the application framework layer includes the media framework, video post-processing service, media codec service, etc. The media framework is used to manage multimedia resources such as photos, pictures, videos, and audio. The media codec service is used to manage the codecs of multimedia resources such as photos, pictures, videos, and audio. The video post-processing service is used to manage and schedule the processing of video files.

[0126] Android runtime includes core libraries and virtual machines. Android runtime is responsible for scheduling and management of the Android system.

[0127] The core library consists of two parts: one part is the functional functions that Java voice needs to call, and the other part is the Android core library.

[0128] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object life cycle management, stack management, thread management, security and exception management, and garbage collection.

[0129] The system library may include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0130] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.

[0131] The media library supports playback and recording of a variety of commonly used audio and video formats, as well as static image files, etc. The media library can support a variety of audio encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0132] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis and layer processing, etc.

[0133] A 2D graphics engine is a drawing engine for 2D drawings.

[0134] The hardware abstraction layer is used to abstract the hardware, encapsulate the kernel layer driver, and provide an interface to the upper layer. Figure 4 As shown, the hardware abstraction layer includes a video pre-processing algorithm unit, a video post-processing algorithm unit, a video encoding and decoding unit, etc.

[0135] The kernel layer provides low-level drivers for various hardware of electronic devices. For example, Figure 4 As shown, the kernel layer may include camera driver, display driver, audio driver, etc.

[0136] The video processing method provided in the embodiment of the present application is described in detail below with reference to the accompanying drawings, taking the electronic device as a mobile phone as an example.

[0137] The mobile phone obtains a video, and by detecting the video stream (image frame) in the video file, the image frame containing the specified action can be determined. An embodiment of the present application provides a video processing method for processing image frames containing specified actions. For example, in one implementation, an interpolation operation is performed on the image frame containing the specified action, and the video file after the interpolation process is played. Slow-motion playback of the specified action is achieved. In another implementation, a frame extraction operation is performed on the image frame containing the specified action, and the video file after the frame extraction process is played. Fast-motion playback of the specified action is achieved. The following embodiment takes the interpolation of image frames containing specified actions to achieve slow-motion playback as an example for detailed introduction.

[0138] In one implementation, the mobile phone obtains a video, that is, obtains a video file at a normal speed. For example, the preset playback frame rate of the mobile phone is 24fps. If the video is shot at a shooting frame rate of 24fps, that is, the shooting frame rate is equal to the playback frame rate, then the video file obtained is a normal speed video file. In one implementation, the mobile phone uses a camera (e.g. Figure 3 camera 193) and an audio module (e.g. Figure 3In one example, the user starts the camera application, and when the mobile phone runs the camera application, the camera is started, and the image is collected through the camera to obtain the video stream; and the microphone is started to collect audio and obtain the audio stream. Figure 5 As shown in (a), in response to the user clicking the "camera" icon 201 in the main screen interface of the mobile phone, the mobile phone displays the following Figure 5 The interface 202 shown in (b) is a preview interface of the mobile phone camera, and the preview interface is used to display the preview image (i.e., the camera preview image) when the mobile phone takes a photo. Figure 5 As shown in (b), the interface 202 also includes a "portrait" option, a "video" option, a "movie" option, etc. Among them, the "video" option and the "movie" option are used to record video files; the "portrait" option is used to take photos. Figure 5 As shown in (b) of FIG. 1 , the user can click on the “Movie” option 203 to enter the video recording mode. Figure 5 As shown in (c), the mobile phone displays a "Movie" interface 204. The "Movie" option in the "Movie" interface 204 also includes a "Normal Speed" sub-option 205, a "Slow Motion" sub-option 206, a "Fast Motion" sub-option 207, etc. The "Normal Speed" sub-option 205 is used for a normal speed video mode, the "Slow Motion" sub-option 206 is used for a slow motion video mode, and the "Fast Motion" sub-option 207 is used for a fast motion video mode. For example, Figure 5 As shown in (d), the user can click on the "slow motion" sub-option 206 to choose to enter the slow motion video mode. The "movie" interface 204 also includes a "record" button 208, and the user can click on the "record" button 208 to start recording the video. In response to the user's click operation on the "record" button 208, the mobile phone starts the camera and captures images through the camera; and starts the microphone to capture audio through the microphone; and displays the captured image in the preview interface. In one implementation, the camera captures image frames at a preset first shooting frame rate (for example, 24fps). Exemplarily, as Figure 5 As shown in (e), the mobile phone displays an interface 209, which is a preview interface of the mobile phone video, and the preview interface is used to display the video preview image (the captured image frame). Optionally, the interface 209 includes a "stop" button 210. In response to the user clicking the "stop" button 210, the mobile phone stops capturing image frames through the camera and stops capturing audio through the microphone.

[0139] For example, please refer to Figure 6In response to the user's operation on the camera application, the camera application starts the camera and microphone. The camera of the mobile phone captures images and generates a video stream through a codec, that is, the image frames displayed on the preview interface, which are called video preview streams; the microphone of the mobile phone captures audio and generates an audio stream through a codec. The video preview stream includes image frames captured by the camera after the mobile phone receives the user's click operation on the "record" button 208 and before the mobile phone receives the user's click operation on the "stop" button 210; the audio stream includes audio frames captured by the microphone after the mobile phone receives the user's click operation on the "record" button 208 and before the mobile phone receives the user's click operation on the "stop" button 210.

[0140] In one implementation, the video pre-processing algorithm unit of the hardware abstraction layer analyzes the video preview stream and obtains the specified action video interval in the video stream. In one implementation, the video pre-processing algorithm unit obtains a video preview stream, performs resolution reduction processing (downsampling) on ​​the video preview stream, generates a low-resolution preview stream, and analyzes the low-resolution preview stream instead of analyzing the full-resolution preview stream. This can increase the processing speed of the video pre-processing algorithm unit and improve performance.

[0141] The scene detection module of the video pre-processing algorithm unit performs designated action detection based on the low-resolution preview stream, for example, the designated action is a sports action such as shooting, long jump, shooting, or swinging a racket in badminton. In some embodiments, the scene detection module can use any motion detection algorithm in conventional technology to perform motion detection, and the embodiments of the present application are not limited to this. In one implementation, the scene detection module needs to determine the motion interval based on the posture changes of the human body in multiple frames of images before and after. After receiving the low-resolution preview stream, the scene detection module caches multiple frames of images for determining the motion interval.

[0142] refer to Figure 6 The scene detection module uses a motion detection algorithm to obtain one or more motion intervals in the low-resolution preview stream. Exemplarily, the scene detection module obtains the 24th to 29th image frames in the low-resolution preview stream as motion intervals. And the motion interval results are transmitted to the motion analysis module.

[0143] It is understandable that different motion detection algorithms may be able to process different color gamut formats. In one implementation, reference Figure 6 The video pre-processing algorithm unit includes an ISP post-processing module, which can convert the color gamut format of the low-resolution preview stream according to the motion detection algorithm preset in the scene detection module; for example, convert the color gamut format of the low-resolution preview stream to the BT.709 format. The converted low-resolution preview stream is transmitted to the scene detection module for motion interval detection.

[0144] After the scene detection module transmits the motion interval result to the motion analysis module, the motion analysis module marks the motion start label and the motion end label according to the motion interval result. Exemplarily, the motion analysis module obtains the 24th image frame to the 29th image frame as the motion interval. In one example, the motion start label is marked as the 24th image frame, and the motion end label is marked as the 29th image frame. In another example, the 24th image frame corresponds to the first moment on the timeline, and the 29th image frame corresponds to the second moment on the timeline. The motion start label is marked as the first moment, and the motion end label is marked as the second moment.

[0145] In one implementation, the video pre-processing algorithm unit sends the marked motion start tag and motion end tag to the camera application through the hardware abstraction layer reporting channel. The camera application sends the marked motion start tag and motion end tag to the media codec service of the framework layer.

[0146] On the other hand, the video preview stream and audio stream of the hardware layer are sent to the media codec service of the framework layer through the reporting channel of the hardware abstraction layer.

[0147] The media codec service synthesizes a normal speed video file (first video file) based on the video preview stream and the audio stream. It also saves the marked motion start tag and motion end tag as metadata in the marker file. The normal speed video file and the corresponding marker file are stored together in a video container. In other words, each video container stores the relevant files of a video (video file, marker file, etc.).

[0148] In one example, the tag file includes the following content in Table 1. The motion start tag of the first motion interval is the 24th image frame, and the motion end tag is the 29th image frame; the motion start tag of the second motion interval is the 49th image frame, and the motion end tag is the 70th image frame.

[0149] Table 1

[0150] Exercise interval number Start Frame End Frame 1 24 29 2 49 70

[0151] In another example, the tag file includes the content shown in Table 2 below.

[0152] Table 2

[0153] Exercise interval number Start tag End Tag 1 First Moment Second Moment

[0154] It should be noted that the above Tables 1 and 2 are only exemplary descriptions of the tag file. The embodiment of the present application does not limit the specific form of the tag file. In a specific implementation, the information in the tag file may not be in the form of a table, and the way the information is recorded in the tag file is not limited. For example, in one example, the tag file includes the following content:

[0155] Frame number 1, TAG = 0

[0156] Frame number 2, TAG = 0

[0157] Frame number 3, TAG = 0

[0158] …

[0159] Frame number 24, TAG = 1

[0160] Frame number 25, TAG = 1

[0161] …

[0162] Frame number 29, TAG = 1

[0163] …

[0164] Frame number 48, TAG = 0

[0165] The tag file includes the frame numbers of all image frames in the preview video stream. A TAG of 0 indicates that the frame does not belong to a motion interval, and a TAG of 1 indicates that the frame belongs to a motion interval.

[0166] After the first video file is generated, the video portion containing the specified action in the video stream can be obtained according to the marker file, and the video portion containing the specified action can be processed (for example, interpolation processing) to generate a video file after interpolation processing. In this way, when playing the video, the video file after interpolation processing can be directly played to achieve automatic playback of the video with slow motion effect.

[0167] In one implementation, after the mobile phone receives the user's operation to start playing the video for the first time, it processes the video portion containing the specified action (for example, interpolation) to generate an interpolation-processed video file, thereby achieving the slow motion effect of automatically playing the specified action. Afterwards, when the user receives the operation to start playing the video again, the interpolation-processed video file can be directly played without the need for interpolation again.

[0168] It is understandable that the embodiment of the present application does not limit the triggering time for processing the video portion containing the specified action. For example, the video portion containing the specified action can also be processed after the first video file is generated. The following takes the processing of the video portion containing the specified action after the mobile phone receives the user's operation of starting to play the video for the first time as an example to describe the implementation method of the embodiment of the present application in detail.

[0169] The user can start playing videos or editing videos on the mobile phone. Taking playing videos as an example, for example, Figure 7 As shown, the mobile phone displays a play interface 701 of video one, and the play interface 701 includes a play button 702. The user can click the play button 702 to start playing video one.

[0170] In one example, in response to the user clicking the play button 702, the mobile phone starts the video player application. Figure 8 As shown, the video player application determines to start video one, and calls the media framework to start playing video one. For example, the video player application sends the identifier of video one to the media framework. The media framework obtains the video file corresponding to video one from the codec through the media codec service. The codec stores the video files recorded or received by the mobile phone. Exemplarily, the video container corresponding to video one includes a first video file and a marker file. The video codec unit decodes the first video file and the marker file in the video container, obtains the decoded first video file and the marker file, and transmits them to the media framework. The media framework obtains the motion start tag and the motion end tag in the marker file, and obtains the image frame corresponding to the motion interval according to the first video file and the motion start tag and the motion end tag.

[0171] Exemplarily, the first video file includes a first video stream and a first audio stream. Figure 2A As shown, the shooting frame rate of the first video stream is 24fps, the duration of the first video file is 2s, and the first video stream includes 48 image frames. The motion start tag corresponds to the 24th image frame, and the motion end tag corresponds to the 29th image frame. It should be noted that Figure 2A Taking the motion interval as a part of the image frames in the first video stream as an example, in other embodiments, the motion interval may also be all the image frames in the first video stream.

[0172] The media framework requests the video post-processing service to perform slow motion image processing on the motion interval. For example, the media framework sends the image frames corresponding to the motion interval to the video post-processing service. The video post-processing service passes the media framework's request (including the image frames corresponding to the motion interval) to the video post-processing algorithm unit of the hardware abstraction layer. The video post-processing algorithm unit uses a preset interpolation algorithm and relevant hardware resources (such as CPU, GPU, NPU, etc.) to perform interpolation processing on the image frames corresponding to the motion interval.

[0173] The video post-processing algorithm unit may use any frame interpolation algorithm in conventional technologies to perform frame interpolation processing, such as a motion estimation and motion compensation (MEMC) algorithm. Fig. 9 As shown in (a), the image frames corresponding to the motion interval are the 24th to 29th frames in the first video stream, a total of 6 frames. The MEMC algorithm is used to perform 4 times interpolation processing on the image frames from the 24th to the 29th frames, that is, 3 frames are inserted between two adjacent frames. The image frames after interpolation processing are shown in FIG. Fig. 9 as shown in (b).

[0174] The video post-processing service returns the interpolated image frames generated by the video post-processing algorithm unit to the media framework. The media framework replaces the image frames corresponding to the motion intervals in the first video stream with the interpolated image frames to obtain a second video stream. The image frames corresponding to the motion intervals in the second video stream are four times the image frames corresponding to the motion intervals in the first video stream. For example, Figure 2A As shown, the image frames corresponding to the motion interval in the first video stream are 6 frames, and the image frames corresponding to the motion interval in the second video stream are 24 frames.

[0175] In one implementation, the media framework calls the video post-processing service to send the second video stream to the display screen for display, that is, to play the second video stream on the display screen. The playback frame rate of the second video stream is a preset value, which is equal to the shooting frame rate; for example, 24fps. In this way, the duration of the first video stream is 2s, and the duration of the second video stream is greater than 2s. The image frames corresponding to the motion interval in the first video stream include 6 frames, and the image frames corresponding to the motion interval in the second video stream include 24 frames. The motion interval playback time is longer, and slow motion playback of the motion interval is achieved.

[0176] In some embodiments, the interpolation multiple used in the interpolation algorithm is a preset value. For example, in the above example, the interpolation multiple used in the MEMC algorithm is 4 times, that is, 4 times slow motion is achieved. In some embodiments, multiple preset values ​​of the interpolation multiple (for example, 4, 8, 16, etc.) can be preset in the mobile phone, and the user can select one of them as the interpolation multiple used in the interpolation algorithm. In other words, the user can select the slow motion multiple.

[0177] In one example, the user can select a slow motion multiple (interpolation multiple) when recording a video. Fig.10 As shown, after the user selects the "slow motion" sub-option 206, the mobile phone displays multiple options 2061, including "4 times", "8 times", "16 times" and other options. "4 times" means that the interpolation multiple is 4 times, corresponding to the preset value 4; "8 times" means that the interpolation multiple is 8 times, corresponding to the preset value 8; "16 times" means that the interpolation multiple is 16 times, corresponding to the preset value 16. For example, when the user selects the "4 times" option, the above-mentioned video post-processing algorithm unit uses the interpolation algorithm to perform interpolation processing, and the interpolation multiple used is 4 times.

[0178] In some embodiments, the media framework processes the first audio stream according to the motion start tag and the motion end tag. For example, the media framework determines the audio stream corresponding to the motion interval in the first audio stream according to the motion start tag and the motion end tag, and eliminates the sound of the audio stream corresponding to the motion interval. The media framework calls the audio module to play the processed first audio stream (i.e., the second audio stream).

[0179] In this way, slow motion playback of the motion interval is achieved, and no sound is played during the motion interval. Fig.11 As shown, the motion interval includes 24 image frames, which are played at a playback frame rate of 24fps to achieve slow motion playback. The sound of the audio stream (original video sound) is not played in the motion interval.

[0180] In one implementation, when entering the motion interval, the volume of the original video sound gradually decreases, and when leaving the motion interval, the volume of the original video sound gradually increases, thus providing users with a smoother sound experience.

[0181] In one implementation, when playing a video file, the soundtrack is also played. For example, the soundtrack is played at a lower volume in the normal speed range, and at a higher volume in the motion range. When entering the motion range, the soundtrack volume gradually increases; when leaving the motion range, the soundtrack volume gradually decreases.

[0182] Exemplary, reference Fig.11 In the normal speed range (non-sports range), the video sound volume is high and the music volume is low. When entering the sports range, the video sound volume gradually decreases and the music volume gradually increases; when leaving the sports range, the video sound volume gradually increases and the music volume gradually decreases.

[0183] In some embodiments, the media framework further generates a second video file based on the second video stream and the second audio stream. Optionally, the media framework also incorporates the audio stream of the soundtrack into the second video file. Figure 8 , the media framework calls the video post-processing service to store the second video file into the video container corresponding to the video one. That is, the video container corresponding to the video one includes the first video file, the marker file, and the second video file. The first video file is the original video file; the marker file is a file that records the motion start tag and the motion end tag after motion detection is performed on the original video file; the second video file is a video file after the original video file is processed in slow motion according to the motion start tag and the motion end tag.

[0184] The second video file can be used for playing, editing, forwarding, etc. In one example, the mobile phone can execute the above only when the user starts playing the video for the first time or starts editing the video for the first time. Figure 8The processing flow shown in the figure is to play the first video file in slow motion according to the first video file and the mark file; and generate the second video file. When the user subsequently starts to play the first video or edit the first video, the second video file can be played or edited directly without repeating the process. Figure 8 In some embodiments, the user can forward the slow motion video to other electronic devices. Fig. 12A As shown in (a), in response to receiving a user's click operation on a blank area of ​​the playback interface 701, the mobile phone displays Fig. 12A (b) shows the playback interface 703. The playback interface 703 includes a "Share" button 704. The user can click the "Share" button 704 to forward the video file corresponding to the interface. Exemplarily, the playback interface 703 is the playback interface of video one. After receiving the operation of the user clicking the "Share" button 704, the mobile phone searches for the video container corresponding to video one. If the video container corresponding to video one includes a second video file, the second video file is forwarded. That is, the video file after slow motion processing is forwarded. In this way, another electronic device receives the second video file and plays the second video file, thereby realizing slow motion playback of the motion interval.

[0185] It should be noted that the above embodiments are described by taking the selection of slow motion video mode when recording a video on a mobile phone as an example. In other embodiments, the mobile phone records a video (a first video file) in a normal speed video mode, or receives a video (a first video file) recorded in a normal speed video mode from another device, and the user can select the slow motion video mode when playing or editing the video. For example, Fig. 12B As shown, the mobile phone displays a start-up playback interface 1301 of video 1, and the start-up playback interface 1301 includes a "play" button 1302. In response to the user clicking the "play" button 1302, the mobile phone displays an interface 1303. Interface 1303 includes a "normal speed" option 1304, a "slow motion" option 1305, a "fast motion" option 1306, etc. Exemplarily, the user selects the "slow motion" option 1305 and clicks the "OK" button 1307. In response to the user clicking the "OK" button 1307, the mobile phone generates a corresponding tag file according to the first video file (for specific steps, please refer to Figure 6 ) and execute Figure 8 The processing flow shown in the figure plays the slow motion video and saves the second video file. The functions and specific implementation steps of each module can be referred to Figure 6 and Figure 8 , which will not be described here. Optional, such as Fig. 12BAs shown, in response to the user selecting the "slow motion" option 1305, the mobile phone interface 1303 also displays slow motion multiple options "4 times", "8 times", "16 times", etc., and the user can select one of them as the interpolation multiple used in the interpolation algorithm.

[0186] In some embodiments, the mobile phone determines the audio stream corresponding to the motion interval in the first audio stream according to the motion start tag and the motion end tag, and uses a speech recognition algorithm to recognize the speech in the audio stream corresponding to the motion interval. If the recognition is successful, the corresponding text is generated, and the text corresponding to the speech is displayed in the form of subtitles when playing the slow motion interval video. It can be understood that the audio stream including the speech can be part or all of the audio stream corresponding to the motion interval.

[0187] It can be understood that, according to the timeline, there is a corresponding relationship between the image frames in the first video stream and the audio frames in the first audio stream. For example, one image frame corresponds to one audio frame, or one image frame corresponds to multiple audio frames, or multiple image frames correspond to one audio frame. The image frame at which the subtitles start to be displayed can be determined based on the audio frame at which the speech starts in the motion interval.

[0188] In one example, the text corresponding to a segment of speech is merged and displayed in the image frame corresponding to the segment of speech. For example, the text corresponding to the speech in the motion interval is "Come on". The duration of the speech "Come on" is the first duration, and the interpolation multiple is N (N>1), then the display duration of the text "Come on" is the first duration * N. That is to say, if the speech "Come on" in the first audio stream corresponds to M image frames in the first video stream, and the interpolation multiple for interpolating the image frames in the motion interval in the first video stream is N (N>1), then the number of frames displayed for the text "Come on" in the second video stream is M*N frames.

[0189] For example, Fig.13A As shown in (a), the image frames corresponding to the motion interval in the first video stream include the 24th to 29th image frames. Voice recognition is performed on the audio stream corresponding to the motion interval in the first audio stream, and the recognized text is "Come on." The audio stream including the voice "Come on" corresponds to the 24th to 28th image frames in the first video stream (i.e., M=5). It can be determined that the image frame that starts displaying subtitles is the 24th frame, and the interpolation multiple for the motion interval is 4 (N=4). As shown Fig.13A As shown in (b), in the second video stream, the number of frames in which the text "Come on" is displayed is 20 (5*4) frames.

[0190] In another example, each word in the voice is displayed in the image frame corresponding to that word. For example, the text corresponding to the voice within the motion interval is "Come on". The word "加 (jiā)" in the voice corresponds to the M1 frame image in the first video stream. If the interpolation multiple for interpolating the image frames within the motion interval of the first video stream is N (N > 1), then the number of frames in which the word "加 (jiā)" is displayed in the second video stream is M1 * N frames. The word "油 (yóu)" in the voice corresponds to the M2 frame image in the first video stream. If the interpolation multiple for interpolating the image frames within the motion interval of the first video stream is N (N > 1), then the number of frames in which the word "油 (yóu)" is displayed in the second video stream is M2 * N frames.

[0191] Exemplarily, as Fig. 13B shown in (a) of, the image frames corresponding to the motion interval in the first video stream include the 24th to 29th frame images. Performing speech recognition on the audio stream corresponding to the motion interval, the recognized text corresponding to the voice is "Come on"; among them, the audio stream including the word "加 (jiā)" corresponds to the 24th and 25th frame images in the first video stream (i.e., M1 = 2); the audio stream including the word "油 (yóu)" corresponds to the 26th to 28th frame images in the first video stream (i.e., M2 = 3). It can be determined that the image frame starting to display the subtitle "加 (jiā)" is the 24th frame, and the image frame starting to display the subtitle "油 (yóu)" is the 26th frame. The interpolation multiple for interpolating the motion interval is 4 (N = 4). As Fig. 13B shown in (b) of, in the second video stream, the number of frames in which the word "加 (jiā)" is displayed is 8 (2 * 4) frames, and the number of frames in which the word "油 (yóu)" is displayed is 12 (3 * 4) frames. It can be seen that the interpolation frames between the 24th frame, the 24th and 25th frames, the 25th frame, and the 25th and 26th frames display the word "加 (jiā)"; the interpolation frames between the 26th frame, the 26th and 27th frames, the 27th frame, the 27th and 28th frames, the 28th frame, and the 28th and 29th frames display the word "油 (yóu)". In this way, the display duration of each word is N (interpolation multiple) times the voice duration corresponding to that word. In the played slow-motion video, the display duration of each word in the subtitle is extended by 4 times (the same as the interpolation multiple), and the subtitle matches the slow-motion image.

[0192] In some embodiments, the user can manually adjust the slow-motion playback interval (i.e., the motion interval). Exemplarily, as Fig.14As shown, the playback interface 703 of the video 1 also includes a "manual" button 705. The user can click the "manual" button 705 to manually adjust the slow motion playback interval. The mobile phone receives the user's click operation on the "manual" button 705 and displays the manual adjustment slow motion playback interval interface 706. Exemplarily, the interface 706 includes a prompt message 707 for prompting the user to manually adjust the slow motion playback interval. A segment of image frames is displayed in the interface 706; in an example, the interface 706 displays the image frames before the interpolation process. The interface 706 also includes a "slow motion start" icon 708 and / or a "slow motion end" icon 709. The image frames between the "slow motion start" icon 708 and the "slow motion end" icon 709 are the slow motion playback interval. The user can adjust the "slow motion start" icon 708 or the "slow motion end" icon 709 separately. The interface 706 also includes an "OK" button 710 and a "Cancel" button 711. The user can click the “OK” button 710 to save the change to the slow motion playback interval; the user can click the “Cancel” button 711 to cancel the change to the slow motion playback interval.

[0193] In one implementation, in response to receiving a user click operation on the "OK" button 710, the mobile phone determines the motion start tag according to the position of the "slow motion start" icon 708, and determines the motion end tag according to the position of the "slow motion end" icon 709; and updates the motion start tag and motion end tag saved in the markup file. Further, the image frames corresponding to the motion interval in the first video stream are interpolated according to the updated motion start tag and motion end tag to generate an updated second video stream. The first audio stream is also processed according to the updated motion start tag and motion end tag to generate an updated second audio stream. Optionally, the music can be re-composed according to the updated motion start tag and motion end tag to generate subtitles. Further, the mobile phone generates an updated second video file according to the updated second video stream and the updated second audio stream, replacing the saved second video file. In this way, when the user plays or forwards the video later, the updated second video file is used.

[0194] For example, Fig.15 A schematic diagram of a process flow of a video processing method provided by an embodiment of the present application is shown. Fig.15 As shown, the method may include:

[0195] S1501: Receive an operation from a user to start recording a video.

[0196] For example, the user adopts Figure 5 The method shown starts recording a video. For example, if the mobile phone receives an operation of the user clicking the "record" button 208, it determines that the user's operation of starting recording a video has been received.

[0197] S1502: In response to a user starting a video recording operation, the electronic device collects image frames at a first frame rate through a camera device and collects audio frames through a recording device.

[0198] In one example, the camera device is a camera of a mobile phone, and the recording device is a microphone. In response to the user starting the operation of recording a video, the mobile phone captures image frames through the camera at a first frame rate, and captures audio frames through the microphone. Exemplarily, the first frame rate is 24fps.

[0199] S1503: Receive the user's operation to end video recording.

[0200] For example, the mobile phone receives a user click Figure 5 If the "Stop" button 210 in (e) is operated, it is determined that the user's operation to end video recording has been received.

[0201] S1504: The electronic device stops collecting image frames through the camera device and stops collecting audio frames through the recording device, and generates a first video file.

[0202] The electronic device obtains the image frames collected by the camera to generate a video preview stream, obtains the audio frames collected by the recording device to generate an audio stream, and synthesizes a normal speed video file, namely the first video file, according to the video preview stream and the audio stream.

[0203] In one implementation, the electronic device may perform resolution reduction processing on the video preview stream, perform designated action detection on the video stream after resolution reduction processing, and generate marking information; and record the marking information in a marking file. Exemplarily, the marking file may include the contents of Table 1 or Table 2 above. For example, the marking information includes designated action start information and designated action end information; the designated action start information is used to indicate an image frame in the image that contains a designated action start action, and the designated action end information is used to indicate an image frame in the image that contains a designated action end action.

[0204] S1505: The electronic device obtains a first video portion containing a designated action according to the markup file and the first video file.

[0205] The video preview stream includes a first video portion consisting of a first image frame and a second video portion consisting of a second image frame. The first image frame is an image frame containing a specified action, and the second image frame is an image frame not containing a specified action. For example, the first image frame is Figure 2A The preview video stream is from frame 24 to frame 29, a total of 6 frames.

[0206] The electronic device determines, according to the tag information in the tag file, an image frame containing a specified action in the video preview stream, that is, obtains the first video portion.

[0207] S1506: The electronic device processes the first video part to generate a third video part.

[0208] In one implementation, the electronic device performs frame interpolation processing on the first video portion to generate a third video portion. In this way, the number of image frames in the third video portion is greater than the number of image frames in the first video portion. Slow motion playback of a specified action can be achieved. For example, Fig. 9 As shown in (a), the first video portion includes frames 24 to 29, a total of 6 image frames; after the interpolation process, as shown in Fig. 9 As shown in (b), the third video portion includes 24 image frames.

[0209] In one implementation, the electronic device performs frame extraction processing on the first video portion, so that the number of image frames of the third video portion is smaller than the number of image frames of the first video portion, so that fast motion playback of a specified action can be achieved.

[0210] In some embodiments, the electronic device obtains the first video file, performs designated action detection on the video preview stream, obtains tag information, and then processes the first video portion.

[0211] In other embodiments, when the electronic device receives an operation from a user to play the video for the first time, the electronic device obtains the first video portion according to the tag information and processes the first video portion.

[0212] S1507: The electronic device replaces the first video portion in the first video file with the third video portion to generate a second video file.

[0213] The second video file includes a third video portion and a second video portion.

[0214] S1508: The electronic device plays the second video file at the first frame rate.

[0215] In one implementation, the audio stream collected by the electronic device includes a first audio portion corresponding to the timeline of the first video portion (the image frame containing the specified action) and a second audio portion corresponding to the timeline of the second video portion (the image frame not containing the specified action). The electronic device performs speech recognition on the first audio portion to generate text. And determines the video sub-portion in the first video portion corresponding to the timeline of the audio (first audio sub-portion) containing the speech.

[0216] When playing the second video file, the text is displayed in the first video subsection of the third video section in the form of subtitles. The first video subsection of the third video section is obtained by inserting frames of the video subsection containing voice in the first video section. That is, the subtitles are matched with the image frames after the inserting frame processing.

[0217] In one example, if Fig.13A As shown, the image frames corresponding to the audio containing speech are from the 24th frame to the 28th frame, a total of 5 frames; the text corresponding to the speech is "Come on". The duration of the audio frame (first audio sub-part) corresponding to the image frames from the 24th frame to the 28th frame is the first duration (the duration corresponding to displaying 5 frames of image frames). After performing N times (interpolation multiple) interpolation processing on the 24th frame to the 28th frame, the image frames displaying the text "Come on" are 20 frames (N times the first duration).

[0218] In one example, if Fig. 13B As shown, the image frames corresponding to the audio containing voice are the 24th to 28th frames, a total of 5 frames; the text corresponding to the voice is "Come on". Among them, the word "plus" (the first text) corresponds to the 24th and 25th image frames, and the duration of the corresponding audio frame is the first duration (the duration corresponding to displaying 2 frames of image frames). After performing N times (interpolation multiples) interpolation processing on the 24th to 28th frames, the image frames displayed with the word "plus" are 8 frames (N times the first duration). In some embodiments, an operation of editing a video by a user is received, and a first interface is displayed. An operation of modifying the range of an image frame interval containing a specified action by a user on the first interface is received, and the electronic device updates the marking information according to the modified range of an image frame interval containing a specified action. Exemplarily, the first interface is Fig.14 The electronic device receives an operation of adjusting the "slow motion start" icon 708 or the "slow motion end" icon 709 on the interface 706, and updates the marking information according to the position of the "slow motion start" icon 708 or the "slow motion end" icon 709. In this way, the electronic device can process the first video file according to the updated marking information to generate an updated second video file.

[0219] The video processing method provided in the embodiment of the present application shoots a first video file at a first shooting frame rate, and performs a specified action (such as a specified motion action) detection on the image frames (first video stream) in the first video file. And an interpolation algorithm is used to perform interpolation processing on the image frames corresponding to the motion interval in the first video stream. The video stream after interpolation processing is played at a first playback frame rate (first playback frame rate = first shooting frame rate). Automatic slow motion playback of the specified action is achieved. In this method, slow motion playback is automatically performed when a specified action is detected, avoiding manual operation by the user, and the captured slow motion is more accurate, improving the user experience. Moreover, the video is recorded at the same shooting frame rate as the playback frame rate, so that advanced capabilities such as DCG and PDAF can be used to record the video, thereby improving the video quality.

[0220] The mobile phone obtains a video, and by detecting the audio stream in the video file, it can determine the audio frames that include voice, and determine the image frames corresponding to the audio frames that include voice (referred to as image frames containing voice in the embodiment of the present application). The embodiment of the present application provides a video processing method, which does not perform frame extraction or frame insertion processing on image frames containing voice to achieve normal speed playback; and processes image frames that do not contain voice. For example, in one implementation, frame extraction processing is performed on image frames that do not contain voice in the video file (i.e., image frames corresponding to audio frames that do not contain voice) to achieve fast motion playback. In another implementation, frame insertion processing is performed on image frames that do not contain voice in the video file (i.e., image frames corresponding to audio frames that do not contain voice) to achieve slow motion playback. The following embodiment takes the example of performing frame extraction processing on image frames that do not contain voice to achieve fast motion playback as a detailed introduction.

[0221] In one implementation, the mobile phone obtains a video file at a normal speed. The mobile phone can obtain a video file at a normal speed by using a method similar to that of a slow motion video. Figure 5 The method shown opens the "Movie" interface 204 of the mobile phone. Fig.16 As shown, the "Movie" option in the "Movie" interface 204 includes a "Quick Motion" sub-option 207. The user can click the "Quick Motion" sub-option 207 to enter the quick motion video mode. The mobile phone receives the user's operation of clicking the "Record" button 208 and starts recording the video. Exemplarily, the camera of the mobile phone captures images according to a preset first shooting frame rate (for example, 24fps), and the microphone of the mobile phone captures audio.

[0222] For example, please refer to Figure 6 In response to the user's operation on the camera application, the camera application starts the camera and microphone. The camera of the mobile phone collects images and generates a video stream through the codec, that is, the image frame displayed on the preview interface, which is called the video preview stream; the microphone of the mobile phone collects audio and generates an audio stream through the codec.

[0223] In one implementation, the video pre-processing algorithm unit of the hardware abstraction layer analyzes the audio stream. Exemplarily, the speech recognition module of the video pre-processing algorithm unit performs speech recognition on the audio stream, obtains the speech therein, and determines the audio frame containing the speech. Optionally, the speech recognition module analyzes the speech and obtains the speech containing valid semantics. Further, the image frame containing the speech is determined based on the audio frame containing the speech. Exemplarily, the speech recognition module obtains the image frames from the 21st frame to the 32nd frame as image frames containing speech. The speech recognition module transmits the image frame results containing speech to the speech analysis module.

[0224] The speech analysis module marks the speech start label and the speech end label according to the image frame result containing the speech. Exemplarily, the speech analysis module obtains that the 21st image frame to the 32nd image frame are the intervals containing speech. In one example, the speech start label is marked as the 21st image frame, and the speech end label is marked as the 32nd image frame. In another example, the 21st image frame corresponds to the first moment on the timeline, and the 32nd image frame corresponds to the second moment on the timeline. The speech start label is marked as the first moment, and the speech end label is marked as the second moment.

[0225] In one implementation, the video pre-processing algorithm unit sends the marked speech start tag and speech end tag to the camera application through the hardware abstraction layer reporting channel. The camera application sends the marked speech start tag and speech end tag to the media codec service of the framework layer.

[0226] On the other hand, the video preview stream and audio stream of the hardware layer are sent to the media codec service of the framework layer through the reporting channel of the hardware abstraction layer.

[0227] The media encoding and decoding service synthesizes a normal speed video file (first video file) according to the video preview stream and the audio stream. The marked speech start tag and speech end tag are also saved as metadata in the marker file. The normal speed video file and the corresponding marker file are stored together in a video container. Among them, the embodiment of the present application does not limit the specific form of storing the speech start tag and the speech end tag in the marker file. As an example, reference can be made to the description in the relevant embodiment of slow motion playback (Table 1, Table 2).

[0228] After the first video file is generated, the video portion containing the voice in the video stream can be obtained according to the marker file, and the video portion not containing the voice can be processed (for example, frame extraction) to generate a frame-extracted video file. In this way, when playing the video, the frame-extracted video file can be directly played to achieve automatic playback of the video with a fast motion effect.

[0229] In one implementation, after the mobile phone receives the user's operation to start playing the video for the first time, it processes the video portion that does not contain the voice (for example, frame extraction) to generate a frame-extracted video file, thereby achieving an automatic fast-motion playback effect. After receiving the user's operation to start playing the video again, the frame-extracted video file can be directly played without the need to perform frame extraction again.

[0230] It is understandable that the embodiment of the present application does not limit the triggering time for processing the video portion that does not contain voice. For example, the video portion that does not contain voice can also be processed after the first video file is generated. The following takes the processing of the video portion that does not contain voice after the mobile phone receives the user's operation of starting to play the video for the first time as an example to describe the implementation method of the embodiment of the present application in detail.

[0231] Users can start playing videos or editing videos on their phones. Take playing videos as an example, please refer to Figure 7 and Figure 7 Description of the video. User starts playing video one.

[0232] In one example, in response to the user clicking the play button 702, the mobile phone starts the video player application. Figure 8 As shown, the video player application determines to start video one, and calls the media framework to start playing video one. The media framework obtains the video file corresponding to video one from the codec through the media codec service. Exemplarily, the video container corresponding to video one includes a first video file and a marker file. The video codec unit decodes the first video file and the marker file in the video container, obtains the decoded first video file and the marker file, and transmits them to the media framework. The media framework obtains the speech start tag and the speech end tag in the marker file, obtains the image frame containing speech according to the first video file and the speech start tag and the speech end tag, and obtains the image frame not containing speech.

[0233] Exemplarily, the first video file includes a first video stream and a first audio stream. Figure 2C As shown, the shooting frame rate of the first video stream is 24fps, the duration of the first video file is 2s, and the first video stream includes 48 image frames. The speech start tag corresponds to the 21st image frame, and the speech end tag corresponds to the 32nd image frame. The image frames that do not contain speech include the 1st to 20th image frames, and the 33rd to 48th image frames.

[0234] The media framework requests the video post-processing service to perform fast-motion image processing on the intervals that do not contain voice. For example, the media framework sends the image frames that do not contain voice to the video post-processing service. The video post-processing service passes the media framework's request (image frames that do not contain voice) to the video post-processing algorithm unit of the hardware abstraction layer. The video post-processing algorithm unit uses a preset frame extraction algorithm and utilizes relevant hardware resources (such as CPU, GPU, NPU, etc.) to perform frame extraction on the image frames corresponding to the intervals that do not contain voice. The video post-processing algorithm unit can use any frame extraction algorithm in conventional technologies to perform frame extraction. For example, the image frames are subjected to 4x frame extraction, that is, one frame is extracted from every 4 image frames and retained. Exemplary, such as Fig.17As shown in (a), the image frame before the frame extraction process is 24 frames. After the frame extraction process is performed by 4 times, as shown in Fig.17 As shown in (b), 6 image frames are retained.

[0235] The video post-processing service returns the image frames generated by the video post-processing algorithm unit after the frame extraction process to the media framework. The media framework replaces the image frames corresponding to the non-voice intervals in the first video stream with the image frames after the frame extraction process to obtain a second video stream. The image frames corresponding to the non-voice intervals in the second video stream are one quarter of the image frames corresponding to the non-voice intervals in the first video stream. For example, Figure 2C As shown, the first video stream does not contain image frames corresponding to the speech interval for 36 frames, and the second video stream does not contain image frames corresponding to the speech interval for 9 frames.

[0236] In one implementation, the media framework calls the video post-processing service to send the second video stream to the display screen for display, that is, to play the second video stream on the display screen. The playback frame rate of the second video stream is a preset value, which is equal to the shooting frame rate; for example, 24fps. In this way, the duration of the first video stream is 2s, and the duration of the second video stream is less than 2s. There are 36 image frames without voice in the first video stream, and 12 image frames with voice in total; there are 9 image frames without voice in the second video stream, and 12 image frames with voice in total; the playback time of the interval without voice is shorter, realizing fast motion playback, and the playback time of the interval containing voice remains unchanged, and plays at normal speed.

[0237] In some embodiments, the frame extraction multiple used in the frame extraction algorithm is a preset value. For example, the frame extraction multiple used in the above example is 4 times, that is, 4 times fast motion is achieved. In some embodiments, multiple preset values ​​of the frame extraction multiple can be preset in the mobile phone (for example, 4, 8, 16, etc.), and the user can select one of them as the frame extraction multiple used in the frame extraction algorithm. In other words, the user can select the fast motion multiple. In one example, the user can select the fast motion multiple (frame extraction multiple) when recording a video. For details, please refer to Fig.10 Example of setting the slow motion multiple in .

[0238] In some embodiments, the media framework processes the first audio stream according to the speech start tag and the speech end tag. For example, the media framework determines the interval containing speech in the first audio stream according to the speech start tag and the speech end tag, retains the interval containing speech in the first audio stream, and eliminates the sound of the interval not containing speech in the first audio stream. The media framework calls the audio module to play the processed first audio stream (i.e., the second audio stream).

[0239] In this way, the video without voice is played in fast motion without playing the original sound of the video; the video with voice is played at normal speed and the original sound of the video is retained. Fig.18 As shown, the interval without speech includes 9 image frames, which are played at a playback frame rate of 24fps to achieve fast motion playback; the sound of the audio stream (original video sound) is not played in the fast motion playback interval. The interval containing speech includes 12 image frames, which are played at a playback frame rate of 24fps and played at normal speed; the original video sound is retained.

[0240] In one implementation, when entering a speech-containing interval, the volume of the original video sound gradually increases; when leaving the speech-containing interval, the volume of the original video sound gradually decreases, thereby providing users with a smoother sound experience.

[0241] In one implementation, when playing a video file, the soundtrack is also played. For example, the soundtrack is played at a lower volume in the normal speed interval (including the speech interval), and at a higher volume in the fast motion playback interval (excluding the speech interval). When entering the normal speed interval, the soundtrack volume gradually decreases; when leaving the normal speed interval, the soundtrack volume gradually increases.

[0242] Exemplary, reference Fig.18 , in the fast motion playback interval (excluding the voice interval), the soundtrack is played. Entering the normal speed interval (including the voice interval), the soundtrack volume gradually decreases, and the video sound volume gradually increases; the video sound volume is high, and the soundtrack volume is low. Leaving the normal speed interval (including the voice interval), the video sound volume gradually decreases, and the soundtrack volume gradually increases.

[0243] In some embodiments, the media framework further generates a second video file based on the second video stream and the second audio stream. Optionally, the media framework also incorporates the audio stream of the soundtrack into the second video file. Figure 8 , the media framework calls the video post-processing service to store the second video file into the video container corresponding to the video one. That is, the video container corresponding to the video one includes the first video file, the marker file, and the second video file. The first video file is the original video file; the marker file is a file that records the motion start tag and the motion end tag after motion detection is performed on the original video file; the second video file is a video file after the original video file is processed in slow motion according to the motion start tag and the motion end tag.

[0244] The second video file can be used for playing, editing, forwarding, etc. In one example, the mobile phone can execute the above only when the user starts playing the video for the first time or starts editing the video for the first time. Figure 8The processing flow shown in the figure performs fast motion playback according to the first video file and the mark file; and generates a second video file. When the user subsequently starts playing the first video or editing the first video, the second video file can be played or edited directly without repeating the process. Figure 8 In some embodiments, the user can forward the fast motion video to other electronic devices. Fig. 12A , the playback interface 703 is the playback interface of video 1, and the user can click the "Share" button 704 to forward the video file corresponding to the interface. After receiving the user's operation of clicking the "Share" button 704, the mobile phone searches for the video container corresponding to video 1. If the video container corresponding to video 1 includes the second video file, the second video file is forwarded. That is, the video file after fast motion processing is forwarded. In this way, another electronic device receives the second video file and plays the second video file, thereby realizing fast motion playback.

[0245] It should be noted that the above embodiments are described by taking the selection of the fast motion video mode when recording a video on a mobile phone as an example. In other embodiments, the mobile phone records a video (a first video file) in a normal speed video mode, or receives a video (a first video file) recorded in a normal speed video mode from another device, and the user can select the fast motion video mode when playing or editing the video. For example, Fig.19 As shown, the mobile phone displays a start-up playback interface 1301 of video 1, and the start-up playback interface 1301 includes a "play" button 1302. In response to the user clicking the "play" button 1302, the mobile phone displays an interface 1303. Interface 1303 includes a "normal speed" option 1304, a "slow motion" option 1305, a "fast motion" option 1306, etc. Exemplarily, the user selects the "fast motion" option 1306 and clicks the "OK" button 1307. In response to the user clicking the "OK" button 1307, the mobile phone generates a corresponding tag file according to the first video file (for specific steps, please refer to Figure 6 ) and execute Figure 8 The processing flow shown in the figure plays the fast motion video and saves the second video file. The functions and specific implementation steps of each module can be referred to Figure 6 and Figure 8 , which will not be described here. Optional, such as Fig.19 As shown, in response to the user selecting the "fast motion" option 1306, the mobile phone interface 1303 also displays fast motion multiple options "4x", "8x", "16x", etc., and the user can select one of them as the frame extraction multiple used in the frame extraction algorithm.

[0246] In some embodiments, the mobile phone determines the audio stream containing voice in the first audio stream based on the voice start tag and the voice end tag, recognizes the voice using a voice recognition algorithm, generates corresponding text, and displays the text corresponding to the voice in the form of subtitles when playing the video in the normal speed range.

[0247] In some embodiments, the user can manually adjust the fast motion playback interval. Fig. 20 As shown, the mobile phone displays an interface 2001 for manually adjusting the fast motion playback interval. Interface 2001 includes prompt information 2002, which is used to prompt the user to manually adjust the fast motion playback interval. Interface 2001 displays a segment of image frames; in one example, interface 2001 displays the image frames before frame extraction. Interface 2001 also includes a "start" icon 2003 and an "end" icon 2004. The image between the "start" icon 2003 and the "end" icon 2004 is the normal speed playback interval, and the image outside the "start" icon 2003 and the "end" icon 2004 is the fast motion playback interval. The user can adjust the "start" icon 2003 or the "end" icon 2004 separately.

[0248] In one example, corresponding subtitles are displayed in the image of the normal speed playback interval. In one implementation, the mobile phone receives an operation of the user moving the "Start" icon 2003 or the "End" icon 2004, determines the voice start tag according to the current position of the "Start" icon 2003, and determines the voice end tag according to the current position of the "End" icon 2004. According to the current voice start tag and voice end tag, the interval containing voice in the first audio stream is re-acquired, and voice recognition is performed on the interval currently containing voice to generate corresponding text. Further, the generated text is displayed in the image of the current normal speed playback interval in the form of subtitles. That is, the normal speed playback interval changes as the user moves the "Start" icon 2003 or the "End" icon 2004, and the interval where the subtitles are displayed and the subtitle content are updated as the normal speed playback interval is updated. Exemplarily, refer to Fig. 20 , the normal speed playback interval includes 6 image frames, subtitles are displayed in the 6 image frames, and the subtitle content is text generated by voice recognition based on the audio frames corresponding to the 6 image frames. For example, the user moves the "Start" icon 2003 backward by one frame, and the normal speed playback interval is updated to include 5 image frames; then it is updated to display subtitles in the 5 image frames, and the subtitle content is text generated by voice recognition based on the audio frames corresponding to the 5 image frames. In this way, it is convenient for users to adjust the normal speed playback interval according to the voice content (subtitles), that is, adjust the fast motion playback interval.

[0249] In one example, the interface 2001 further includes a "play" button 2005, and the user can click the "play" button 2005 to preview the video after adjusting the "start" icon 2003 and / or the "end" icon 2004. The mobile phone receives the user's click operation on the "play" button 2005, and in response to the user's click operation on the "play" button 2005, updates the voice start tag and the voice end tag, and performs fast motion processing on the first video file according to the updated voice start tag and the voice end tag (including frame extraction, eliminating the original sound of the video in the fast motion playback interval, soundtrack, generating subtitles, etc.), and plays the video after fast motion processing.

[0250] In one example, the interface 2001 further includes an "OK" button 2006 and a "Cancel" button 2007. The user can click the "OK" button 2006 to save the changes to the fast motion playback interval; the user can click the "Cancel" button 2007 to cancel the changes to the fast motion playback interval. In one implementation, in response to receiving the user's click operation on the "OK" button 2006, the mobile phone determines the voice start tag according to the position of the "Start" icon 2003, and determines the voice end tag according to the position of the "End" icon 2004; and updates the voice start tag and voice end tag saved in the tag file. Further, according to the updated voice start tag and voice end tag, the image frame that does not contain voice in the first video stream is subjected to frame extraction processing to generate an updated second video stream. The first audio stream is also processed according to the updated voice start tag and voice end tag to generate an updated second audio stream. Optionally, the music can also be re-composed according to the updated voice start tag and voice end tag to generate subtitles. Further, the mobile phone generates an updated second video file according to the updated second video stream and the updated second audio stream, and replaces the stored second video file. In this way, when the user plays or forwards the video later, the updated second video file is used.

[0251] For example, Fig.21 A schematic diagram of a process flow of a video processing method provided by an embodiment of the present application is shown. Fig.21 As shown, the method may include:

[0252] S2101. The electronic device obtains a first video file. The first video file includes a first image frame and a first audio frame. The shooting frame rate of the first image frame is a first frame rate.

[0253] In one implementation, the electronic device obtains the first video file from another device.

[0254] In one implementation, the electronic device records a first video file. The electronic device receives an operation from a user to start recording a video. For example, the user uses Fig.16The method shown starts recording a video. For example, if the mobile phone receives an operation of the user clicking the "record" button, it is determined that the user has started recording a video. In response to the user starting the video recording operation, the electronic device collects image frames at a first frame rate through a camera device and collects audio frames through a recording device. In one example, the camera device is a camera of a mobile phone, and the recording device is a microphone. In response to the user starting the video recording operation, the mobile phone collects image frames through the camera at a first frame rate; and collects audio frames through the microphone. Exemplarily, the first frame rate is 24fps. The user ends the video recording operation. For example, if the mobile phone receives an operation of the user clicking the "stop" button, it is determined that the user has ended the video recording operation. The electronic device stops collecting image frames through the camera device, stops collecting audio frames through the recording device, and generates a first video file. The electronic device obtains the image frames collected by the camera to generate a video preview stream (first image frame), and obtains the audio frames collected by the recording device to generate an audio stream (first audio frame). A video file of normal speed is synthesized according to the video preview stream and the audio stream, that is, the first video file.

[0255] The first audio frame includes a first audio portion consisting of a second audio frame and a second audio portion consisting of a third audio frame, wherein the second audio frame is an audio frame not containing speech, and the third audio frame is an audio frame containing speech.

[0256] The first image frame includes a first video part composed of second image frames and a second video part composed of third image frames; wherein the second image frame (image frame not containing voice) corresponds to the second audio frame on the timeline, and the third image frame (image frame containing voice) corresponds to the third audio frame on the timeline.

[0257] In one implementation, the electronic device may perform speech recognition on the first audio frame, mark the interval range of the third audio frame and the interval range of the third image frame, and generate marking information; and record the marking information in a marking file.

[0258] S2102: The electronic device obtains a first video portion that does not contain speech in a first image frame according to the mark file and the first video file.

[0259] That is, an image frame that does not contain voice in the preview video stream (the first image frame) is obtained.

[0260] S2103: The electronic device processes the first video portion to generate a third video portion.

[0261] In one implementation, the electronic device performs frame insertion processing on the first video portion to generate a third video portion. In this way, the number of image frames of the third video portion is greater than the number of image frames of the first video portion, so that slow motion playback can be achieved.

[0262] In one implementation, the electronic device performs frame extraction processing on the first video portion, so that the number of image frames of the third video portion is smaller than the number of image frames of the first video portion, so that fast motion playback can be achieved.

[0263] In some embodiments, the electronic device obtains the first video file, performs voice recognition on the first audio stream, obtains tag information, and then processes the first video portion.

[0264] In other embodiments, when the electronic device receives an operation from a user to play the video for the first time, the electronic device obtains the first video portion according to the tag information and processes the first video portion.

[0265] S2104: The electronic device replaces the first video portion in the first video file with the third video portion to generate a second video file.

[0266] The second video file includes a third video portion and a second video portion.

[0267] S2105: The electronic device plays the second video file at the first frame rate.

[0268] That is, the image frames containing voice (second video part) are played at normal speed to retain the original sound. The image frames not containing voice are processed and the processed video (third video part) is played to achieve fast motion playback or slow motion playback.

[0269] In some embodiments, upon receiving a user's operation of editing a video, the first interface is displayed. Upon receiving a user's operation of modifying the image frame interval range corresponding to the audio frame containing speech on the first interface, the electronic device updates the marking information according to the modified image frame interval range corresponding to the audio frame containing speech. Fig. 20 The electronic device receives an operation of adjusting the "start" icon 2003 or the "end" icon 2004 on the interface 2001 by the user, and updates the marking information according to the position of the "start" icon 2003 or the "end" icon 2004. In this way, the electronic device can process the first video file according to the updated marking information to generate an updated second video file.

[0270] The video processing method provided in the embodiment of the present application shoots a first video file at a first shooting frame rate, performs voice recognition on the audio frames (first audio stream) in the first video file to obtain image frames containing voice and image frames not containing voice. The image frames not containing voice are subjected to frame extraction processing, and the video stream after frame extraction processing is played at a first playback frame rate (first playback frame rate = first shooting frame rate); fast motion playback of the video not containing voice is realized. No frame extraction processing is performed on the image frames containing voice, and normal speed playback is realized, and the original sound of the video is played normally, thereby improving the user experience.

[0271] The above embodiments are respectively introduced for the two scenarios of slow motion playback and fast motion playback. It should be noted that the above two scenarios of slow motion playback and fast motion playback can be combined. For example, the motion interval in a video is played in slow motion, the interval containing voice is played at normal speed, and the rest is played in fast motion. The various embodiments in the above two scenarios can also be combined arbitrarily. The specific implementation method can refer to the description of the slow motion playback scenario and the fast motion playback scenario in the above embodiments, which will not be repeated here.

[0272] In other embodiments, in the above slow motion playback and fast motion playback scenarios, the mobile phone can perform atmosphere detection based on the video file to determine the atmosphere type of the video file, and match the corresponding soundtrack according to the atmosphere type.

[0273] Exemplarily, a plurality of atmosphere types are preset in the mobile phone, for example, the atmosphere types include childishness, pets, Spring Festival, Christmas, birthday, wedding, graduation, food, art, travel, sports, nature, etc. In one implementation, the mobile phone performs image analysis on each image frame in the video file to determine the atmosphere type corresponding to each image frame. For example, if it is detected that the image frame includes a birthday cake, it is determined that the image frame corresponds to the "birthday" atmosphere type. The mobile phone determines the atmosphere type of the video in combination with the atmosphere types of all image frames in the video file; for example, the video file includes 48 image frames, of which 35 frames have an atmosphere type of "birthday", then it is determined that the video corresponds to the "birthday" atmosphere type. In another implementation, the mobile phone performs voice analysis on the audio stream in the video file to determine the atmosphere type corresponding to the audio stream. For example, according to the voice analysis, the audio stream includes the voice "Happy Birthday", then it is determined that the video corresponds to the "birthday" atmosphere type. It can be understood that in another implementation, the mobile phone can determine the atmosphere type of the video in combination with the image and audio in the video file.

[0274] In one implementation, a plurality of music pieces are preset in the mobile phone, and a correspondence between the atmosphere type and the music pieces is preset. In one example, the correspondence between the atmosphere type and the music pieces is shown in Table 3.

[0275] Table 3

[0276] Atmosphere Type music Children's fun, pets, Spring Festival, Christmas, graduation, travel Music 1 Birthdays, weddings, food, art, nature Music 2 sports Music 3

[0277] In one implementation, multiple music types are preset in the mobile phone, and a correspondence between atmosphere types and music types is preset; wherein each music type includes one or more music. Exemplarily, the music types include cheerful, warm, intense, etc. In one example, the correspondence between atmosphere types and music is shown in Table 4.

[0278] Table 4

[0279] Atmosphere Type Music Type music Children's fun, pets, Spring Festival, Christmas, graduation, travel Cheerful Music 1, Music 4, Music 5 Birthdays, weddings, food, art, nature Warmth Music 2 sports fierce Music 3

[0280] It is understandable that the electronic device provided in the embodiment of the present application includes a hardware structure and / or software module corresponding to each function in order to realize the above functions. Those skilled in the art should easily realize that, in conjunction with the units and algorithm steps of each example described in the embodiment disclosed herein, the embodiment of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiment of the present application.

[0281] The embodiment of the present application can divide the functional modules of the above-mentioned electronic device according to the above-mentioned method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0282] In one example, see Fig. 22 , which shows a possible structural diagram of the electronic device involved in the above embodiment. The electronic device 2200 includes: a processing unit 2210, a storage unit 2220 and a display unit 2230.

[0283] The processing unit 2210 is used to control and manage the actions of the electronic device 2200. For example, it acquires a first video file, generates a marker file, performs frame insertion or frame extraction on image frames, processes an audio stream, generates a second video file, and the like.

[0284] The storage unit 2220 is used to store program codes and data of the electronic device 2200. For example, the first video file, the markup file, the second video file, etc. are stored.

[0285] The display unit 2230 is used to display the interface of the electronic device 2200. For example, slow motion video, fast motion video, normal speed video, etc. are displayed.

[0286] Of course, the unit modules in the above-mentioned electronic device 2200 include but are not limited to the above-mentioned processing unit 2210, storage unit 2220 and display unit 2230.

[0287] Optionally, the electronic device 2200 may further include an image acquisition unit 2240. The image acquisition unit 2240 is used to acquire images.

[0288] Optionally, the electronic device 2200 may further include an audio unit 2250. The audio unit 2250 is used to collect audio, play audio, and the like.

[0289] Optionally, the electronic device 2200 may further include a communication unit 2260. The communication unit 2260 is used to support the electronic device 2200 to communicate with other devices, for example, to obtain video files from other devices.

[0290] Among them, the processing unit 2210 can be a processor or a controller, for example, a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The storage unit 2220 can be a memory. The display unit 2230 can be a display screen, etc. The image acquisition unit 2240 can be a camera, etc. The audio unit 2250 can include a microphone, a speaker, etc. The communication unit 2260 can include a mobile communication unit and / or a wireless communication unit.

[0291] For example, the processing unit 2210 is a processor (such as Figure 3 The processor 110 shown in FIG. 10A ) and the storage unit 2220 may be a memory (such as Figure 3 The internal memory 121 shown in FIG. 1 ), the display unit 2230 may be a display screen (such as Figure 3 The image acquisition unit 2240 may be a camera (eg, Figure 3 The audio unit 2250 may be an audio module (eg, Figure 3 The communication unit 2260 may include a mobile communication unit (such as Figure 3The mobile communication module 150 shown) and the wireless communication unit (such as Figure 3 The wireless communication module 160 shown in FIG. 1 ). The electronic device 2200 provided in the embodiment of the present application may be Figure 3 The electronic device 100 shown in FIG. The processor, memory, display screen, camera, audio module, mobile communication unit, wireless communication unit, etc. may be connected together, for example, via a bus.

[0292] The embodiment of the present application also provides a chip system, which includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected by lines. For example, the interface circuit can be used to receive signals from other devices (such as a memory of an electronic device). For another example, the interface circuit can be used to send signals to other devices (such as a processor). Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can perform the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, which are not specifically limited in the embodiment of the present application.

[0293] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned electronic device, the electronic device executes each function or step executed by the mobile phone in the above-mentioned method embodiment.

[0294] The embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute each function or step executed by the mobile phone in the above method embodiment.

[0295] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0296] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0297] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0298] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0299] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0300] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A video processing method, applied to an electronic device, wherein the electronic device comprises a camera device and a recording device, wherein: The method comprises: In response to starting the operation of recording a video, the electronic device collects image frames at a first frame rate through the camera device, and collects audio frames through the audio recording device; Upon receiving an operation to end recording the video, the electronic device stops collecting image frames and audio frames, and generates a first video file; the first video file includes a first video portion consisting of first image frames and a second video portion consisting of second image frames, wherein the first image frames include a specified action; the first video file also includes a first audio frame, wherein the first audio frame includes a first audio portion corresponding to a timeline of the first video portion and a second audio portion corresponding to a timeline of the second video portion; The electronic device performs resolution reduction processing on the image frames captured by the camera device to obtain corresponding low-resolution image frames; The electronic device caches the multiple low-resolution image frames before and after, marks the first image frame in the first video file according to the multiple low-resolution image frames before and after, and generates marking information; the marking information includes designated action start information and designated action end information; The electronic device processes the first video file to generate a second video file; the second video file includes a third video portion and the second video portion, wherein the third video portion is obtained by performing frame insertion processing on the first video portion, and the first video portion is determined from the first video file according to the tag information; The electronic device performs speech recognition on the first audio part to generate text corresponding to the first audio sub-part containing speech in the first audio part; the electronic device plays the second video file at the first frame rate; the playing time of the second video file is longer than the shooting time of the first video file; When the electronic device plays the second video file, the text is displayed in the first video subsection of the third video section in the form of subtitles; wherein the first video subsection of the third video section is obtained by interpolating the second video subsection of the first video section, and the second video subsection is an image frame corresponding to the timeline of the first audio subsection; The duration of the first audio sub-part is a first duration, the display duration of the text is N times of the first duration; N is the interpolation multiple of the interpolation process; or, The duration of the audio frame corresponding to the first character in the text is the first duration, and the display duration of the first character is N times the first duration; N is the interpolation multiple of the interpolation processing.

2. The method according to claim 1, characterized in that After the electronic device processes the first video file, the method further includes: Upon receiving the operation of sharing the video, the electronic device forwards the second video file.

3. The method according to claim 1, characterized in that The method further comprises: The electronic device obtains the first video portion according to the tag information and the first video file.

4. The method according to claim 3, characterized in that The method further comprises: Upon receiving an operation to edit the video, the electronic device displays a first interface; the first interface includes part or all of the image frames of the first video file; Upon receiving an operation by the user to modify the image frame interval range including the specified action on the first interface, the electronic device updates the marking information according to the modified image frame interval range including the specified action.

5. The method according to any one of claims 1 to 4, characterized in that: Before the electronic device processes the first video file, the method further includes: An operation of playing the video is received.

6. The method according to any one of claims 1 to 4, characterized in that: The first frame rate is 24 frames per second.

7. The method according to claim 1, characterized in that The method further comprises: The electronic device acquires a third video file, where the third video file includes a first-category image frame and a first-category audio frame; the shooting frame rate of the first-category image frame is a first frame rate; The first type of audio frame includes a first sub-audio portion composed of the second type of audio frames and a second sub-audio portion composed of the third type of audio frames; the third type of audio frames contains speech; The first type of image frames includes a first sub-video portion composed of second type of image frames and a second sub-video portion composed of third type of image frames; the second type of image frames correspond to the second type of audio frames on a timeline, and the third type of image frames correspond to the third type of audio frames on a timeline; The electronic device processes the third video file to generate a fourth video file; the fourth video file includes a third sub-video portion and the second sub-video portion, wherein the third sub-video portion is obtained by performing frame insertion processing or frame extraction processing on the first sub-video portion; The electronic device plays the fourth video file at the first frame rate; The electronic device acquiring the third video file includes: In response to starting the operation of recording a video, the electronic device collects image frames at a first frame rate through the camera device and collects audio frames through the audio recording device; Upon receiving an operation to end recording the video, the electronic device stops capturing image frames and stops capturing audio frames, and generates the third video file.

8. The method according to claim 7, characterized in that After the electronic device processes the third video file, the method further includes: Upon receiving the operation of sharing the video, the electronic device forwards the fourth video file.

9. The method according to claim 7, characterized in that: The method further comprises: When the electronic device plays the third sub-video portion in the fourth video file, the audio frame stops playing; When the electronic device plays the second sub-video part in the fourth video file, it plays the third type of audio frame.

10. The method according to claim 9, characterized in that The method further comprises: When the electronic device plays the third sub-video portion of the fourth video file, the soundtrack is played at a first volume; When the electronic device plays the second sub-video part in the fourth video file, the soundtrack is played at a second volume; the second volume is smaller than the first volume, and the second volume is smaller than the playback volume of the third type of audio frame.

11. The method according to any one of claims 7 to 10, characterized in that: Before the electronic device processes the third video file, the method further includes: The electronic device marks the image frame corresponding to the third type of audio frame timeline to generate second marking information; The electronic device obtains the first sub-video portion according to the second tag information and the third video file.

12. An electronic device, characterized in that: The electronic device comprises: a processor and a memory; the processor is coupled to the memory; the memory is used to store computer program code; the computer program code comprises computer instructions, and when the processor executes the above-mentioned computer instructions, the electronic device executes the method as described in any one of claims 1-11.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer instructions, and when the computer instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Intelligent video recording method, electronic equipment and computer readable storage medium

    CN112532903A

  • Video recording method and electronic equipment

    CN113067994A