A video processing method and related device

By performing scene and transition analysis on the videos recorded by users, deleting invalid clips and editing wonderful clips, the problem of too much meaningless content in the recorded video is solved, and the viewing of videos is improved.

CN115529378BActive Publication Date: 2025-06-13HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210193721.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-06-13
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

When users use smartphones and other electronic devices to record videos, the recorded videos usually contain too much meaningless content, which leads to users feeling tired when watching the videos and having a poor viewing experience.

Method used

By performing scene analysis and transition analysis on the videos recorded by the user, invalid clips are deleted, multiple wonderful video clips are edited, and these clips are fused into one wonderful video to improve the viewing of the video.

Benefits of technology

It achieves the viewing of the video recorded by users, makes the video content more compact and interesting, and enhances the viewing experience of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115529378B_ABST
    Figure CN115529378B_ABST
Patent Text Reader

Abstract

The present application provides a video processing method and related devices, which can analyze scenes and transitions in videos recorded by users, delete invalid clips in the recorded videos, edit multiple wonderful video clips in the recorded videos, and merge these multiple wonderful video clips into one wonderful video. In this way, the viewing quality of videos recorded by users can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a video processing method and related devices. Background Art

[0002] Since the development of smartphones to date, taking photos and videos has become one of the most important features. As the photo-taking and video-recording functions of electronic devices such as smartphones become more and more powerful, more and more people use electronic devices such as smartphones to replace professional cameras for taking photos.

[0003] When a user uses an electronic device such as a smartphone to record a video, the electronic device needs to synthesize an image stream and an audio stream continuously acquired within a period of time into a video stream. Since there is a lot of content in the video recorded by the user, when the user looks back at the recorded video, it is easy to feel tired due to too much uninteresting content in the video, and the viewing experience of the user is poor. Summary of the Invention

[0004] This application provides a video processing method and related devices, which realizes deleting meaningless segments in the recorded video, editing multiple wonderful video segments in the recorded video, and fusing these multiple wonderful video segments into a wonderful video by analyzing the scenes and transitions in the video recorded by the user. The viewing property of the video recorded by the user is improved.

[0005] In a first aspect, this application provides a video processing method, including: an electronic device displays a shooting interface, the shooting interface includes a preview frame and a recording start control, and a picture captured in real time by a camera of the electronic device is displayed in the preview frame; the electronic device detects a first input to the recording start control; in response to the first input, the electronic device starts to record a first video; the electronic device displays a recording interface, the recording interface includes a recording end control and a video picture of the first video recorded in real time by the electronic device; the electronic device detects a second input to the recording end control; in response to the second input, the electronic device ends the recording of the first video; the electronic device saves the first video and a second video; wherein, the first video includes a first video segment, a second video segment and a third video segment, the end time of the first video segment is earlier than or equal to the start time of the second video segment, and the end time of the second video segment is earlier than or equal to the start time of the third video segment; the second video includes the first video segment and the third video segment, and does not include the second video segment.

[0006] The present application provides a video processing method, which can analyze the scenes in the video recorded by the user, delete invalid clips in the recorded video (for example, scene switching, screen zooming, fast camera movement, severe screen shaking, etc.), edit multiple wonderful video clips of designated shooting scenes (for example, people, Spring Festival, Christmas, ancient buildings, beaches, fireworks, plants or snow scenes, etc.) in the recorded video, and merge these multiple wonderful video clips into one wonderful video. In this way, the viewing quality of the video recorded by the user can be improved.

[0007] In one possible implementation, the duration of the first video is greater than that of the second video; or, the duration of the first video is less than that of the second video; or, the duration of the first video is equal to that of the second video.

[0008] In a possible implementation, before the electronic device saves the second video, the method further includes: the electronic device splicing the first video segment and the third video segment in the first video together to obtain the second video.

[0009] In a possible implementation, the electronic device splices the first video segment and the third video segment in the first video together to obtain the first video, specifically including: the electronic device splices the end position of the first video segment and the start position of the third video segment together to obtain the second video; or, the electronic device splices the end position of the first video segment and the start position of the first special effects segment together, and splices the end position of the first special effects segment and the start position of the third video segment together to obtain the second video.

[0010] In a possible implementation manner, the first video segment and the third video segment are wonderful video segments, and the second video segment is an invalid video segment.

[0011] In a possible implementation, the first video also includes a fourth video clip; if the fourth video clip is a wonderful video clip, the second video includes the fourth video clip; if the fourth video clip is an invalid video clip, the second video does not include the fourth video clip.

[0012] In a possible implementation, the wonderful video clip includes a video clip in which the shooting scene in the first video is a wonderful scene and does not include a transition clip.

[0013] In a possible implementation, the wonderful video clip includes a video clip in which the shooting scene in the first video is a designated wonderful scene and does not include a transition clip with noise or no sound.

[0014] Among them, the wonderful scenes include one or more of the scenes such as people, scenery, food, Spring Festival, Christmas, buildings, beaches, fireworks, plants, snow scenes or travels, and so on.

[0015] In a possible implementation manner, the recording interface further includes a snapshot control. When the electronic device displays the recording interface, the method further includes: the electronic device receives a third input from the user for the snapshot control; in response to the third input, the electronic device saves the first video frame of the first video when receiving the third input as a first picture.

[0016] In a possible implementation manner, after the electronic device finishes recording the first video, the method further includes: the electronic device saves a third video, where the first video includes a fifth video segment and a sixth video segment; the end time of the fifth video segment is earlier than or equal to the start time of the sixth video segment, the third video includes the fifth video segment and the sixth video segment, and the same shooting subject is included in both the fifth video segment and the sixth video segment.

[0017] In a possible implementation manner, after the electronic device saves the first video and the second video, the method further includes: the electronic device displays a video album interface, and the video album interface includes a first option corresponding to the first video; after the electronic device detects a fourth input for the first option, it displays a first video display interface of the first video. The first video display interface of the first video includes a first display area for the first video and a second display area for the second video. The first display area is used to display the video frame of the first video, and the second display area is used to display the video frame of the second video. In this way, classifying the first video and the second video in one video display interface facilitates the user to find the first video and the second video.

[0018] In a possible implementation manner, after the electronic device saves the first video and the second video, the method further includes: the electronic device displays a video album interface, and the video album interface includes a first option corresponding to the first video and a second option corresponding to the second video; when the electronic device detects a fourth input for the first option, it displays a first video display interface of the first video. The first display interface of the first video includes a first display area for the first video, and the first display area is used to display the video frame of the first video; when the electronic device detects a fifth input for the second option, it displays a second video display interface of the second video. The second display interface of the second video includes a second display area for the second video, and the second display area is used to display the video frame of the second video. In this way, displaying the options of the first video and the second video in parallel in a video album can facilitate the user to quickly open the display interface of the first video or the display interface of the second video.

[0019] In a possible implementation, after the electronic device saves the first video and the second video, the method further includes: the electronic device displays the shooting interface and displays a first prompt on the shooting interface, the first prompt being used to prompt the user that the electronic device has generated and saved the second video from the recorded first video. In this way, the user can see the generated second video in a timely manner.

[0020] In a possible implementation, after the first input to the recording start control is detected, the method further includes: the electronic device collects the image stream of the first video in real time through a camera, and collects the audio stream of the first video in real time through a microphone; the electronic device performs scene detection on the image stream of the first video, and determines the scene category of each picture frame in the image stream of the first video; the electronic device performs transition detection on the image stream of the first video, and determines the transition position and transition category of the scene transition in the image stream of the first video; the electronic device divides the image stream of the first video into the following categories based on the scene category of each picture frame in the image stream of the first video and the transition position and transition category of the scene transition in the image stream of the first video. The electronic device is a device for recording a plurality of screen segments, and determining a segment theme of each of the plurality of screen segments; the electronic device determines a plurality of wonderful screen segments under a wonderful theme from the screen segments based on the segment themes of the plurality of screen segments, and records the positions of the plurality of wonderful screen segments in the image stream in the first video; after the electronic device detects a second input for the recording end control, the electronic device mixes the image stream of the first video and the audio stream of the first video into the first video; the electronic device extracts the first video segment and the third video segment from the original video based on the positions of the plurality of wonderful screen segments in the image stream in the first video; the electronic device generates the second video based on the first video segment and the third video segment.

[0021] In this way, by performing scene analysis and scene transition analysis on the recorded video during the user's video recording process, invalid clips (such as scene switching, screen zooming, fast camera movement, severe screen shaking, etc.) can be deleted, multiple wonderful video clips in the recorded video can be edited, and these multiple wonderful video clips can be merged into one wonderful video. In this way, the viewing quality of the user's recorded video can be improved.

[0022] In a possible implementation, after the first input to the recording start control is detected, the method further includes: the electronic device collects the image stream of the first video in real time through a camera and collects the audio stream of the first video in real time through a microphone; the electronic device performs scene detection on the image stream of the first video to determine the scene category of each frame in the image stream of the first video; the electronic device performs transition detection on the image stream of the first video to determine the transition position and transition category of the scene transition in the image stream of the first video; the electronic device performs sound activation detection on the audio stream of the first video to identify the start and end time points of the voice signal in the audio stream of the first video, and divides the audio stream of the first video into multiple audio segments based on the start and end time points of the voice signal; the electronic device performs audio event classification on the multiple audio segments in the audio stream of the first video to determine the audio event category of each audio segment in the multiple audio segments; the electronic device divides the image stream of the first video into multiple frames based on the scene category of each frame in the image stream of the first video and the transition position and transition category of the scene transition in the image stream of the first video. The electronic device determines a plurality of audio event image segments corresponding to the plurality of audio segments in the image stream of the first video, and an audio event category corresponding to each audio event image segment, based on the start and end time points of the plurality of audio segments and the segment theme of each of the plurality of image segments; the electronic device divides the image stream of the first video into a plurality of image segments based on the scene category of each frame in the image stream of the first video, the transition position and transition category of the scene transition in the image stream of the first video, and the audio event category of the plurality of audio event image segments, and determines a segment theme of each of the plurality of image segments; after the electronic device detects a second input for the recording end control, the electronic device mixes the image stream of the first video and the audio stream of the first video into the first video; the electronic device extracts the first video segment and the third video segment from the first video based on the positions of the plurality of wonderful image segments in the image stream of the first video; the electronic device generates the second video based on the first video segment and the third video segment.

[0023] In this way, the scene analysis, transition analysis and audio event analysis of the recorded video can be performed while the user is recording the video, meaningless clips in the recorded video can be deleted, multiple wonderful video clips in the recorded video can be edited, and these multiple wonderful video clips can be merged into one wonderful video. In this way, the viewing quality of the video recorded by the user can be improved.

[0024] In a possible implementation, after the electronic device generates the second video, the method further includes: the electronic device adding background music to the second video; the electronic device saving the second video, specifically including: the electronic device saving the second video with the added background music.

[0025] In a possible implementation, the first input includes one or more of the following: gesture input, click input, double-click input, and so on.

[0026] In a second aspect, the present application provides an electronic device, including a display screen, a camera, one or more processors, and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, and the computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device is caused to execute the video processing method in any possible implementation of any aspect above.

[0027] In a third aspect, the present application provides a chip system, applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the video processing method in any possible implementation of any aspect above.

[0028] In a fourth aspect, the present application provides a computer storage medium, including computer instructions. When the computer instructions run on an electronic device, the electronic device is caused to execute the video processing method in any possible implementation of any aspect above.

[0029] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a computer, the computer is caused to execute the video processing method in any possible implementation of any aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;

[0031] Figure 2 It is a schematic diagram of the software architecture of an electronic device provided by an embodiment of the present application;

[0032] Figures 3A - 3I It is a set of schematic diagrams of video recording interfaces provided by an embodiment of the present application;

[0033] Figure 3J It is a schematic diagram of wonderful video splicing provided by an embodiment of the present application;

[0034] Figures 4A - 4G It is a set of schematic diagrams of wonderful video display interfaces provided by an embodiment of the present application;

[0035] Figures 5A - 5E Schematic diagram of a set of wonderful video setting interfaces provided by an embodiment of the present application;

[0036] Figures 6A - 6F Schematic diagram of a set of interfaces for generating wonderful videos provided by an embodiment of the present application;

[0037] Figures 7A - 7H Another set of schematic diagrams of interfaces for generating wonderful videos provided by an embodiment of the present application;

[0038] Figures 8A - 8C Schematic diagram of an interface for generating wonderful videos in a video call scenario provided by an embodiment of the present application;

[0039] Figure 9 Schematic diagram of a process flow of a video processing method provided by an embodiment of the present application;

[0040] Figure 10 Schematic diagram of a timing sequence for generating wonderful videos provided by an embodiment of the present application;

[0041] Figure 11 Schematic diagram of the splicing of wonderful videos provided by an embodiment of the present application;

[0042] Figure 12 Schematic diagram of modules of a video processing system provided by an embodiment of the present application;

[0043] Figure 13 Schematic diagram of a process flow of a video processing method provided by another embodiment of the present application;

[0044] Figure 14 Schematic diagram of a timing sequence for generating wonderful videos provided by another embodiment of the present application;

[0045] Figure 15 Schematic diagram of modules of a video processing system provided by another embodiment of the present application. Detailed implementation manners

[0046] Next, the technical solutions in the embodiments of the present application will be clearly and elaborately described in conjunction with the accompanying drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0047] Hereinafter, the terms "first" and "second" are for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0048] In the following embodiments of the present application, the term "user interface (UI)" is a media interface for interaction and information exchange between an application or an operating system and a user, and it realizes the conversion between the internal form of information and the form acceptable to the user. The user interface is source code written in a specific computer language such as Java, Extensible Markup Language (XML), etc. The interface source code is parsed and rendered on an electronic device and finally presented as content recognizable by the user. The common manifestation form of the user interface is the graphical user interface (GUI), which refers to the user interface related to computer operations displayed in a graphical manner. It can be visual interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, Widgets, etc. displayed on the display screen of an electronic device.

[0049] Figure 1 The structural schematic diagram of the electronic device 100 is shown.

[0050] Hereinafter, the embodiments will be specifically described by taking the electronic device 100 as an example. It should be understood that Figure 1 the shown electronic device 100 is only an example, and the electronic device 100 may have more or fewer components than those Figure 1 shown, may combine two or more components, or may have different component configurations. The various components shown in the figure can be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.

[0051] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0052] It can be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0053] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0054] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.

[0055] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the said memory. This avoids repeated accesses and reduces the waiting time of the processor 110, thus improving the efficiency of the system.

[0056] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, and / or a universal serial bus (USB) interface, etc.

[0057] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change the display information.

[0058] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel may adopt a liquid crystal display (LCD). The display panel may also be made of an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniled, a microled, a micro-oled, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0059] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0060] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the algorithms for image noise, brightness, and skin color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.

[0061] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0062] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0063] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0064] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.

[0065] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0066] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area.

[0067] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.

[0068] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0069] The speaker 170A, also called a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A.

[0070] The receiver 170B, also called a "handset", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be listened to by bringing the receiver 170B close to the human ear.

[0071] The microphone 170C, also called a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C.

[0072] The earphone interface 170D is used to connect a wired earphone and can be a USB interface 130 or a 3.5 mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0073] The pressure sensor 180A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc.

[0074] The gyro sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyro sensor 180B detects the angle of the electronic device 100 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device 100 through reverse movement to achieve anti-shake. The gyro sensor 180B can also be used for navigation and somatosensory game scenes.

[0075] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in all directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0076] The distance sensor 180F is used to measure the distance. The electronic device 100 can measure the distance by infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure the distance to achieve fast focusing.

[0077] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.

[0078] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy.

[0079] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 together form a touch screen, also known as the "touch control screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a position different from that of the display screen 194.

[0080] Figure 2 An exemplary schematic diagram of the software architecture of the electronic device according to an embodiment of the present application is shown.

[0081] As Figure 2 shown, the layered architecture divides the system into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom, namely the application layer, the application framework layer, the hardware abstraction layer, the kernel layer, and the hardware layer.

[0082] The application layer may include a series of application packages.

[0083] The application packages may include a camera, a gallery, etc.

[0084] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0085] In some embodiments, the application framework layer may include a camera access interface. Among them, the camera access interface may include camera management and camera devices. The camera access interface is used to provide application programming interfaces and programming frameworks for the camera application.

[0086] The hardware abstraction layer is an interface layer located between the application framework layer and the kernel layer, and provides a virtual hardware platform for the operating system.

[0087] In the embodiments of the present application, the hardware abstraction layer may include a camera hardware abstraction layer and a camera algorithm library.

[0088] Among them, the camera hardware abstraction layer can provide virtual hardware for camera device 1 (the first camera) and camera device 2 (the second camera). It can also obtain attitude data and transmit it to the camera algorithm library. The camera hardware abstraction layer can also be used to calculate the number N of images to be stitched. And, obtain information from the camera algorithm library.

[0089] The camera algorithm library may include an algorithm module and a motion detection module.

[0090] Among them, the algorithm module includes several algorithms for processing images, which can be used to implement the stitching and other processing of N frames of images to be stitched.

[0091] The motion detection module can be used to calculate whether the current shooting scene of the electronic device is moving.

[0092] The kernel layer is the layer between the hardware and the software. The kernel layer includes drivers for various hardware.

[0093] In some embodiments, the kernel layer may include a camera device driver, a digital signal processor driver, an image processor driver, etc.

[0094] Among them, the camera device driver is used to drive the sensor of the camera to collect images and drive the image signal processor to preprocess the images.

[0095] The digital signal processor driver is used to drive the digital signal processor to process images.

[0096] The image processor driver is used to drive the graphics processor to process images.

[0097] Next, in combination with the above hardware structure and system structure, the method in the embodiments of the present application will be specifically described:

[0098] 1. The electronic device 100 starts the video recording function and obtains an image stream and an audio stream.

[0099] This step 1 is continuously performed. In response to the user's operation on the video recording start control in the shooting interface (such as a click operation), the camera application calls the camera access interface of the application framework layer to start the camera application. Then, by calling the camera device 1 (the first camera) in the camera hardware abstraction layer to send an instruction to start video recording, the camera hardware abstraction layer sends this instruction to the camera device driver in the kernel layer. This camera device driver can start the sensor of the first camera of the camera (sensor 1), collect image optical signals through sensor 1, and transmit this image optical signal to the image signal processor for preprocessing to obtain an image stream (an image sequence composed of at least 2 frames of original image frames). Then, this original stream is transmitted to the camera hardware abstraction layer through the camera device driver. The camera application also sends an instruction to start video recording through the audio input unit in the audio hardware abstraction layer. The audio hardware abstraction layer sends this instruction to the audio driver in the kernel layer, and this audio driver can start the microphone to collect audio signals to obtain an audio stream.

[0100] 2. The electronic device 100 obtains a processing stream according to the image stream.

[0101] This step 2 is carried out continuously. The camera hardware abstraction layer can send the original stream to the camera algorithm library. With the support of the digital signal processor and the image processor, the camera algorithm library can first downsample the original stream to obtain a processed stream with low resolution.

[0102] 3. The electronic device 100 performs scene detection and transition detection on the image frames in the processed stream to determine the wonderful picture segments.

[0103] This step 3 is carried out continuously. With the support of the digital signal processor and the image processor, the camera algorithm library can call scene detection algorithms, transition detection algorithms, etc. to detect the scene category of each frame in the image stream, the transition position and transition category where a scene transition occurs, etc., and then determine the wonderful picture segments.

[0104] 4. The electronic device 100 mixes the image stream and the audio stream into the original video.

[0105] With the support of the digital signal processor and the image processor, the image stream and the audio stream can be mixed into the original video based on the same time track.

[0106] 5. The electronic device 100 can extract multiple wonderful video segments from the original video based on the positions of the wonderful picture segments, and fuse the multiple wonderful video segments into one wonderful video.

[0107] The camera algorithm library can call the clipping algorithm and the fusion algorithm to extract multiple wonderful video segments from the video stream based on the positions of the wonderful picture segments, and fuse the multiple wonderful video segments into one wonderful video. Among them, the wonderful video segment can include the video segment in the original video where the shooting scene is a wonderful scene and does not include the transition segment. Or, the wonderful video segment can include the video segment in the original video where the shooting scene is a wonderful scene and does not include the transition segment with noise or no sound. Among them, the wonderful scene includes one or more scenes such as people, scenery, food, Spring Festival, Christmas, architecture, beach, fireworks, plants, snow scene or travel, etc.

[0108] 6. The electronic device 100 can save the wonderful video and the original video.

[0109] The camera algorithm library can send the wonderful video to the camera hardware abstraction layer. Then, the camera hardware abstraction layer can save it.

[0110] In an embodiment of the present application, a video processing method is provided, which can analyze the scenes in the video recorded by the user, delete invalid clips in the recorded video (for example, scene switching, screen zooming, fast camera movement, severe screen shaking, etc.), edit multiple wonderful video clips of designated shooting scenes (for example, people, Spring Festival, Christmas, ancient buildings, beaches, fireworks, plants or snow scenes, etc.) in the recorded video, and merge these multiple wonderful video clips into one wonderful video. In this way, the viewing quality of the video recorded by the user can be improved.

[0111] A video processing method provided by an embodiment of the present application is introduced below in conjunction with an application scenario.

[0112] In some application scenarios, the user can record a video in a normal video recording mode in the camera application of the electronic device 100. During the process of recording a video by the electronic device 100, the electronic device 100 can identify and edit multiple wonderful video clips of wonderful scenes in the recorded original video, and merge these multiple wonderful frequency bands into one wonderful video. After the video recording is completed, the electronic device 100 can save the original video and the wonderful video. In this way, the viewing quality of the video recorded by the user can be improved.

[0113] For example, Figure 3A As shown, the electronic device 100 can display a desktop 310, in which a page with application icons is displayed, and the page includes multiple application icons (for example, weather application icon, stock application icon, calculator application icon, setting application icon, mail application icon, gallery application icon 312, music application icon, video application icon, browser application icon, etc.). A page indicator is also displayed below the multiple application icons to indicate the positional relationship between the currently displayed page and other pages. There are multiple tray icons (for example, dialing application icon, information application icon, contact application icon, camera application icon 311) below the page indicator, and the tray icon remains displayed when the page is switched. In some embodiments, the above-mentioned page may also include multiple application icons and page indicators, and the page indicator may not be part of the page and exist separately. The above-mentioned tray icon is also optional, and the embodiments of the present application are not limited to this.

[0114] The electronic device 100 may receive an input operation (eg, a single click) from a user on the camera application icon 311, and in response to the input operation, the electronic device 100 may display a Figure 3B The shooting interface 320 is shown.

[0115] like Figure 3BAs shown in the figure, the shooting interface 320 may include an echo control 321, a shooting control 322, a camera switching control 323, a preview frame, a setting control 325, a zoom ratio control 326, and one or more shooting mode controls (for example, "night scene mode" control 327A, "portrait shooting mode" control 327B, "large aperture mode" control 327C, "ordinary shooting mode" control 327D, "video recording mode" control 327E, "professional mode" control 327F, "more modes" control 327G). Among them, a preview image 324 is displayed in the preview frame. The echo control 321 can be used to display the captured images. The shooting control 322 is used to trigger the saving of the images captured by the camera. The camera switching control 323 can be used to switch the camera for taking pictures. The setting control 325 can be used to set the shooting function. The zoom ratio control 326 can be used to set the zoom ratio of the camera. The shooting mode control can be used to trigger the image processing process corresponding to the shooting mode. For example, the "night scene mode" control 327A can be used to trigger an increase in the brightness and color richness in the captured image. The "portrait mode" control 327B can be used to trigger the beautification process of the portrait in the captured image. As Figure 3B shown, the currently selected shooting mode by the user is the "ordinary shooting mode".

[0116] The electronic device 100 can receive an input (such as a click) from the user to select the "video recording mode" control 327E, as Figure 3C shown. In response to this input, the electronic device 100 can switch from the "ordinary shooting mode" to the "video recording mode", and replace the above shooting control 322 with a recording start control 331. The electronic device 100 can also display the video recording time information 332.

[0117] As Figure 3CAs shown, when the electronic device 100 is recording video, there are person A, person B, and a Ferris wheel in the direction where the camera of the electronic device 100 is aimed. The electronic device 100 can receive a first input (such as a click) from the user on the recording start control 331. In response to this first input, the electronic device 100 can start recording a video. For example, after starting to record the video, the user can capture person A within the time period of 0 to 3 seconds after starting the recording. Transition from person A to capture person B within the time period of 3 seconds to 5 seconds after starting the recording. Capture person B within the time period of 5 seconds to 7 seconds after starting the recording. Transition from person B to capture the Ferris wheel within the time period of 7 seconds to 9 seconds after starting the recording. Capture the Ferris wheel within the time period of 9 seconds to 11 seconds after starting the recording. Zoom in on the details of the Ferris wheel in the video within the time period of 11 seconds to 13 seconds after starting the recording, but the picture is blurred. Continue to capture the Ferris wheel within the time period of 13 seconds to 15 seconds after starting the recording. Transition from the Ferris wheel to capture a panoramic view including person A, person B, and the Ferris wheel within the time period of 15 seconds to 17 seconds after starting the recording. Capture a panoramic view including person A, person B, and the Ferris wheel within the time period of 17 seconds to 20 seconds after starting the recording.

[0118] As Figure 3D shown, after the electronic device 100 starts recording a video, it can display a video recording interface 330. Among them, the video recording interface 330 includes a recording end control 333, a shooting control 334, video recording time information 332, and a video recording screen. The recording end control 333 can be used to trigger the electronic device 100 to end the video recording. The shooting control 334 can be used to, in response to a third input from the user, trigger the electronic device 100 to save the first video frame of the first video captured by the camera of the electronic device 100 when receiving the third input as a first picture.

[0119] Among them, the original video recorded by the electronic device 100 can include multiple exciting video segments of exciting scenes. Among them, the exciting scenes can include one or more of scenes such as people, landscapes, food, Spring Festival, Christmas, buildings, beaches, fireworks, plants, snow scenes, or travels, etc.

[0120] For example, as Figure 3D shown, in the video frame 341 at about the 4th second in the original video recorded by the electronic device 100, there is person A. The electronic device 100 can determine that the scene category of the video segment at about the 4th second in the original video is "person". As Figure 3E shown, in the video frame 342 at about the 9th second in the original video recorded by the electronic device 100, there is person B. The electronic device 100 can determine that the scene category of the video segment at about the 9th second in the original video is "person". As Figure 3FAs shown, in the video frame 343 around the 12th second of the original video recorded by the electronic device 100, there is a building (e.g., a Ferris wheel). The electronic device 100 can determine that the scene category of the video segment around the 12th second in the original video is "building". As Figure 3G shown, in the video frame 344 around the 16th second of the original video recorded by the electronic device 100, the zoom is being increased, so the building (e.g., a Ferris wheel) in the video frame 344 is relatively blurred. The electronic device 100 can determine that the video segment around the 16th second in the original video is an invalid segment. As Figure 3H shown, in the video frame 345 around the 20th second of the original video recorded by the electronic device 100, there is a panoramic view of a building, person A, and person B. The electronic device 100 can determine that the scene category of the video segment around the 20th second in the original video is "travel". Optionally, the transition part between each shooting scene can be considered invalid. For example, the transition from person A to person B within the time period of 3s - 5s after starting the recording, the transition from person B to the Ferris wheel within the time period of 7s - 9s after starting the recording, the zooming in of the picture to shoot the details of the Ferris wheel (e.g., increasing the zoom ratio) within the time period of 11s - 13s after starting the recording, and the transition from the Ferris wheel to the panoramic view including person A, person B, and the Ferris wheel within the time period of 15s - 17s after starting the recording can all be considered invalid segments.

[0121] As Figure 3H shown, the electronic device 100 can receive a second input from the user acting on the recording end control 333 (e.g., at the 20th second of starting the video recording, clicking the recording end control 333). In response to this second input, the electronic device 100 can end the video recording and save the recorded original video and the wonderful video cropped from the original video.

[0122] Among them, during the process of recording the original video, the electronic device 100 can continuously identify and crop multiple wonderful video segments in the original video that are in the above - specified scenes. After the electronic device 100 ends the recording of the original video, the electronic device 100 can fuse the multiple wonderful video segments in the original video into one wonderful video. The electronic device 100 can save the original video and the wonderful video.

[0123] Optionally, as Figure 3I shown, after the electronic device 100 ends the video recording, it can display the shooting interface 340. Among them, for the text description in the shooting interface 340, reference can be made to the foregoing Figure 3CThe written description shown above will not be elaborated here. After the electronic device 100 generates and saves an exciting video, the electronic device 100 may display a prompt message 335 (which may be referred to as the first prompt in the embodiments of the present application) on the shooting interface 340. The prompt message 335 is used to prompt the user that the electronic device 100 has generated and saved an exciting video from the originally recorded video. Among them, the prompt message 335 may be a text prompt (for example, "An exciting video has been generated from the video you shot. Please view it in the gallery"), a pattern prompt, an animation prompt, and so on.

[0124] In a possible implementation manner, after the electronic device 100 finishes recording the original video, it may save the original video, and then identify and crop multiple exciting video segments in the original video that are in the above-mentioned specified scene. After cropping multiple exciting video segments, the electronic device 100 may fuse these multiple exciting video segments into an exciting video. After generating the exciting video, the electronic device 100 may save the exciting video.

[0125] Exemplarily, such as Figure 3JAs shown, the segment from 0 to 3 seconds in the original video is shooting person A, the segment from 3 to 5 seconds transitions from person A to shooting person B, the segment from 5 to 7 seconds is shooting person B, the segment from 7 to 9 seconds transitions from person B to shooting the Ferris wheel, the segment from 9 to 11 seconds is shooting the Ferris wheel, the segment from 11 to 13 seconds zooms in on the Ferris wheel for details, but the picture is blurred, the segment from 13 to 15 seconds is shooting the Ferris wheel, and the segment from 15 to 17 seconds transitions from the Ferris wheel to shooting a panoramic view including person A, person B, and the Ferris wheel. The segment from 17 to 20 seconds is shooting the panoramic view including person A, person B, and the Ferris wheel. Among them, the segments from 3 to 5 seconds, from 7 to 9 seconds, from 11 to 13 seconds, and from 15 to 17 seconds are all during transitions or picture zooms, so they can all be determined as invalid segments. The remaining segments from 0 to 3 seconds, from 5 to 7 seconds, from 9 to 11 seconds, from 13 to 15 seconds, and from 17 to 20 seconds can all be determined as wonderful video segments. Among them, the segment from 0 to 3 seconds is wonderful video segment 1, the segment from 5 to 7 seconds is wonderful video segment 2, the segment from 9 to 11 seconds is wonderful video segment 3, the segment from 13 to 15 seconds is wonderful video segment 4, and the segment from 17 to 20 seconds is wonderful video segment 5. The electronic device 100 can splice wonderful video segment 1, wonderful video segment 2, wonderful video segment 3, wonderful video segment 4, and wonderful video segment 5 together in chronological order from start to end to obtain a wonderful video. For example, the end of wonderful video segment 1 can be spliced with the start of wonderful video segment 2, the end of wonderful video segment 2 can be spliced with the start of wonderful video segment 3, the end of wonderful video segment 3 can be spliced with the start of wonderful video segment 4, and the end of wonderful video segment 4 can be spliced with the start of wonderful video segment 5.

[0126] Optionally, if Figure 3J wonderful video segment 4 in it is an invalid segment, for example, a blurred picture or no shooting object, or the user's hand shakes or a passerby passes by and blocks or other obstacles block, resulting in this wonderful video segment 4 being an invalid segment; then, the wonderful video is wonderful video segment 1, wonderful video segment 2, wonderful video segment 3, and wonderful video segment 5 spliced together. Among them, the specific process of the electronic device 100 identifying and cropping multiple wonderful video segments in the original video and fusing multiple wonderful video segments into one wonderful video can refer to the subsequent embodiments of this application and will not be elaborated here.

[0127] In the embodiments of this application, the above first input, second input, and other inputs include, but are not limited to, gesture input, click operation input, voice input, and so on.

[0128] In some embodiments, after the electronic device 100 saves the original video and the highlight video generated from the original video, it can simultaneously display the display area of the highlight video in the display interface of the original video. When the electronic device 100 receives an input (such as a click) from the user for the display area of the highlight video, the electronic device 100 can play the highlight video.

[0129] Exemplarily, as Figure 4A shown, the electronic device 100 can display the desktop 310. Among them, the textual description of the desktop 310 can refer to the foregoing Figure 3A shown embodiments and will not be elaborated here.

[0130] The electronic device 100 can receive an input (such as a click) from the user acting on the gallery application icon 312. In response to this input, the electronic device 100 can display a gallery application interface 410 as Figure 4B shown.

[0131] As Figure 4B shown, the gallery application interface 410 can display one or more albums (for example, all photos album, video album 416, camera album, portrait album, WeChat album, Weibo album, etc.). The electronic device 100 can display a gallery menu 411 below the gallery application interface 410. Among them, the gallery menu 411 includes a photo control 412, an album control 413, a moment control 414, and a discovery control 415. Among them, the photo control 412 is used to trigger the electronic device 100 to display all local pictures in the form of picture thumbnails. The album control 413 is used to trigger the electronic device 100 to display the albums to which the local pictures belong. As Figure 4B shown, when the current album control 413 is in the selected state, the electronic device 100 displays the gallery application interface 410. The moment control 414 can be used to trigger the electronic device 100 to display the selected pictures stored locally. The discovery control 415 can be used to trigger the electronic device 100 to display the classified albums of pictures.

[0132] The electronic device 100 can receive an input (such as a click) from the user for the video album 416. In response to this input, the electronic device 100 can display a video album interface 420 as Figure 4C shown.

[0133] As Figure 4CAs shown, the video album interface 420 may include options for one or more video files, such as option 421 corresponding to the original video recorded by the user in the above embodiment (which may be referred to as the first option in the embodiments of the present application). Thumbnails of specified frame images in the video file and video time length information may be displayed on the option of the video file. For example, the option 421 may display a thumbnail of the first frame image in the original video recorded by the user and video time length information (e.g., 20 seconds).

[0134] The electronic device 100 may receive a fourth input (such as a click) from the user on the above option 421. In response to this fourth input, the electronic device 100 may display a video display interface 430 as Figure 4D shown.

[0135] In a possible implementation, the electronic device 100 may also receive and respond to an input (such as a click) from the user on the echo control 321 in the above Figure 3I shown, and display the video display interface 430 as Figure 4D shown (which may be referred to as the first video display interface in the embodiments of the present application).

[0136] As Figure 4D shown, the video display interface 430 may include a display area 431 for the original video (which may be referred to as the first display area in the embodiments of the present application), a display area 433 for the highlight video generated from the original video (which may be referred to as the second display area in the embodiments of the present application), a menu 436, and so on. Among them, frame images and time information 432 in the original video (e.g., the time length is 20 seconds) may be displayed on the display area 431 for the original video. When the display area 431 for the original video receives an input from the user (such as a click), the electronic device 100 may play or pause the original video. Frame images and highlight video time information 435 (e.g., the time length is 12 seconds) may be displayed on the display area 433 for the highlight video. Optionally, a highlight mark 434 may be displayed on the display area 433 for the highlight video, and the highlight mark 434 can be used to prompt the user that the video displayed in the display area 433 is the highlight video generated from the original video. The menu 436 may include a share button, a favorite button, an edit button, a delete button, and a more button. The share button can be used to trigger the sharing of the original video and / or the highlight video. The favorite button can be used to trigger the collection of the original video and / or the highlight video into a favorite folder. The edit button can be used to trigger editing functions such as rotating, trimming, adding filters, and blurring for the original video and / or the highlight video. The delete button can be used to trigger the deletion of the original video and / or the highlight video. The more button can be used to trigger the opening of more functions related to the original video and / or the highlight video.

[0137] The electronic device 100 can receive an input (such as a click) from the user for the display area 433 of the highlight video. In response to this input, as Figure 4E shown, the electronic device 100 can reduce the display of the display area 431 of the original video and enlarge the display of the display area 433 of the highlight video in the video display interface 430. After the display area 433 of the highlight video is enlarged, the electronic device 100 can receive and respond to an input (such as a click) from the user for the display area 433 of the highlight video. In response to this input, the electronic device 100 can play the highlight video.

[0138] In some embodiments, after the electronic device 100 saves the above-mentioned original video and the highlight video generated from the original video, it can display the option of the original video and the option of the highlight video side by side in the video album. When the electronic device 100 receives an input from the user for the option of the original video, the electronic device 100 can display the display interface of the original video. When the electronic device 100 receives an input from the user for the option of the highlight video, the electronic device 100 can display the display interface of the highlight video.

[0139] Exemplarily, when the electronic device 100 can receive an input (such as a click) from the user for the above-mentioned Figure 4B shown video album 416, the electronic device 100 can display a video album interface 440 as Figure 4F shown.

[0140] As Figure 4F shown, the video album interface 440 may include options for multiple video files. Among them, the options for the multiple video files include the option 421 of the original video (which can be referred to as the first option in the embodiments of the present application) and the option 423 of the highlight video generated based on the original video (which can be referred to as the second option in the embodiments of the present application). Among them, the option 421 may display a thumbnail of a specified frame in the original video recorded by the user and video time length information (for example, 20 seconds). The option 423 may display a thumbnail of a specified frame in the highlight video, video time length information (for example, 12 seconds), and a highlight mark 425. Among them, the highlight mark 425 can be used to prompt the user that the video file corresponding to the option 423 is a highlight video generated from the original video.

[0141] The electronic device 100 can receive a fifth input (such as a click) from the user for the option 423 of the highlight video. In response to this fifth input, the electronic device 100 can display a video display interface 450 as Figure 4G shown (which can be referred to as the second video display interface in the embodiments of the present application).

[0142] As Figure 4GAs shown, the video display interface 450 may include a display area 451 of a wonderful video, and a frame picture and time information 452 (for example, a time length of 12 seconds) of the wonderful video may be displayed in the display area 451. Optionally, a wonderful mark 453 may be displayed on the display area 451 of the wonderful video, and the wonderful mark 453 may be used to prompt the user that the wonderful video generated from the original video is displayed in the display area 451. The menu 454 may include a share button, a favorite button, an edit button, a delete button, and more buttons. The share button may be used to trigger the sharing of the wonderful video. The favorite button may be used to trigger the collection of the wonderful video to a collection folder. The edit button may be used to trigger the editing functions such as rotation, trimming, adding filters, and blurring of the wonderful video. The delete button may be used to trigger the deletion of the wonderful video. The more buttons may be used to trigger the opening of more functions related to the wonderful video. When the display area 451 of the wonderful video receives the user's input (for example, click), the electronic device 100 may play or pause the wonderful video.

[0143] In some application scenarios, the user can record a video using a special video recording mode (e.g., wonderful video recording) in the camera application of the electronic device 100. During the process of recording a video by the electronic device 100, the electronic device 100 can identify and edit multiple wonderful video clips of a specified shooting scene in the recorded original video, and merge these multiple wonderful frequency bands into one wonderful video. After the video recording is finished, the electronic device 100 can save the wonderful video. Optionally, the electronic device 100 can also save the original video. In this way, the viewing quality of the video recorded by the user can be improved.

[0144] For example, Figure 5A As shown, the electronic device 100 can have a shooting interface 510. The shooting interface 510 may include an echo control 511, a shooting control 512, a camera conversion control 513, a preview box, a setting control 515, a zoom ratio control 516, and one or more shooting mode controls (for example, a "night scene mode" control 517A, a "portrait shooting mode" control 517B, a "large aperture mode" control 517C, a "normal shooting mode" control 517D, a "video recording mode" control 517E, a "wonderful recording mode" control 517H, a "professional mode" control 517F, a "more modes" control, etc.). The electronic device 100 can receive input (for example, a single click) of the user's option wonderful recording mode control 517H, and in response to the input, the electronic device 100 can switch from the "normal shooting mode" to the "wonderful recording mode". For the text description of the controls in the shooting interface 510, please refer to the aforementioned Figure 3B The shooting interface 320 shown in FIG. 3 is not described in detail here.

[0145] like Figure 5BAs shown, after switching to the "wonderful video recording mode", the electronic device 100 can replace the above-mentioned shooting control 512 with a video recording start control 521. The electronic device 100 can also display video recording time information 522.

[0146] The electronic device 100 can receive an input (such as a click) from the user on the video recording start control 521. In response to this input, the electronic device 100 can start recording a video. Among them, in the wonderful video recording mode, the electronic device 100 can continuously identify and crop multiple wonderful video segments in the original video that are in the above-mentioned specified scene during the process of recording the original video. After the electronic device 100 finishes recording the original video, the electronic device 100 can merge the multiple wonderful video segments in the original video into a wonderful video. The electronic device 100 can save the wonderful video. Optionally, the electronic device 100 can also save the original video.

[0147] In a possible implementation manner, as Figure 5C shown, the electronic device 100 can display a prompt message 523 on the shooting interface when switching to the wonderful video recording mode. Among them, the prompt message 523 can be used to prompt the user about the mode introduction of this wonderful video recording mode (for example, it will identify the wonderful video segments in your video recording process and generate a wonderful video).

[0148] In a possible implementation manner, the electronic device 100 can preset the wonderful scenes required by the user during video recording. After the user sets the wonderful scenes, during the video recording process of the electronic device 100, the electronic device 100 can identify multiple wonderful video segments corresponding to the wonderful scenes set by the user from the original video, and merge these multiple wonderful video segments into a wonderful video.

[0149] Exemplarily, as Figure 5C shown, the electronic device 100 can receive an input (such as a click) from the user for the setting control 515. In response to this input, the electronic device 100 can display a setting window 530 as Figure 5D shown on the shooting interface 510.

[0150] As Figure 5D shown, the setting window 530 can display including a window closing control 531, one or more setting items. For example, a resolution setting bar 532 and a wonderful scene setting bar. Among them, the wonderful scene setting bar can include setting items for one or more wonderful scenes. For example, a "character scene" setting item 533, a "scenery scene" setting item 534, an "architecture scene" setting item 535, a "food scene" setting item 536, and a "travel scene" setting item 537, etc.

[0151] As Figure 5EAs shown, the electronic device 100 can receive the user's input for the wonderful scene setting column, and select a character scene, a landscape scene, a food scene, a travel scene, etc. as a wonderful scene when generating a wonderful video. After the user sets the wonderful scene, during the recording process, the electronic device 100 can identify multiple wonderful video clips corresponding to the wonderful scene set by the user from the original video, and merge the multiple wonderful video clips into one wonderful video.

[0152] In some application scenarios, after the electronic device 100 has recorded the original video and saved the original video to the video album, the user can trigger the generation of a wonderful video from the original video in the display interface of the original video in the video album. After the user triggers the generation of a wonderful video from the original video, the electronic device 100 can identify and edit multiple wonderful video clips with wonderful scenes in the original video, and merge these multiple wonderful frequency bands into one wonderful video. After the electronic device 100 finishes recording the video, it can save the wonderful video. In this way, the viewing quality of the video recorded by the user can be improved.

[0153] For example, Figure 6A As shown, the electronic device 100 can display a gallery application interface 410. For a text description of the gallery application interface 410, please refer to the aforementioned Figure 4B The text part of the illustrated embodiment will not be repeated here.

[0154] The electronic device 100 may receive a user input (eg, a click) for the video album 416, and in response to the input, the electronic device 100 may display the video album 416. Figure 6B Video album interface 420 is shown.

[0155] like Figure 6B As shown, the video album interface 420 may include one or more video file options, such as option 421 corresponding to the original video recorded by the user in the above embodiment. For a detailed text description of the video album interface 420, please refer to the above Figure 4C The text part of the illustrated embodiment will not be repeated here.

[0156] The electronic device 100 may receive an input (eg, a click) from a user on the option 421 of the original video. In response to the input, the electronic device 100 may display the following: Figure 6C The video display interface 610 is shown.

[0157] like Figure 6CAs shown, the video display interface 610 may include a display area 611 for the original video, a menu 613, a highlight video generation control 614, and so on. Among them, frame images and time information 612 (such as a time length of 20 seconds) in the original video may be displayed on the display area 611 of the original video. When the display area 611 of the original video receives user input (such as a click), the electronic device 100 may play or pause the original video. The highlight video generation control 614 can be used to trigger the electronic device 100 to generate a highlight video from the original video displayed in the display area 611. The menu 613 may include a share button, a favorite button, an edit button, a delete button, and a more button. The share button can be used to trigger sharing of the original video. The favorite button can be used to trigger favoriting the original video to a favorite folder. The edit button can be used to trigger editing functions such as rotating, trimming, adding filters, and blurring the original video. The delete button can be used to trigger deletion of the original video. The more button can be used to trigger opening more functions related to the original video.

[0158] The electronic device 100 may receive user input (such as a click) for the highlight video generation control 614. In response to this input, the electronic device 100 may identify and crop multiple highlight video segments in the original video in highlight scenes, and fuse these multiple highlight video segments into a highlight video.

[0159] Optionally, as Figure 6D shown, during the process of the electronic device 100 generating a highlight video, the electronic device 100 may display the generation progress 615 of the highlight video on the video display interface 610 of the original video. Among them, the specific processes of the electronic device 100 identifying and cropping multiple highlight video segments in the original video and fusing multiple highlight video segments into a highlight video may refer to subsequent embodiments of this application and will not be elaborated here.

[0160] As Figure 6E shown, after the electronic device 100 generates a highlight video, the electronic device 100 may display a display area 616 corresponding to the highlight video generated from the original video on the video display interface 610 of the original video. Among them, frame images in the highlight video and time information 618 of the highlight video (such as a time length of 12 seconds) may be displayed on the display area 616 of the highlight video. Optionally, a highlight mark 617 may be displayed on the display area 616 of the highlight video, and the highlight mark 617 can be used to prompt the user that the content displayed in the display area 616 is a highlight video generated from the original video.

[0161] The electronic device 100 may receive user input (such as a click) for the display area 616 of the highlight video. In response to this input, as Figure 6FAs shown, the electronic device 100 may reduce the display area 611 of the original video and enlarge the display area 616 of the highlight video in the video display interface 610. After the display area 616 of the highlight video is enlarged, the electronic device 100 may receive and respond to an input (such as a click) from the user for the display area 616 of the highlight video. In response to this input, the electronic device 100 may play the highlight video.

[0162] In a possible implementation, when the user confirms to generate a highlight video from the original video in the display interface of the original video on the electronic device 100, the electronic device 100 may receive the highlight scene set by the user. The electronic device 100 may identify and clip multiple highlight video segments in the original video under the highlight scene based on the highlight scene set by the user, and merge these multiple highlight video segments into a highlight video. Among them, for the same original video, when the user selects different highlight scenes, the electronic device 100 may generate different highlight videos.

[0163] Exemplarily, as Figure 7A shown, the electronic device 100 may display the video display interface 610. For the textual description of the video display interface 610, reference may be made to the foregoing Figure 6C shown embodiments, which will not be elaborated herein.

[0164] The electronic device 100 may receive an input (such as a click) from the user for the highlight video generation control 614. In response to this input, the electronic device 100 may display a scene setting window 710 as Figure 7B shown.

[0165] As Figure 7B shown, the scene setting window 710 may display a confirmation control 716, a cancellation control 717, and setting items for one or more highlight scenes. For example, a "person scene" setting item 711, a "scenery scene" setting item 712, a "building scene" setting item 713, a "food scene" setting item 714, and a "travel scene" setting item 715, and so on.

[0166] As Figure 7C shown, the electronic device 100 may receive an input from the user for the setting item of the highlight scene, and select a person scene, a scenery scene, a food scene, a travel scene, etc. as the highlight scene when generating the highlight video. After the user sets the highlight scene, the electronic device 100 may receive an input (such as a click) from the user for the confirmation control 716. In response to this input, the electronic device 100 may identify multiple highlight video segments corresponding to the set highlight scene type set a from the original video, and merge these multiple highlight video segments into a highlight video (for example, highlight video a).

[0167] Optionally, as Figure 7D shown, during the process of the electronic device 100 generating an exciting video, the electronic device 100 may display the generation progress 615 of the exciting video a and the scene type of the exciting video a on the video display interface 610 of the original video. The specific processes of the electronic device 100 identifying and cropping multiple exciting video clips in the original video and fusing the multiple exciting video clips into one exciting video may refer to the subsequent embodiments of this application and will not be elaborated here.

[0168] As Figure 7E shown, after the electronic device 100 generates the exciting video a, the electronic device 100 may display a display area 616 corresponding to the exciting video a generated from the original video on the video display interface 610 of the original video. Among them, frame images in the exciting video a, time information 618 of the exciting video a (for example, the time length is 12 seconds), and scene information 619 of the exciting video a (for example, people, scenery, food, and travel) may be displayed on the display area 616 of the exciting video a. Optionally, an exciting mark 617 may be displayed on the display area 616 of the exciting video, and the exciting mark 617 may be used to prompt the user that the content displayed in the display area 616 is the exciting video a generated from the original video.

[0169] Among them, for the same original video, when the user selects different exciting scenes, the electronic device 100 may generate different exciting videos. Therefore, while and after the electronic device 100 generates the exciting video a from the original video, the electronic device 100 may continue to display an exciting video generation control 614 on the above-mentioned video display interface 610.

[0170] After the electronic device 100 generates the exciting video a from the original video, the electronic device 100 may continue to receive an input (such as a click) from the user for the exciting video generation control 614. In response to this input, the electronic device 100 may display a scene setting window 710 as Figure 7F shown. For the textual description of the scene setting window 710, reference may be made to the textual part of the foregoing Figure 7B shown embodiment and will not be elaborated here.

[0171] As Figure 7FAs shown, when the set b of wonderful scene types selected by the user in the scene setting window 710 is the same as the set a of wonderful scene types of the already generated wonderful video a, the electronic device 100 can output a prompt 718 and disable the determination control 716. Among them, after the determination control 716 is disabled, the determination control 716 cannot respond to the user's input to execute the corresponding wonderful video generation function. The prompt 718 can be used to prompt the user that the wonderful scene type selected by the user is the same as the wonderful scene of the already generated wonderful video a. For example, the prompt 718 can be a text prompt "You have already generated a wonderful video with the same wonderful scene. Please reselect."

[0172] As Figure 7G shown, when the set b of wonderful scene types selected by the user in the scene setting window 710 is different from the set a of wonderful scene types of the already generated wonderful video a, the electronic device 100 can enable the determination control 716.

[0173] The electronic device 100 can receive the user's input (such as a click) for the determination control 716. In response to this input, the electronic device 100 can identify multiple wonderful video segments corresponding to the set b of wonderful scenes set by the user from the original video, and fuse these multiple wonderful video segments into a wonderful video (for example, wonderful video b).

[0174] As Figure 7H shown, after the electronic device 100 generates the wonderful video b, the electronic device 100 can display a display area 721 corresponding to the wonderful video b generated from the original video on the video display interface 610 of the original video. Among them, on the display area 721 of the wonderful video b, frame images in the wonderful video b, time information 723 of the wonderful video b (for example, the time length is 8 seconds), and scene information 724 of the wonderful video b (for example, people, scenery, food, and travel) can be displayed. Optionally, a wonderful mark 722 can be displayed on the display area 721 of the wonderful video, and the wonderful mark 722 can be used to prompt the user that the content displayed in the display area 721 is the wonderful video b generated from the original video.

[0175] In some application scenarios, during a video call, the electronic device 100 can identify and crop multiple wonderful video segments with wonderful scenes in the video stream during the video call, and fuse these multiple wonderful video segments into a wonderful video. After the video call ends, the electronic device 100 can save the wonderful video. Optionally, the electronic device 100 can also share the generated wonderful video with the other party of the video call. In this way, during the video call, multiple wonderful video segments in the video stream can be fused into a wonderful video, which is convenient for the user to review the content of the video call.

[0176] Exemplarily, as Figure 8AAs shown, the electronic device 100 can display a video call answering interface 810. The video call answering interface 810 can include a reject control 811, a video-to-audio control 812, and an answer control 813.

[0177] The electronic device 100 can receive an input (such as a click) from the user for the answer control 813. In response to this input, the electronic device 100 can display Figure 8B the video call interface 820 as shown, and receive the video stream sent by the call counterpart, while simultaneously receiving the video stream captured by the camera and microphone in real time.

[0178] As Figure 8B shown, the video call interface 820 can include a picture 821 in the video stream captured by the electronic device 100 through the camera and microphone in real time, a picture 822 in the video stream sent by the call counterpart, a hang-up control 823, a video-to-audio control 824, a lens switching control 825, a highlight recording control 826, and a picture switching control 827. Among them, the hang-up control 823 can be used to trigger the electronic device 100 to hang up the video call with the other party. The video-to-audio control 824 can be used to trigger the electronic device 100 to convert the video call into a voice call. The lens switching control 825 can be used to trigger the camera of the electronic device 100 to capture video pictures in real time (for example, switch the front camera to the rear camera or the rear camera to the front camera). The highlight recording control 826 can be used to trigger the electronic device 100 to generate a highlight video based on the call video stream. The picture switching control 827 can be used to trigger the electronic device 100 to switch the display positions of the pictures 821 and 822.

[0179] The electronic device 100 can receive an input (such as a click) from the user for the highlight recording control 826. In response to this input, the electronic device 100 can identify multiple highlight video segments with highlight scenes in the video stream captured by the electronic device 100 through the camera and microphone in real time and / or the video stream sent by the call counterpart, and fuse these multiple highlight video segments into a highlight video. After the recording ends or the video call ends, the electronic device 100 can save the highlight video.

[0180] As Figure 8C shown, when the electronic device 100 starts highlight recording, the electronic device 100 can replace the highlight recording control 826 with an end recording control 828. The end recording control 828 can be used to trigger the electronic device 100 to end the recording of the highlight video.

[0181] In some application scenarios, during a video live stream, the electronic device 100 can identify and crop multiple exciting video clips from the video stream during the video live stream, and fuse these multiple exciting video clips into one exciting video. After the live stream ends, the electronic device 100 can save the exciting video. Optionally, the electronic device 100 can also synchronize the generated exciting video to the server of the live streaming application, bind it to the live streaming account, and share it to a public viewing area for other accounts following the live streaming account to watch. In this way, during the video live stream, multiple exciting video clips in the video live stream can be fused into one exciting video, facilitating users and other users following the live streaming account to review the content of the video call.

[0182] In a possible implementation, during the video live stream, the live streaming server can obtain the video stream live streamed by the electronic device 100. The live streaming server can identify multiple exciting video clips in the video stream live streamed by the electronic device 100, and fuse the multiple exciting video clips into one exciting video, and save it to the storage space associated with the video live streaming account logged in by the electronic device 100. The user can also use the electronic device 100 to share the exciting video with other users through the live streaming server. In this way, it is convenient for the user and other users following the live streaming account to review the content of the video call.

[0183] In the embodiments of the present application, the original video can be referred to as the first video, and the exciting video can be referred to as the second video. The second video may include some video clips in the first video. For example, the first video includes a first video clip, a second video clip (exciting video clip), a second video clip (invalid video clip), and a third video clip (exciting video clip). Among them, the end time of the first video clip is earlier than or equal to the start time of the second video clip, and the end time of the second video clip is earlier than or equal to the start time of the third video clip. Since the second video clip is an invalid clip, the second video includes the first video clip and the third video clip, and does not include the second video clip.

[0184] Among them, the first video further includes a fourth video clip. If the fourth video clip is an exciting video clip, the second video includes the fourth video clip; if the fourth video clip is an invalid video clip, the second video does not include the fourth video clip.

[0185] The duration of the first video is greater than the duration of the second video; or, the duration of the first video is less than the duration of the second video; or, the duration of the second video is equal to the duration of the second video.

[0186] The wonderful video clip includes a video clip in the first video where the shooting scene is a wonderful scene and does not include a transition clip. Alternatively, the wonderful video clip includes a video clip in the first video where the shooting scene is a specified wonderful scene and does not include a transition clip with noise or no sound. Among them, the wonderful scene includes one or more scenes such as people, scenery, food, Spring Festival, Christmas, architecture, beach, fireworks, plants, snow scene or travel, etc.

[0187] The following introduces a video processing method provided in an embodiment of the present application in combination with a flowchart and a functional module diagram.

[0188] Figure 9 The flowchart of a video processing method provided in an embodiment of the present application is shown.

[0189] As Figure 9 shown, the method may include the following steps:

[0190] S901. The electronic device 100 acquires an audio stream and an image stream collected in real time during video recording.

[0191] During video recording, the electronic device 100 can collect an image stream in real time through a camera, and collect an audio stream in real time through a microphone and an audio circuit. Among them, the time stamps of the audio stream and the image stream collected in real time are the same.

[0192] Among them, the interface for the video recording process can refer to the foregoing Figures 3A - 3I shown embodiment or Figures 5A - 5E shown embodiment, which will not be elaborated here.

[0193] S902. The electronic device 100 performs scene detection on the image stream to determine the scene category of each frame in the image stream.

[0194] Among them, the scene category may include people, Spring Festival, Christmas, ancient architecture, beach, fireworks, plants, snow scene, food and travel, etc.

[0195] The electronic device 100 can use a trained scene classification model to identify the scene category of each frame in the image stream. Among them, for the training of the scene classification model, a data set can be established in advance through a large number of image data with labeled scene categories. Then, the data set is input into the classification model to train the neural network classification model. Among them, the neural network used by the scene classification model is not limited. For example, it can be a convolutional neural network, a fully convolutional neural network, a deep neural network, a BP neural network, etc.

[0196] In a possible implementation, in order to improve the recognition speed of the scene categories of the frames in the image stream, before inputting the image stream into the scene classification model, the electronic device 100 can first perform interval sampling on the real-time acquired image stream (for example, take 1 frame every 3 frames) to obtain a sampled image stream, record the frame numbers of the sampled image frames in the sampled image stream in the real-time image stream, and input the sampled image stream into the neural network classification model to identify the scene categories of each sampled image frame in the sampled image stream. After identifying the scene categories of each sampled image frame in the sampled image stream, the electronic device 100 can label multiple frames in the image stream that have the same and adjacent frame numbers as the sampled image frame with the scene category corresponding to the sampled image frame. For example, the electronic device 100 can take 1 frame from every 3 frames of the image stream as the sampled image frame. Among them, the 77th frame in the image stream is the sampled image frame, and the scene category of the sampled image frame with the frame number 77 is "person". Then, the electronic device 100 can label the scene categories of the 77th frame, the 76th frame, and the 78th frame in the image stream as "person".

[0197] In a possible implementation, in order to improve the recognition speed of the scene categories of the frames in the image stream, the resolution of the image stream can also be reduced (for example, from 4K to 640*480 resolution) before inputting it into the scene classification model.

[0198] In a possible implementation, in order to improve the recognition speed of the scene categories of the frames in the image stream, the resolution of the image stream can also be reduced (for example, from 4K to 640*480 resolution) and interval sampling can be performed before inputting it into the scene classification model.

[0199] S903. The electronic device 100 performs transition detection on the image stream to determine the transition positions and transition categories of the scene transitions in the image stream.

[0200] Among them, the conversion categories of the scene transitions can include video subject conversion (for example, it can be specifically divided into the video subject changing from scenery to people, people to scenery, people to food, food to people, people to ancient buildings, ancient buildings to scenery, etc.), picture scaling, rapid camera movement, etc.

[0201] The electronic device 100 can use the trained transition recognition model to identify the transition positions and transition categories of the scene transitions in the recognition image stream. Among them, for the training of the transition recognition model, a data set can be established in advance through a large number of image streams with labeled transition positions and transition categories. Then, the data set is input into the transition recognition model to train the transition recognition model. The neural network used by the transition recognition model is not limited. For example, it can be a 3D convolutional neural network, etc.

[0202] In a possible implementation, in order to improve the recognition speed of the transition position and transition category where a scene transition occurs in the image stream. Before inputting the image stream into the transition recognition model, the electronic device 100 may first perform downsampling on the real-time captured image stream (for example, reducing the resolution from 4K to 640*480 resolution) to obtain a low-resolution image stream. Then, the low-resolution image stream is input into the transition recognition model for transition detection to identify the transition position and transition category in the low-resolution image stream. The electronic device 100 may determine the corresponding transition position and transition category in the real-time acquired image stream based on the transition position and transition category in the low-resolution image stream.

[0203] In the embodiments of the present application, the execution order of the above steps S902 and S903 is not limited. Step S902 may be executed first, step S903 may be executed first, or steps S902 and S903 may be executed in parallel.

[0204] S904. The electronic device 100 divides the image stream into multiple video clips based on the scene category of each frame in the image stream and the transition position and transition category where a scene transition occurs in the image stream, and determines the clip theme of each video clip.

[0205] S905. The electronic device 100 determines multiple exciting video clips under the exciting theme from the multiple video clips based on the clip themes of the multiple video clips, and records the positions of the multiple exciting video clips in the image stream.

[0206] Exemplarily, as Figure 10 shown, the time length of the image stream may be 0 to t14. Among them, the recognition result of the scene category in the image stream may be: the scene category of the 0 to t2 segment in the image stream is "person (person A)", the scene category of the t2 to t5 segment in the image stream is "person (person B)", the scene category of the t5 to t10 segment in the image stream is "food", and the scene category of the t10 to t14 segment in the image stream is "scenery".

[0207] The recognition result of the transition in the image stream may be: the transition category of the t1 to t3 segment in the image stream is "person to person", the transition category of the t4 to t6 segment in the image stream is "person to food", the transition category of the t7 to t8 segment in the image stream is "fast camera movement", and the transition category of the t9 to t11 segment in the image stream is "image zoom".

[0208] The division of the video segments in the video stream and the segment themes can be as follows: in the video stream, there can be video segments from t0 to t1, from t1 to t3, from t3 to t4, from t4 to t6, from t6 to t7, from t7 to t8, from t8 to t9, from t9 to t11, from t11 to t12, from t12 to t13, and from t13 to t14. Among them, the segment theme of the t0 - t1 video segment is "person", the segment theme of the t1 - t3 video segment is "invalid", the segment theme of the t3 - t4 video segment is "person", the segment theme of the t4 - t6 video segment is "invalid", the segment theme of the t6 - t7 video segment is "food", the segment theme of the t7 - t8 video segment is "invalid", the segment theme of the t8 - t9 video segment is "food", the segment theme of the t9 - t11 video segment is "invalid", the segment theme of the t11 - t12 video segment is "scenery", the segment theme of the t12 - t13 video segment is "invalid", and the segment theme of the t13 - t14 video segment is "scenery".

[0209] The electronic device 100 can remove the invalid - theme segments from multiple video segments and retain the remaining wonderful video segments. For example, as Figure 10 shown, the remaining wonderful video segments can include the video segments from t0 to t1, from t3 to t4, from t6 to t7, from t8 to t9, from t11 to t12, and from t13 to t14.

[0210] S906. At the end of recording, the electronic device 100 mixes the video stream and the audio stream into the original video.

[0211] Among them, at the end of recording, the electronic device 100 can, based on the time axis of the video stream and the time axis of the audio stream, mix the video stream and the audio stream into the original video. Among them, the electronic device 100 can receive the user's input to trigger the end of video recording, or it can be that the electronic device 100 automatically ends recording when recording for a specified duration.

[0212] S907. The electronic device 100 extracts multiple wonderful video segments from the original video based on the positions of multiple wonderful video segments in the video stream.

[0213] For example, multiple wonderful video segments may include video segments from t0 to t1, from t3 to t4, from t6 to t7, from t8 to t9, from t11 to t12, and from t13 to t14. The electronic device 100 may extract the video segment with the time line from t0 to t1 in the original video as the wonderful video segment 1, extract the video segment with the time line from t3 to t4 as the wonderful video segment 2, extract the video segment with the time line from t6 to t7 as the wonderful video segment 3, extract the video segment with the time line from t8 to t9 as the wonderful video segment 4, extract the video segment with the time line from t11 to t12 as the wonderful video segment 5, and extract the video segment with the time line from t13 to t14 as the wonderful video segment 6.

[0214] S908. The electronic device 100 combines multiple wonderful video segments into one wonderful video.

[0215] Among them, the electronic device 100 can directly splice multiple wonderful video segments together in chronological order as one wonderful video. For example, when the original video includes a first video segment, a second video segment, and a third video segment, and the wonderful video segments include the first video segment and the third video segment, the electronic device can splice the end position of the first video segment and the start position of the third video segment together to obtain the wonderful video.

[0216] In a possible implementation manner, the electronic device 100 can add video effects in the splicing area of the wonderful video segments during the splicing process for video transition. Among them, the video effects may include picture effects. Optionally, the video effects may also include audio effects. For example, when the original video includes a first video segment, a second video segment, and a third video segment, and the wonderful video segments include the first video segment and the third video segment, the electronic device can splice the end position of the first video segment and the start position of the first special effect segment together, and splice the end position of the first special effect segment and the start position of the third video segment together to obtain the second video.

[0217] Among them, the splicing area can add a time region between the end position of the previous wonderful video segment and the start position of the next wonderful video segment in two wonderful video segments. For example, as Figure 10As shown, there can be splicing area 1 between the end position of wonderful video clip 1 and the start position of wonderful video clip 2, splicing area 2 between the end position of wonderful video clip 2 and the start position of wonderful video clip 3, splicing area 3 between the end position of wonderful video clip 3 and the start position of wonderful video clip 4, splicing area 4 between the end position of wonderful video clip 5 and the start position of wonderful video clip 6, and splicing area 5 between the end position of wonderful video clip 5 and the start position of wonderful video clip 6.

[0218] In a possible implementation, the splicing area can be an area composed of the end part area (e.g., the last 500 ms part) of the previous wonderful video clip and the start part area (e.g., the first 500 ms part) of the next wonderful video clip among two wonderful video clips. For example, as Figure 11 shown, the end part area of wonderful video clip 1 and the start part area of wonderful video clip 2 can be splicing area 1, the end part area of wonderful video clip 2 and the start part area of wonderful video clip 3 can be splicing area 2, the end part area of wonderful video clip 3 and the start part area of wonderful video clip 4 can be splicing area 3, the end part area of wonderful video clip 4 and the start part area of wonderful video clip 5 can be splicing area 4, and the end part area of wonderful video clip 5 and the start part area of wonderful video clip 6 can be splicing area 5.

[0219] Among them, the video special effects of the splicing area can include flying in, flying out, and the video fusion of the previous and the next wonderful video clips, etc. For example, in the splicing area of two wonderful video clips, the video of the previous wonderful video clip can gradually fly out of the video display window from the left, and at the same time, the video of the next wonderful video clip can gradually fly into the video display window from the right.

[0220] Among them, the audio special effects of the splicing area can include pure music, songs, etc. In a possible implementation, when the splicing area can be an area composed of the end part area of the previous wonderful video clip and the start part area (e.g., the first 500 ms part) of the next wonderful video clip among two wonderful video clips, the electronic device 100 can gradually reduce the audio volume of the previous wonderful video clip and gradually increase the audio volume of the next wonderful video clip in the splicing area.

[0221] In a possible implementation, the electronic device 100 may select a video special effect for the splicing area based on the segment themes corresponding to the two wonderful video segments before and after the splicing area. For example, the segment theme corresponding to the wonderful video segment 1 before the splicing area 1 is "person", and the segment theme corresponding to the wonderful video segment 2 after the splicing area 1 is "person", so the video special effect 1 can be used in the splicing area 1. The segment theme corresponding to the wonderful video segment 2 before the splicing area 2 is "person", and the segment theme corresponding to the wonderful video segment 3 after the splicing area 2 is "food", so the video special effect 2 can be used in the splicing area 2. The segment theme corresponding to the wonderful video segment 3 before the splicing area 3 is "food", and the segment theme corresponding to the wonderful video segment 4 after the splicing area 3 is "food", so the video special effect 3 can be used in the splicing area 3. The segment theme corresponding to the wonderful video segment 4 before the splicing area 4 is "food", and the segment theme corresponding to the wonderful video segment 5 after the splicing area 4 is "scenery", so the video special effect 4 can be used in the splicing area 4. The segment theme corresponding to the wonderful video segment 5 before the splicing area 5 is "scenery", and the segment theme corresponding to the wonderful video segment 6 after the splicing area 5 is "scenery", so the video special effect 5 can be used in the splicing area 5.

[0222] In a possible implementation, after splicing multiple wonderful video segments together in chronological order to form a wonderful video, the electronic device 100 may add background music to the wonderful video. Optionally, the electronic device 100 may select the background music based on the segment themes in these multiple wonderful video segments. For example, the electronic device 100 may select the segment theme that appears for the longest time among the segment themes in these multiple wonderful video segments as the theme of the wonderful video, and based on the theme of the wonderful video, select the music corresponding to the theme of the wonderful video as the background music and add it to the wonderful video.

[0223] In a possible implementation, the electronic device 100 may score each of the multiple wonderful video segments based on their segment themes respectively. Then, splice the scored multiple wonderful video segments together in chronological order to form a wonderful video. For example, the segment theme corresponding to the wonderful video segment 1 is "person", so the music 1 can be used as the theme music for the wonderful video segment 1. The segment theme corresponding to the wonderful video segment 2 is "person", so the music 1 can be used as the theme music for the wonderful video segment 1. The segment theme corresponding to the wonderful video segment 3 is "food", so the music 2 can be used as the theme music for the wonderful video segment 1. The segment theme corresponding to the wonderful video segment 4 is "food", so the music 2 can be used as the theme music for the wonderful video segment 1. The segment theme corresponding to the wonderful video segment 5 is "scenery", so the music 3 can be used as the theme music for the wonderful video segment 1. The segment theme corresponding to the wonderful video segment 6 is "scenery", so the music 3 can be used as the theme music for the wonderful video segment 1.

[0224] S909. The electronic device 100 stores the original video and the highlight video.

[0225] After the electronic device 100 stores the original video and the highlight video, for the schematic diagram of the interface showing the stored original video and highlight video, reference may be made to the foregoing Figures 4A - 4G illustrated embodiments and will not be elaborated herein.

[0226] In some embodiments, the electronic device 100 may generate a highlight video from the already captured original video in the gallery application. At this time, the electronic device 100 may first split the original video into an image stream and an audio stream. Then, based on the image stream, perform the above steps S902 to S905, and steps S907 to S908 to generate a highlight video.

[0227] In a possible implementation, the electronic device 100 may store a third video, where the original video may include a fifth video segment and a sixth video segment. The end time of the fifth video segment is earlier than or equal to the start time of the sixth video segment, and the third video includes the fifth video segment and the sixth video segment. Both the fifth video segment and the sixth video segment include the same shooting subject. For example, both the fifth video segment and the sixth video segment include a human shooting subject, etc. In this way, the segments of the same type of shooting subject in the original video can be extracted to generate a highlight video, improving the viewing experience of the video recorded by the user.

[0228] Through a video processing method provided by an embodiment of the present application, by analyzing the scene and transition in the video recorded by the user, invalid segments (such as scene switching, picture zooming, rapid camera movement, severe picture jitter, etc.) in the recorded video can be deleted, multiple highlight video segments in the recorded video can be clipped out, and these multiple highlight video segments can be merged into one highlight video. In this way, the viewing performance of the video recorded by the user can be improved.

[0229] Figure 12 Fig. shows a functional module diagram of a video processing system provided in an embodiment of the present application.

[0230] As Figure 12 shown, the video processing system 1200 may include: a data module 1201, a perception module 1202, a fusion module 1203, and a video processing module 1204. Among them,

[0231] The data module 1201 is used to obtain the image stream and the audio stream during video recording. The data module 1201 may transfer the image stream to the perception module 1202 and transfer the image stream and the audio stream to the video processing module 1204.

[0232] The perception module 1202 can perform video understanding on the image stream. Among them, video understanding includes transition detection and scene detection. Specifically, the perception module 1202 can perform scene detection on the image stream to identify the scene category of each frame in the image stream. The perception module 1202 can perform transition detection on the image stream to identify the transition position and transition category where the scene conversion occurs in the image stream. Among them, for the specific content of the transition detection and scene detection of the image stream, reference can be made to the steps S902 and S903 in the foregoing Figure 9 illustrated embodiment, which will not be elaborated here.

[0233] The perception module 1202 can transmit the scene category of each frame, as well as the transition position and transition category where the scene conversion occurs in the image stream, to the fusion module 1203.

[0234] The fusion module 1203 can divide the image stream into multiple video clips based on the transition position where the scene conversion occurs in the image stream. The fusion module 1203 can determine the clip theme of each video clip in the multiple video clips based on the transition position and transition category where the scene conversion occurs, as well as the scene category of each frame. For specific content, reference can be made to the step S905 in the foregoing Figure 9 illustrated embodiment, which will not be elaborated here.

[0235] The fusion module 1203 can present the positions and clip themes of the multiple video clips to the video processing module 1204.

[0236] The video processing module 1204 can mix the audio stream and the image stream into the original video. The video processing module 1204 can remove the video clips with invalid themes in the original video based on the positions and clip themes of the multiple video clips, so as to extract multiple wonderful video clips. For specific content, reference can be made to the steps S906 to S907 in the foregoing Figure 9 illustrated embodiment, which will not be elaborated here.

[0237] The video processing module 1204 can fuse the multiple wonderful video clips into one wonderful video. Among them, the fusion process includes: splicing the wonderful video clips, adding special effects, adding background music, etc. For specific content, reference can be made to the step S908 in the foregoing Figure 9 illustrated embodiment, which will not be elaborated here.

[0238] The video processing module 1204 can output the original video and the wonderful video.

[0239] Figure 13 shows a schematic flowchart of a video processing method provided in another embodiment of the present application.

[0240] As Figure 13 shown, the video processing method includes:

[0241] S1301. The electronic device 100 acquires the audio stream and the image stream collected in real time during the video recording process.

[0242] For the specific content, reference can be made to step S901 in the foregoing Figure 9 illustrated embodiment, which will not be elaborated herein.

[0243] S1302. The electronic device 100 performs scene detection on the image stream to determine the scene category of each frame in the image stream.

[0244] For the specific content, reference can be made to step S902 in the foregoing Figure 9 illustrated embodiment, which will not be elaborated herein.

[0245] S1303. The electronic device 100 performs transition detection on the image stream to determine the transition position and transition category where a scene transition occurs in the image stream.

[0246] For the specific content, reference can be made to step S903 in the foregoing Figure 9 illustrated embodiment, which will not be elaborated herein.

[0247] In the embodiments of the present application, the execution order of the above steps S1302 and S1303 is not limited. Step S1302 can be executed first, step S1303 can be executed first, or steps S1302 and S1303 can be executed in parallel.

[0248] S1304. The electronic device 100 performs voice activation detection on the audio stream, identifies the start and end time points of the voice signal in the audio stream, and divides the audio stream into multiple audio segments.

[0249] Among them, the electronic device 100 can perform sliding window slicing on the voice signal to detect the audio features of the voice signal within the sliding window. The electronic device 100 can identify the start and end time points of the voice signal in the image stream based on the audio features in the image stream. The electronic device 100 can divide the audio stream into multiple audio segments based on the start and end time points of the voice signal in the audio stream. Among them, the audio features can include features such as spectral slope, correlation coefficient, log likelihood ratio, cepstral, and weighted cepstral.

[0250] S1305. The electronic device 100 classifies audio events for multiple audio segments in the audio stream.

[0251] Among them, the electronic device 100 can use the trained audio event classification model to identify the audio event category of the audio segment. Among them, for the training of the audio event classification model, a data set can be established in advance through a large number of data with labeled audio signals and audio event categories. Then, the data set is input into the audio event classification model to train the audio event classification model. Among them, the neural network used by the transition recognition model is not limited. For example, it can be a recurrent neural network (RNN) classification model, a long short-term memory (LSTM) artificial neural network classification model, and so on.

[0252] Among them, the audio event categories can include human voices, laughter, music, noise, and so on. Optionally, the noise can be further subdivided into vehicle driving sounds, animal calls, bird calls, dog barks, wind sounds, and so on.

[0253] S1306. The electronic device 100 determines multiple audio event image segments corresponding to multiple audio segments in the image stream based on the start and end time points of the multiple audio segments, and the audio event category corresponding to each audio event image segment.

[0254] S1307. The electronic device 100 divides the image stream into multiple picture segments based on the scene category of each frame in the image stream, the transition position and transition category where the scene transition occurs in the image stream, and the position and audio event category of the multiple audio event image segments, and determines the segment theme of each picture segment.

[0255] Specifically, the electronic device 100 can divide the image stream into multiple picture segments based on the position of the audio event image segment and the transition position where the scene transition occurs in the image stream. Among them, the union of the position of the audio event image segment and the transition position can be taken to divide the image stream into multiple picture segments.

[0256] Then, the electronic device 100 can determine the theme of each picture segment based on the scene category, transition category, and audio event category corresponding to each picture segment.

[0257] Exemplarily, as Figure 14 shown, the time length of the original video can be 0 to t20. Among them, the recognition results of the scene categories in the image stream can be: the scene category of the 0 to t3 segment in the image stream is "person (person A)", the scene category of the t3 to t7 segment in the image stream is "person (person B)", the scene category of the t7 to t13 segment in the image stream is "food", the scene category of the t13 to t16 segment in the image stream is "no scene", and the scene category of the t16 to t20 segment in the image stream is "scenery".

[0258] The recognition results of the transitions in the image stream can be as follows: the transition category of the t2 - t4 segment in the image stream is "person to person", the transition category of the t6 - t8 segment in the image stream is "person to food", the transition category of the t10 - t11 segment in the image stream is "rapid camera movement", the transition category of the t12 - t14 segment in the image stream is "food to no scene", and the transition category of the t17 - t19 segment in the image stream is "image zoom".

[0259] The position of the audio event image segments and the recognition results of the audio event categories in the image stream can be as follows: the audio event category of the t0 - t1 segment in the image stream is "speaking", the audio event category of the t1 - t5 segment in the image stream is "laughter", the audio event category of the t5 - t9 segment in the image stream is "music", the audio event category of the t9 - t11 segment in the image stream is "no sound", the audio event category of the t11 - t18 segment in the image stream is "noise", and the audio event category of the t18 - t20 segment in the image stream is "no sound".

[0260] The division of the picture segments and the segment themes in the image stream can be as follows: in the image stream, there can be divided t0 - t1 picture segments, t1 - t2 picture segments, t2 - t4 picture segments, t4 - t5 picture segments, t5 - t6 picture segments, t6 - t8 picture segments, t8 - t9 picture segments, t9 - t10 picture segments, t10 - t11 picture segments, t11 - t12 picture segments, t12 - t14 picture segments, t14 - t16 picture segments, t16 - t17 picture segments, t17 - t18 picture segments, t18 - t19 picture segments, and t19 - t20 picture segments. Among them, the segment theme of the t0 - t1 picture segment is "person", the segment theme of the t1 - t2 picture segment is "person", the segment theme of the t2 - t4 picture segment is "person to person + laughter", the segment theme of the t4 - t5 picture segment is "person", the segment theme of the t5 - t6 picture segment is "person", the segment theme of the t6 - t8 picture segment is "person to food + music", the segment theme of the t8 - t9 picture segment is "food", the segment theme of the t9 - t10 picture segment is "food", the segment theme of the t10 - t11 picture segment is "rapid camera movement", the segment theme of the t11 - t12 picture segment is "food", the segment theme of the t12 - t14 picture segment is "food to no scene + noise", the segment theme of the t14 - t16 picture segment is "noise", the segment theme of the t16 - t17 picture segment is "scenery", the segment theme of the t17 - t18 picture segment is "image zoom + noise", the segment theme of the t18 - t19 picture segment is "image zoom", and the segment theme of the t19 - t20 picture segment is "scenery".

[0261] S1308. The electronic device 100 determines multiple wonderful video clips under a wonderful theme from multiple video clips based on the clip themes of the multiple video clips, and records the positions of the multiple wonderful video clips in the video stream.

[0262] Among them, the electronic device 100 may determine the video clips under a preset wonderful theme among the multiple video clips as wonderful video clips.

[0263] The electronic device 100 may determine the video clips with only transitions but no effective sound (such as speech, laughter, music, etc.) and the video clips with no effective sound, no transitions, and no scene categories as invalid clips. The video clips among the multiple video clips other than the invalid clips are determined as wonderful video clips.

[0264] For example, as Figure 14 shown, the electronic device 100 may determine the video clips from t0 to t1, from t1 to t2, from t2 to t4, from t4 to t5, from t5 to t6, from t6 to t8, from t8 to t9, and from t9 to t10 as wonderful video clips, determine the video clip from t10 to t11 as an invalid clip, determine the video clip from t11 to t12 as a wonderful video clip, determine the video clips from t12 to t14 and from t14 to t16 as invalid clips, determine the video clip from t16 to t17 as a wonderful video clip, determine the video clips from t17 to t18 and from t18 to t19 as invalid clips, and determine the video clip from t19 to t20 as a wonderful video clip.

[0265] S1309. At the end of recording, the electronic device 100 mixes the video stream and the audio stream into an original video.

[0266] S1310. The electronic device 100 extracts multiple wonderful video clips from the original video based on the positions of the multiple wonderful video clips in the video stream.

[0267] For example, as Figure 14As shown, since the video clips from t0 to t1, from t1 to t2, from t2 to t4, from t4 to t5, from t5 to t6, from t6 to t8, from t8 to t9, and from t9 to t10 are continuous and all wonderful video clips, the electronic device 100 can determine the video clip from t0 to t10 in the original video as the wonderful video clip 1. Since the video clip from t11 to t12 is a wonderful video clip, the electronic device 100 determines the video clip from t11 to t12 in the original video as the wonderful video clip 2. Since the video clip from t16 to t17 is a wonderful video clip, the electronic device 100 determines the video clip from t16 to t17 in the original video as the wonderful video clip 3. Since the video clip from t19 to t20 is a wonderful video clip, the electronic device 100 determines the video clip from t19 to t20 in the original video as the wonderful video clip 4.

[0268] S1311. The electronic device 100 fuses multiple wonderful video clips into one wonderful video.

[0269] S1312. The electronic device 100 saves the original video and the wonderful video.

[0270] Among them, the electronic device 100 can directly splice multiple wonderful video clips together in chronological order as one wonderful video.

[0271] In a possible implementation, when splicing, the electronic device 100 can add video effects to the splicing area of the wonderful video clips for video transition. Among them, the video effects can include picture effects. Optionally, the video effects can also include audio effects.

[0272] Among them, the splicing area can be a time area added between the end position of the previous wonderful video clip and the start position of the next wonderful video clip among two wonderful video clips. For example, as Figure 14 shown, there can be a splicing area 1 between the end position of the wonderful video clip 1 and the start position of the wonderful video clip 2, a splicing area 2 between the end position of the wonderful video clip 2 and the start position of the wonderful video clip 3, and a splicing area 3 between the end position of the wonderful video clip 3 and the start position of the wonderful video clip 4.

[0273] In a possible implementation, the splicing area can be an area composed of the end part area (for example, the last 500ms part) of the previous wonderful video clip and the start part area (for example, the first 500ms part) of the next wonderful video clip among two wonderful video clips. For details, reference can be made to the foregoing Figure 11 shown embodiments, which will not be elaborated here.

[0274] The screen effects of the splicing area may include flying in, flying out, fusion of the screens of the two wonderful video clips, etc. For example, in the splicing area of ​​two wonderful video clips, the screen of the previous wonderful video clip may gradually fly out of the video display window from the left, and the screen of the next wonderful video clip may gradually fly into the video display window from the right.

[0275] The audio effects in the splicing area may include pure music, songs, etc. In a possible implementation, when the splicing area may be an area consisting of the end area of ​​the first wonderful video segment and the beginning area (e.g., the beginning 500ms part) of the second wonderful video segment in two wonderful video segments, the electronic device 100 may gradually reduce the audio volume of the first wonderful video segment in the splicing area, and gradually increase the audio volume of the second wonderful video segment from a small volume.

[0276] In a possible implementation, the electronic device 100 may select a video effect to be used in the splicing area based on the clip themes corresponding to the two wonderful video clips before and after the splicing area.

[0277] In a possible implementation, the electronic device 100 can add background music to the wonderful video after splicing the multiple wonderful video clips together in chronological order as a wonderful video. Optionally, the electronic device 100 can select background music based on the clip themes in the multiple wonderful video clips. For example, the electronic device 100 can select the clip theme with the longest appearance time among the clip themes in the multiple wonderful video clips as the theme of the wonderful video, and based on the theme of the wonderful video, select the music corresponding to the theme of the wonderful video as the background music and add it to the wonderful video.

[0278] In a possible implementation, the electronic device 100 can respectively match music to the multiple wonderful video clips based on the clip themes of the multiple wonderful video clips, and then splice the multiple wonderful video clips after the music match together in chronological order as a wonderful video.

[0279] Through a video processing method provided by an embodiment of the present application, invalid segments in a recorded video can be deleted by analyzing scenes, transitions, and audio events in a video recorded by a user, multiple wonderful video segments in the recorded video can be edited, and the multiple wonderful video segments can be merged into one wonderful video. In this way, the viewing quality of the video recorded by the user can be improved.

[0280] Figure 15 A functional module diagram of a video processing system provided in an embodiment of the present application is shown.

[0281] like Figure 15As shown, the video processing system 1500 may include: a data module 1501, a perception module 1502, a fusion module 1503, and a video processing module 1504. Among them,

[0282] The data module 1501 is used to obtain the image stream and audio stream during video recording. The data module 1501 can transfer the image stream and audio stream to the perception module 1502 and transfer the image stream and audio stream to the video processing module 1504.

[0283] The perception module 1502 can perform video understanding on the image stream. Among them, video understanding includes transition detection and scene detection. Specifically, the perception module 1502 can perform scene detection on the image stream to identify the scene category of each frame in the image stream. The perception module 1502 can perform transition detection on the image stream to identify the transition position and transition category where the scene conversion occurs in the image stream. Among them, for the specific content of the transition detection and scene detection of the image stream, reference can be made to steps S1302 and S1303 in the foregoing Figure 13 illustrated embodiment, which will not be elaborated here.

[0284] The perception module 1502 can also perform audio understanding on the audio stream. Among them, audio understanding includes voice activation detection and audio event classification. Specifically, the perception module 1502 can perform voice activation detection on the audio stream to identify the start and end time points of the voice signal in the audio stream and divide the audio stream into multiple audio segments. The perception module 1502 can perform audio event classification on the multiple audio segments in the audio stream. Among them, for the specific content of the voice activation detection and audio event classification of the audio stream, reference can be made to steps S1304 and S1305 in the foregoing Figure 13 illustrated embodiment, which will not be elaborated here.

[0285] The perception module 1502 can transfer the scene category of each frame, the transition position and transition category where the scene conversion occurs in the image stream, the position of the audio segment, and the audio event category to the fusion module 1503.

[0286] The fusion module 1503 can divide the image stream into multiple picture segments based on the position of the audio event image segment corresponding to the audio segment and the transition position where the scene conversion occurs in the image stream. The fusion module 1503 can determine the theme of each picture segment based on the scene category, transition category, and audio event category corresponding to each picture segment. For the specific content, reference can be made to step S1307 in the foregoing Figure 13 illustrated embodiment, which will not be elaborated here.

[0287] The fusion module 1503 can present the positions and segment themes of the multiple picture segments to the video processing module 1504.

[0288] The video processing module 1504 can mix the audio stream and the image stream into the original video. The video processing module 1504 can remove the picture segments with invalid themes in the original video based on the positions and segment themes of multiple picture segments, so as to extract multiple wonderful video segments. For specific content, reference can be made to steps S1308 to S1310 in the foregoing Figure 13 illustrated embodiment, which will not be elaborated herein.

[0289] The video processing module 1504 can fuse multiple wonderful video segments into a wonderful video. The fusion process includes: splicing of wonderful video segments, adding special effects, adding background music, etc. For specific content, reference can be made to Figure 13 step S1311 in the foregoing illustrated embodiment, which will not be elaborated herein.

[0290] The video processing module 1504 can output the original video and the wonderful video.

[0291] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.

Claims

1. A video processing method, characterized in that, it includes: The electronic device determines the scene category of each frame in the image stream of the first video, the transition position and transition category where a scene transition occurs in the image stream of the first video, and divides the audio stream of the first video into multiple audio segments; the first video includes a first video segment, a second video segment, and a third video segment, the end time of the first video segment is earlier than or equal to the start time of the second video segment, the end time of the second video segment is earlier than or equal to the start time of the third video segment, and the duration of the first video segment is greater than the duration of the third video segment; The electronic device determines multiple audio event image segments corresponding to the multiple audio segments in the image stream of the first video, and the audio event category corresponding to each audio event image segment; The electronic device divides the image stream of the first video into multiple picture segments based on the scene category of each frame in the image stream of the first video, the transition position and transition category where a scene transition occurs in the image stream of the first video, and the audio event category of the multiple audio event image segments, and determines the segment theme of each picture segment in the multiple picture segments; The electronic device determines multiple wonderful picture segments under the wonderful theme from the multiple picture segments based on the segment themes of the multiple picture segments, and records the positions of the multiple wonderful picture segments in the image stream of the first video, where the wonderful picture segments are the picture segments other than the invalid segments in the multiple picture segments, and the invalid segments are the picture segments with only transitions but no effective sound and the picture segments with no effective sound, no transitions, and no scene categories; The electronic device extracts the first video segment and the third video segment from the first video based on the positions of the multiple wonderful picture segments in the image stream of the first video; The electronic device generates the second video based on the first video segment and the third video segment; During the generation of the second video, the electronic device determines the background music of the second video based on the segment theme of the first video segment; the second video includes the first video segment, the third video segment, and the background music, and does not include the second video segment.

2. The method according to claim 1, characterized in that, before the electronic device determines the scene category of each frame in the image stream of the first video, the transition position and transition category where a scene transition occurs in the image stream of the first video, and divides the audio stream of the first video into multiple audio segments, the method includes: The electronic device displays a video album interface, the video album interface includes one or more video files, and the one or more video files include the first video; The video album interface displays a thumbnail of the first video; The electronic device detects a first input to the first video; In response to the first input, the electronic device displays a first interface, which includes a display area of the first video and a first control for generating the second video from the first video. The electronic device detects a second input to the first control. The electronic device determines the scene category of each frame in the image stream of the first video, the transition position and transition category where a scene transition occurs in the image stream of the first video, and divides the audio stream of the first video into multiple audio segments, specifically including: In response to the second input, the electronic device determines the scene category of each frame in the image stream of the first video, the transition position and transition category where a scene transition occurs in the image stream of the first video, and divides the audio stream of the first video into multiple audio segments.

3. The method according to claim 1 or 2, wherein, the method further includes: The electronic device displays a second interface, which includes a display area of the first video and a first window for displaying the generation progress of the second video. After the second video is generated, the electronic device displays a third interface, which includes a display area of the second video.

4. The method according to claim 3, wherein, the third interface further includes a sharing control and an editing control; The sharing control is used to share the second video, and the editing control is used to edit the second video.

5. The method according to claim 3 or 4, wherein, the third interface further includes a display area of the first video; In response to a third input to the display area of the second video by the user, the display area of the first video is reduced, and the display area of the second video is enlarged.

6. The method according to claim 1, wherein, the duration of the first video is greater than the duration of the second video; or, the duration of the first video is equal to the duration of the second video.

7. The method according to claim 3, wherein, before the electronic device displays the third interface, the method further includes: The electronic device stitches together the first video segment and the third video segment in the first video to obtain the second video.

8. The method according to claim 7, wherein, the electronic device stitches together the first video segment and the third video segment in the first video to obtain the second video, specifically including: The electronic device stitches together the end position of the first video segment and the start position of the third video segment to obtain the second video; or, The electronic device stitches together the end position of the first video segment and the start position of a first special effect segment, and stitches together the end position of the first special effect segment and the start position of the third video segment to obtain the second video.

9. The method according to claim 1, wherein, The first video segment and the third video segment are wonderful video segments, and the second video segment is an invalid video segment.

10. The method according to claim 1, wherein, the first video further includes a fourth video segment; if the fourth video segment is a wonderful video segment, the second video includes the fourth video segment; if the fourth video segment is an invalid video segment, the second video does not include the fourth video segment.

11. The method according to claim 9, wherein, the wonderful video segment includes a video segment in the first video where the shooting scene is a wonderful scene and does not include a transition segment.

12. The method according to claim 9, wherein, the wonderful video segment includes a video segment in the first video where the shooting scene is a wonderful scene and does not include a transition segment with noise or no sound.

13. The method according to claim 11 or 12, wherein, the wonderful scene includes one or more of a person, a landscape, food, the Spring Festival, Christmas, a building, a beach, fireworks, plants, a snow scene, or a trip.

14. The method according to claim 2, wherein, before the electronic device displays the video album interface, the method further includes: the electronic device displays a shooting interface, the shooting interface includes a preview frame and a recording start control, and the preview frame displays the image captured in real time by the camera of the electronic device; the electronic device detects a fourth input to the recording start control; in response to the fourth input, the electronic device displays a recording interface and starts recording the first video; the electronic device displays a recording interface, the recording interface includes a recording end control and the video image of the first video recorded in real time by the electronic device; the electronic device detects a fifth input to the recording end control; in response to the fifth input, the electronic device ends the recording of the first video; the electronic device saves the first video.

15. The method according to claim 14, wherein, the recording interface further includes a capture control, and when the electronic device displays the recording interface, the method further includes: the electronic device receives a sixth input from the user for the capture control; in response to the sixth input, the electronic device saves the first video image captured by the camera of the electronic device when receiving the sixth input as a first picture.

16. The method according to claim 1, wherein, after the electronic device generates the second video, the method further includes: the electronic device saves the second video.

17. An electronic device, wherein, including a camera, a display screen, one or more processors, and one or more memories; wherein, the one or more memories are coupled to the one or more processors, and the one or more memories are used for storing computer program code, the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the electronic device is caused to execute the method according to any one of claims 1-16.

18. A chip system, the chip system is applied to an electronic device, the chip system includes one or more processors, and the processors are used for calling computer instructions to cause the electronic device to execute the method according to any one of claims 1-16.

19. A computer-readable storage medium, including instructions, characterized in that, when the instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Video transition processing method and device

    CN105245810A

  • Method for processing video file and electronic equipment

    CN111061912A

  • Video segmentation method, device, equipment, system and storage medium

    CN113766314A

  • Video collection generation method and display device

    CN113973216A