Video slicing method and electronic equipment
By performing image analysis and segmentation processing on video clips, the problem that highlighting actions cannot be retained in the prior art is solved, the complete retention of highlighting actions and the consistency of videos is achieved, and the presentation effect of videos is improved.
Patent Information
- Application Number
- CN202410046018.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-01-10
AI Technical Summary
When existing electronic devices use the one-click film-forming function to generate film-forming videos, they may not be able to effectively retain the highlights in the material videos, resulting in poor presentation of film-forming videos.
By performing image analysis on image frames in video clips, multiple consecutive image frames whose image scores are greater than the threshold and belong to the same pointer are divided into the same video clip based on the preset frame skipping strategy and image score threshold, and these clips are spliced into pieces of video to ensure the retention of highlight motion.
It realizes that all highlight actions are retained when generating a film video, ensuring that the content of the film video is coherent and smooth, and improving the presentation effect of the film video.
Smart Images

Figure CN120343185A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of video data processing, and particularly to a method for generating a video clip and an electronic device. Background Art
[0002] Currently, some electronic devices can provide a one-click video clip generation function. Specifically, in response to a user's operation of using the one-click video clip generation function after selecting a source video, the electronic device can automatically analyze and extract highlight segments from the source video through an algorithm, and combine the highlight segments into a video clip. Currently, after processing a source video including highlight actions using the one-click video clip generation function, the video clip may not include highlight actions or may include fewer highlight actions, resulting in a poor final presentation effect of the video clip. Summary of the Invention
[0003] Embodiments of the present application provide a method for generating a video clip and an electronic device, which are used to retain as many highlight actions in the original video segments as possible when generating a video clip using the one-click video clip generation function, and improve the final presentation effect of the video clip.
[0004] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0005] In a first aspect, a method for generating a video clip is provided, which is applied to an electronic device. The electronic device includes at least one first video segment, and one or more highlight actions are included in the first video segment. Specifically, the method may include:
[0006] First, the electronic device acquires the first video segment, and performs image analysis on the first image frames in the first video segment according to a preset frame skipping strategy to obtain image scores. Then, the electronic device divides multiple consecutive first image frames in the first video segment whose image scores are greater than the score threshold and belong to the same shot into the same video segment to obtain at least one second video segment. Finally, the electronic device obtains a first video clip based on at least one second video segment among all the first video segments. Among them, the image scores of the first image frames including highlight actions are greater than the score threshold.
[0007] As can be seen from the above, the image scores of the first image frames including highlight actions are greater than the score threshold. Therefore, all the first image frames including highlight actions in the first video segment will be retained in the second video segment, so that all the highlight actions will be retained in the first video clip. In addition, each of the first image frames retained in the second video segment belongs to the same shot. In this way, the content of each second video segment is more coherent and smooth. Based on this, the finally obtained first video clip can include the highlight actions in all the first video segments and the content is coherent and smooth.
[0008] In a possible implementation of the first aspect, the above-mentioned process of dividing multiple consecutive first image frames in the first video segment with image scores greater than the score threshold and belonging to the same scene into the same video segment to obtain at least one second video segment may include:
[0009] First, the electronic device filters out multiple second image frames with image scores greater than the score threshold from the first image frames in the first video segment. Then, the electronic device divides consecutive second image frames among all the filtered second image frames into a third video segment to obtain at least one third video segment. Next, the electronic device divides the second image frames belonging to the same scene in the third video segment into a second video segment to obtain at least one second video segment.
[0010] In another possible implementation of the first aspect, the above-mentioned process of obtaining the first finished video based on at least one second video segment among all the first video segments may include:
[0011] First, the electronic device divides the second video segment into multiple fourth video segments with durations less than the first duration at the first splitting point based on the duration of the second video segment being greater than the first duration. Then, the electronic device filters out fifth video segments with durations greater than or equal to the third duration from all the fourth video segments. Next, the electronic device splices all the fifth video segments to obtain the first finished video.
[0012] Among them, the first splitting point includes at least one of the target moment, the first candidate moment, and the second candidate moment. The duration between the target moment and the start moment of the second video segment is an integer multiple of the first duration. The candidate moment can be the starting moment of the second duration before the third image frame, or the ending moment of the second duration after the third image frame. The third image frame is a second image frame including a highlight action. Thus, it can be seen that if an image frame with a highlight action is included in a fifth video segment, the duration of this image frame is at least the second duration from the start moment of this fifth video segment, or the duration of this image frame is at least the second duration from the ending moment of the fifth video segment.
[0013] Among them, the third duration is twice the second duration, and the second duration is less than the first duration. For example, the second duration is 1 second, the third duration is 2 seconds, and the first duration is 8 seconds.
[0014] In another possible implementation of the first aspect, the fourth video segment does not include a highlight action. At this time, the first splitting point is the target moment.
[0015] In another possible implementation of the first aspect, the fourth video segment includes a highlight action. At this time, the first splitting point includes at least one of the target moment, the first candidate moment, and the second candidate moment.
[0016] In another possible implementation of the first aspect, obtaining the image scores by performing image analysis on the first image frames in the first video segment includes:
[0017] First, the electronic device obtains first analysis information for each first image frame in the first video segment. The first analysis information may include the brightness, sharpness, aesthetic score, and first action information of the corresponding image frame. Among them, the first action information includes an inaction identifier or a target action type. The inaction identifier is used to indicate that the first image frame does not include a highlight action, and the target action type is the action type of the highlight action included in the first image frame.
[0018] Then, when the first action information indicates that the first image frame does not include a highlight action, based on the sharpness of the first image frame being within a preset sharpness range and the brightness of the first image frame being greater than the brightness threshold, the electronic device uses the aesthetic score of the first image frame as the image score of the first image frame. Otherwise, the electronic device uses a third value as the image score of the first image frame.
[0019] Next, when the first action information indicates that the first image frame includes a highlight action, the electronic device determines that the image score of the first image frame is a second value, and the image scores of all first image frames within a third time period before the first image frame are greater than or equal to a score threshold, and the image scores of all first image frames within a third time period after the first image frame are greater than or equal to the score threshold. Among them, the second value (such as 1) is greater than the score threshold (such as 0.25), and the score threshold is greater than the third value (such as 0).
[0020] As can be seen from the above, in the process of obtaining the image score of an image frame, the image frame including a highlight action is forced to be scored highly, and the image frames within a certain range of this image frame are also forced to be scored highly. Based on this, it is ensured that the image frames containing highlight actions will be retained in the first compiled video.
[0021] In another possible implementation of the first aspect, obtaining the first analysis information for each first image frame in the first video segment includes:
[0022] First, the electronic device obtains second analysis information for the first image frame. The second analysis information includes brightness, sharpness, aesthetic score, and second action information. Among them, the second action information includes an inaction identifier or at least one predicted action information, and each predicted action information includes a corresponding predicted action type and the confidence level of this predicted action type.
[0023] Then, the electronic device determines the predicted action type with the highest confidence level in each predicted action information as the target action type of the predicted action information, and obtains the first analysis information.
[0024] In another possible implementation of the first aspect, the above-mentioned first analysis information further includes shot information. The shot information includes a no-shot identifier or a shot identifier. Among them, the no-shot identifier is used to indicate that there is no shot point between the corresponding image frame and the adjacent image frame, and the shot identifier is used to indicate that there is a shot point between the corresponding image frame and the adjacent image frame;
[0025] The above-mentioned dividing the multiple second image frames belonging to the same shot in the third video segment into a second video segment includes: the electronic device divides the third video segment based on the shot information indicating that there is a shot point between the second image frame and the adjacent image frame, and uses the shot point as the first splitting point to obtain the second video segment.
[0026] In another possible implementation of the first aspect, the above-mentioned obtaining the second analysis information of the first image frame may include:
[0027] First, the electronic device performs image recognition on the first image frame to obtain the image information of the first image frame. The image recognition may include at least one of clarity recognition, brightness recognition, or aesthetic scoring, as well as action recognition and shot recognition. The image information may include at least one of clarity, brightness, or aesthetic score, as well as second action information. At the same time, the type of image recognition corresponds one-to-one with the image information. For example, action recognition corresponds to second action information and shot information.
[0028] Then, when the image recognition includes action recognition, clarity recognition, brightness recognition, shot recognition, and aesthetic scoring, the image information of the first image frame is the second analysis information of the first image frame. When the image recognition does not include clarity recognition, brightness recognition, or aesthetic scoring, the second analysis information of the Nth first image frame inherits part of the image information in the second analysis information of the (N - 1)th first image frame. N is a positive integer;
[0029] Among them, the partial image information is obtained by recognizing the first image frame based on some recognition types among clarity recognition, brightness recognition, and aesthetic scoring. The partial image information includes at least one of clarity, brightness, and aesthetic score.
[0030] In another possible implementation of the first aspect, the above-mentioned preset frame skipping strategy may indicate: starting from the first image frame, extract one image frame from the first video segment as the first image frame every L frames; or, all the image frames in the first video segment are used as the first image frames. The preset frame skipping strategy may also indicate: the target recognition type of the first image frame. The target recognition type includes: one or more of clarity recognition, brightness recognition, aesthetic scoring, face recognition, scene recognition, and human body recognition, as well as action recognition and shot recognition.
[0031] In another possible implementation of the first aspect, the above-mentioned splicing of all the fifth video segments to obtain the first finished video includes: First, the electronic device obtains a theme template corresponding to the content in each fifth video segment. Then, the electronic device splices the respective fifth video segments and applies the theme template to obtain the first finished video.
[0032] In another possible implementation of the first aspect, the above-mentioned splicing of the respective fifth video segments and applying the theme template to obtain the first finished video may include: First, the electronic device splices the respective fifth video segments to obtain a second finished video. Then, the electronic device performs video speed change processing (such as slow processing or slow-motion playback) on the video segments including multiple highlight actions in the second finished video to obtain a third finished video. Next, the electronic device applies the theme template to the third finished video to obtain the first finished video.
[0033] In another possible implementation of the first aspect, the electronic device may further include a target picture. At this time, the above-mentioned splicing of the respective fifth video segments may include: splicing the respective fifth video segments and the target picture.
[0034] In a second aspect, an electronic device is provided, which includes a memory, a display screen, and one or more processors. The memory and the display screen are coupled to the processor. Computer program code is stored in the memory, and the computer program code includes instructions. When the instructions are executed by the processor, the electronic device is caused to execute the method described in any one of the above-mentioned first aspects.
[0035] In a third aspect, a computer-readable storage medium is provided, in which instructions are stored. When it runs on an electronic device, the electronic device can be caused to execute the method described in any one of the above-mentioned first aspects.
[0036] In a fourth aspect, a computer program product containing instructions is provided. When it runs on an electronic device, the electronic device can be caused to execute the method described in any one of the above-mentioned first aspects.
[0037] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor. The processor is used to call a computer program in the memory to execute the method as in the first aspect.
[0038] It can be understood that the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect can refer to the beneficial effects in the first aspect and any of its possible design manners, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 One of the interface schematic diagrams provided by the embodiments of this application;
[0040] Figure 2 Another interface schematic diagram provided by the embodiments of this application;
[0041] Figure 3 The third interface schematic diagram provided by the embodiments of this application;
[0042] Figure 4 The hardware structure schematic diagram of an electronic device provided by the embodiments of this application;
[0043] Figure 5 The software architecture schematic diagram of an electronic device provided by the embodiments of this application;
[0044] Figure 6 One of the process schematic diagrams of a video production method provided by the embodiments of this application;
[0045] Figure 7 Another process schematic diagram of a video production method provided by the embodiments of this application;
[0046] Figure 8 The schematic diagram of the scoring strategy provided by the embodiments of this application;
[0047] Figure 9 The third process schematic diagram of a video production method provided by the embodiments of this application;
[0048] Figure 10 The schematic diagram of screening corresponding video segments provided by the embodiments of this application;
[0049] Figure 11 The fourth process schematic diagram of a video production method provided by the embodiments of this application;
[0050] Figure 12 The fifth process schematic diagram of a video production method provided by the embodiments of this application;
[0051] Figure 13 The fourth interface schematic diagram provided by the embodiments of this application;
[0052] Figure 14 The sixth process schematic diagram of a video production method provided by the embodiments of this application;
[0053] Figure 15 The seventh process schematic diagram of a video production method provided by the embodiments of this application;
[0054] Figure 16 The eighth process schematic diagram of a video production method provided by the embodiments of this application;
[0055] Figure 17 It is the ninth flowchart diagram of a video production method provided by an embodiment of the present application;
[0056] Figure 18 It is a schematic structural diagram of a chip system provided by an embodiment of the present application. Detailed implementation manners
[0057] The terms "first", "second", etc. involved in the embodiments of the present application are only used for the purpose of distinguishing the same type of features, and cannot be understood as indicating relative importance, quantity, order, etc.
[0058] The term "exemplary" or "for example" etc. involved in the embodiments of the present application is used to represent an example, illustration or explanation. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using the words "exemplary" or "for example" etc. aims to present relevant concepts in a specific manner.
[0059] The terms "coupled" and "connected" involved in the embodiments of the present application should be understood in a broad sense. For example, it can refer to a direct physical connection, or an indirect connection realized through electronic devices, such as a connection realized through resistors, inductors, capacitors or other electronic devices.
[0060] First, some nouns or terms involved in the embodiments of the present application are explained.
[0061] A high-light action refers to actions such as a smiling face of a person, splashing water, throwing an object, running, table tennis smashing, badminton smashing, and a smiling glance back recorded in materials (such as material videos, material pictures).
[0062] A high-light segment, also known as a wonderful segment, refers to an image extracted from materials that records wonderful moments such as high-light actions, moments of winning the championship, an airplane landing, sunrise or sunset. Among them, high-light actions can include: actions such as a person jumping, splashing water, throwing an object, running, table tennis smashing, badminton smashing, and a smiling glance back.
[0063] The one-key video production function means that after selecting at least one material, the high-light segments in the materials are automatically analyzed through an algorithm, and all the high-light segments are spliced into a finished video.
[0064] An electronic device can provide a function of generating a ready-made video with one click. For example, an image processing application (such as a photo gallery) is installed on the electronic device. The image processing application can support processing operations for editing materials such as pictures and videos. The electronic device can obtain the image information of each frame of the image in the material through the image processing application, and obtain at least one highlight segment according to the image information of each frame of the image. Finally, the electronic device stitches all the highlight segments obtained from multiple materials into a ready-made video and displays it on the display interface.
[0065] Currently, even if the material includes highlight actions, in the ready-made video generated by the electronic device through the one-click video generation function, the highlight actions are not included, or there are few highlight actions, resulting in the ready-made video not meeting the user's needs, that is, the final presentation effect of the ready-made video is poor.
[0066] Therefore, an embodiment of the present application provides a method for generating a ready-made video. When the electronic device responds to processing at least one first video segment including highlight actions using the one-click video generation function, it divides multiple consecutive image frames in each first video segment with an image score greater than a score threshold and belonging to the same shot into a second video segment, and the duration of the second video segment is less than the first duration and greater than the second duration. Then, the second video segments corresponding to all the first video segments are stitched together to form a first ready-made video.
[0067] Among them, the image score of the first image frame containing the highlight action is greater than the score threshold, so all the first image frames containing the highlight action will be retained in the second video segment, so that all the highlight actions will be retained in the first ready-made video. In addition, each of the first image frames retained in the second video segment belongs to the same shot, so the content of each second video segment is more coherent and smooth.
[0068] The electronic device involved in the embodiments of the present application may also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The electronic device may be a mobile phone, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city or wireless terminal in smart home, etc. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the electronic device.
[0069] Figure 1 One of the interface schematic diagrams provided by the embodiments of the present application is shown. Figure 2 Another interface schematic diagram provided by the embodiments of the present application is shown. Figure 3 A third interface schematic diagram provided by the embodiments of the present application is shown.
[0070] Taking the electronic device as a mobile phone and the mobile phone storing multiple pre-shot videos and multiple pictures as an example, the application scenario of the video compilation method provided by the embodiments of the present application will be exemplarily described in combination with Figures 1 - 3 the following.
[0071] In one embodiment, as shown at A in Figure 1 , the application icon 111 of the gallery is displayed on the desktop 110 of the mobile phone. In response to a click operation on the application icon 111 of the gallery, as shown at B in Figure 1 , the mobile phone displays the first gallery interface 120. The first gallery interface 120 includes a "One-Click Blockbuster" control 121. In response to a click operation on the "One-Click Blockbuster" control 121, as shown at C in Figure 1 , the mobile phone displays the second gallery interface 130. The second gallery interface 130 may include multiple videos and multiple pictures. In response to a click operation on any video or any picture in the second gallery interface 130, as shown at D in Figure 1As shown by D in [description], the mobile phone displays the third gallery interface 140. The third gallery interface 140 includes all videos and all pictures. In response to a selection operation on at least one video, or at least one video and picture in the third gallery interface 140, for example, a selection operation on one video, as Figure 1 shown by D in [description], a first pop-up window 141 is displayed in the third gallery interface 140. The first pop-up window 141 includes the selected video and a video generation control 142. For example, the video generation control can be a one-click blockbuster control. It should be understood that the one-click blockbuster control in the first pop-up window 141 is different from the "one-click blockbuster" control 121 included in the first gallery interface 120. In response to a click operation on the video generation control 142, the mobile phone analyzes the selected video, or the video and picture, filters out the highlight segments from the selected video, or the video and picture, and generates a finished video based on the highlight segments. During this process, as Figure 1 shown by E in [description], a second pop-up window 143 can be displayed in the third gallery interface 140. The second pop-up window 143 includes the analysis progress so that the user can intuitively view the analysis progress.
[0072] In another embodiment, as Figure 2 shown by A in [description], after obtaining the finished video, the mobile phone displays the fourth gallery interface 150. The fourth gallery interface 150 includes the finished video (such as the first finished video). The mobile phone can automatically play the finished video. In addition, the fourth gallery interface 150 can also include a video export control 151. In response to a click operation on the video export control 151, the finished video is exported. During this process, as Figure 2 shown by B in [description], a third pop-up window 152 can be displayed in the fourth gallery interface 150. The third pop-up window 152 includes the export progress so that the user can intuitively view the export progress.
[0073] In another embodiment, continuing as Figure 2 shown by A in [description], the fourth gallery interface 150 can also include other function controls so that the user can perform operations such as editing, adding special effects, and sharing on the finished video based on these function controls. For example, as Figure 2 shown by A in [description], the other function controls can include but are not limited to controls such as templates, music, segments, and sharing.
[0074] In another embodiment, if the user selects more videos, or videos and pictures in the third gallery interface 140, the mobile phone can give a prompt to the user. As Figure 3 shown by B in [description], a first prompt message can be displayed in the first pop-up window 141. For example, the first prompt message is "Up to 30 materials can be selected".
[0075] In another embodiment, if the video selected by the user in the third gallery interface 140, or the number of videos and pictures is small (such as 1), the mobile phone can give a prompt to the user. For example, Figure 3 as shown in A of Figure 3 , a second prompt message can be displayed in the first pop-up window 141. For example, the second prompt message is "The best effect is achieved with more than 6 materials."
[0076] It should be noted that the videos and pictures displayed in the second gallery interface 130 and the third gallery interface 140 can be taken by the mobile phone or stored in the gallery after being sent to the mobile phone by other electronic devices. The embodiments of the present application do not limit the sources of the videos and pictures displayed in the second gallery interface 130 and the third gallery interface 140.
[0077] Figure 4 shows a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. The following combines Figure 4 to introduce the hardware structure of the electronic device.
[0078] Taking the electronic device as a mobile phone as an example. As Figure 4 shown, the electronic device 400 may include: a processor 410, a memory 420, a universal serial bus (USB) interface 430, a power management module 440, an antenna, a communication module 450, a display screen 460, an audio module 470, a camera 480, a sensor module 490, etc.
[0079] The processor 410 may include one or more processing units. For example, the processor 410 may include an application processor (AP), a modulation and demodulation processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The controller may be the nerve center and command center of the electronic device 400. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.
[0080] The memory 420 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 410 executes various functional applications and data processing of the electronic device by running the instructions stored in the memory 420. The memory 420 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function (such as a sound playback function, an interface display function, etc.). The data storage area can store data created during the use of the electronic device (such as notification messages). In addition, the memory 420 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0081] The power management module 440 is used to connect the battery to the processor 410. The power management module 440 receives battery and / or power input and supplies power to the processor 410, the memory 420, the communication module 450, the display screen 460, the display screen 460, and the camera 480, etc. The power management module 440 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 440 can also be disposed in the processor 410.
[0082] The communication module 450 can provide wireless communication solutions applied to the electronic device 400, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The communication module 450 can be one or more devices integrating at least one communication processing module. The communication module 450 receives electromagnetic waves via an antenna, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 410. The communication module 450 can also receive signals to be sent from the processor 410, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna and radiate them out.
[0083] In some embodiments, the antenna of the electronic device 400 is coupled to the communication module 450, enabling the electronic device 400 to communicate with the network and other devices via wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, Global Navigation Satellite System (GNSS), WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include Global Positioning System (GPS), Beidou Navigation Satellite System (BDS), Global Navigation Satellite System (GLONASS), and / or Galileo Satellite Navigation System (GALILEO).
[0084] The electronic device 400 implements the display function through the GPU, the display screen 460, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 460 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 410 may include one or more GPUs, which execute program instructions to generate or change display information.
[0085] The display screen 460 is used to display images, videos, etc. The display screen 460 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In the embodiments of the present application, the display screen 460 can be used to display a finished video.
[0086] The electronic device 400 can implement the shooting function through an ISP, a camera 480, a video codec, a GPU, a display screen 460, an application processor, etc.
[0087] The audio module 470 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 470 can also be used to encode and decode audio signals. In some embodiments, the audio module 470 can be disposed in the processor 410, or some functional modules of the audio module 470 can be disposed in the processor 410.
[0088] The camera 480 is used to capture static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats.
[0089] The sensor module 490 can include a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, and a bone conduction sensor, etc.
[0090] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 400. In other embodiments, the electronic device 400 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0091] Generally speaking, in addition to hardware support, the implementation of the one - click video - making function in an electronic device also requires software cooperation. The software system of the electronic device can adopt a layered architecture, an event - driven architecture, a micro - kernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of this application, the Android operating is taken as an example, combined with Figure 5 to introduce the software architecture of the electronic device involved in the embodiments of this application.
[0092] Figure 5 FIG. shows a schematic diagram of the software architecture of an electronic device provided by the embodiments of this application.
[0093] As Figure 5 shown, the layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android operating is divided into four layers, from top to bottom in sequence: the application layer, the media middle - platform framework layer, the application framework (FWK) layer, and the hardware abstraction layer (HAL).
[0094] The application layer may include a series of application packages, such as cameras, galleries, third - party video - editing software, calendars, maps, and navigation, etc. When these application packages are run, they can access the various service modules provided by the media middle - platform framework layer and the application framework layer through the application programming interface (API), and execute the corresponding intelligent services.
[0095] In one embodiment, the camera is used to respond to user operations to capture photos, videos, slow - motion images, panoramic images, etc. After these images are captured by the camera, or after the user triggers a mobile phone screenshot, or after the user triggers a mobile phone screen recording, or after the terminal device downloads images from other devices, the terminal device can save these images in the gallery, so that the user can perform video - editing operations on the images in the gallery, such as one - click video - making operations.
[0096] In the embodiments of this application, the gallery is divided into: the business layer, the application function layer, and the basic function layer from top to bottom.
[0097] Among them, the service layer, also known as the video editing service layer. The service layer includes multiple modules such as multi-shot video automatic video creation, AI music short film with multiple benefits from one recording, one-click video creation, and wonderful moments, which are respectively used to provide corresponding services. These services are presented in the form of controls in the user interface (UI) of the gallery, or the gallery interface. By operating on a certain control, the user can trigger the camera to perform corresponding video processing actions. For example, after the user selects one or more materials (materials include material videos and / or material pictures) and clicks the one-click video creation control in the gallery, the gallery will automatically analyze and extract the highlight segments (such as the fifth video segment mentioned below) from the materials through algorithms, and then combine the highlight segments into a clipped video (such as the first video mentioned below).
[0098] Among them, the application function layer includes an automatic editing framework. Each service in the service layer can call the automatic editing framework to provide automatic editing services for the materials. Exemplarily, the automatic editing framework may include function modules such as segment optimization, story line organization, layout splicing, and special effect beautification. Segment optimization is used to call the highlight segment analysis interface, the highlight segment selection strategy interface, and the video frame analysis interface to extract highlight segments from the materials. Story line organization is used to sequentially splice multiple materials in the form of a story line based on the content of the materials. Layout splicing is used to adjust the interface layout of the materials. Special effect beautification is used to adjust the beautification effect of the materials, such as adjusting the picture brightness and beautifying the human face, etc.
[0099] Among them, the basic function layer is used to perform basic function processing on the clipped video segments after the automatic editing framework clips multiple materials. Exemplarily, the basic function layer may include basic function modules such as video splicing, synthesis and saving, video speed change, and audio-visual effect processing. Among them, video splicing is used to splice the multiple extracted highlight segments. Synthesis and saving is used to store the video obtained after splicing. Video speed change is used to perform speed change processing on the video. For example, video speed change is used to add a slow-play effect to the part of the video including the highlight action after splicing. The audio-visual effect processing module is used to add video effects and sound effects to the video obtained after splicing. For example, adding a style filter and theme to the video, adding background music, etc.
[0100] The media middle platform framework layer is a software layer set between the application layer and the application framework. The media middle platform framework layer may include an analysis performance query interface, a highlight segment selection strategy interface, a policy monitoring module, a pipeline interface, a theme summary interface, an initialization interface, a video frame analysis interface, and a highlight segment analysis interface.
[0101] Among them, the analysis performance query interface is used to calculate the total duration of each video according to the analysis speed of the algorithm module.
[0102] Among them, the policy monitoring module can be used to extract the first video segment from each material video jointly with the channel interface, service interface, and algorithm module. The policy monitoring module can also allocate available duration for each material video.
[0103] Among them, the channel interface is used to decode, convert the format, reduce the resolution, etc. of each first video segment, and forward the data address of the processed material video to the hardware abstraction layer through the application framework layer, and then report the analysis completion message (including the highlight segments of all material videos) returned by the video frame analysis interface to the application function layer.
[0104] The channel interface is also used to find each first video segment from the material video according to the file descriptor and analysis policy of each first video segment sent by the policy monitoring module.
[0105] The channel interface is also used to perform frame dropping processing on the decoded first video segment, and retain some video frames in each first video segment. In this way, the time for the electronic device to perform operations such as format conversion and resolution reduction can be reduced, and the efficiency can be improved.
[0106] Among them, the highlight segment selection policy interface is used to screen out all highlight segments (such as the fifth video segment involved below) from the corresponding each first video segment based on the image scores of each first video frame in each first video segment.
[0107] Among them, the theme summary interface is used to call the theme algorithm to obtain the theme type that conforms to all highlight segments (such as the fifth video segment involved below) returned by the video frame analysis interface, and send the theme type to the application function layer. The application function layer obtains the theme template corresponding to the theme type.
[0108] Among them, the initialization interface is used to initialize the aesthetic algorithm, face detection algorithm, scene recognition algorithm, clarity algorithm, subject detection algorithm, action recognition algorithm, brightness algorithm, storyboarding algorithm, etc. of the hardware abstraction layer.
[0109] It should be noted that this application is described by taking the gallery providing the one - click video composition function as an example, which does not limit the embodiments of this application. In actual implementation, third - party video editing software can adopt the video segment processing method provided by the embodiments of this application to synthesize multiple pictures and videos selected by the user into a single video in one click.
[0110] The application framework layer, simply referred to as the framework layer, can be used to support the operation of each module in the media middle platform framework layer. For example, the framework layer can include service interfaces such as one-click video creation interfaces, parameter management interfaces, image data transmission interfaces, theme analysis interfaces, and performance analysis interfaces.
[0111] The hardware abstraction layer can include a capability query interface and algorithm modules. The video frame analysis interface is used to call the capability query interface to obtain the algorithm capabilities supported by the hardware abstraction layer. The algorithm modules can include: aesthetic algorithms, face detection algorithms, scene recognition algorithms, clarity algorithms, subject detection algorithms, action recognition algorithms, and brightness algorithms, etc.
[0112] Among them, the aesthetic algorithm is used to score the aesthetics of a frame of image to obtain the aesthetic score of the frame of image. The aesthetic score can be used as a basis for evaluating whether a frame of image is a highlight segment. For example, when the aesthetic score of a frame of image is greater than a score threshold (such as 0.25), the frame of image can be used as a highlight segment.
[0113] Among them, the face detection algorithm is used to detect faces in a frame of image, the human body detection algorithm is used to detect limbs in a frame of image, and the action recognition algorithm is used to recognize the actions in a frame of image to obtain the predicted action type of the person and the confidence level of the predicted action type, or a no-action flag. The no-action flag is used to indicate that there is no action in the image.
[0114] When there is an action in a frame of image, the action recognition algorithm can recognize at least one predicted action type and the confidence level of each predicted action type. Usually, the predicted action type with the highest confidence level is used as the target action type of the frame of image.
[0115] Whether there is an action in a frame of image can be used as a basis for evaluating whether the frame of image belongs to a highlight segment. For example, when it is recognized that there is an action in a frame of image, the frame of image can belong to a highlight segment. When it is recognized that there is no action in a frame of image, whether the frame of image belongs to a highlight segment needs to be judged in combination with the aesthetic score, brightness, and clarity of the frame of image.
[0116] In one embodiment, the action recognition algorithm is bound to the face detection algorithm and the human body detection algorithm. That is to say, when performing action recognition on a frame of image, face detection and human body detection are also performed on it.
[0117] Among them, the scene recognition algorithm is used to recognize the scene of a frame of image to obtain the scene information of the frame of image. The scene information can include a scene parent label and a scene sub-label. The scene parent label is used to indicate the scene type of a frame of image, such as scenery, food, architecture, etc. The scene sub-label is used to indicate the specific scene (or object) of a frame of image, such as sunrise, noodles, skyscraper, etc.
[0118] Among them, the clarity algorithm is used to identify the clarity of a frame of image, and the brightness algorithm is used to identify the brightness of a frame of image. The clarity and brightness can be used as the basis for evaluating whether a frame of image is a high-light segment. For example, if the clarity of a frame of image is within the clarity range and the brightness is greater than the brightness threshold, this frame of image may belong to a high-light segment.
[0119] Among them, the shot segmentation algorithm is used to identify whether there is a shot segmentation point in a frame of image.
[0120] It should be noted that Figure 5 The layers in the shown software structure and the components included in each layer do not constitute a specific limitation on the terminal device. In some other embodiments, the terminal device may include more layers than shown, such as the system library (FWK LIB) layer and the kernel layer. Each layer may include more or fewer components than shown. In addition, the above-mentioned various functional modules may also be combined into one functional module, and each layer may also be combined into one layer. For example, the media middle platform framework layer may be set in the application framework layer.
[0121] It can be understood that in order for the electronic device to implement the method in the embodiments of the present application, it includes corresponding hardware and / or software modules for executing various functions. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments.
[0122] The following combines Figures 6 - 13 , taking the electronic device shown in Figure 4 as an example of the execution subject, to introduce the video composition method provided in the embodiments of the present application.
[0123] Figure 6 FIG. shows one of the flow diagrams of a video composition method provided in the embodiments of the present application. As Figure 6 shown, the method may include:
[0124] S601. The electronic device displays at least one candidate video.
[0125] Exemplarily, as shown in C in Figure 1 , the second gallery interface 130 displayed on the mobile phone. The second gallery interface 130 includes at least one video and at least one picture. This video may be referred to as a candidate video, and this picture may be referred to as a candidate picture.
[0126] Among them, the candidate video can be obtained by the electronic device through shooting, screen recording, or received from other electronic devices. The embodiments of the present application do not limit this.
[0127] S602. The electronic device obtains a first video segment based on the candidate video.
[0128] Among them, the first video segment is related to the candidate video displayed on the electronic device. The first video segment includes highlight actions. For example, the highlight actions may include: actions such as a person jumping, splashing water, throwing an object, running, table tennis smashing, badminton smashing, and a smiling glance back. The first video segment can be obtained in the following different ways:
[0129] In one embodiment, the first video segment can be a candidate video selected by the user. Herein, the candidate video selected by the user is referred to as the material video.
[0130] Exemplarily, as Figure 1 shown in D of, in response to a selection operation on at least one candidate video in the third gallery interface 140 displayed on the mobile phone, the mobile phone uses the selected candidate video as the first video segment.
[0131] In another embodiment, the first video segment can be a partial video segment of the corresponding material video. In this embodiment, the electronic device preprocesses the corresponding material video and extracts the first video segment from the material video. The first video segment can be the highlight segment of the corresponding material video.
[0132] The process of the electronic device preprocessing the material video will be introduced in detail below, and will not be elaborated herein.
[0133] S603. The electronic device performs image analysis on the first image frames in each first video segment according to a preset frame skipping strategy to obtain the image scores of each first image frame.
[0134] Among them, the preset frame skipping strategy can indicate which image frames in the first video segment are the first image frames and the target recognition types for each first image frame. Among them, the target recognition types include: one or more of clarity recognition, brightness recognition, aesthetic scoring, face recognition, scene recognition, and human body recognition, as well as action recognition and shot recognition.
[0135] In one embodiment, the electronic device starts from the first image frame in the first video segment and extracts one image frame as the first image frame from the first video segment every L frames. Wherein, L is a positive integer.
[0136] Exemplarily, the first image frame is the image frame with frame data addresses of 1, AL, 2AL, and 4AL in the first video segment. Wherein, A is an integer greater than 1 (such as 3). That is to say, the first image frame, the AL-th image frame, the 2AL-th image frame, and the 4AL-th image frame in the first video segment are the first image frames.
[0137] In another embodiment, the electronic device uses all the image frames in the first video segment as the first image frames.
[0138] It should be noted that the time to process one image frame is fixed. Therefore, within a certain duration, the number of image frames that can be processed is limited. Therefore, the selection method of the first image frame can also include the following two:
[0139] In another embodiment, within the available duration of the first video segment, starting from the first image frame in the first video segment, one image frame is extracted from the first video segment every L frames as the first image frame.
[0140] Exemplarily, assume that the available duration of a first video segment is 20 seconds. Starting from the first image frame in the first video segment, 100 image frames can be extracted from the first video segment every L frames, and 50 image frames can be processed within 20 seconds. Then the first image frames are the first 50 image frames with frame data addresses of 1, AL, 2AL, and 4AL within 20 seconds in the first video segment.
[0141] In another embodiment, all the image frames within the available duration of the first video segment are used as the first image frames.
[0142] Exemplarily, assume that the available duration of a first video segment is 20 seconds, and the first video segment includes 100 image frames, and 50 image frames can be processed within 20 seconds. Then the first image frames are the first 50 image frames within 20 seconds in the first video segment.
[0143] Figure 7 Fig. 2 shows the second schematic flowchart of a video composition method provided by an embodiment of the present application. The following will be combined with Figure 7 , and how to obtain the image scores of each first image frame will be introduced in detail. Specifically, as Figure 7 shown, obtaining the image scores of each first image frame may include:
[0144] S701. The electronic device obtains the first image analysis information of each first image frame in each first video segment.
[0145] Among them, the first image analysis information may include the brightness, clarity, aesthetic score, second action information, face information, limb information, scene information, shot information, etc. of the corresponding image frame. The first image analysis information may be referred to as the second analysis information.
[0146] Among them, the second action information may include an inaction identifier or at least one predicted action information. Each predicted action information includes the corresponding predicted action type and the confidence of the predicted action type. The inaction identifier may be used to indicate that there is no highlight action in the first image frame. For example, the inaction identifier may be represented by a first value (such as -1).
[0147] Among them, the scene information may include a scene parent label and a scene child label. The scene child label is a child node of the scene parent label. The scene parent label is used to indicate the scene type of the image frame, and the scene child label is used to indicate the shooting object in the image frame. For example, the scene parent label is "Food", and the scene child label is "Noodles".
[0148] Among them, the shot information includes an unshot identifier or a shot identifier. Among them, the unshot identifier is used to indicate that there is no shot point between the corresponding image frame and the adjacent image frame, and the shot identifier is used to indicate that there is a shot point between the corresponding image frame and the adjacent image frame.
[0149] The following specifically introduces how the electronic device obtains the first image analysis information of each first image frame:
[0150] In one embodiment, the electronic device performs a full-scale image analysis on all the first image frames to directly obtain the first image analysis information of each first image frame. In this way, it is more convenient. Among them, the full-scale image analysis means that the electronic device calls all algorithms to analyze each first image frame.
[0151] Exemplarily, the full-scale image analysis is that the electronic device calls an aesthetic algorithm to perform an aesthetic score on the first image frame to obtain the aesthetic score of the first image frame; calls a face detection algorithm to perform face detection on the first image frame to obtain the face information of the first image frame; calls a subject detection algorithm to perform limb detection on the first image frame to obtain the limb information of the first image frame; calls an action recognition algorithm to perform action recognition on the first image frame to obtain the second action information of the first image frame; calls a clarity algorithm to perform clarity recognition on the first image frame to obtain the clarity of the first image frame; calls a brightness algorithm to perform brightness recognition on the first image frame to obtain the brightness of the first image frame; calls a scene recognition algorithm to perform scene recognition on the first image frame to obtain the scene information of the first image frame; calls a shot algorithm to perform shot point recognition on the first image frame to obtain the shot information of the first image frame. But it is not limited to this.
[0152] It should be noted that when the first image frame is all the image frames in the first video segment, the electronic device performs a full-image analysis on all the image frames in the first video segment. When the first image frame is the image frames in the first video segment with frame data addresses of 1, AL, 2AL, and 4AL, a full-image analysis is performed on this part of the image frames in the first video segment.
[0153] In another embodiment, the electronic device performs a first image analysis on the first image frame with a frame data address of AL, a second image analysis on the first image frame with a frame data address of 2AL, and a full-image analysis on the first first image frame and the first image frame with a frame data address of 4AL. Based on this, the analysis efficiency can be improved, the analysis time can be saved, and the finished video can be obtained quickly.
[0154] Among them, the first image analysis refers to the algorithm module calling the first part of the algorithm to analyze a first image frame to obtain the first image information of the first image frame. The second image analysis refers to the algorithm module calling the second part of the algorithm to analyze a first image frame to obtain the second image information of the first image frame.
[0155] Specifically, the first part of the algorithm and the second part of the algorithm are different, so the first image information is different from the second image information. The first image analysis information can include the first image information and the second image information.
[0156] Exemplarily, the first part of the algorithm can include an action recognition algorithm and a storyboarding algorithm, etc. Specifically, the first image analysis is that the algorithm module calls the action recognition algorithm to perform action recognition on the first image frame to obtain the second action information of the first image frame, and calls the storyboarding algorithm to perform storyboarding recognition on the first image frame to obtain the storyboarding information of the first image frame, but not limited to this.
[0157] Exemplarily, the second part of the algorithm can include an action recognition algorithm, a clarity algorithm, and a storyboarding algorithm, etc. Specifically, the second image analysis is that the algorithm module calls the action recognition algorithm to perform action recognition on the first image frame to obtain the second action information of the first image frame, calls the storyboarding algorithm to perform storyboarding recognition on the first image frame to obtain the storyboarding information of the first image frame, and calls the clarity algorithm to perform clarity recognition on the first image frame to obtain the clarity of the first image frame, but not limited to this.
[0158] It should be noted that for an action in an image frame, the action recognition algorithm may recognize at least one predicted action type, and each predicted action type has a corresponding confidence level. The greater the confidence level of a predicted action type, the closer the action corresponding to the predicted action type is to the highlight action in the image frame. In the embodiments of the present application, the predicted action type with the greatest confidence level is called the target action type.
[0159] Exemplarily, take an image frame recording badminton playing as an example. The predicted action types recognized by the action recognition algorithm may include: running, jumping, and playing the ball. Among them, the confidence level of running is 0.3, the confidence level of jumping is 0.5, and the confidence level of playing the ball is 0.7. Based on this, the action corresponding to the target action type in this image frame is playing the ball.
[0160] In summary, the image information of the first image frame after full-image analysis is the first image analysis information of this first image frame. The image information of the first image frame that has not undergone full-image analysis is part of the first image analysis information of this first image frame.
[0161] The following introduces the method for obtaining the first image analysis information of the first image frame that has not undergone full-image analysis:
[0162] In one embodiment, the (N + 1)-th first image frame can obtain the first image analysis information of the (N + 1)-th first image frame by inheriting part of the image information of the N-th first image frame. Among them, this part of the inherited image information is the image information that is not included in the image information of the (N + 1)-th first image frame. Where N is a positive integer.
[0163] Exemplarily, assume A = 3. Then the first image frame with a frame data address of 3 is the second first image frame. The image information of the second first image frame includes second action information and shot information. At this time, the second first image frame can inherit this part of the image information such as clarity, brightness, scene information, and aesthetics score of the first first image frame (i.e., the first image frame with a frame data address of 1).
[0164] S702. The electronic device processes each first image analysis information to obtain corresponding second image analysis information.
[0165] Among them, the second image analysis information may include: aesthetics score, face information, limb information, first action information, clarity, brightness, scene sub-label, shot information, etc. The second image analysis information can be referred to as the first analysis information.
[0166] In one embodiment, after the electronic device obtains each first image analysis information, it processes the second action information and scene information in each first image analysis information to obtain the second image analysis information.
[0167] Exemplarily, when the second action information includes multiple predicted action types and the confidence level corresponding to each predicted action type, the electronic device determines the predicted action type with the highest confidence level in the second action information as the target action type to obtain the first action information. That is to say, the first action information may include a no-action identifier or the target action type. In addition, the electronic device can screen out the scene sub-label from the scene information.
[0168] Exemplarily, when the second action information includes a predicted action type and a confidence level corresponding to the predicted action type, the electronic device directly determines the second action information as the first action information.
[0169] S703. The electronic device obtains the image scores of the corresponding first image frames according to each piece of second image analysis information.
[0170] After the electronic device obtains the second image analysis information of each first image frame in a first video segment, it scores each first image frame in the first video segment based on a scoring strategy to obtain the image scores of the corresponding first image frames. The detailed scoring strategy is as follows:
[0171] In one embodiment, as Figure 8 shown, taking the example that the first action information of a first image frame indicates that the first image frame includes a highlight action. The electronic device uses a second value (such as 1) as the image score of the first image frame. At the same time, the electronic device sets the image scores of all first image frames within a second time period (such as 1 second) before the first image frame to the second value, and the electronic device sets the image scores of all first image frames within the second time period after the first image frame to the second value. Wherein, the second value is greater than the score threshold. That is to say, the image scores of all first image frames within the second time period before and after the first image frame containing the highlight action are greater than the score threshold.
[0172] In another embodiment, continuing as Figure 8 shown, taking the example that the first action information of a first image frame indicates that the first image frame does not include a highlight action. Based on the clarity of the first image frame being within a preset clarity range and the brightness of the first image frame being greater than the brightness threshold, the electronic device uses the aesthetic score of the first image frame as the image score of the first image frame. Conversely, the electronic device uses a third value (such as 0) as the image score of the first image frame. Wherein, the third value is less than the score threshold. That is to say, if the clarity of the first image frame is not within the preset clarity range and the brightness of the first image frame is less than the brightness threshold, then the image score of the first image frame is less than the score threshold.
[0173] S604. The electronic device divides multiple consecutive first image frames in the first video segment whose image scores are greater than the score threshold and belong to the same shot into the same video segment to obtain at least one second video segment.
[0174] Among them, the electronic device can determine whether an image frame belongs to a high-light picture according to the size relationship between the image score of an image frame and the score threshold. The high-light pictures in the first video segment can be retained. Generally, if the image score of an image frame is greater than or equal to the score threshold, then the image frame belongs to a high-light picture. Otherwise, the image frame does not belong to a high-light picture.
[0175] It should be noted that S604 is executed for each first video frame.
[0176] Figure 9 Fig. 3 shows a schematic flow chart of a video production method provided by an embodiment of the present application. Figure 10 Fig. 4 shows a schematic diagram of screening corresponding video segments provided by an embodiment of the present application. The following combines Figure 9 and Figure 10 to introduce how the electronic device obtains at least one second video segment from the first video segment. Specifically, as Figure 9 shown, S604 may include:
[0177] S901. The electronic device screens out multiple second image frames with image scores greater than the score threshold from the first image frames in the first video segment.
[0178] Exemplarily, assume that the score threshold is 0.25. As shown in A of Figure 10 , the first video segment A includes 9 first image frames. At the same time, the image scores of the first 3 first image frames in the first video segment are 0.3, the image score of the 4th first image frame is 0.2, the image scores of the 5th and 6th first image frames are 0.4, and the image scores of the 7th to 9th first image frames are 0.5. At this time, the first 3 first image frames among the 9 first image frames are second image frames, and the 5th to 9th first image frames among the 9 first image frames are second image frames.
[0179] S902. The electronic device divides multiple consecutive second image frames among all the screened second image frames into one third video segment to obtain at least one third video segment.
[0180] Exemplarily, continuing with A of Figure 10 , since the first 3 second image frames are consecutive, the first 3 second image frames are one third video segment in the first video segment. At the same time, since the 4th to 8th second image frames are consecutive, the 4th and 5th second image frames are one third video segment in the first video segment.
[0181] It should be understood that the 4th to 8th second image frames are the above-mentioned 5th to 9th first image frames.
[0182] S903. The electronic device divides the second image frames within the same shot in a third video clip into a second video clip, obtaining at least one second video clip.
[0183] In one embodiment, the electronic device may indicate, based on the shot information, that a second image frame has a shot point, and divides the third video clip at this shot point as the first segmentation point to obtain the second video clip. Among the second image frames between two adjacent shot points in a third video clip, they belong to the same shot.
[0184] Exemplarily, as shown in B of Figure 10 assuming that the 3rd second image frame in a third video clip includes a shot point (i.e., the second segmentation point), the electronic device divides the third video clip into two second video clips based on this shot point.
[0185] Among them, the second image frame including the segmentation point may be located in the previous second video clip or the subsequent second video clip, and the embodiments of the present application do not limit this.
[0186] It should be noted that after the electronic device finishes executing the above S901 - S903 on a first video clip, if there is a next first video clip, it continues to execute the above S901 - S903 on the next first video clip until all first video clips are processed.
[0187] S605. The electronic device obtains a first finished video based on at least one second video clip among all first video clips.
[0188] In one embodiment, after the electronic device obtains all second video clips, it directly splices all second video clips to obtain the first finished video.
[0189] In another embodiment, after the electronic device obtains all second video clips, the electronic device may execute the steps shown in Figure 11 Specifically, as shown in Figure 11 S605 may include the following steps:
[0190] S1101. The electronic device determines whether the duration of the second video clip exceeds a first duration.
[0191] Among them, the first duration is greater than the above - mentioned second duration. For example, the first duration is 8 seconds and the second duration is 1 second.
[0192] S1102. Based on the fact that the duration of the second video clip exceeds the first duration, the electronic device divides the second video clip into at least two fourth video clips with a duration less than the first duration at the first segmentation point.
[0193] In one embodiment, the first segmentation point may be the target moment in the second video segment related to the first duration. Alternatively, the first segmentation point may be the candidate moment in the second video segment related to the second image frame including the highlight action. The second image frame including the highlight action may be referred to as the third image frame.
[0194] Among them, the duration between the target moment in a second video segment and the start moment of the second video segment is an integer multiple of the first duration. For example, assuming the first duration is 8 seconds, the target moment in the second video segment is the 8th second, 16th second, 24th second, etc. in the second video segment.
[0195] Among them, the candidate moment may include a first candidate moment and a second candidate moment. The first candidate moment is the start moment of the second duration before the third image frame in the second video segment. The second candidate moment is the end moment of the second duration after the third video frame image frame in the second video segment. For example, assuming the third image frame is located at the 10th second of the second video segment and the second duration is 1 second, the first candidate moment in the second video segment may be the 9th second in the second video segment, and the second candidate moment in the second video segment may be the 11th second in the second video segment.
[0196] The following is an example to introduce the segmentation method for a second video segment longer than the first duration:
[0197] In one embodiment, take the example that a second video segment does not include a highlight action. At this time, the electronic device segments the second video segment with the target moment as the first segmentation point to obtain at least two fourth video segments.
[0198] Exemplarily, assume that the duration of a second video segment is 10 seconds and the second video segment does not include a highlight action. The target moment in the second video segment is the 8th second of the second video segment. At this time, the electronic device segments with the 8th second of the third highlight segment as the first segmentation point to obtain a fourth video segment with a duration of 8 seconds and a fourth video segment with a duration of 2 seconds.
[0199] In another embodiment, take the example that a second video segment includes a highlight action. At this time, the electronic device segments the second video segment with at least one of the target moment and the candidate moment as the first segmentation point to obtain at least two fourth video segments.
[0200] Exemplarily, assume that the duration of a second video segment is 20 seconds and the second video segment includes the third image frame. The target moments in the second video segment are the 8th second and 16th second of the second video segment. At this time, how to segment the second video segment depends on the position (i.e., moment) of the third image frame in the second video segment.
[0201] For example, assume that the first duration is 8 seconds, the second duration is 1 second, and the third image frame is before the 8th second in the second video segment, such as the 7th second. At this time, the first candidate moment is the 6th second, and the second candidate moment is the 8th second. The duration corresponding to the first candidate moment does not exceed the first duration, and the second candidate moment is the first target moment. At this time, the first split point is the second candidate moment (i.e., the first target moment of 8 seconds) and the second target moment (i.e., 16 seconds).
[0202] In this example, the electronic device splits the second video segment into two fourth video segments with a duration of 8 seconds each and one fourth video segment with a duration of 4 seconds.
[0203] Again, for example, assume that the first duration is 8 seconds, the second duration is 1 second, and the third image frame is before the 8th second in the second video segment, such as the 6th second. At this time, the first candidate moment is the 5th second, and the second candidate split point is the 7th second. The durations corresponding to both candidate moments do not exceed the first duration. At this time, the first split point is the two target moments.
[0204] Again, for example, assume that the first duration is 8 seconds, the second duration is 1 second, and the third image frame is after the 8th second and before the 16th second in the second video segment, such as the 10th second. At this time, the first candidate moment is the 9th second, and the second candidate moment is the 11th second. At this time, both candidate split points are between the two target moments, so the first split point is the two target moments.
[0205] In this example, the electronic device splits the third highlight segment into two fourth video segments with a duration of 8 seconds each and one fourth video segment with a duration of 4 seconds.
[0206] Again, for example, as Figure 10 shown in C of, assume that the first duration is 8 seconds, the second duration is 1 second, and the third image frame is at the 8th second of the second video segment. At this time, the first candidate moment is the 7th second, and the second candidate moment is the 9th second. At this time, the two candidate split points are on both sides of the first target moment. At this time, the first split point can include the first candidate moment and the second candidate moment.
[0207] At this time, in order to retain the video segment related to the highlight action in one fourth video segment as much as possible and make the fourth video segment less than the first duration, in this example, the first candidate moment, the second candidate moment, and the second target moment are used as the first split point.
[0208] In this example, the electronic device splits the third highlight segment into a fourth video segment with a duration of 7 seconds, a fourth video segment with a duration of 2 seconds, a fourth video segment with a duration of 8 seconds, and a fourth video segment with a duration of 4 seconds.
[0209] It should be noted that if the duration of a second video segment is less than the first duration, the electronic device does not need to split the second video segment. At this time, the second video segment is the fourth video segment.
[0210] S1103. The electronic device screens out fifth video segments with a duration greater than or equal to the third duration from all the fourth video segments of the second video segment.
[0211] Among them, the third duration is twice the second duration. For example, assuming the second duration is 1 second, the third duration is 2 seconds. The fifth video segment is a video segment with a duration greater than the third duration and less than the first duration.
[0212] In addition, combining the principle of obtaining the fourth video segment above, if a fourth video segment includes a highlight action, the duration of the fourth video segment is greater than or equal to the third duration. That is to say, all the highlight actions in all the fourth video segments are included in all the fifth video segments. At the same time, all the highlight actions in the first video segment are retained in the fourth video segment. Therefore, all the highlight actions in all the first video segments are included in all the fifth video segments.
[0213] It should be noted that when the electronic device obtains a fifth video segment, it is equivalent to obtaining the start position and the end position of this fifth video segment.
[0214] It should be noted that after the electronic device finishes executing the above S1101 - S1103 on a second video segment, if there is a next second video segment, it continues to execute the above S1101 - S1103 on the next second video segment until all the second video segments are processed. Then, the electronic device executes the following S1104.
[0215] S1104. The electronic device splices all the fifth video segments.
[0216] In one embodiment, after the electronic device obtains all the fifth video segments, it can obtain a theme template corresponding to the content of each fifth video segment. Then, after the electronic device splices the respective fifth video segments, it directly applies the theme template to obtain the first finished video.
[0217] In another embodiment, on the basis of the above embodiment, after the electronic device splices the respective fifth video segments, it can generate a second finished video. Then, the electronic device performs video speed change processing on the video segments including multiple actions in the second finished video to obtain a third finished video. Finally, the electronic device applies the theme template to the third finished video to obtain the first finished video.
[0218] In some other embodiments, the electronic device further includes a target picture. At this time, the splicing of all the fifth video segments by the electronic device may be the splicing of the target picture and all the fifth video segments by the electronic device. The target picture may be the candidate picture (i.e., the material picture) selected by the user as described above, or the target picture may be a highlight picture with an image score greater than the score threshold screened out by the electronic device after image analysis of multiple material pictures.
[0219] S606. The electronic device displays the first finished video.
[0220] Exemplarily, as Figure 2 shown in A of, the finished video displayed on the fourth picture gallery interface 150 is the first finished video.
[0221] Figure 12 FIG. 5 shows a schematic flowchart of a video finishing method provided by an embodiment of the present application, Figure 13 FIG. 4 shows a schematic interface diagram provided by an embodiment of the present application. As Figure 12 shown, after the above S606, the video finishing provided by the embodiment of the present application may further include:
[0222] S1201. In response to an editing operation on the first finished video, the electronic device generates a fourth finished video corresponding to the first finished video.
[0223] Among them, the editing operation may include: custom speed change processing and other speed change processing, etc. For example, other speed change processing may include: montage speed change processing, key frame speed change processing, bullet time speed change processing, etc.
[0224] Exemplarily, as Figure 2 shown in A of, the mobile phone displays the fourth picture gallery interface 150. The fourth picture gallery interface 150 includes the first finished video and an editing control. In response to a click operation on the editing control, as Figure 13 shown in A of, the mobile phone displays the fifth picture gallery interface 1301. The fifth picture gallery interface 1301 may include the first finished video and a speed change control 1302. In response to a click operation on the speed change control 1302, as Figure 13 shown in B of, the mobile phone displays a first pop-up window 1303 on the fifth picture gallery interface 1301. The first pop-up window 1303 may include a custom control 1304 corresponding to the curve speed change function. In response to a click operation on the custom control 1304, as Figure 13As shown by C in the figure, the mobile phone displays a second pop-up window 1305 on the fifth gallery interface 1301. The second pop-up window 1305 displays a custom variable speed curve 1306 and a confirmation control 1307 (such as a tick). The custom variable speed curve 1306 can change accordingly in response to the user's adjustment operation. In response to a click operation on the confirmation control 1307, the mobile phone performs variable speed processing on the first generated video according to the custom curve to obtain a target finished video.
[0225] S1202. The electronic device displays a fourth finished video.
[0226] The above introduced the video finishing method provided by the embodiments of the present application with the electronic device as the execution subject. The following will combine Figures 14 - 17 , with the relevant modules in the software architecture diagram shown in Figure 5 as an example to introduce the video finishing method provided by the embodiments of the present application.
[0227] In one scenario, the first video segment involved in the embodiments of the present application is a candidate video selected by the user, that is, the first video segment is a complete material video. Based on this, the following introduces the video finishing method provided by the embodiments of the present application.
[0228] Figure 14 FIG. shows the sixth flow chart of a video finishing method provided by the embodiments of the present application. Figure 15 FIG. shows the seventh flow chart of a video finishing method provided by the embodiments of the present application. Figure 16 FIG. shows the eighth flow chart of a video finishing method provided by the embodiments of the present application.
[0229] Specifically, as Figure 14 shown, the video finishing method provided by the embodiments of the present application may include:
[0230] S1401. The one-click video finishing module in the service layer receives the operation of the user enabling the one-click video finishing function.
[0231] In one embodiment, as Figure 1 shown by B in the figure, the operation of enabling the one-click video finishing function may be: a click operation on the one-click blockbuster control 121 on the first gallery interface 120.
[0232] S1402. The one-click video finishing module in the service layer loads and displays candidate pictures and candidate videos.
[0233] In one embodiment, as Figure 1As shown by C in [the figure], in response to a click operation on the one - key blockbuster control 121 in the first gallery interface 120, the one - key video creation module loads and displays multiple pictures and multiple videos in the second gallery interface 130. The videos in the second gallery interface 130 are candidate videos. The pictures in the second gallery interface 130 are candidate pictures.
[0234] In another embodiment, when there are only videos in the gallery, S1402 loads and displays candidate videos for the one - key video creation module in the business layer. Or, when there are only pictures in the gallery, S1402 loads and displays candidate pictures for the one - key video creation module in the business layer.
[0235] S1403. The one - key video creation module in the business layer receives the operation of the user selecting the material video and determines the operation of executing the one - key video creation function.
[0236] Among them, the material video is the candidate video selected by the user. At least one material video can be a video including highlight actions or a video not including highlight actions. In the embodiments of the present application, the case where at least one material video includes highlight actions is taken as an example for introduction.
[0237] In some embodiments, in this step, the one - key video creation module in the business layer receives the operation of the user selecting the material picture. The material picture is the candidate picture selected by the user.
[0238] Exemplarily, as Figure 1 shown by D in [the figure], the operations of selecting the material video and the material picture can be: click operations on the candidate video and the candidate picture in the third gallery interface 140, etc. As Figure 1 shown by D in [the figure], the operation of determining to execute the one - key video creation function can be: click operation on the video generation control 142 in the third gallery interface 140.
[0239] S1404. The one - key video creation module in the business layer sends the file descriptor (FD) of each material video to the channel interface of the media middle - platform framework layer through the application function layer.
[0240] Among them, a file descriptor can be used to uniquely identify an opened file (such as Video 1). Generally, the file descriptor is a non - negative integer and is an index value.
[0241] After S1404, for each material video (such as Video 1, Video 2, etc.), the following S1405 - S1418 are executed.
[0242] S1405. The channel interface of the media middle - platform framework layer decodes, converts the format, reduces the resolution, etc. of Video 1 and stores the processed Video 1.
[0243] Specifically, in this step, the channel interface first obtains Video 1 based on the file descriptor of Video 1, then decodes, converts the format, reduces the resolution, etc. of Video 1, and stores the processed Video 1.
[0244] Generally, in order to save the analysis time of the video by the electronic device, before analyzing the video, the electronic device can perform downsampling on the video.
[0245] Exemplarily, taking a video as an example, this step may include: First, the electronic device decodes all the image frames in the video. Next, the electronic device converts the decoded video from the first format to the second format. Next, the electronic device reduces the resolution of the video with the changed format from the first resolution to the second resolution. Next, the electronic device converts the video with the reduced resolution from the second format back to the first format. Finally, the electronic device stores the video with the format converted again.
[0246] Among them, the first format may be the nv12 format, and the second format may be the i420 format, but not limited thereto. The first resolution may be 1080p, and the second resolution may be 480p, but not limited thereto.
[0247] It should be noted that since the downsampling algorithm only supports the second format, it is necessary to first convert the video from the first format to the second format, then reduce the resolution of the video, and then convert the video with the reduced resolution back to the original first format and store it in the memory.
[0248] Of course, in some embodiments, if the downsampling algorithm supports the first format, then the electronic device does not need to perform format conversion on the video and can directly reduce the resolution after decoding. The embodiments of the present application do not limit this.
[0249] In one embodiment, after the electronic device decodes the video, it can perform frame dropping on the decoded video and retain some video frames in the video. In this way, the time for the electronic device to perform operations such as format conversion and resolution reduction can be reduced, and the efficiency can be improved.
[0250] Exemplarily, the electronic device equally-spaced retains 30 video frames from the video within a certain time period (such as 1 second).
[0251] It should be noted that in the embodiments of the present application, the first video segment to be segmented is the first video segment after the above-mentioned decoding, format conversion, resolution reduction, etc.
[0252] S1406. The channel interface of the media middle platform framework layer sends the frame data address of Video 1 to the video frame analysis interface of the media middle platform framework layer.
[0253] S1407. The video frame analysis interface of the media middle platform framework layer sends the frame data address of Video 1 and the second analysis strategy to the algorithm module of the hardware abstraction layer through the service interface of the application framework layer.
[0254] Among them, the service interface in this step can be the one - key video composition interface. The second analysis strategy can be called the preset frame - skipping strategy.
[0255] Specifically, the second analysis strategy in this step can refer to the relevant introduction in S603 above.
[0256] S1408. The algorithm module of the hardware abstraction layer analyzes each first image frame of Video 1 according to the second analysis strategy to obtain the image information of each first image frame.
[0257] Among them, this step can refer to the relevant introduction in S701 above and will not be elaborated here.
[0258] S1409. The algorithm module of the hardware abstraction layer sends the image information of each first image frame in Video 1 to the video frame analysis interface of the media middle platform framework layer through the service interface of the application framework layer.
[0259] S1410. The video frame analysis interface of the media middle platform framework layer obtains the first image analysis information of each first image frame according to the image information of each first image frame of Video 1.
[0260] Among them, this step can refer to the relevant introduction in S701 above and will not be elaborated here.
[0261] S1411. The video frame analysis interface of the media middle platform framework layer processes the first image analysis information of each first image frame of Video 1 to obtain the corresponding second image analysis information of each one.
[0262] Among them, this step can refer to the relevant introduction in S702 above and will not be elaborated here.
[0263] S1412. The video frame analysis interface of the media middle platform framework layer sends the second image analysis information of each first image frame of Video 1 to the highlight segment selection strategy interface of the media middle platform framework layer.
[0264] S1413. The highlight segment selection strategy interface of the media middle platform framework layer obtains the image score of each corresponding first image frame according to the second image analysis information of each first image frame in Video 1.
[0265] Among them, this step can refer to the relevant introduction in S703 above and will not be elaborated here.
[0266] The highlight segment selection strategy interface of the media middle platform framework layer filters out at least one third video segment according to the image scores of all the first image frames in Video 1.
[0267] For this step, the relevant introductions in the above S901 and S902 can be referred to and will not be elaborated here.
[0268] S1415. The highlight segment selection strategy interface of the media middle platform framework layer divides a third video segment into two second video segments based on the shot points in the third video segment.
[0269] For this step, the relevant introduction in the above S903 can be referred to and will not be elaborated here.
[0270] S1416. The highlight segment selection strategy interface of the media middle platform framework layer divides a second video segment with a duration longer than the first duration to obtain multiple fourth video segments.
[0271] For this step, the relevant introductions in the above S1101 and S1102 can be referred to and will not be elaborated here.
[0272] S1417. The highlight segment selection strategy interface of the media middle platform framework layer filters out a fifth video segment with a duration longer than the third duration from all the fourth video segments.
[0273] For this step, the relevant introduction in the above S1103 can be referred to and will not be elaborated here.
[0274] S1418. The highlight segment selection strategy interface of the media middle platform framework layer sends an analysis completion message of Video 1 to the channel interface of the media middle platform framework layer.
[0275] An analysis completion message of a material video includes: the start position and end position of each fifth video segment in the material video, and the second image analysis information of each image frame in each fifth video segment.
[0276] After the channel interface of the media middle platform framework layer obtains an analysis completion message of a material video (such as Video 1), if there are other material videos (such as Video 2), the electronic device can continue to execute the above S1405 - S1418 to obtain the analysis completion messages of other material videos. After obtaining the analysis completion messages of all the material videos.
[0277] Further, after the above S1403 and before S1405, as Figure 15 shown, the video compilation method provided by the embodiment of the present application may include:
[0278] S1501. The one - click video creation module in the service layer calls the initialization interface of the media middleware framework layer through the segment optimization module in the application function layer to initialize the relevant algorithms in the hardware abstraction layer.
[0279] Among them, the relevant algorithms refer to the algorithms for the service functions to be implemented in the service layer. Here, the service function is the one - click video creation function. Therefore, the relevant algorithms may include aesthetic algorithms, face detection algorithms, scene recognition algorithms, clarity algorithms, subject detection algorithms, action recognition algorithms, brightness algorithms, etc.
[0280] S1502. The initialization interface of the media middleware framework layer sequentially sends the corresponding initialization parameters to the algorithm module in the hardware abstraction layer through the channel interface in the media middleware framework layer and the service interface corresponding to the application framework layer.
[0281] Among them, the service interface in this step can be the parameter management interface. The parameter management interface is used for data transparent transmission between the channel interface and the algorithm module.
[0282] S1503. The algorithm module in the hardware abstraction layer initializes each algorithm using the initialization parameters.
[0283] S1504. The algorithm module in the hardware abstraction layer sends an initialization success message to the video frame analysis interface of the media middleware framework layer through the service interface of the application framework layer.
[0284] Among them, the service interface in this step can be the parameter management interface involved in S1502.
[0285] S1505. The video frame analysis interface of the media middleware framework layer calls the capability query interface in the hardware abstraction layer through the service interface of the application framework layer to query the algorithm capabilities supported by the algorithm module.
[0286] Among them, the service interface in this step can be the performance analysis interface. The algorithm capabilities supported by the algorithm module can be the chip analysis speed of the chip corresponding to the algorithm module (such as an image signal processor). Different chips have different chip analysis speeds. For an already - launched electronic device, the chip in the electronic device is fixed. Therefore, the algorithm capabilities of the algorithm module in the electronic device are fixed.
[0287] In one embodiment, the algorithm capability can be a multiple. The larger the multiple, the stronger the algorithm capability and the faster the analysis speed of the image.
[0288] Exemplarily, taking a 60 - second material video as an example. If the multiple is 1, the analysis time of this material video is 60 seconds. If the multiple is 2, the analysis time of this material video is 30 seconds.
[0289] The ability query interface of the hardware abstraction layer sends algorithm capabilities to the initialization interface of the media middleware framework layer through the service interface of the application framework layer, the video frame analysis interface of the media middleware framework layer, and the channel interface.
[0290] Among them, the service interface in this step is the performance analysis interface involved in S1505.
[0291] In one embodiment, the channel interface of the media middleware framework layer can also return other relevant performance parameters of various algorithms to the initialization interface of the media middleware framework layer.
[0292] S1507. The initialization interface of the media middleware framework layer sends an initialization success message to the one-click video creation module of the service layer through the application function layer.
[0293] Among them, the initialization success message can include the algorithm capabilities of the algorithm module.
[0294] It should be noted that after S1507, for each material video (such as Video 1), the following S1508 - S1510 can be executed. Hereinafter, the material video is taken as an example of Video 1 for introduction.
[0295] S1508. The one-click video creation module of the service layer sends a query message to the analysis performance query interface of the media middleware framework layer through the application function layer.
[0296] Among them, the query message can include the file descriptor of a material video (such as Video 1) and the algorithm capabilities of the algorithm module.
[0297] S1509. The analysis performance query interface of the media middleware framework layer queries the resolution of Video 1 according to the file descriptor of Video 1.
[0298] Specifically, the resolution of a video is an attribute of the video. The resolution of a video is related to the hardware device that captured the video. When a video is stored in an electronic device, the resolution of the video can be obtained by querying the attribute information of the video. Specifically, the resolution of a video can be queried through the analysis performance query interface.
[0299] S1510. The analysis performance query interface of the media middleware framework layer calculates the analysis speed of Video 1 according to the resolution and algorithm capabilities of Video 1.
[0300] Specifically, the analysis speed of a material video depends on the video decoding speed and algorithm capabilities. Among them, the video decoding speed of a material video can be calculated according to the resolution and frame rate of the material video. Therefore, the analysis speed of the material video can be calculated according to the algorithm capabilities of the algorithm module, the resolution and frame rate of a material video.
[0301] It can be seen from this that when the algorithm capabilities are fixed: the analysis speeds of two material videos with different resolutions are different, or the analysis speeds of two material videos with different frame rates are different, or the analysis speeds of two material videos with different resolutions and different frame rates are different.
[0302] S1511. The analysis performance query interface of the media middle platform framework layer sends the analysis speed of Video 1 to the one-click video creation module in the service layer through the application function layer.
[0303] It should be noted that as Figure 15 shown, after the one-click video creation module in the service layer obtains the analysis speed of a material video (such as Video 1), it can continue to execute S1508 - S1511 to obtain the analysis speed of the next material video (such as Video 2), and so on, until the analysis speeds of all material videos are obtained.
[0304] S1512. The one-click video creation module in the service layer calculates the total target analysis duration and the analysis time consumption according to the durations and analysis speeds of all material videos.
[0305] Specifically, the total target analysis duration is the sum of the durations of each highlight segment in all material videos. The duration of a highlight segment is related to the maximum duration and the minimum duration of the highlight segment. Among them, the maximum duration of a highlight segment refers to the expected maximum duration of a highlight segment. The minimum duration of a highlight segment refers to the expected minimum duration of a highlight segment.
[0306] In one embodiment, the duration of a highlight segment can be represented by the following formula (1):
[0307] T 高光 =min(1.2*T 短高光 , T 长高光 ) (1)
[0308] Wherein, T 高光 is the duration of a highlight segment; T 短高光 is the minimum duration of the highlight segment; T 长高光 is the maximum duration of the highlight segment. The min() function is used to obtain the minimum value. That is to say, the minimum value in 1.2*T 短高光 ,T 长高光 is determined as the duration of a highlight segment. The maximum duration of the highlight segment and the minimum duration of the highlight segment can both be set according to the actual situation.
[0309] In one embodiment, the analysis time consumption of a material video is the duration required for the algorithm module to analyze each frame image of the material video. Dividing the duration of a material video by the analysis speed of the material video can obtain the analysis time consumption of the material video.
[0310] S1513. The one - click video composition module at the service layer sends the file descriptors of all material videos, the total target analysis duration, and the analysis time consumption to the policy monitoring module in the media middle - platform framework layer through the application function layer.
[0311] S1514. The policy monitoring module in the media middle - platform framework layer assigns available analysis durations to each material video according to the total target analysis duration and the analysis time consumption.
[0312] In one embodiment, the policy monitoring module can assign available analysis durations to each material video according to the ratio of the analysis time consumption of each material video to the total target analysis duration. For example, the longer the analysis time consumption of a material video, the longer the available analysis duration assigned to this material video. The shorter the analysis time consumption of a material video, the shorter the available analysis duration assigned to this material video. Of course, the policy monitoring module can also use other methods to assign available analysis durations to each video, which is not limited in the embodiments of this application.
[0313] S1515. The policy monitoring module in the media middle - platform framework layer sends the file descriptor, the available analysis duration, and the analysis time consumption of video 1 to the channel interface in the media middle - platform framework layer.
[0314] It should be noted that in this embodiment, for the above - mentioned S1406, the channel interface in the media middle - platform framework layer can send the frame data address of video 1, as well as the available analysis duration and the analysis time consumption of video 1 to the video frame analysis interface in the media middle - platform framework layer, so that the video frame analysis interface can determine the first image frame within the available analysis duration of video 1.
[0315] Exemplarily, if the available analysis duration of a material video is greater than or equal to the analysis time consumption of this material video, then within the available analysis duration assigned to this material video, the first image frame in this material video can be completely analyzed.
[0316] Exemplarily, if the available analysis duration of a material video is less than the analysis time consumption of this material video, then within the available analysis duration assigned to this material video, some of the first image frames in this material video can be analyzed. For example, assume that a material video includes 100 first image frames, then the first 50 first image frames can be analyzed.
[0317] Further, on the basis of any of the above - mentioned embodiments, the electronic device obtains all the fifth video segments. Then, after the above - mentioned S1418, as Figure 16 shown, the video composition method provided by the embodiments of this application may include:
[0318] S1601. The channel interface of the media middle platform framework layer sends the analysis completion message of all material videos to the application function layer through the highlight segment analysis interface of the media middle platform framework layer.
[0319] S1602. The application function layer clips and filters all material videos according to all the analysis completion messages to obtain all the fifth video segments.
[0320] Specifically, first, the application function layer can obtain the start position and end position of each fifth video segment in a material video according to the analysis completion message of the material video. Then, the application function layer clips the material video according to the start position and end position of each fifth video segment, and filters out the fifth video segments in the material video. Based on this, by processing each material video in this way, the application function layer can obtain all the fifth video segments.
[0321] S1603. The application function layer calls the theme summary interface of the media middle platform framework to request to obtain the theme type.
[0322] Among them, the theme summary interface includes a theme algorithm. The request may include the second image analysis information of each image frame in each fifth video segment.
[0323] S1604. The theme summary interface of the media middle platform framework obtains the theme type.
[0324] S1605. The theme summary interface of the media middle platform framework layer sends the theme type to the application function layer.
[0325] S1606. The application function layer obtains the theme template corresponding to the theme type.
[0326] In one embodiment, the application function layer includes a template set, and the template set includes multiple templates. After the application function layer obtains the theme type, it obtains the template corresponding to the theme type from the template set as the theme template.
[0327] S1607. The application function layer sends the theme template and all the fifth video segments to the basic capability layer.
[0328] S1608. The basic capability layer splices all the fifth video segments and applies the theme template to obtain the first finished video.
[0329] Among them, S1608 can refer to the relevant introduction in the above S1104, as well as other relevant introductions in this part above, which will not be elaborated here.
[0330] S1609. The basic capability layer sends an instruction message to play the finished video to the service layer.
[0331] S1610. The business layer plays the first finished video.
[0332] In another scenario, the first video segment involved in the embodiments of the present application is a partial video segment of the candidate videos selected by the user, that is, the first video segment is a partial video segment of the source video. Based on this, after performing the above S1508 on the first video segment and before S1405, the video composition method provided by the embodiments of the present application may include Figure 17 the steps shown in
[0333] Figure 17 FIG. 9 shows a schematic flowchart of a video composition method provided by an embodiment of the present application.
[0334] Specifically, as Figure 17 shown, the video composition method provided by the embodiments of the present application may include:
[0335] S1701. The analysis performance query interface of the media middle platform framework layer queries the estimated analysis duration of Video 1 according to the file descriptor of Video 1.
[0336] Among them, the estimated analysis duration of a source video may include the following situations:
[0337] In one embodiment, the estimated analysis duration of a source video may be the duration corresponding to analyzing 1 video frame in the source video (that is, the minimum analysis duration corresponding to analyzing video frames).
[0338] In another embodiment, the estimated analysis duration of a source video may be the duration corresponding to analyzing all video frames in the source video.
[0339] In another more specific embodiment, the estimated analysis duration of a source video may be the duration corresponding to analyzing M video frames in the source video. Wherein, M is a positive integer greater than 1 and less than the number of all video frames of the source video.
[0340] S1702. The analysis performance query interface of the media middle platform framework layer returns the estimated analysis duration of Video 1 to the business layer through the application function layer.
[0341] S1703. The one-key video composition module of the business layer obtains analysis parameters according to the estimated analysis durations of all source videos and all source pictures.
[0342] Among them, the analysis parameters may include the total analysis duration, the upper limit recommended value of the total analysis duration, the maximum duration of the highlight segment, the minimum duration of the highlight segment, the total recommended duration of the highlight segment, the recommended duration of the highlight segment, whether to force each source video to output a highlight segment, the selected highlight segments, etc.
[0343] Among them, whether to force each material video to output a highlight segment is default to yes. That is to say, in this embodiment, a highlight segment needs to be output for each material video.
[0344] Among them, the total analysis duration represents the sum of the durations required to complete the analysis of all materials selected by the user.
[0345] In one embodiment, the materials only include material pictures. At this time, the processing duration of processing a material picture can be directly determined according to the algorithm ability. The product of the number of material pictures and the processing duration of processing a material picture is used to determine the sum of the estimated analysis durations of all material pictures. At this time, the total analysis duration is the sum of the estimated analysis durations of all material pictures.
[0346] In another embodiment, the materials only include material videos. At this time, the estimated analysis durations of all material videos can be accumulated to obtain the sum of the estimated analysis durations of all material videos. At this time, the total analysis duration is the sum of the estimated analysis durations of all material videos.
[0347] In another embodiment, the materials include material videos and material pictures. The sum of the estimated analysis durations of all material videos and the sum of the estimated analysis durations of all material pictures are accumulated to obtain the total analysis duration of all materials.
[0348] Among them, the recommended value of the total analysis duration upper limit represents the recommended value of the maximum duration for completing the analysis of all material pictures and all material videos. The recommended value of the total analysis duration upper limit can be calculated according to the estimated analysis duration of each material video and the estimated analysis duration of the pictures. Generally, the recommended value of the total analysis duration upper limit is greater than the total analysis duration.
[0349] Among them, the maximum duration of the highlight segment represents the expected maximum duration of a highlight segment. The minimum duration of the highlight segment represents the expected minimum duration of a highlight segment. The maximum duration of the highlight segment and the minimum duration of the highlight segment can be preset values.
[0350] Among them, the total recommended duration of the highlight segment represents the recommended value of the sum of the durations of the highlight segments of all material videos. The total recommended duration of the highlight segment can be determined according to the number of material videos, the maximum duration of the highlight segment, and the minimum duration of the highlight segment.
[0351] Among them, the recommended duration of the highlight segment represents the recommended value of the duration of the highlight segment of a material video. The recommended duration of the highlight segment can be determined according to the maximum duration of the highlight segment and the minimum duration of the highlight segment.
[0352] S1704. The business layer, through the application function layer, sends the file descriptors and analysis parameters of all material videos and material pictures to the policy monitoring module in the media middle platform framework layer.
[0353] It should be noted that when the material does not include material pictures, the content related to material pictures is not included in the above steps.
[0354] Specifically, after obtaining the total analysis duration of all materials, the number of material pictures and the number of material videos to be actually analyzed can be determined according to the total analysis duration and the recommended upper limit value of the total analysis duration.
[0355] S1705. The strategy monitoring module of the media middle platform framework layer determines the first analysis strategy according to the total analysis duration and the recommended upper limit value of the total analysis duration.
[0356] In one embodiment, if the total analysis duration is less than or equal to the recommended upper limit value of the total analysis duration, the first analysis strategy may be to analyze all material pictures and all material videos.
[0357] In another embodiment, if the total analysis duration is greater than the recommended upper limit value of the total analysis duration, the first analysis strategy may be a random sampling strategy. Among them, the random sampling strategy refers to randomly selecting some materials from all materials for analysis.
[0358] The random sampling strategy is related to the category of materials, and the following is introduced by cases:
[0359] In one embodiment, the materials may include material videos and material pictures. At this time, the random sampling strategy may include the following two methods:
[0360] Exemplarily, material pictures take precedence over material videos. At this time, the random sampling strategy may be: select P material pictures and then select Q material videos each time until the total analysis duration of the selected materials reaches the recommended value of the analysis duration upper limit or all material pictures have been taken. Among them, Q < P, and both P and Q are positive integers.
[0361] Exemplarily, material videos take precedence over material pictures. At this time, the random sampling strategy may be: select A material videos and then select B material pictures each time until the total analysis duration of the selected materials reaches the recommended value of the analysis duration upper limit or all material videos have been taken. Among them, B < A, and both A and B are positive integers.
[0362] In another embodiment, the materials may include material videos. At this time, the random sampling strategy may include: randomly selecting C material videos from all material videos.
[0363] In another embodiment, the materials may include material pictures. At this time, the random sampling strategy may include: randomly selecting D material pictures from all material pictures.
[0364] It should be noted that in the embodiments of the present application, the extracted material videos and / or material pictures in this step will be further analyzed, and the material videos and / or material pictures not extracted in this step will not be analyzed again. For the convenience of description, the material videos extracted in this step will be referred to as the first material videos, and the material pictures extracted in this step will be referred to as the first material pictures in the following text.
[0365] When the material includes material pictures and material videos, after the policy monitoring module determines the first analysis policy, the policy monitoring module interacts with the channel interface, the service interface, and the algorithm module to execute the following S1802.
[0366] S1706. Obtain all target pictures from all the first material pictures, and obtain the first video segments from each of the first material videos.
[0367] Among them, the target picture can be the first material picture whose image score is greater than the score threshold. The first video segment is the highlight segment selected from the corresponding first material video. The first video segment may or may not include highlight actions. In the embodiments of the present application, the case where at least one first video segment includes highlight actions is taken as an example for introduction. That is to say, at least one of the first material videos includes highlight actions.
[0368] In one embodiment, when the electronic device (such as a mobile phone) analyzes each of the first material videos, it can first perform an overview analysis on each of the first material videos to obtain an overview analysis result. The overview analysis result may be an overview analysis score, and the overview analysis score is an aesthetic score obtained by performing an aesthetic score on the image frames. Among them, during the overview analysis, an image frame in a first material video can be randomly selected, or image frames of different segments can be extracted according to the storyboard points of the first material video. The number of extracted image frames can represent the number of image frames that can be analyzed within the remaining analysis duration. Through the overview analysis results of each image frame, the target area of the material video is determined. Then, frame-by-frame analysis is performed on the image frames in the target area to obtain the analysis results of each frame of the image frames in the target area. Based on the analysis results of each frame of the image frames in the target area. The analysis result may be the score of each frame of the image frame. Based on the segments covered by the continuous image frames with scores greater than the first threshold, the first video segment of the first material video is obtained.
[0369] It can be seen that in the embodiments of the present application, when extracting the first video segment from the first material video, only the target area of the first material video is analyzed frame by frame, rather than the entire first material video. Based on this, the workload of the electronic device for image frame analysis can be reduced, thereby reducing the time-consuming for the electronic device to implement the one-click video generation function, improving the efficiency of the electronic device in processing material videos, and further improving the user experience of the one-click video generation function.
[0370] After S1706, for each first video segment in each first material video, the following S1707 and the steps after S1405 above are executed.
[0371] S1707. The policy monitoring module of the media middle platform framework layer sends the file descriptor of the first video segment in each first material video to the channel interface of the media middle platform framework layer.
[0372] As Figure 18 shown, an embodiment of the present application also provides a chip system. The chip system 1900 includes at least one processor 1901 and at least one interface circuit 1902. The at least one processor 1901 and the at least one interface circuit 1902 can be interconnected by a line. The processor 1901 is used to support the electronic device to implement each step in the above method embodiments, and the at least one interface circuit 1902 can be used to receive signals from other devices (such as a memory), or send signals to other devices (such as a communication interface). The chip system may include chips and may also include other discrete devices.
[0373] An embodiment of the present application also provides a computer storage medium. The computer storage medium includes instructions that, when running on the above electronic device, cause the electronic device to execute each step in the above method embodiments.
[0374] An embodiment of the present application also provides a computer program product including instructions that, when running on the above electronic device, cause the electronic device to execute each step in the above method embodiments.
[0375] For the technical effects of the chip system, computer storage medium, and computer program product, refer to the technical effects of the foregoing method embodiments.
[0376] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0377] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0378] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0379] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical, or other forms.
[0380] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they can be located in one device or distributed to multiple devices. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0381] In addition, in each embodiment of the present application, the functional modules can be integrated in one device, or each module can exist physically separately, or two or more modules can be integrated in one device.
[0382] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer storage medium or transmitted from one computer storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0383] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for generating a video finished product, characterized in that, Applied to an electronic device, the electronic device includes at least one first video segment, and one or more highlight actions are included in the first video segment; the method includes: Obtain the first video segment, and perform image analysis on the first image frames in the first video segment according to a preset frame skipping strategy to obtain image scores; among them, the image scores of the first image frames including highlight actions are greater than a score threshold; Divide multiple consecutive first image frames in the first video segment whose image scores are greater than the score threshold and belong to the same shot into the same video segment to obtain at least one second video segment; Based on at least one second video segment among all the first video segments, obtain a first assembled video.
2. The method according to claim 1, wherein The step of dividing multiple consecutive first image frames in the first video segment whose image scores are greater than the score threshold and belong to the same shot into the same video segment to obtain at least one second video segment includes: Screen out multiple second image frames with image scores greater than the score threshold from the first image frames in the first video segment; Divide multiple consecutive second image frames among all the screened second image frames into a third video segment to obtain at least one third video segment; Divide the second image frames belonging to the same shot in the third video segment into a second video segment to obtain at least one second video segment.
3. The method according to claim 1 or 2, characterized in that, The step of obtaining a first assembled video based on at least one second video segment among all the first video segments includes: Based on the fact that the duration of the second video segment is greater than a first duration, divide the second video segment into multiple fourth video segments with durations less than the first duration at a first splitting point; among them, the first splitting point includes at least one of a target moment, a first candidate moment, and a second candidate moment; the duration between the target moment and the start moment of the second video segment is an integer multiple of the first duration; the candidate moment is the starting moment of a second duration before a third image frame, or the candidate moment is the ending moment of a second duration after the third image frame; among them, the second duration is less than the first duration, and the third image frame is a second image frame including a highlight action; Screen out fifth video segments with durations greater than or equal to a third duration from all the fourth video segments; among them, the third duration is twice the second duration; Stitch all the fifth video segments to obtain the first assembled video.
4. The method according to claim 3, wherein The fourth video segment does not include a highlight action; the first splitting point is the target moment.
5. The method according to claim 3, characterized in that, The fourth video segment includes a highlight action; the first splitting point includes at least one of the target moment, the first candidate moment, and the second candidate moment.
6. The method according to any one of claims 3 to 5, characterized in that The step of performing image analysis on the first image frames in the first video segment to obtain image scores includes: Obtain first analysis information for each of the first image frames in the first video segment; wherein, the first analysis information includes the brightness, clarity, aesthetics score, and first action information of the corresponding image frame, and the first action information includes a no-action flag or a target action type; wherein, the no-action flag is used to indicate that the first image frame does not include a highlight action, and the target action type is the action type of the highlight action included in the first image frame; In the case where the first action information indicates that the first image frame does not include a highlight action, based on the clarity of the first image frame being within a preset clarity range and the brightness of the first image frame being greater than a brightness threshold, use the aesthetics score of the first image frame as the image score of the first image frame; otherwise, use a third value as the image score of the first image frame; wherein, the score threshold is greater than the third value; In the case where the first action information indicates that the first image frame includes a highlight action, determine that the image score of the first image frame is a second value, and the image scores of all first image frames within a third time period before the first image frame are greater than or equal to the score threshold, and the image scores of all first image frames within the third time period after the first image frame are greater than or equal to the score threshold; wherein, the second value is greater than the score threshold.
7. The method according to claim 6, wherein The obtaining of the first analysis information for each of the first image frames in the first video segment includes: Obtain second analysis information for the first image frame; wherein, the second analysis information includes the brightness, the clarity, the aesthetics score, and second action information, and the second action information includes the no-action flag or at least one predicted action information, and each predicted action information includes a corresponding predicted action type and the confidence level of the predicted action type; Determine the predicted action type with the highest confidence level among each of the predicted action information as the target action type of the predicted action information to obtain the first analysis information.
8. The method according to claim 6 or 7, characterized in that The first analysis information further includes shot information, and the shot information includes a no-shot flag or a shot flag; wherein, the no-shot flag is used to indicate that there is no shot point between the corresponding image frame and the adjacent image frame, and the shot flag is used to indicate that there is a shot point between the corresponding image frame and the adjacent image frame; The dividing of the multiple second image frames belonging to the same shot in the third video segment into one second video segment includes: Based on the shot information indicating that there is a shot point between the second image frame and the adjacent image frame, use the shot point as a first segmentation point to segment the third video segment to obtain the second video segment.
9. The method according to claim 8, wherein The obtaining of the second analysis information for the first image frame includes: Perform image recognition on the first image frame to obtain the image information of the first image frame; wherein, the image recognition includes at least one of clarity recognition, brightness recognition, or aesthetic scoring, as well as action recognition and shot recognition; the image information includes at least one of clarity, brightness, or aesthetic score, as well as the second action information and the shot information; the types of image recognition correspond one-to-one with the image information. In the case where the image recognition includes the action recognition, the clarity recognition, the brightness recognition, the shot recognition, and the aesthetic scoring, the image information of the first image frame is the second analysis information of the first image frame. In the case where the image recognition does not include the clarity recognition, the brightness recognition, or the aesthetic scoring, the second analysis information of the Nth first image frame inherits part of the image information in the second analysis information of the (N - 1)th first image frame; N is a positive integer. Wherein, the part of the image information is obtained by recognizing the first image frame based on some of the recognition types among the clarity recognition, the brightness recognition, and the aesthetic scoring; wherein, the part of the image information includes at least one of the clarity, the brightness, and the aesthetic score.
10. The method according to claim 9, characterized in that The preset frame skipping strategy indicates: starting from the first image frame, extract one image frame as the first image frame from the first video segment every L frames; or, all the image frames in the first video segment are used as the first image frame. The preset frame skipping strategy also indicates: the target recognition type of the first image frame. The target recognition type includes: one or more of the clarity recognition, the brightness recognition, the aesthetic scoring, face recognition, scene recognition, and human body recognition, as well as the action recognition and the shot recognition.
11. The method according to any one of claims 6-10, characterized in that, The splicing of all the fifth video segments to obtain the first finished video includes: Obtain a theme template corresponding to the content in each of the fifth video segments. Splice each of the fifth video segments and apply the theme template to obtain the first finished video.
12. The method according to claim 11, wherein The splicing of each of the fifth video segments and applying the theme template to obtain the first finished video includes: Splice each of the fifth video segments to obtain a second finished video. Perform video speed change processing on the video segments in the second finished video that include the multiple highlight actions to obtain a third finished video. Apply the theme template to the third finished video to obtain the first finished video.
13. The method according to claim 11 or 12, characterized in that The electronic device further includes a target picture; the splicing of each of the fifth video segments includes: Splice each of the fifth video segments and the target picture.
14. An electronic device, characterized in that, Includes one or more processors, as well as a display screen and a memory coupled to the processor; computer program code is stored in the memory, the computer program code includes instructions, and the display screen is used to display a third video; when the processor executes the instructions, the electronic device executes the method according to any one of claims 1 - 13.
15. A computer-readable storage medium, characterized in that, Comprising instructions that, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1-13.
16. A computer program product, characterized in that, Comprising instructions that, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1-13.
Citation Information
Patent Citations
Video editing method and device, equipment and medium
CN115052188A
Video editing method, device and equipment, computer readable storage medium and product
CN115278355A
Video Format for Digital Video Recorder
US20110058792A1