Video generation method and electronic device
By integrating one-click blockbuster-related algorithm modules and a chip computing platform into electronic devices, images and videos are automatically detected and edited to generate video templates that match the theme. This solves the problem of wasted editing time for users and achieves the effect of quickly generating edited videos.
Patent Information
- Application Number
- CN202311682479.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-24
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-02-24
AI Technical Summary
In existing technologies, photos and videos taken by users require manual editing, which is time-consuming and requires editing skills, and cannot meet the needs of rapid sharing.
By integrating one-click blockbuster-related algorithm modules and a chip computing platform into electronic devices, and utilizing the private pathways between the media platform, hardware abstraction layer, and chip computing platform, images and videos are automatically detected and edited to generate video templates that match the theme.
It enables automatic editing of images and videos, meeting users' needs for quickly generating edited videos, improving processing speed, and satisfying personalized requirements.
Smart Images

Figure CN118555444B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed with the Chinese Patent Office with application number 202310209599.9, application date February 24, 2023, entitled "Video Generation Method and Electronic Device". Technical Field
[0002] This application relates to the field of terminal equipment, and more particularly to a video generation method and electronic device. Background Technology
[0003] Currently, the camera function has become one of the most important functions of electronic devices. Using the camera function of electronic devices, people can take pictures anytime, anywhere, providing convenience to their lives.
[0004] People often want to share photos and videos that capture precious moments in life with friends or on social media platforms. However, not all photos and videos taken with the camera function of electronic devices are what users want. Furthermore, users need to manually edit the photos and videos they want to share, which is both time-consuming and requires a certain level of editing skills. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a video generation method and electronic device that can automatically edit images and videos to quickly generate edited videos, meeting users' editing needs.
[0006] Firstly, this application provides a video generation method. This method is applied to an electronic device, which includes a media platform, a Hardware Abstraction Layer (HAL), a one-click blockbuster-related algorithm module located in the HAL, and a chip computing platform. A first private path is established between the media platform and the HAL. The one-click blockbuster-related algorithm module has a first interface connected to the chip computing platform. The method includes: the media platform, in response to receiving a video generation request, displays a first interface showing multiple images and / or videos; the media platform receives a user's selection of a target image and / or target video on the first interface, and sends the target image and / or target video to the one-click blockbuster-related algorithm module through the first private path; the one-click blockbuster-related algorithm module sends the target image and / or target video to the chip computing platform through the first interface; the one-click blockbuster-related algorithm module performs a first preset detection algorithm on each image or video frame in the target image and / or target video. The system performs detection using a first detection method to obtain a first detection result. The chip computing platform then performs detection on each image or video frame in the target image and / or target video using a second preset detection algorithm to obtain a second detection result. The chip computing platform sends the second detection result to the "One-Click Blockbuster" related algorithm module via a first interface. The "One-Click Blockbuster" related algorithm module then sends the first and second detection results to the media platform via a first private path. The "One-Click Blockbuster" related algorithm module performs scene detection on the target image and / or target video, determines the target theme corresponding to the target image and / or target video based on the scene detection result, and sends the target theme to the media platform via the first private path. The media platform then generates and displays a first video based on the target image and / or target video frames whose detection results meet the first condition, according to a target video template that matches the target theme. The detection results include both the first and second detection results. This allows for automatic editing of images and videos, quickly generating edited videos to meet users' editing needs.
[0007] According to the first aspect, before generating and displaying the first video based on a target video template that matches the target theme, the method further includes: the media platform searching for a first video template that matches the target theme from a preset video template library, and using the first video template as the target video template that matches the target theme. The video template library stores the correspondence between video templates and themes. In this way, a video template with the corresponding visual style can be automatically selected according to the target theme.
[0008] According to the first aspect, before generating and displaying the first video using target images and / or target video frames that meet the first condition based on a target video template that matches the target theme, the method further includes: the media platform receiving the user's operation of selecting a second video template from a preset video template library, and using the second video template as the target video template that matches the target theme. In this way, users can choose video templates with their preferred visual style according to their own needs, satisfying their personalized requirements.
[0009] According to the first aspect, before generating and displaying the first video based on the target image and / or target video frame whose detection result meets the first condition, according to the target video template that matches the target theme, the method further includes: the one-click blockbuster related algorithm module determining the first evaluation score of the image or video frame based on the detection result; the one-click blockbuster related algorithm module determining whether the detection result meets the first condition based on the first evaluation score, wherein the first condition is that the image or video frame corresponding to the detection result is a set number of images and / or video frames with the highest first evaluation score.
[0010] According to the first aspect, the first evaluation score of an image or video frame is equal to the weighted sum of the second evaluation scores corresponding to each detection algorithm in the first preset detection algorithm and the second preset detection algorithm.
[0011] According to the first aspect, the first preset detection algorithm and the second preset detection algorithm include a quality detection algorithm and an image perception algorithm.
[0012] According to the first aspect, the quality detection algorithm includes any one or more of the following detection algorithms: jitter detection algorithm; color detection algorithm; brightness detection algorithm; sharpness detection algorithm.
[0013] According to the first aspect, the quality detection algorithm includes any one or more of the following detection algorithms: aesthetic scoring algorithm; action detection algorithm; face detection algorithm; smile detection algorithm; child detection algorithm; scene detection algorithm; and storyboard detection algorithm.
[0014] According to the first aspect, the target theme corresponding to the target image and / or target video is determined based on the scene detection results, including: the one-click blockbuster related algorithm module performs theme inference based on the scene detection results to obtain the target theme corresponding to the target image and / or target video.
[0015] According to the first aspect, in response to receiving a video generation request, a first interface is displayed, on which multiple images and / or videos are displayed, including: receiving a user's selection operation of a first function option in the gallery, generating a video generation request; and displaying the first interface, which is an interface displaying a list of images and / or videos in the gallery.
[0016] In a second aspect, this application provides an electronic device, including: a memory and a processor, wherein the memory is coupled to the processor; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device performs the video generation method of any one of the first aspects.
[0017] Thirdly, this application provides a computer-readable storage medium including a computer program that, when run on an electronic device, causes the electronic device to perform the video generation method described in any of the first aspects above. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the structure of an electronic device 100 as an example.
[0019] Figure 2 This is a software structure block diagram of an electronic device 100 according to an embodiment of this application, which is an example shown.
[0020] Figure 3 This is a timing diagram illustrating the video generation method in this embodiment as an example.
[0021] Figure 4 This is a schematic diagram illustrating the processing of images and video frames in the video generation process of this embodiment, which is an example of such a process.
[0022] Figure 5 This is a schematic flowchart illustrating the video generation method in this embodiment as an example. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0025] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.
[0026] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0027] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.
[0028] The video generation method in this embodiment can be applied to electronic devices such as mobile phones and tablets. The structure of the electronic device can be as follows: Figure 1 As shown.
[0029] Figure 1 This is a schematic diagram illustrating the structure of an electronic device 100 as an example. It should be understood that... Figure 1 The electronic device 100 shown is merely an example of an electronic device, and the electronic device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. Figure 1 The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0030] Please see Figure 1 The electronic device 100 may include: a processor 110, internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, an indicator 192, a camera 193, etc.
[0031] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0032] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0033] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory.
[0034] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0035] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0036] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0037] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0038] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0039] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0040] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0041] The software system of the electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to illustrate the software structure of the electronic device 100.
[0042] Figure 2 The above is a software structure block diagram of an electronic device 100 according to an embodiment of this application.
[0043] The layered architecture of the electronic device 100 divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system may include an application layer, an application framework layer, system libraries, and a kernel layer, etc.
[0044] The application layer can include a series of application packages.
[0045] like Figure 2 As shown, the application package can include applications such as camera, calendar, SMS, gallery, call, and video.
[0046] The gallery application includes a video editing app, which in turn includes a media platform module. The video editing app, the media platform module, and the subsequent one-click blockbuster-related algorithm module and chip computing platform module in the hardware abstraction layer are used to implement the video generation method of this embodiment. For detailed descriptions of the functions of these modules, please refer to the subsequent embodiments in this document.
[0047] like Figure 2 As shown, the application framework layer may include a window manager, content provider, resource manager, phone manager, view system, etc.
[0048] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0049] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0050] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0051] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).
[0052] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0053] The Android Runtime consists of core libraries and a virtual machine. The Android Runtime is responsible for scheduling and managing the Android system.
[0054] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0055] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0056] like Figure 2 As shown, the hardware abstraction layer may include a one-click blockbuster related algorithm module and a chip computing platform module.
[0057] The kernel layer is the layer between hardware and software.
[0058] like Figure 2 As shown, the kernel layer can include modules such as audio driver, display driver, Bluetooth driver, camera driver, and sensor driver.
[0059] Understandable, Figure 2 The layers in the illustrated software structure and the components contained in each layer do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer layers than illustrated, and each layer may include more or fewer components; this application does not impose any limitations.
[0060] The following example uses a mobile phone as an electronic device, combined with... Figure 3 , Figure 4 and Figure 5 The embodiments of this application will be described in detail. It will be understood that the following examples are also applicable to scenarios of other types of electronic devices (e.g., tablets).
[0061] Figure 3 This is a timing diagram illustrating the video generation method in this embodiment as an example. Figure 4 This is a schematic diagram illustrating the processing of images and video frames in the video generation process of this embodiment, which is an example of such a process. Figure 5 This is a schematic flowchart illustrating the video generation method in this embodiment as an example.
[0062] Please see Figure 3 In this embodiment:
[0063] The media platform runs within the video editing app;
[0064] A private path (referred to as CameraResource in this article) is set up between HAL and the media middle platform. Using the private interface, the chip computing platform, which is independent of the camera system architecture, implements the upper-level capability interface (init, process, deinit, registerListener, unregisterListener).
[0065] HAL integrates algorithm pathways and adds new algorithm interfaces (init, process, deinit). The process interface can include functions such as setOpt, imgAnalyse, Framepush, Reset, getTheme, Stop, getPerformance, and calFramesSameShotProb. These interfaces are used for communication with the chip computing platform. These new algorithm interfaces can be set in algorithms related to one-click large-scale data processing. In this way, this embodiment can utilize the chip computing platform to execute algorithms, significantly improving processing speed, as the chip computing platform's processing speed is much faster than the software processing speed within HAL.
[0066] The functions of each interface can be summarized as follows:
[0067] setOpt is used to set video analysis parameters;
[0068] imgAnalyse is used for image / single-frame analysis and scoring, outputting the location and number of faces, whether there are smiling faces, scene category, etc.
[0069] Framepush is used to send frame information from images and videos to the algorithm for processing;
[0070] Reset is used to reset the video scoring algorithm and clear the algorithm's historical state. It is called when the user uses the one-click video creation function again without exiting the function.
[0071] getTheme is used to process all user-uploaded videos or images using algorithms and then report the current topic.
[0072] The Stop function is used to forcibly stop the current calculation and return the current calculation result.
[0073] Getperformance is used to obtain the processing speed of the single-frame / multi-frame analysis interface;
[0074] calFramesSameShotProb is used to calculate whether images are the same and to analyze the confidence level of the images.
[0075] Please see Figure 4 The overall video generation process in this embodiment may include:
[0076] The first step involves the user inputting multiple videos and images. If it's a video, frame extraction is also required.
[0077] The second step involves performing basic quality checks (e.g., jitter detection, color detection, brightness detection, and sharpness detection) and image perception algorithm analysis (e.g., aesthetic scoring, motion detection, face detection, smile detection, child detection, scene detection, and scene segmentation detection) on the frames (i.e., the extracted video frames) or images. Figure 4 Not shown in the image. Figure 4 Image perception in the process can be further included in segment detection.
[0078] The third step is to score each frame based on the basic quality detection results and the image perception algorithm results.
[0079] The fourth step is to select highlight segments based on the constraints, and finally output the highlight segments.
[0080] Fifth, unlike videos, images do not undergo shake detection, motion detection, or scene segmentation detection.
[0081] The sixth step is to perform topic inference based on the scene detection results of video frames and images to obtain the topics of all inputs.
[0082] Then, you can select the corresponding video template based on the theme and generate the video. Figure 4 In the "topic" processing following topic reasoning, the corresponding video template can be selected based on the topic.
[0083] For example, in one instance, the media platform can have a pre-set video template library, which stores the correspondence between video templates and themes. The media platform can search for a video template from the pre-set video template library that matches the inferred theme, and use that first video template as the target video template that conforms to the inferred theme.
[0084] For example, in another example, a user can choose a video template from a preset video template library. The media platform receives the user's selection and uses it as the target video template that matches the inferred theme. Furthermore, after automatically selecting a target video template based on the previous example, the user can also manually change the video template.
[0085] Different video templates offer different visual styles. By selecting a video template based on the theme, you can adjust the visual style (e.g., transitions, filters). Visual styles can include effects, music, and more.
[0086] The video generation method of this embodiment will be further described in detail below through an example. It should be noted that, in the following... Figure 5 In the description, HAL communicates with the media platform through the aforementioned private channel, such as by transmitting messages. The algorithm modules related to the one-click blockbuster video service communicate with the chip computing platform through the aforementioned newly added algorithm interface, also for the same purpose. The chip computing platform is a hardware chip within an electronic device.
[0087] In this embodiment, the algorithm module related to the one-click blockbuster video executes the algorithm with a smaller data volume, while the chip computing platform executes the algorithm with a larger data volume. This leverages the powerful computing capabilities of the chip computing platform to improve the overall speed of the video generation process. Of course, Figure 5 The allocation of algorithms in the One-Click Blockbuster related algorithm module and the chip computing platform is an illustrative example and is not intended to limit this embodiment. In other embodiments, other allocation methods may be adopted, for example, all algorithms may be executed by the chip computing platform.
[0088] Please see Figure 5 In this embodiment, the video generation method may include the following steps:
[0089] S501, Users select the One-Click Blockbuster feature in the Gallery application.
[0090] When an electronic device detects that a user has selected the "One-Click Movie" feature in the gallery application, the media platform considers that it has received a video generation request.
[0091] Upon receiving a video generation request, the various modules used to execute the video generation method are initialized to request resources.
[0092] Furthermore, upon receiving a video generation request, the electronic device will switch from the interface displaying the "One-Click Blockbuster" function options to the photo interface, which displays images, videos, etc. The photo interface can be displayed after step S504 is completed.
[0093] S502, the media platform is initialized.
[0094] This initialization step allows you to allocate memory resources for the media platform.
[0095] S503 and H AL are initialized.
[0096] This initialization step allocates memory resources for the HAL.
[0097] S504, the algorithm modules related to one-click blockbuster movies are initialized.
[0098] This initialization step allocates memory resources for the algorithm modules related to the one-click blockbuster feature.
[0099] S505: The user selects image a and video b from the images and videos in the gallery.
[0100] The media platform receives the selection operation for image a and video b, and confirms that a VLOG will be generated based on image a and video b.
[0101] This step corresponds to Figure 3 The timing sequence ①.
[0102] It should be noted that, Figure 3 The time sequences ① to ⑤ in the diagram correspond to the processing of one image or one video frame. In this embodiment, the processing of image a is taken as an example. Figure 3 The timing sequence ① to ⑤ in the diagram will be explained. The timing sequence of the video frame processing is the same as that of the image processing, and will not be repeated in this embodiment.
[0103] S506, the media middle platform decodes image a and sends the decoded image a to HAL.
[0104] In the steps that follow this step, "image a" refers to the decoded image a.
[0105] S507 and HAL send image 'a' to the relevant algorithm module for one-click blockbuster videos.
[0106] S5081 and the related algorithm module for one-click blockbuster movies perform brightness detection, color detection and aesthetic analysis on image a to obtain the first detection result.
[0107] Steps S506 to S5081 correspond to Figure 3 The timing sequence in ②.
[0108] Please see Figure 4 The detection of images and video frames can include quality detection and image perception.
[0109] Quality inspection may include any one or more of the following tests:
[0110] Jitter detection, color detection, brightness detection, and sharpness detection are all performed according to their respective detection algorithms. For example, jitter detection is performed using a jitter detection algorithm, color detection using a color detection algorithm, brightness detection using a brightness detection algorithm, and sharpness detection using a sharpness detection algorithm.
[0111] Image perception can include any one or more of the following:
[0112] Aesthetic scoring, motion detection, face detection, smile detection, child detection, scene detection, and scene detection are represented by the corresponding algorithms: aesthetic scoring algorithm, motion detection algorithm, face detection algorithm, smile detection algorithm, child detection algorithm, scene detection algorithm, and scene detection algorithm.
[0113] It should be noted that, although Figure 4 Various quality detection algorithms and image perception algorithms are listed; however, image or video frames do not need to undergo [further processing / processing]. Figure 4 All quality inspections and image perception listed in the document.
[0114] For example, in this embodiment, image quality detection does not include jitter detection. Similarly, image perception does not include motion detection.
[0115] In this embodiment, when performing quality inspection, the image perception of video frames extracted from the video does not include scene detection.
[0116] S5082, the one-click blockbuster related algorithm module sends a calculation request to the chip computing platform to perform sharpness detection, face detection, smile detection, child detection, and scene detection on image a. This request may include image a.
[0117] Step S5082 corresponds to Figure 3 The timing sequence in ③.
[0118] S5083, the chip computing platform returns the second detection result of image a to the one-click blockbuster related algorithm module, including sharpness detection, face detection, smile detection, child detection, and scene detection.
[0119] Step S5083 corresponds to Figure 3 The timing sequence in the text ④.
[0120] S509, the one-click blockbuster related algorithm module returns all detection results of image a to HAL.
[0121] Steps S509 to S510 correspond to Figure 3 The timing sequence in the text is ⑤.
[0122] In this step, the detection results returned to HAL by the one-click blockbuster related algorithm module include the detection results of brightness detection, color detection, aesthetic analysis, sharpness detection, face detection, smile detection, child detection, and scene detection of image a, which is the combination of the aforementioned first detection results and second detection results.
[0123] S510 and HAL return all detection results for image a to the media platform.
[0124] S511, the media middleware decodes video b and sends the decoded video b to HAL.
[0125] S512 and HAL extract video frames from video b and send the extracted video frames to the relevant algorithm module for one-click blockbuster.
[0126] For each of the extracted video frames, the processing steps 5131-51331 can be performed separately.
[0127] The S5131 and One-Click Blockbuster related algorithm modules perform brightness detection, color detection, aesthetic analysis, and jitter detection on video frames to obtain the third detection result.
[0128] S5132, the One-Click Blockbuster related algorithm module sends a calculation request to the chip computing platform to perform clarity detection, face detection, smile detection, child detection, scene detection, and motion detection on video frames.
[0129] S5133, the chip computing platform returns the fourth detection result of the video frame to the one-click blockbuster related algorithm module, including sharpness detection, face detection, smile detection, child detection, scene detection, and motion detection.
[0130] S514, the One-Click Blockbuster related algorithm module returns all detection results of video frames to HAL.
[0131] In this step, the detection results returned to HAL by the one-click blockbuster related algorithm module include the detection results of brightness detection, color detection, aesthetic analysis, jitter detection, sharpness detection, face detection, smile detection, child detection, scene detection, and motion detection of video frames, which is the combination of the aforementioned third and fourth detection results.
[0132] S515 and HAL return all detection results of video frames to the media center.
[0133] The S516 and One-Click Blockbuster related algorithm modules score all processed images and video frames and determine highlight segments based on their scores.
[0134] This step corresponds to Figure 4 The processing includes "frame scoring" and "highlight segment selection". It should be noted that... Figure 4 The term "frame scoring" in this context refers not only to scoring video frames but can also refer to scoring images.
[0135] The scoring criteria can be pre-set. For example, each test result can be scored, and the sum or average of the scores of all test results can be used as the final score.
[0136] Highlight segments refer to images or video frames with high scores. For example, a set number of images and / or video frames with the highest scores can be selected as highlight segments. For example, the set number can be 1 or greater than 1.
[0137] S517, the one-click blockbuster related algorithm module returns highlight segments to HAL.
[0138] S518 and HAL returned highlight footage to the media center.
[0139] S519, The media platform sends a request to HAL to retrieve the topic.
[0140] S520 and HAL forward requests to the relevant algorithm modules for obtaining topics to the One-Click Blockbuster.
[0141] S521, the algorithm module related to one-click blockbuster movies votes based on the scene detection results of each obtained image and video frame.
[0142] Here, voting means determining the theme of image a and video b based on the scene detection results.
[0143] S522, the algorithm module related to one-click blockbuster movies returns the topic to HAL.
[0144] S523 and HAL return to the topic at the media center.
[0145] The above steps S519 to S523 correspond to Figure 4 The "topic reasoning" process in the text.
[0146] In this way, after acquiring the highlight clips and themes, the media platform can generate video logs (VLOGs) according to the target video templates that match the themes, and display the VLOGs on electronic device screens.
[0147] The target video templates can be obtained in the following ways: either through automatic recommendation or through user selection.
[0148] In this embodiment, video templates can be preset and stored in a video template library, which stores the correspondence between video templates and themes.
[0149] In this way, once a theme is obtained, the system can search for the first video template that matches the theme from a pre-set video template library, and use it as the target video template that matches the obtained theme. Therefore, this embodiment can automatically recommend video templates that match the content based on the scene theme (people, scenery, food, children, pets, sports, travel).
[0150] In this embodiment, the internal interface after entering the one-click blockbuster can display the video template function option. Users can also actively select their favorite video template from the video template library through this function option and use the selected template as the target video template that matches the obtained theme.
[0151] S524. The media platform will generate VLOGs from the highlight clips according to the target video template that matches the theme and then display them.
[0152] S525, the media platform sends a request to HAL to reset parameters.
[0153] The parameter reset step is used to clear the parameters from the previous steps, preparing for the next execution of the video generation method.
[0154] S526 and HAL forward the request to reset parameters to the relevant algorithm modules for one-click blockbuster movies.
[0155] S5271, the algorithm module related to one-click blockbuster movies resets its own global variables.
[0156] S5272, One-Click Blockbuster related algorithm modules reset global variables of the chip computing platform.
[0157] S528, When the user executes the operation to exit the one-click movie, the media platform receives the instruction to exit the one-click movie.
[0158] S529, The media platform will send the command to exit the one-click movie feature to HAL.
[0159] Upon receiving the instruction to exit the one-click movie feature, S530 and HAL send a destruction instruction to the relevant algorithm modules of the one-click movie feature.
[0160] S5311, the algorithm module related to one-click blockbuster movies destroys related data.
[0161] Here, the data to be destroyed could be, for example, images or videos from the aforementioned steps, in preparation for the next execution of the video generation method.
[0162] S5312, the algorithm module related to one-click blockbuster movies destroys the relevant data in the chip computing platform.
[0163] Here, the data to be destroyed could be, for example, images or videos from the aforementioned steps, in preparation for the next execution of the video generation method.
[0164] The video generation method of this embodiment can be applied to various editing needs of users. For example, in the following scenarios:
[0165] Scenario 1: After recording a video longer than 1 minute, the user wants to cut out the highlights and add effects and music to generate a short video.
[0166] Scenario 2: Users select multiple images and videos to generate a short video with music, a title, and special effects.
[0167] The video generation method of this embodiment can automatically edit images and videos to meet users' editing needs. Furthermore, this embodiment can leverage the computing power of the chip computing platform to significantly improve processing speed, enabling rapid generation of edited videos and providing users with a better user experience.
[0168] This application also provides an electronic device, which includes a memory and a processor. The memory is coupled to the processor and stores program instructions. When the program instructions are executed by the processor, the electronic device performs a video generation method as described above.
[0169] It is understood that, in order to achieve the above-mentioned functions, electronic devices include hardware and / or software modules that perform the respective functions. Based on the algorithmic steps of the examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of this application.
[0170] This embodiment also provides a computer storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the video generation method described above.
[0171] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the video generation method described above.
[0172] In addition, this application also provides an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory. The memory is used to store computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory to cause the chip to execute the video generation methods in the above method embodiments.
[0173] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.
[0174] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0175] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0176] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0177] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0178] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.
[0179] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0180] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0181] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.
[0182] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0183] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A video generation method, characterized in that, Applied to electronic devices, the electronic devices including a Hardware Abstraction Layer (HAL), a one-click blockbuster related algorithm module located in the HAL, and a chip computing platform, the method includes: In response to the received first operation, a first interface is displayed, which displays multiple images and / or videos. The first operation is the operation of selecting the one-click blockbuster function. In response to the received second operation on the first interface, select the target image and / or the target video; The one-click blockbuster related algorithm module is invoked to detect each image or video frame in the target image and / or the target video according to the first preset detection algorithm, and a first detection result is obtained; The computing platform is invoked to detect each image or video frame in the target image and / or the target video according to the second preset detection algorithm to obtain a second detection result. The data volume of the second preset detection algorithm is greater than the data volume of the first preset detection algorithm. The second detection result includes scene detection results. Based on the scene detection results, the target theme corresponding to the target video frame in the target image and / or the target video is determined; If the first detection result and the second detection result meet the first condition, the first video is generated from the target image and / or the target video frame according to the target video template corresponding to the target theme.
2. The method according to claim 1, characterized in that, The electronic device also includes a media platform, and a first privatization path is provided between the media platform and the HAL. Before invoking the one-click blockbuster related algorithm module to detect each image or video frame in the target image and / or the target video according to the first preset detection algorithm and obtaining the first detection result, the method further includes: The media platform sends the target image and / or the target video to the one-click blockbuster related algorithm module through the first private channel; The one-click blockbuster related algorithm module sends the target image and / or the target video to the chip computing platform; After invoking the computing platform to detect each image or video frame in the target image and / or the target video according to the second preset detection algorithm and obtaining the second detection result, the method further includes: The chip computing platform sends the second detection result to the one-click blockbuster related algorithm module; The one-click blockbuster related algorithm module sends the first detection result and the second detection result to the media platform through the first private channel; When the first detection result and the second detection result meet the first condition, generating a first video from the target image and / or the target video frame according to the target video template corresponding to the target theme includes: If the first detection result and the second detection result meet the first condition, the media platform generates a first video from the target image and / or the target video frame according to the target video template corresponding to the target theme.
3. The method according to claim 2, characterized in that, The one-click blockbuster related algorithm module interacts with the chip computing platform through the first interface; The one-click blockbuster related algorithm module sends the target image and / or the target video to the chip computing platform, including: The one-click blockbuster related algorithm module sends the target image and / or the target video to the chip computing platform through the first interface; The chip computing platform sends the second detection result to the one-click blockbuster related algorithm module, including: The chip computing platform sends the second detection result to the one-click blockbuster related algorithm module through the first interface.
4. The method according to claim 1, characterized in that, When the first detection result and the second detection result meet the first condition, generating a first video from the target image and / or the target video frame according to the target video template corresponding to the target theme includes: From the preset video template library, find the first video template that matches the target theme; According to the first video template, the target image and / or the target video frame whose detection results meet the first condition are used to generate a first video.
5. The method according to claim 1, characterized in that, When the first detection result and the second detection result meet the first condition, generating a first video from the target image and / or the target video frame according to the target video template corresponding to the target theme includes: In response to the user's third action in the preset video template library, select the second video template; The second video template is determined as the target video template corresponding to the target theme; According to the target video template, the target images and / or target video frames whose detection results meet the first condition are used to generate the first video.
6. The method according to claim 1, characterized in that, When the first detection result and the second detection result meet the first condition, generating a first video from the target image and / or the target video frame according to the target video template corresponding to the target theme includes: Based on the first detection result and the second detection result of each image or video frame in the target image and / or the target video, a first evaluation score is determined for each image or video frame. Based on the first evaluation score, determine whether the detection result meets the first condition, wherein the first condition is that the image or video frame corresponding to the detection result is a set number of images and / or video frames with the highest first evaluation score. The first video is generated from the target images and / or target video frames whose detection results meet the first condition.
7. The method according to claim 6, characterized in that, The first evaluation score of an image or video frame is equal to the weighted sum of the second evaluation scores corresponding to each detection algorithm in the first preset detection algorithm and the second preset detection algorithm.
8. The method according to claim 1, characterized in that, When the second operation selects a target image, the first preset detection algorithm includes any one or more of the following algorithms: Color detection algorithm; Brightness detection algorithm; Aesthetic scoring algorithm; The second preset detection algorithm includes any one or more of the following algorithms: Face detection algorithm; Smile detection algorithm; Child detection algorithm; Scene detection algorithm.
9. The method according to claim 1, characterized in that, When the second operation selects a target video, the first preset detection algorithm includes any one or more of the following algorithms: Color detection algorithm; Brightness detection algorithm; Aesthetic scoring algorithm; jitter detection algorithm; The second preset detection algorithm includes any one or more of the following algorithms: Face detection algorithm; Smile detection algorithm; Child detection algorithm; Scene detection algorithm; Action detection algorithm; Storyboard detection algorithm.
10. The method according to claim 4, characterized in that, The step of determining the target theme corresponding to the target image and / or the target video based on the scene detection results includes: Based on the scene detection results, topic inference is performed to obtain the target topic corresponding to the target image and / or the target video.
11. The method according to claim 1, characterized in that, In response to the received first operation, a first interface is displayed, on which multiple images and / or videos are displayed, including: Receive the user's first operation on the first function option in the gallery, and generate a video generation request; Based on the video generation request, a first interface is displayed, which is the interface that displays a list of pictures and / or videos in the gallery.
12. An electronic device, characterized in that, include: A memory and a processor, wherein the memory is coupled to the processor; The memory stores program instructions that, when executed by the processor, cause the electronic device to perform the video generation method as described in any one of claims 1 to 11.
13. A computer-readable storage medium comprising a computer program, characterized in that, When the computer program is run on an electronic device, the electronic device causes the electronic device to perform the video generation method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Method for quickly generating short video based on template
CN112732977A
Video clip selector used in image creation and diagnosis for medial care
JP2020024478A