Video generation method, electronic device, and computer readable storage medium
By using the event camera to detect the motion details of the target object and insert it into the image captured by the first camera, the problem that electronic devices are difficult to capture slow motion details is solved, and high-quality video generation is achieved and user experience is improved.
Patent Information
- Application Number
- PCT/CN2024/110742
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-08-08
- Publication Date
- 2025-05-22
AI Technical Summary
When shooting moving target objects, existing electronic devices are difficult to capture the slow motion details of the target objects, resulting in poor continuity, poor fluency and poor user experience.
The event camera is used to detect the pixel light brightness changes caused by the movement of the target object, obtain the slow motion detail images during the movement of the multi-frame target object, and insert it between the RGB images collected by the first camera to generate high-quality video.
It improves the continuity and fluency of video generation, improves the picture quality of the video and the user's visual experience.
Smart Images

Figure CN2024110742_22052025_PF_FP_ABST
Abstract
Description
Video generation method, electronic device, and computer-readable storage medium
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 13, 2023, with application number 202311508927.1 and invention name “A video generation method, electronic device and computer-readable storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of image processing, and in particular to a video generation method, an electronic device, and a computer-readable storage medium. Background Art
[0003] As camera functionality in electronic devices develops, it becomes increasingly powerful. As you can understand, videos are generally composed of multiple frames, and the content presented in videos can enhance the user's viewing experience. Beyond simple camera functions like taking photos and recording videos, electronic devices now also have the ability to create videos from captured photos. For example, electronic devices can create videos of wonderful moments captured. This eliminates the need for users to spend time and effort editing videos themselves, while allowing them to immersively recall the moments.
[0004] However, currently, when electronic devices capture photos, due to the relative motion between the target object and the device, factors such as camera exposure time and sampling frequency often result in the capture of discrete, continuous moments of the target object's non-action. Furthermore, electronic devices can generate a continuous video of multiple discrete moments—that is, moments with significantly different image content—from one another. This video generation method results in poor continuity and smoothness, resulting in a poor user experience.
[0005] For example, Figure 1 shows a schematic diagram of two frames of images captured by an electronic device. The electronic device captures a scene of user A jumping over a hurdle. During the capture process, the electronic device uses a camera to capture frame A and frame C. Frame A shows user A just beginning to jump over the hurdle, while frame C shows the user completing the jump. The electronic device can generate video 1 from frame A and frame B.
[0006] In the example shown in Figure 2, a schematic diagram of the interface display during the playback of Video 1 is shown. When an electronic device begins playing Video 1, it can display all images in the order in which they appear in Video 1. For example, the electronic device can display interface 111, which includes Frame A, and displays Frame A from the 1st second to the 5th second. At the 5th second, the electronic device can display interface 112, which includes Frame C, and displays Frame C from the 5th second to the 15th second. Thus, during the playback of Video 1, after Frame A shows User A just beginning to jump over a hurdle, Frame C shows the user having completed the jump and running towards their destination. Because Frame A and Frame C represent discrete moments, the user's movements presented in these two images are not a relatively continuous process. The slow-motion details of User A's multiple hurdle jumps are missing, resulting in poor continuity and fluidity, and a poor user experience.
[0007] In addition, some existing solutions use algorithms to generate content between slow-motion video frames, which also have problems such as low image quality, many pseudo-textures, poor continuity, poor smoothness, and poor user experience.
[0008] Summary of the Invention
[0009] The embodiments of the present application provide a video generation method, an electronic device, and a computer-readable storage medium for improving the continuity and smoothness of video generation, thereby enhancing the user experience.
[0010] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0011] In a first aspect, a video generation method is provided. The method is applied to an electronic device, the electronic device including an event camera and a first camera. The method comprises: in response to a user operation of turning on a camera function, turning on the first camera and displaying a first interface; wherein the first interface includes a preview window for displaying a preview stream image, the preview stream image being an image captured in real time by the first camera; when an automatic capture mode in the camera function is turned on, turning on the event camera; wherein the event camera is used to capture and obtain event camera stream information; if it is detected based on the preview stream information corresponding to the preview stream image and the event camera stream information that a target object satisfies a preset motion posture, determining N highlight frames for generating a video based on the preview stream information; wherein the N highlight frames include a start frame and an end frame of the video to be generated, and N is an integer greater than 1; obtaining event camera stream information between the capture time of the start frame and the capture time of the end frame, and determining, based on the event camera stream information, a first image to be inserted between at least one set of adjacent two highlight frames in the N highlight frames; wherein the first image includes slow-motion details of the target object between at least one set of adjacent two highlight frames in the N highlight frames; and generating a target highlight video based on the N highlight frames and the first image.
[0012] In the present application, when the electronic device uses the first camera to shoot the target object, if the target object is in motion, the brightness change of the pixel points of the target object detected by the event camera due to the motion can be used to obtain multiple frames of slow-motion detail images of the target object during motion. Then, the electronic device inserts the multiple frames of slow-motion detail images of the target object during motion between the multiple RGB images collected by the first camera to generate a video. The video generated according to the above scheme has higher image quality, less pseudo-texture, better continuity, and higher smoothness. And the electronic device can play each frame of the image according to the arranged image, which is convenient for users to watch the wonderful moments of slow motion at a slow speed, and the user's viewing experience is better.
[0013] In a possible implementation manner of the first aspect, the number of first images inserted between at least two adjacent groups of highlight frames in the N highlight frames is the same.
[0014] That is, the number of first images inserted between every two adjacent wonderful frames in N frames is the same, or the number of first images inserted between not all groups of adjacent two wonderful frames in N frames is the same, but the number of first images inserted between two or more groups of adjacent two wonderful frames in N frames is the same.
[0015] In a possible implementation manner of the first aspect, the numbers of first images inserted between at least two adjacent groups of highlight frames in the N highlight frames are different.
[0016] That is, the number of first images inserted between each two adjacent highlight frames in N highlight frames is different. The number of first images inserted between not all groups of adjacent highlight frames in N highlight frames is different, but rather the number of first images inserted between two or more groups of adjacent highlight frames in N highlight frames is different. For example, the number of first images inserted between the first group of two adjacent highlight frames in N highlight frames is 2, the number of first images inserted between the second group of two adjacent highlight frames in N highlight frames is 2, and the number of first images inserted between the third group of two adjacent highlight frames in N highlight frames is 0, meaning no images are inserted between the third group of two adjacent highlight frames in N highlight frames. For another example, the number of first images inserted between the first group of two adjacent highlight frames in N highlight frames is 3, the number of first images inserted between the second group of two adjacent highlight frames in N highlight frames is 2, and the number of first images inserted between the third group of two adjacent highlight frames in N highlight frames is 2, meaning no images are inserted between the third group of two adjacent highlight frames in N highlight frames.
[0017] In a possible implementation of the first aspect, determining, based on event camera stream information, a first image to be inserted between at least one group of adjacent two highlight frames in N highlight frames includes:
[0018] The number of first images inserted between at least one group of adjacent two wonderful frames in the N wonderful frames is determined based on the motion speed of the target object between two adjacent wonderful frames in the N wonderful frames.
[0019] In the present application, the speed at which the target object moves at different moments may be different. When the target object moves faster, within a certain period of time, the target object has more movements, the richness of the movements is higher, and these movements are difficult to be captured clearly by the first camera. Therefore, when generating a video, the electronic device can compensate for the movement details of the target object that are not captured by the first camera based on the event stream information captured by the event camera. Specifically, the electronic device can insert more images when the target object moves faster according to the movement speed of the target object. In this way, the continuity and smoothness of the video generation can be improved, thereby enhancing the user experience.
[0020] When the target object moves slowly, within a certain period of time, the target object's movements are relatively few and the richness of the movements is low, and these movements can be easily captured more clearly by the first camera. Therefore, when generating a video, the electronic device can basically obtain the target object's movement details without compensating for them. Specifically, the electronic device can insert fewer images when the target object moves slowly, depending on the target object's movement speed. In this way, a video with higher continuity and smoothness can be generated, and the power consumption of the electronic device when inserting images can be reduced, thereby improving the user experience.
[0021] In a possible implementation of the first aspect, the preview stream image is at a first resolution, and the N highlight frames of the generated video are at a second resolution; wherein the second resolution is greater than the first resolution.
[0022] That is, the resolution of the image displayed on the preview interface of the electronic device is lower than the resolution of the image of the video to be generated. When the electronic device determines the wonderful frame, it generates an image with a higher resolution, which can improve the user's viewing experience.
[0023] In a possible implementation of the first aspect, determining a first image to be inserted between at least one group of adjacent two highlight frames in N highlight frames based on event camera stream information includes: determining motion information based on the event camera stream information, and determining the first image to be inserted between at least one group of adjacent two highlight frames in the N highlight frames based on the motion information; wherein the motion information is used to represent the motion of a moving object captured by the event camera, and the moving object includes a target object.
[0024] The motion of a moving object includes the motion of the target object and the motion of objects in the target object's background. Based on the motion of the moving object captured by the event camera, the electronic device can accurately obtain a clearer image of the target object even when the target object is in motion. Because the generated inset image is clearer, the electronic device can produce higher-quality video, presenting a higher-quality video to the user and providing a better user experience.
[0025] In a possible implementation manner of the first aspect, the motion information includes optical flow information, where the optical flow information is used to represent changes in pixel brightness in an image captured by the event camera.
[0026] Optical flow refers to the change in pixel brightness during the capture time caused by light obstruction or other factors in the object's surroundings. Based on this change in pixel brightness, electronic devices can accurately and clearly capture images of the target object's motion. Because the resulting interpolated image is clearer, electronic devices can produce higher-quality videos, presenting higher-quality videos to users and providing a superior user experience.
[0027] In a possible implementation of the first aspect, determining motion information based on event camera stream information includes: denoising the event camera stream information using a preset denoising algorithm to obtain first event camera stream information; and determining motion information based on the first event camera stream information.
[0028] By denoising the event camera stream, the electronic device can accurately determine the target object's motion and generate a clearer image to be inserted. This clearer image allows the electronic device to produce a higher-quality video, presenting a higher-quality video to the user and providing a better user experience.
[0029] In a possible implementation of the first aspect, determining a first image to be inserted between two adjacent wonderful frames in N wonderful frames based on motion information includes: the two adjacent wonderful frames include the i-th wonderful frame and the i+1-th wonderful frame, and inserting M frames of images between the i-th wonderful frame and the i+1-th wonderful frame; wherein i takes an integer value from 1 to N-1; for the n-th frame of the M frames, wherein n takes an integer value from 1 to H in sequence: determining first motion difference information based on the i-th frame motion information corresponding to the shooting moment of the i-th wonderful frame and the n-th frame motion information of the event camera corresponding to the n-th frame of the image at the shooting moment; wherein the first motion difference information represents the n-th frame of the M frames of images. The change between the i-th frame motion information and the n-th frame motion information; determining the second motion difference information based on the n-th frame motion information of the event camera corresponding to the n-th frame image at the shooting moment, and the i+1-th frame motion information of the event camera corresponding to the i+1-th frame wonderful frame at the shooting moment; wherein the second motion difference represents the change between the first i-frame motion information and the n-th frame motion information; determining the first n-th frame image to be fused based on the i-th frame wonderful frame and the first motion difference information; determining the second n-th frame image to be fused based on the i+1-th frame wonderful frame and the second motion difference information; generating the n-th frame image based on the first n-th frame image to be fused and the second n-th frame image to be fused.
[0030] In a possible implementation of the first aspect, a target wonderful video is generated based on N wonderful frames and a first image, including: generating a target wonderful video based on N wonderful frames, a first image and a second image; wherein the second image is an image captured within a preset time period before the start frame shooting moment.
[0031] The electronic device generates a video together with the images captured within a preset time period before the start frame shooting moment, N frames of wonderful frames and the first image. While presenting the wonderful movement moments of the target object to the user, it can also present the target object's behavior before the movement to the user, so that a video with richer picture content can be obtained, thereby improving the user experience.
[0032] In a possible implementation of the first aspect, if before detecting that the target object satisfies a preset motion posture based on the preview stream information corresponding to the preview stream screen, the method further includes: detecting whether the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information; detecting that the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream screen includes: detecting that the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information.
[0033] Event cameras can accurately track the changes in the target object's posture during movement. The event camera stream information captured by the event camera can assist in previewing the stream information, more accurately determining whether the target object meets the preset motion posture.
[0034] In a second aspect, an electronic device is provided, which includes a processor and a memory; the memory is used to store code instructions; the processor is used to run the code instructions to execute the audio signal adjustment method in any possible design method as in the first aspect.
[0035] In a third aspect, a computer-readable storage medium is provided, in which instructions are stored. When the instructions are executed on a computer, the computer executes the method for adjusting the audio signal in any possible design manner as in the first aspect.
[0036] In a fourth aspect, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the method in any possible design manner in the first aspect.
[0037] Among them, the technical effects brought about by any design method in the second, third and fourth aspects can refer to the technical effects brought about by different design methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] FIG1 is a schematic diagram showing two frames of images captured by an electronic device;
[0039] FIG2 is a schematic diagram showing the display of the interface during video playback;
[0040] FIG3 shows a schematic diagram of three frames of images obtained by shooting with an event camera and a first camera;
[0041] FIG4 shows a schematic structural diagram of a mobile phone 100;
[0042] FIG5 shows a schematic flow chart of a video generation method;
[0043] FIG6 shows a schematic diagram of the operation interface changes when the mobile phone 100 turns on the camera and the automatic snapshot mode in the camera;
[0044] FIG7 shows a schematic diagram of the operation interface changes when the mobile phone 100 turns on the camera and the automatic snapshot mode in the camera;
[0045] FIG8 shows a schematic diagram of a wonderful frame recognition algorithm;
[0046] FIG9 shows a schematic diagram of another wonderful frame recognition algorithm;
[0047] FIG10 is a schematic diagram showing a comparison of the shooting frequencies of the first camera and the event camera;
[0048] FIG11 is a schematic diagram showing a process of the mobile phone 100 determining N frames of highlight frames to generate a video from a preview stream;
[0049] FIG12 shows a schematic diagram of inserting a slow-motion detail image;
[0050] FIG13 shows a schematic diagram of an image frame insertion;
[0051] FIG14 shows a schematic diagram of yet another image frame insertion method;
[0052] FIG15 shows a schematic diagram of a process for generating an insert image;
[0053] FIG16 is a schematic diagram showing a process of generating an interpolated image;
[0054] FIG17 is a schematic diagram showing changes in the interface of a gallery;
[0055] FIG18 shows a schematic diagram of a process of a video generation method. DETAILED DESCRIPTION
[0056] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0057] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.
[0058] In the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.
[0059] In order to solve the technical problems in the background technology, an embodiment of the present application proposes a video generation method.
[0060] Understandably, with the advancement of photography technology, people's demands for photo quality, particularly clarity, are becoming increasingly stringent. Currently, when mobile phones and other electronic devices capture moving objects, the relative motion between the object and the device affects the camera's exposure time and sampling frequency. This often results in discrete, continuous moments of the object's non-moving activity.
[0061] The image sensor corresponding to the event-based vision (EVS) camera of the electronic device can detect the brightness changes of the pixel points caused by the movement of the target object in real time (this change is referred to as a motion event). Even if the relative movement speed of the camera (first camera) and the target object is fast, the posture changes of the target object during the movement can be determined more accurately.
[0062] Therefore, the electronic device can use the event camera to compensate for the problem that the first camera cannot capture the slow-motion details of the target object when the target object moves faster. If the target object is moving, the brightness changes of the pixels of the target object caused by the movement detected by the event camera can be used to obtain multiple frames of slow-motion detail images of the target object during the movement. The electronic device then inserts the multiple frames of slow-motion detail images of the target object during the movement between the multiple frames of images (such as RGB images) captured by the first camera to generate a video.
[0063] In some embodiments, the posture of the target object during the motion process meets certain conditions, and the photos of the target object taken by the electronic device at these moments are called highlight frames of the target object.
[0064] In the embodiments of the present application, the video generated according to the above scheme has high image quality, less pseudo-texture, better continuity, and higher smoothness. Furthermore, the electronic device can play each frame of the image according to the arranged image, making it convenient for users to watch the wonderful slow-motion moments at a slow speed, and providing a better user viewing experience.
[0065] For example, Figure 3 shows a schematic diagram of three frames of images captured by an event camera and a first camera. As shown in Figure 3, based on the A-frame image and the C-frame image captured by the camera (the first camera), the electronic device can capture the event stream information between the capture time of the A-frame image and the capture time of the C-frame image. Based on the event stream information and the A-frame image and the C-frame image, the electronic device generates a B-frame image corresponding to the motion details (slow-motion details) of the target object between the two capture times. Specifically, the B-frame image presents the slow-motion details of user A jumping over a hurdle. The electronic device can then insert the B-frame image between the A-frame image and the C-frame image, generating Video 2 from the A-frame image, the B-frame image, and the C-frame image.
[0066] In the example of Figure 2, the display of the interface during the playback of Video 2 is shown. The electronic device can display all images in sequence according to the order of the images included in Video 1. Exemplarily, when the electronic device starts playing Video 2, the electronic device can display interface 113, which includes an A-frame image, and displays the A-frame image from the 1st second to the 5th second. When the display reaches the 5th second, the electronic device can display interface 114, which includes a B-frame image, and displays the B-frame image from the 5th second to the 15th second. When Video 1 is about to finish playing, for example, at the 15th second, the electronic device can display interface 115, which includes a C-frame image.
[0067] In this way, during the playback of video 2, frame A shows the user A just starting to jump over the hurdle, frame B shows the user A about to complete the hurdle, and frame C shows the user having completed the hurdle and running towards the destination. The user's actions presented by these three frames are a relatively continuous process, with high image quality, good continuity, less pseudo-texture, and a better user experience.
[0068] For example, the electronic device in the embodiments of the present application may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) or virtual reality (VR) device, and other devices with a wireless network connection function. The embodiments of the present application do not impose any special restrictions on the specific form of the electronic device.
[0069] FIG4 shows a schematic structural diagram of a mobile phone 100. As shown in FIG4, the mobile phone 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0070] It should be understood that the illustrated structure of the embodiment of the present invention does not constitute a specific limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may include more or fewer components than shown, or some components may be combined or separated, or arranged differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0071] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP),
[0072] Baseband processor, and / or neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors.
[0073] The controller may be the nerve center and command center of the mobile phone 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0074] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or is reusing. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. The processor is used to execute the video generation method provided in the embodiments of the present application.
[0075] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0076] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present application is merely an illustrative illustration and does not constitute a structural limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may also adopt a different interface connection method from the above embodiment, or a combination of multiple interface connection methods.
[0077] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the mobile phone 100. While charging the battery 142, the charging management module 140 can also provide power to the mobile phone 100 via the power management module 141.
[0078] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0079] The wireless communication function of the mobile phone 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0080] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0081] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied on the mobile phone 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0082] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0083] The wireless communication module 160 can provide wireless communication solutions applied to the mobile phone 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0084] In some embodiments, the antenna 1 of the mobile phone 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the mobile phone 100 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0085] Mobile phone 100 implements display functionality through a GPU, display screen 194, and an application processor. The GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0086] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), or an active matrix organic light-emitting diode (AMOLED).
[0087] (active-matrix organic light emitting diode, AMOLED), flexible light-emitting diode (flex light-emitting diode, FLED), Miniled, MicroLed, Micro-oLed, quantum dot light emitting diodes (quantum dot light emitting diodes, QLED), etc. In some embodiments, the mobile phone 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0088] The mobile phone 100 can realize the shooting function through the ISP, camera 193, video codec, GPU, display 194 and application processor.
[0089] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0090] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the mobile phone 100 may include 1 or N cameras 193, where N is a positive integer greater than 1. In an embodiment of the present application, the camera 193 may include a first camera and an event camera, and the first camera may be an RGB camera or a black and white camera.
[0091] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the mobile phone 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0092] Video codecs are used to compress or decompress digital video. Mobile phone 100 may support one or more video codecs. This allows mobile phone 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0093] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0094] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the mobile phone 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the mobile phone 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0095] The mobile phone 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0096] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0097] The speaker 170A, also called a "horn," is used to convert audio electrical signals into sound signals. The mobile phone 100 can listen to music or make hands-free calls through the speaker 170A.
[0098] The receiver 170B, also called the "earpiece", is used to convert audio electrical signals into sound signals. When the mobile phone 100 receives a call or a voice message, the voice can be heard by placing the receiver 170B close to the ear.
[0099] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The mobile phone 100 can be provided with at least one microphone 170C. In other embodiments, the mobile phone 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the mobile phone 100 can also be provided with three, four or more microphones 170C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.
[0100] The headphone jack 170D is used to connect a wired headphone and can be a USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0101] Keys 190 include a power button, a volume button, etc. Keys 190 may be mechanical keys or touch keys. Mobile phone 100 may receive key inputs and generate key signal inputs related to user settings and function control of mobile phone 100.
[0102] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0103] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.
[0104] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to and disconnected from the mobile phone 100 by inserting it into or removing it from the SIM card interface 195. The mobile phone 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The mobile phone 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the mobile phone 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the mobile phone 100 and cannot be separated from the mobile phone 100.
[0105] FIG5 shows a schematic flow chart of a video generation method. As shown in FIG5 , the process includes the following steps:
[0106] S501. The mobile phone 100 responds to the user's operation of turning on the camera function, turns on the first camera, and displays the first interface; wherein the first interface includes a preview window, the preview window is used to display a preview stream image, and the preview stream image is an image captured in real time by the first camera.
[0107] It is understandable that when the mobile phone 100 turns on the camera function, the picture taken by the first camera in real time, ie, the preview stream picture, will be displayed for the user to watch. The first camera can be an RGB camera or a black and white camera.
[0108] If the first camera is an RGB camera, the preview stream image is a color image. Exemplarily, the format of the preview stream image can be an RGB image, a RAW image, or a YUV image.
[0109] Figure 6 illustrates a schematic diagram of the interface changes when activating the camera and the automatic capture mode on the mobile phone 100. As shown in Figure 6(a), a camera icon 4011 is displayed in the interface 401 of the mobile phone 100. In response to a user clicking on the camera icon 4011, the mobile phone 100 activates the first camera and displays the interface 402. For example, as shown in Figure 6(b), the interface 402 includes a preview window 4022 for displaying images captured by the first camera.
[0110] S502 : The mobile phone 100 turns on the automatic snapshot mode in the camera function, and then turns on the event camera after turning on the first camera; wherein the event camera is used to capture and obtain event camera stream information.
[0111] It can be understood that the first camera (frame-based image sensor) outputs an overall image containing the object and background, such as an RGB stream, according to the frame rate and at a certain time interval. The event camera (event-based visual sensor) detects brightness changes in an asynchronous manner, and outputs the changed pixel data in combination with the coordinate and time information, that is, it captures the trajectory information of the moving object with a very high temporal resolution and outputs the EVS stream. The event camera can capture the movement details of the moving target object during the movement process. In this way, when the mobile phone 100 turns on the first camera and the event camera, the mobile phone 100 is aimed at the current scene, continuously outputs the preview stream (RGB stream + EVS stream), and uses the EVS stream to guide the insertion of frames. In this way, the mobile phone 100 can capture the movement details of the moving target object during the movement process, which facilitates the subsequent generation of a video with high continuity of image content.
[0112] As shown in FIG6(b), interface 402 also includes a settings control 4021. In response to a user clicking on settings control 4021, the mobile phone 100 displays the settings interface. For example, as shown in FIG6(c), the settings interface includes a smart photo option. In response to a user clicking on an expansion control 4031 corresponding to the smart photo option, the mobile phone 100 displays the smart photo interface. As shown in FIG6(d), the mobile phone 100 displays a smart photo interface 404. Smart photo interface 404 includes an automatic snapshot option and a switch 4041 for toggling automatic snapshot mode on or off. In response to a user turning on switch 4041, the mobile phone 100 enters automatic snapshot mode. This automatic snapshot mode can be used to automatically capture a photo when automatic snapshot is enabled, intelligently identifying a scene containing a target object that meets a preset motion pose. For example, it can automatically capture photos when intelligently identifying a person smiling, jumping, running, or a cat or dog.
[0113] As a possible implementation, after turning on the automatic snapshot mode from the settings, the mobile phone 100 turns off the camera and then turns it on again, the camera automatically enters the automatic snapshot mode. If the automatic snapshot mode is turned on in the above manner, the mobile phone 100 can automatically enter the automatic snapshot mode without executing S502 after executing S501.
[0114] In some other embodiments, the mobile phone 100 opens the camera and sets an automatic snapshot control in the camera interface. The mobile phone 100 can activate the automatic snapshot mode in response to the user's operation of the automatic snapshot control. For example, the mobile phone 100 activates the automatic snapshot mode in response to the user clicking the automatic snapshot control in the camera interface. For example, FIG7 illustrates a schematic diagram of the changes in the operation interface of the mobile phone 100 when the camera is opened and the automatic snapshot mode in the camera is activated. As shown in FIG7(a), the interface 402 also includes an automatic snapshot control 4023. When not operated, the automatic snapshot control 4023 is displayed in a first display mode, for example, only an outline is displayed. In response to the user clicking the automatic snapshot control 4023, the automatic snapshot control 4023 changes from the first display mode to the second display mode. For example, as shown in FIG7(b), the outline area of the automatic snapshot control 4023 is filled with color. Of course, the first display mode may also display the automatic snapshot control 4023 in a dark color, and the second display mode may display the automatic snapshot control 4023 in a bright color.
[0115] As a possible implementation, after turning on the automatic snapshot mode from the camera interface, the mobile phone 100 turns off the camera, and then turns it on again, it is necessary to operate the automatic snapshot control again to put the camera into the automatic snapshot mode again. If the automatic snapshot mode is turned on in the above manner, the mobile phone 100 can also execute S502 to enter the automatic snapshot mode after executing S501.
[0116] S503: The mobile phone 100 detects whether the target object satisfies a preset motion posture based on the preview stream information corresponding to the preview stream image.
[0117] If the mobile phone 100 detects that the target object does not meet the preset motion posture based on the preview stream information corresponding to the preview stream screen, it is not necessary to generate a picture in the shooting mode, and the process ends. If the mobile phone 100 detects that the target object meets the preset motion posture based on the preview stream information corresponding to the preview stream screen, it can generate a picture in the shooting mode as the material for generating the wonderful frame video.
[0118] If a preview frame in the preview stream contains a target object and the target object meets a preset motion pose, the preview frame is determined to be a highlight frame. Specifically, an image frame is one with meaningful image content. Meaningful image content includes images with clear image content, such as people or animals. It is understood that in some embodiments, clear image content also includes a person smiling with eyes open, or a person smiling, laughing, smiling with eyes closed, blinking, or pouting. Furthermore, meaningful image content also includes images performing specific preset motions. Examples include a person leaping, an athlete hurdling, a puppy running, and so on.
[0119] The target object can be a person, an animal, a vehicle (such as a vehicle in a racing scene), etc.
[0120] The preset motion gestures can be motion gestures in various sports scenes, for example, the motions of people in various sports games. For example, in a basketball game scene, an athlete may wave, jump, shoot, or fly. Another example is in a racing scene, a vehicle may move quickly or turn. Another example is in an animal sports scene, a horse or dog may run quickly.
[0121] It should be noted that the above examples are merely examples and are not exhaustive. In other embodiments of the present application, the target object may include other objects, and the preset motion gesture may also be other gestures.
[0122] In some embodiments, the mobile phone 100 can use a wonderful frame recognition algorithm to detect whether there is a picture in the preview stream in which the target object meets a preset motion posture.
[0123] Figure 8 shows a schematic diagram of a highlight frame recognition algorithm. As shown in Figure 8, the highlight frame recognition algorithm includes a target object motion posture recognition network 1, which is used to identify whether a target object exists in an image and whether the target object meets a preset motion posture. If the target object motion posture recognition network identifies the presence of a target object in a preview frame of the preview stream and the target object meets a preset motion posture, the preview frame is determined to be a highlight frame.
[0124] The mobile phone 100 may input the preview stream information corresponding to the preview stream screen into the target object motion posture recognition network 1 , and the target object motion posture recognition network 1 may output a recognition result of whether it is a wonderful frame.
[0125] It can be understood that the event camera can more accurately track the changes in the posture of the target object during the movement. The event camera stream information captured by the event camera can assist the preview stream information to more accurately determine whether there is a target object that meets the preset motion posture. Therefore, in some other embodiments, the mobile phone 100 can also detect whether the target object meets the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information. If the mobile phone 100 detects that the target object does not meet the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information, there is no need to generate a picture in the shooting mode, and the process ends. If the mobile phone 100 detects that the target object meets the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information, it can generate a picture in the shooting mode as material for generating a wonderful frame video.
[0126] As a possible implementation, mobile phone 100 may also input the preview stream information corresponding to the preview stream image and the event camera stream information into a target object motion posture recognition network. The target object motion posture recognition network is used to identify whether a target object exists in the image and whether the target object satisfies a preset motion posture. If the target object motion posture recognition network identifies the presence of a target object in a preview frame of the preview stream image and the target object satisfies the preset motion posture, the preview frame is determined to be a highlight frame.
[0127] For example, another schematic diagram of a wonderful frame recognition algorithm is shown in Figure 9. As shown in Figure 9, the wonderful frame recognition algorithm includes a target object motion posture recognition network 2, which is used to identify whether a target object exists in an image and whether the target object meets a preset motion posture.
[0128] The mobile phone 100 may input the preview stream information and the event camera stream information corresponding to the preview stream screen into the target object motion posture recognition network 2. The target object motion posture recognition network 2 may output a recognition result of whether it is a wonderful frame.
[0129] Generally, the shooting frequency of the event camera is greater than the shooting frequency of the first camera. For example, the shooting frequency of the event camera is 1000 frames per second, and the shooting frequency of the first camera is 30 frames per second. 1000 frames per second is greater than 30 frames per second.
[0130] Figure 10 shows a schematic diagram comparing the shooting frequencies of the first camera and the event camera. As shown in Figure 10, the shooting frequency of the first camera is 30 frames per second, about once every 33ms, while the shooting frequency of the event camera is 1000 frames per second, about once every 1ms.
[0131] As can be seen from FIG9 and FIG10 , the mobile phone 100 can input the preview information corresponding to the first shot of the first camera and the frame of event camera information corresponding to the event camera, for example, the frame of event camera information captured by the event camera for the first time, to the target object motion posture recognition network 2. The target object motion posture recognition network 2 can output a recognition result indicating whether the preview information corresponding to the first shot of the first camera is a highlight frame.
[0132] The mobile phone 100 can input the preview information corresponding to the second shot of the first camera and the frame of event camera information corresponding to the event camera, for example, the frame of event camera information captured by the event camera for the 33rd time, to the target object motion posture recognition network 2. The target object motion posture recognition network 2 can output a recognition result indicating whether the preview information corresponding to the second shot of the first camera is a highlight frame.
[0133] The mobile phone 100 can input the preview information corresponding to the third shot by the first camera and the frame of event camera information corresponding to the shot by the event camera, for example, the frame of event camera information shot by the event camera 66, to the target object motion posture recognition network 2. The target object motion posture recognition network 2 can output a recognition result indicating whether the preview information corresponding to the third shot by the first camera is a highlight frame.
[0134] If the preview information corresponding to the 1st to 3rd shots of the first camera are all identified as highlight frames, the mobile phone 100 can insert an image between the highlight frame corresponding to the 1st shot of the first camera and the highlight frame corresponding to the 2nd shot, and insert an image between the highlight frame corresponding to the 2nd shot of the first camera and the highlight frame corresponding to the 2nd shot, to generate a highlight frame video.
[0135] If the preview information corresponding to the first and third shots of the first camera are both identified as highlight frames, the mobile phone 100 may insert an image between the highlight frame corresponding to the first shot of the first camera and the highlight frame corresponding to the third shot to generate a highlight frame video.
[0136] It can be understood that the multiple frames of preview information obtained by the first camera during multiple consecutive shots can be referred to as preview stream information. The multiple frames of event camera information obtained by the event camera during multiple consecutive shots can be referred to as event camera stream information. For example, as shown in FIG10 , the two frames of preview information corresponding to the first and second shots of the first camera can be referred to as preview stream information. The three frames of preview information corresponding to the first to third shots of the first camera can also be referred to as preview stream information. The 33 frames of event camera information corresponding to the first to 33rd shots of the event camera can be referred to as event camera stream information. The 66 frames of event camera information corresponding to the first to 66th shots of the event camera can be referred to as event camera stream information.
[0137] S504 : The mobile phone 100 determines to generate N highlight frames of the video based on the preview stream information; wherein the N highlight frames include the start frame and the end frame of the video to be generated.
[0138] That is, the mobile phone 100 can determine N frames of the video to be generated based on the preview stream information of the target object meeting the preset motion posture. The number of frames of the video to be generated can be multiple frames, and N is an integer greater than 1, for example, N is 2, 3, 4, 5, etc.
[0139] Figure 11 illustrates a schematic diagram of the process by which mobile phone 100 determines N highlight frames from preview stream information to generate a video. As shown in Figure 11 , if the target object motion posture recognition network identifies the presence of a target object in a preview frame of the preview stream information, and if the target object satisfies a preset motion posture, the preview frame is determined to be a highlight frame. If the preview frame is a highlight frame, the preview frame is extracted and subjected to image quality processing to obtain a higher-quality image. Higher-quality image can include higher image clarity and / or higher image resolution.
[0140] Specifically, the mobile phone 100 can input the preview frame into a neural network, and the neural network can be used for image quality processing to obtain an image with higher image quality.
[0141] It is understood that image quality indicators include but are not limited to clarity, noise, color, etc.
[0142] In this way, when the mobile phone 100 determines a wonderful frame, it generates an image with higher resolution and / or clarity, which can improve the user's viewing experience.
[0143] On the other hand, the mobile phone 100 can also extract the event camera stream information corresponding to the shooting moment of the preview frame, denoise the event camera stream information, and obtain the denoised event camera stream information. It should be noted that since the shooting frequency of the event camera is greater than the shooting frequency of the first camera, between the preview frames at the two adjacent shooting moments, there are also multiple frames of event camera stream information shot by the event camera. The mobile phone 100 can extract the event camera stream information corresponding to the shooting time period between the two adjacent highlight frame preview frames, denoise all the event camera stream information within this time period, and obtain the denoised event camera stream information. The denoising method can be referred to in related technologies and will not be described in detail here. The denoising algorithm of event camera information or event camera stream information can be referred to in related technologies and will not be described in detail here.
[0144] For example, as shown in Figure 10, if the preview information corresponding to the 1st to 3rd shots of the first camera are all identified as wonderful frames, the mobile phone 100 can denoise the event camera information obtained by the 1st shot of the event camera, the event camera information obtained by the 66th shot of the event camera, and the event camera stream information obtained from the 1st to 66th shots of the event camera.
[0145] If the preview information corresponding to the first and second shots of the first camera are both identified as highlight frames, the mobile phone 100 may denoise the event camera information obtained by the first shot of the event camera and the event camera information obtained by the 33rd shot of the event camera.
[0146] The mobile phone 100 denoises the event camera stream to accurately obtain a clearer image of the target object's motion. Since the generated inserted image is clearer, the mobile phone 100 can obtain a higher-quality video, presenting a higher-quality video to the user, and improving the user experience.
[0147] S505 : The mobile phone 100 obtains event camera stream information between the shooting time of the start frame and the shooting time of the end frame, and determines motion information based on the event camera stream information.
[0148] The motion information is used to represent the motion of the moving object captured by the event camera. The moving object includes at least the target object and may also include objects in the background of the target object.
[0149] The motion of a moving object includes the motion of the target object and the motion of objects within the target object's background. Based on the motion of the moving object captured by the event camera, mobile phone 100 can accurately capture a clearer image of the target object even when the target object is in motion. Because the generated inset image is clearer, mobile phone 100 can produce higher-quality video, presenting a higher-quality video to the user and providing a better user experience.
[0150] Motion information can include optical flow information, which represents changes in pixel brightness within images captured by the event camera. Optical flow refers to changes in pixel brightness during the capture time caused by light obstruction in the object's surroundings or other factors. Based on these changes in pixel brightness, mobile phone 100 can accurately and clearly capture images of the target object's motion. Because the resulting interpolated image is clearer, mobile phone 100 can produce higher-quality videos, presenting higher-quality videos to users and providing a better user experience.
[0151] The mobile phone 100 may refer to related technologies to determine motion information based on event camera stream information, which will not be described in detail here.
[0152] It can be understood that the event camera stream information is used to determine the image to be inserted. For example, FIG12 shows a schematic diagram of inserting a slow motion detail image. As shown in FIG12, if an image is inserted at the intermediate time T / 2 between time 0 and time T of the shooting time, then the image to be played I0 corresponding to time 0 and the image to be played I corresponding to time T can be inserted. t , the event camera stream information from time 0 to the intermediate time T / 2 and the event camera stream information from the intermediate time T / 2 to time T are obtained to obtain the image inserted at the intermediate time T / 2.
[0153] Specifically, the mobile phone 100 can obtain the event camera stream information between the capture time of the start frame and the capture time of the end frame, and determine the motion information within the aforementioned time period based on the event camera stream. In this way, the mobile phone 100 can insert a slow-motion detailed image of the target object moving within the aforementioned time period based on the motion information within the aforementioned time period.
[0154] S506. The mobile phone 100 determines a first image to be inserted between each two adjacent wonderful frames in the N wonderful frames based on the motion information; wherein the first image includes slow motion details of the target object between each two adjacent wonderful frames in the N wonderful frames.
[0155] The mobile phone 100 may insert a slow-motion detail image, ie, a first image, of the target object moving within the aforementioned time period based on the motion information within the aforementioned time period.
[0156] It can be understood that, in some embodiments, the number of first images to be inserted between every two adjacent highlight frames in the N second highlight frames is the same.
[0157] For example, taking N as 4, Figure 13 shows a schematic diagram of image frame insertion. As shown in Figure 13, Highlight Frame 1, Highlight Frame 2, Highlight Frame 3, and Highlight Frame 4 are four specific Highlight Frames. To create a video with continuous and smooth content, M frames of images can be inserted between Highlight Frame 1 and Highlight Frame 2, M frames of images can be inserted between Highlight Frame 2 and Highlight Frame 3, and M frames of images can be inserted between Highlight Frame 3 and Highlight Frame 4.
[0158] Among them, the M frame images inserted between wonderful frame 1 and wonderful frame 2 are slow-motion details of the target object's movement from shooting moment 1 of wonderful frame 1 to shooting moment 2 of wonderful frame 2, and the content of each frame of the M frame images can be different. The specific content is the actual movement picture of the target object from shooting moment 1 to shooting moment 2.
[0159] The M frames of images inserted between highlight frame 2 and highlight frame 3 are slow-motion details of the target object's movement from shooting moment 2 of highlight frame 2 to shooting moment 3 of highlight frame 3, and the content of each frame of the M frames can be different. The specific content is the actual movement picture of the target object from shooting moment 2 to shooting moment 3.
[0160] The M frames of images inserted between wonderful frame 3 and wonderful frame 4 are slow-motion details of the target object's movement from shooting moment 3 of wonderful frame 3 to shooting moment 3 of wonderful frame 4, and the content of each frame of the M frames can be different. The specific content is the actual movement picture of the target object from shooting moment 3 to shooting moment 4.
[0161] In addition to the above-mentioned method in which the number of first images to be inserted between every two adjacent wonderful frames in N frames is the same, in some other embodiments, the number of first images to be inserted between not all groups of two adjacent wonderful frames in N frames may be the same, but the number of first images to be inserted between two or more groups of two adjacent wonderful frames in N frames may be the same.
[0162] In some embodiments, the number of first images to be inserted between each two adjacent wonderful frames in the M-frame second wonderful frame is not the same. Specifically, the number of first images to be inserted between each two adjacent wonderful frames in the M-frame second wonderful frame is determined based on the motion speed of the target object between the two adjacent wonderful frames in the M-frame second wonderful frame. If the motion speed of the target object between the two adjacent wonderful frames in the M-frame second wonderful frame is faster, then the target object moves more and has more motion details between the two adjacent wonderful frames. To improve the continuity and smoothness of the image playback content in the video, the number of first images to be inserted between the two adjacent wonderful frames in the M-frame second wonderful frame can be determined to be larger. If the motion speed of the target object between the two adjacent wonderful frames in the M-frame second wonderful frame is slower, then the target object moves less and has less motion details between the two adjacent wonderful frames. Determining the number of first images to be inserted between the two adjacent wonderful frames in the M-frame second wonderful frame can also ensure the continuity and smoothness of the image playback content in the video.
[0163] In an embodiment of the present application, the target object may move at different speeds at different moments. When the target object moves at a faster speed, the target object may make more movements within a certain period of time, with a higher richness of movements, and these movements are difficult to be captured by the first camera. Therefore, when generating a video, the mobile phone 100 can compensate for the movement details of the target object that are not captured by the first camera based on the event stream information captured by the event camera. Specifically, the mobile phone 100 can insert more images when the target object moves at a faster speed based on the movement speed of the target object. In this way, the continuity and smoothness of the video generation can be improved, thereby enhancing the user experience.
[0164] When the target object moves slowly, within a certain period of time, the target object's movements are relatively few, the richness of the movements is low, and these movements are easily captured by the first camera. Therefore, when generating a video, the mobile phone 100 can basically not compensate for the motion details of the target object and still obtain the motion details of the target object. Specifically, the mobile phone 100 can insert fewer images when the target object moves slowly based on the speed of the target object. In this way, a video with higher continuity and smoothness can be generated, and the power consumption of the mobile phone 100 when inserting images can be reduced, thereby improving the user experience.
[0165] In addition to the aforementioned method of having the number of first images to be inserted between each two adjacent highlight frames in N highlight frames differ, in other embodiments, the number of first images to be inserted between not all groups of two adjacent highlight frames in N highlight frames differs, but rather the number of first images to be inserted between two or more groups of two adjacent highlight frames in N highlight frames differs. For example, the number of first images to be inserted between the first group of two adjacent highlight frames in N highlight frames is 2, the number of first images to be inserted between the second group of two adjacent highlight frames in N highlight frames is 2, and the number of first images to be inserted between the third group of two adjacent highlight frames in N highlight frames is 0, i.e., no images are inserted between the third group of two adjacent highlight frames in N highlight frames. For another example, the number of first images to be inserted between the first two adjacent wonderful frames in the N-frame wonderful frames is 3 frames, the number of first images to be inserted between the second two adjacent wonderful frames in the N-frame wonderful frames is 2 frames, and the number of first images to be inserted between the third two adjacent wonderful frames in the N-frame wonderful frames is 2 frames, that is, no image is inserted between the third two adjacent wonderful frames in the N-frame wonderful frames.
[0166] It is understood that the playback time interval between each two frames of the generated video can be the same. Thus, if the generated video contains a large number of images, the video playback duration can be longer; if the generated video contains a small number of images, the video playback duration can be shorter. Therefore, the mobile phone 100 can control the playback duration of the target object's action between two adjacent highlight frames by adjusting the number of images inserted between each two adjacent highlight frames.
[0167] Specifically, the mobile phone 100 may default to inserting the same number of images between every two adjacent highlight frames. For example, taking N as 5, FIG14 shows another schematic diagram of image insertion. As shown in FIG14 , highlight frames 1, 2, 3, 4, and 5 are the five confirmed highlight frames.
[0168] The playing time between wonderful frame 1 and wonderful frame 2 is T, the playing time between wonderful frame 2 and wonderful frame 3 is T, the playing time between wonderful frame 3 and wonderful frame 4 is T, and the playing time between wonderful frame 4 and wonderful frame 5 is also T.
[0169] Of course, in order to improve the user's experience, the play duration between each two adjacent wonderful frames can also be adjusted under the user's operation, that is, by adjusting the insertion of different numbers of images between each two adjacent wonderful frames, so as to achieve the purpose of making the play duration between each two adjacent wonderful frames different after adjustment. As shown in Figure 14, the play duration between wonderful frame 1 and wonderful frame 2 is T12, the play duration between wonderful frame 2 and wonderful frame 3 is T23, the play duration between wonderful frame 3 and wonderful frame 4 is T34, and the play duration between wonderful frame 4 and wonderful frame 5 is also T45. T12, T23, T34 and T45 can all be different.
[0170] The mobile phone 100 can determine the first image to be inserted between every two adjacent wonderful frames in N wonderful frames based on motion information by using relevant technical means. An implementation scheme is briefly introduced below as an example.
[0171] Two adjacent highlight frames include the i-th highlight frame and the i+1-th highlight frame. The i-th highlight frame and the i+1-th highlight frame can also be referred to as a set of two adjacent highlight frames. M frames are inserted between the i-th highlight frame and the i+1-th highlight frame; where i is an integer between 1 and N-1 and has a value of N frames, and M is an integer greater than or equal to 0.
[0172] For example, if N is 5, the five wonderful frames include the first wonderful frame, the second wonderful frame, the third wonderful frame, the fourth wonderful frame, and the fifth wonderful frame. The first wonderful frame, the second wonderful frame, the third wonderful frame, the fourth wonderful frame, and the fifth wonderful frame are arranged in sequence according to the chronological order of the electronic device's shooting time, or the first wonderful frame, the second wonderful frame, the third wonderful frame, the fourth wonderful frame, and the fifth wonderful frame are arranged in a reverse chronological order of the chronological order of the electronic device's shooting time. Then the first wonderful frame and the second wonderful frame are two adjacent wonderful frames, the second wonderful frame and the third wonderful frame are two adjacent wonderful frames, the third wonderful frame and the fourth wonderful frame are two adjacent wonderful frames, and the fourth wonderful frame and the fifth wonderful frame are two adjacent wonderful frames.
[0173] For example, still taking N as 5, if M is 2, the mobile phone 100 can insert 2 frames of images between the 1st wonderful frame and the 2nd wonderful frame, insert 2 frames of images between the 2nd wonderful frame and the 3rd wonderful frame, insert 2 frames of images between the 3rd wonderful frame and the 4th wonderful frame, and insert 2 frames of images between the 4th wonderful frame and the 5th wonderful frame.
[0174] Of course, in addition to the above-mentioned method of inserting the same number of images between the i-th wonderful frame and the (i+1)-th wonderful frame, in some other embodiments, it is also possible that not all groups of adjacent two wonderful frames in N frames have the same number of images (i.e., first images) inserted between them, but the number of images inserted between two or more groups of adjacent two wonderful frames in N frames is the same.
[0175] For example, still taking N as 5, M can be 2, 2, 0, and 1. Then the mobile phone 100 can insert 2 frames of images between the 1st and 2nd wonderful frames, insert 2 frames of images between the 2nd and 3rd wonderful frames, insert 0 frames of images between the 3rd and 4th wonderful frames, and insert 1 frame of image between the 4th and 5th wonderful frames.
[0176] For example, still taking N as 5, M can be 2, 2, 2, and 1. Then the mobile phone 100 can insert 2 frames of images between the 1st wonderful frame and the 2nd wonderful frame, insert 2 frames of images between the 2nd wonderful frame and the 3rd wonderful frame, insert 2 frames of images between the 3rd wonderful frame and the 4th wonderful frame, and insert 1 frame of image between the 4th wonderful frame and the 5th wonderful frame.
[0177] In some other embodiments, the number of images inserted between every two adjacent highlight frames in the N highlight frames is different.
[0178] For example, still taking N as 5, M can be 2, 0, 1, and 3. Then the mobile phone 100 can insert 2 frames of images between the 1st and 2nd wonderful frames, insert 0 frames of images between the 2nd and 3rd wonderful frames, insert 1 frame of image between the 3rd and 4th wonderful frames, and insert 3 frames of image between the 4th and 5th wonderful frames.
[0179] In some other embodiments, the number of images inserted between two adjacent frames in not all groups of N highlight frames is different, but the number of images inserted between two or more groups of two adjacent frames in the N highlight frames is different.
[0180] For example, still taking N as 5, M can be 2, 1, 0, and 0. Then the mobile phone 100 can insert 2 frames of images between the 1st and 2nd wonderful frames, insert 1 frame of image between the 2nd and 3rd wonderful frames, insert 0 frames of image between the 3rd and 4th wonderful frames, and insert 0 frames of image between the 4th and 5th wonderful frames.
[0181] For the n-th frame image in M frames, where n takes integers from 1 to H in sequence, FIG15 shows a schematic flow chart of generating an inserted image.
[0182] S1501. The mobile phone 100 determines first motion difference information based on the i-th frame motion information corresponding to the shooting moment of the i-th wonderful frame and the n-th frame motion information corresponding to the n-th frame image at the shooting moment of the event camera.
[0183] The first motion difference information is used to represent the change between the motion information of the i-th frame and the motion information of the n-th frame.
[0184] For example, as shown in FIG10 , taking i as 1, if the preview information corresponding to the first and second shots of the first camera are both identified as highlight frames, the highlight frame corresponding to the first shot of the first camera is the first highlight frame. The highlight frame corresponding to the second shot of the first camera is the second highlight frame.
[0185] Taking i as 2 as an example, if the preview information corresponding to the first to third shots of the first camera are all identified as highlight frames, the highlight frame corresponding to the second shot of the first camera is the second highlight frame. The highlight frame corresponding to the third shot of the first camera is the third highlight frame.
[0186] Continuing with Figure 10, assuming i is 1, M is 2, and n is 1, the mobile phone 100 can use the motion information of the first frame corresponding to the moment of the event camera's first shot and the motion information of the first frame corresponding to the moment of the event camera's 22nd shot to determine the motion difference between the two frames. Similarly, if n is 2, the same logic can be applied.
[0187] The motion information is determined based on the event camera information. For example, the first frame of motion information is determined based on the first frame of event camera information corresponding to the first shooting moment of the event camera.
[0188] Figure 16 shows a schematic diagram of a process for generating an inserted image. As shown in Figure 16, if an image is inserted at any time t between time t0 and time t1 of the shooting time, the image to be played corresponding to time t0 can be The image to be played corresponding to time t1 Event camera stream information from time t0 to time t And the event camera stream information from time t to time t1 Get the image I inserted at time t t .
[0189] Specifically, the mobile phone 100 can store the event camera stream information from time t0 to time t And the event camera stream information from time t to time t1 Input to the optical flow estimation network to generate motion information from time t0 to time t And the motion information from time t to time t1
[0190] The mobile phone 100 can calculate the difference between the motion information at time t0 and time t to obtain motion difference information
[0191] S1502: The mobile phone 100 determines second motion difference information based on the nth frame motion information corresponding to the insertion moment of the nth frame image and the i+1th frame motion information corresponding to the i+1th wonderful frame at the shooting moment of the event camera.
[0192] The second motion difference is used to represent the change between the motion information of the first i frames and the motion information of the nth frame.
[0193] For example, as shown in FIG10 , taking i as 1, if the preview information corresponding to the first to third shots of the first camera are all identified as highlight frames, the highlight frame corresponding to the first shot of the first camera is the first highlight frame. The highlight frame corresponding to the second shot of the first camera is the second highlight frame.
[0194] Continuing with Figure 10, taking i as 1, M as 2, and n as 1 as an example, mobile phone 100 can use the motion information of the second frame corresponding to the second shot of the event camera and the motion information of the first frame corresponding to the 22nd shot of the event camera to determine the motion difference between the two frames. Similarly, if n is 2, the same logic can be applied.
[0195] The motion information is determined based on the event camera information. For example, the first frame of motion information is determined based on the first frame of event camera information corresponding to the first shooting moment of the event camera.
[0196] For example, as shown in FIG16 , the mobile phone 100 can perform difference calculation on the motion information at time t1 and time t to obtain motion difference information
[0197] S1503: The mobile phone 100 determines the nth frame image to be first fused based on the i-th wonderful frame and the first motion difference information.
[0198] For example, as shown in FIG16 , the mobile phone 100 can be based on the image and motion difference information Get the image at time t
[0199] S1504: The mobile phone 100 determines the second n-th frame image to be fused based on the (i+1)-th wonderful frame and the second motion difference information.
[0200] For example, as shown in FIG16 , the mobile phone 100 can be based on the image I t and motion difference information Get the image g at time t
[0201] S1505 : The mobile phone 100 generates an n-th frame image based on the first n-th frame image to be fused and the second n-th frame image to be fused.
[0202] For example, as shown in FIG16 , the mobile phone 100 can be based on the image at time t and images The fusion is performed to obtain the image with the wonderful frame inserted.
[0203] S507 : The mobile phone 100 generates a target wonderful video based on the N wonderful frames and the first image.
[0204] It can be understood that the target wonderful video is used to record fleeting wonderful moments, and the target wonderful video can facilitate users to watch slow-motion wonderful moments at a slow speed.
[0205] In some embodiments, to enhance the continuity of the target subject's environment when playing the target highlight frame video, the mobile phone 100 may include several frames of preparatory footage of the target subject before the start frame. Specifically, the mobile phone 100 may generate the target highlight video based on N highlight frames, a first image, and a second image; wherein the second image is an image captured within a preset period of time before the start frame.
[0206] The mobile phone 100 generates a video together with the images captured within a preset time period before the start frame shooting moment, N frames of wonderful frames, and the first image. While presenting the wonderful movement moments of the target object to the user, it can also present the behavior of the target object before movement to the user, and obtain a video with richer picture content, thereby improving the user experience.
[0207] In some embodiments, the playback duration of each frame of the target wonderful video is the same.
[0208] The target wonderful video refers to a video frame with meaningful image content. Meaningful image content includes clear image content in the image, such as people, animals, etc. Preferably, clear image content also includes a person smiling with eyes open, or a person smiling, laughing, smiling with eyes closed, blinking, or pouting. Furthermore, meaningful image content also includes the image content performing preset specific actions. For example, a person jumping up, an athlete hurdling, a puppy running, etc. For images that do not include people and animals that can perform specific actions, meaningful image content also includes plants and buildings that are intact and upright.
[0209] In some embodiments, if the photos taken by the mobile phone 100 are to be automatically generated into videos, the "geolocation" switch in the settings interface of the camera application of the mobile phone 100 and the "network connection" switch in the gallery application can be turned on. In this way, the mobile phone 100 can place the video data generated by the mobile phone 100 in the gallery application.
[0210] In some embodiments, the generated target wonderful video can be placed in a gallery. For example, Figure 17 shows a schematic diagram of interface changes in a gallery. As shown in Figure 17 (a), the figure is an interface 405 corresponding to the moment option in the gallery. Interface 405 includes a thumbnail area 4051 for the target wonderful video. It is understandable that there can be multiple target wonderful videos, each corresponding to a thumbnail area of multiple target wonderful videos. To briefly introduce this solution, only the thumbnail area of one target wonderful video is shown in interface 405.
[0211] The mobile phone 100 can respond to the user's operation on the thumbnail area 4051 of the target wonderful video. For example, the mobile phone 100 displays the interface 406 shown in Figure 17 (b) in response to the user's click operation on the thumbnail area 4051 of the target wonderful video. The interface 406 includes a thumbnail window 4061 for dynamically playing the target wonderful video. The thumbnail window 4061 for dynamically playing the target wonderful video includes a video play button 4062. In response to the user clicking the video play button 4062, the mobile phone 100 plays the target wonderful video in a larger window. Exemplarily, the mobile phone 100 can play the target wonderful video in full screen mode.
[0212] FIG18 is a schematic diagram showing the process of a video generation method. As shown in FIG18 , the mobile phone 100 is ready to take a photo, aims at the target scene, and displays the preview stream. The mobile phone 100 then performs highlight frame recognition on the preview stream information, automatically identifying a series of continuous highlight frames, which are a highlight frame segment. The mobile phone 100 then performs an automatic interpolation operation: using two adjacent highlight frames to synthesize N dense images. All interpolated images are compressed into a short video with continuous action. The mobile phone can then synthesize a video based on the aforementioned images: all interpolated images are compressed into a short video with continuous action.
[0213] It is understood that the mobile phone 100 may generally have a default order for the highlight frames. For example, the highlight frames may be arranged in the order in which they were captured. Furthermore, the mobile phone 100 may generally have a default interval length between two adjacent highlight frames. Of course, to facilitate user flexibility and generate more satisfying highlight videos, the mobile phone 100 may rearrange the order of the highlight frames and readjust the interval length between two adjacent highlight frames at the user's discretion.
[0214] In the embodiment of the present application, when the target object and the camera are in relative motion, the first camera cannot capture the slow-motion details of the target object's motion due to factors such as exposure time. However, the data obtained by the event camera can obtain the slow-motion details of the target object. The video generation method provided by the embodiment of the present application generates a video with high image quality, less pseudo-texture, better continuity, and higher smoothness. Moreover, the mobile phone 100 can play each frame of the image according to the arranged image, making it convenient for the user to watch the wonderful slow-motion moments at a slow speed, and the user's viewing experience is better.
[0215] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned mobile terminal, the mobile terminal executes the various functions or steps executed by the mobile phone 100 in the above-mentioned method embodiment.
[0216] The present application also provides a computer program product, which, when executed on a computer, enables the computer to execute the functions or steps executed by the mobile phone 100 in the above method embodiment. The computer may be the above mobile terminal (such as the mobile phone 100).
[0217] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0218] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0219] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0220] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, instructions can be distributed over a network or through other computer-readable storage media. Therefore, a machine-readable storage medium can include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including but not limited to, floppy disks, optical disks, optical discs, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROM), random access memories (RAM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) based on the Internet in electrical, optical, acoustic or other forms of propagation signals. Accordingly, machine-readable storage media include any type of machine-readable storage media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).
[0221] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.
[0222] It can be understood that the various units / modules mentioned in the various device embodiments of this application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.
[0223] It will be understood that in the examples and description of this patent, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0224] Although the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the application.
Claims
1. A video generation method, characterized in that: The method is applied to an electronic device, the electronic device comprising an event camera and a first camera, and the method comprises: In response to the user's operation of turning on the camera function, the first camera is turned on and a first interface is displayed; wherein the first interface includes a preview window, the preview window is used to display a preview stream image, and the preview stream image is an image captured by the first camera in real time; When the automatic snapshot mode in the camera function is turned on, the event camera is turned on; wherein the event camera is used to capture and obtain event camera stream information; If it is detected based on the preview stream information corresponding to the preview stream screen that the target object satisfies the preset motion posture, N frames of wonderful frames of the generated video are determined based on the preview stream information; wherein the N frames of wonderful frames include the start frame and the end frame of the video to be generated, and N is an integer greater than 1; Acquire event camera stream information between the shooting time of the start frame and the shooting time of the end frame, and determine a first image inserted between at least one group of adjacent two wonderful frames in the N wonderful frames based on the event camera stream information; wherein the first image includes slow motion details of the target object between at least one group of adjacent two wonderful frames in the N wonderful frames; A target wonderful video is generated based on the N wonderful frames and the first image.
2. The method according to claim 1, characterized in that The number of first images inserted between at least two adjacent groups of wonderful frames in the N wonderful frames is the same.
3. The method according to claim 1, characterized in that The numbers of first images inserted between at least two adjacent groups of wonderful frames in the N wonderful frames are different.
4. The method according to claim 3, characterized in that Before determining the first image to be inserted between at least one group of adjacent two wonderful frames in the N wonderful frames based on the event camera stream information, the method includes: Based on the motion speed of the target object between two adjacent wonderful frames in the N wonderful frames, the number of first images inserted between at least one group of two adjacent wonderful frames in the N wonderful frames is determined.
5. The method according to any one of claims 1 to 4, characterized in that The preview stream image is of a first resolution, and the N highlight frames of the generated video are of a second resolution; wherein the second resolution is greater than the first resolution.
6. The method according to any one of claims 1 to 5, characterized in that The determining, based on the event camera stream information, a first image inserted between at least one group of adjacent two wonderful frames in the N wonderful frames comprises: Motion information is determined based on the event camera stream information, and a first image inserted between at least one group of adjacent two wonderful frames in the N wonderful frames is determined based on the motion information; wherein the motion information is used to represent the motion of a moving object photographed by the event camera, and the moving object includes a target object.
7. The method according to claim 6, characterized in that The determining motion information based on the event camera stream information comprises: De-noising the event camera stream information using a preset de-noising algorithm to obtain first event camera stream information; Motion information is determined based on the first event camera stream information.
8. The method according to claim 6, characterized in that The step of determining the first image to be inserted between two adjacent wonderful frames in the N wonderful frames based on the motion information comprises: The two adjacent wonderful frames include the i-th wonderful frame and the i+1-th wonderful frame. and inserting M frames of images into the i+1th wonderful frame; wherein i is an integer between 1 and N-1; For the nth frame image in the M frames, where n is an integer from 1 to H: Determine first motion difference information based on the i-th frame motion information corresponding to the shooting moment of the i-th wonderful frame and the n-th frame motion information corresponding to the n-th frame image at the shooting moment of the event camera; wherein the first motion difference information represents the change between the first i-frame motion information and the n-th frame motion information; Determine second motion difference information based on the nth frame motion information of the event camera corresponding to the nth frame image at the shooting moment and the i+1th frame motion information of the event camera corresponding to the i+1th frame image at the shooting moment; wherein the second motion difference represents the change between the first i-frame motion information and the n-frame motion information; Determine the first n-th frame image to be fused based on the i-th wonderful frame and the first motion difference information; Determine the nth frame image to be merged based on the i+1th wonderful frame and the second motion difference information; An n-th frame image is generated based on the first n-th frame image to be fused and the second n-th frame image to be fused.
9. The method according to claim 1, characterized in that: The step of generating a target wonderful video based on the N wonderful frames and the first image includes: A target wonderful video is generated based on the N wonderful frames, the first image and the second image; wherein the second image is an image captured within a preset time period before the shooting moment of the start frame.
10. The method according to claim 1, characterized in that If the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream picture, the method further comprises: detecting whether the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream picture and the event camera stream information; Detecting that the target object satisfies a preset motion posture based on the preview stream information corresponding to the preview stream picture includes: Based on the preview stream information corresponding to the preview stream picture and the event camera stream information, it is detected that the target object satisfies a preset motion posture.
11. An electronic device, characterized in that: The electronic device comprises a processor and a memory; the memory is used to store code instructions; the processor is used to run the code instructions, so that the electronic device executes the method as described in any one of claims 1-10.
12. A computer-readable storage medium, characterized in that: The method comprises computer instructions, and when the computer instructions are executed on an electronic device, the electronic device executes the method as claimed in any one of claims 1 to 10.
Citation Information
Patent Citations
Bimodal image signal processor and image sensor
CN112714301A
Video processing method, electronic equipment and readable medium
CN114827342A
Video processing method, electronic equipment and readable medium
CN116916149A
Method and apparatus for capturing dynamic images
US20200014836A1
Image processing method and device
WO2022141477A1