Video generation method, electronic equipment and computer readable storage medium
By using event cameras to capture the motion details of the target object and insert them between RGB images, the problem of difficulty in shooting slow motion details by electronic devices is solved, improving the continuity and fluency of the video, and improving the user experience.
Patent Information
- Application Number
- CN202311508927.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-23
AI Technical Summary
When existing electronic devices shoot moving target objects, it is difficult to capture the slow motion details of the target objects, resulting in poor continuity, poor fluency and poor user experience.
The event camera is used to detect the pixel light brightness changes caused by the movement of the target object, obtain the slow motion detail image during the movement of the multi-frame target object, and insert it between the RGB images collected by the first camera to generate video.
It improves the continuity and fluency of video generation, has high picture quality, fewer pseudo textures, and significantly improves the user's visual experience.
Smart Images

Figure CN120034728A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing, and in particular to a video generation method, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the development of camera functions in electronic devices, the camera functions in electronic devices are becoming more and more powerful. It can be understood that videos are generally composed of multiple frames of images, and the content presented in the video can improve the user's viewing experience. In addition to simple camera functions such as taking pictures and recording videos, electronic devices also have the function of making videos from the captured photos. For example, electronic devices can make videos of the wonderful moments captured. In this way, users do not need to spend time and energy on editing, and can also allow users to recall the wonderful moments in an immersive way.
[0003] However, currently, when electronic devices take photos, due to the relative motion between the target object and the electronic device, affected by factors such as camera exposure time and sampling frequency, the captured results are generally the non-motion continuous (discrete) wonderful moments of the target object. Moreover, electronic devices can generate a continuous video of multiple discrete wonderful moments, that is, wonderful moments with large differences in image content. The video generated by this video generation method has poor continuity, poor fluency, and poor user experience.
[0004] For example, in Figure 1 In the example of , a schematic diagram of two frames of images captured by an electronic device is shown. The electronic device captures a picture of user A jumping over a hurdle. During the shooting process, the electronic device uses a camera to capture an A-frame image and a C-frame image. The A-frame image presents a picture of user A just starting to jump over a hurdle, and the C-frame image presents a picture of the user having completed the hurdle. The electronic device can generate a video 1 from the A-frame image and the B-frame image.
[0005] In such Figure 2In the example, a schematic diagram of the display of the interface during the playback of video 1 is shown. When the electronic device starts to play video 1, the electronic device can display all the images in sequence according to the order of the images included in the video 1. Exemplarily, the electronic device can display interface 111, which includes an A frame image, and displays the A frame image from the 1st second to the 5th second. When the display reaches the 5th second, the electronic device can display interface 112, which includes a C frame image, and displays the C frame image from the 5th second to the 15th second. In this way, during the playback of video 1, after the A frame image still presents the picture of user A just starting to jump over the hurdle, the C frame image presents the picture of the user having completed the hurdle and running to the destination. Since the A frame image and the C frame image are discrete wonderful moments, that is, the user's actions presented by these two frames are not a relatively continuous process, there are no multiple slow-motion details of the hurdle during the hurdle process of user A, the continuity is not good, the fluency is poor, and the user experience is poor.
[0006] In addition, some existing solutions generate the content between the gaps in slow-motion video frames through algorithms, which also have problems such as low image quality, many pseudo-textures, poor continuity, poor smoothness, and poor user experience. Summary of the invention
[0007] The embodiments of the present application provide a video generation method, an electronic device, and a computer-readable storage medium, which are used to improve the continuity and smoothness of video generation, thereby improving the user experience.
[0008] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0009] In a first aspect, a video generation method is provided, which is applied to an electronic device, the electronic device comprising an event camera and a first camera, the method comprising: in response to a user's operation of turning on a camera function, turning on the first camera and displaying a first interface; wherein the first interface comprises a preview window, the preview window is used to display a preview stream image, and the preview stream image is an image captured by the first camera in real time; when the automatic capture mode in the camera function is turned on, turning on the event camera; wherein the event camera is used to capture and obtain event camera stream information; if it is detected based on the preview stream information corresponding to the preview stream image and the event camera stream information that the target object meets a preset motion posture, determining N frames of a generated video based on the preview stream information; wherein the N frames of the wonderful frames include a start frame and an end frame of the video to be generated, and N is an integer greater than 1; obtaining the event camera stream information between the shooting time of the start frame and the shooting time of the end frame, and determining a first image inserted between at least one group of adjacent two wonderful frames in the N frames of wonderful frames based on the event camera stream information; wherein the first image includes slow motion details of the target object between at least one group of adjacent two wonderful frames in the N frames of wonderful frames; and generating a target wonderful video based on the start frame, the end frame and the first image.
[0010] In the present application, when the electronic device uses the first camera to shoot the target object, if the target object is moving, the brightness change of the pixel points of the target object detected by the event camera due to the movement can be used to obtain multiple frames of slow-motion detail images of the target object during the movement. Then, the electronic device inserts the slow-motion detail images of the target object during the movement of multiple frames between the multiple RGB images collected by the first camera to generate a video. The video generated according to the above scheme has higher image quality, less pseudo-texture, better continuity, and higher smoothness. And the electronic device can play each frame of the image according to the arranged image, which is convenient for users to watch the slow-motion wonderful moments slowly, and the user's viewing experience is better.
[0011] In a possible implementation manner of the first aspect, the number of first images inserted between at least two adjacent groups of highlight frames in the N highlight frames is the same.
[0012] That is, the number of first images inserted between every two adjacent wonderful frames in N frames is the same, or the number of first images inserted between not all groups of two adjacent wonderful frames in N frames is the same, but the number of first images inserted between two or more groups of two adjacent wonderful frames in N frames is the same.
[0013] In a possible implementation manner of the first aspect, the numbers of first images inserted between at least two groups of adjacent two highlight frames in the N highlight frames are different.
[0014] That is, the number of first images inserted between each adjacent two wonderful frames in N wonderful frames is different; the number of first images inserted between all groups of adjacent two wonderful frames in N wonderful frames is not different, but the number of first images inserted between two groups and more than two adjacent wonderful frames in N wonderful frames is different. For example, the number of first images inserted between the first group of adjacent two wonderful frames in N wonderful frames is 2 frames, the number of first images inserted between the second group of adjacent two wonderful frames in N wonderful frames is 2 frames, and the number of first images inserted between the third group of adjacent two wonderful frames in N wonderful frames is 0 frames, that is, no images are inserted between the third group of adjacent two wonderful frames in N wonderful frames. For another example, the number of first images inserted between the first group of adjacent two wonderful frames in N wonderful frames is 3 frames, the number of first images inserted between the second group of adjacent two wonderful frames in N wonderful frames is 2 frames, and the number of first images inserted between the third group of adjacent two wonderful frames in N wonderful frames is 2 frames, that is, no images are inserted between the third group of adjacent two wonderful frames in N wonderful frames.
[0015] In a possible implementation manner of the first aspect, determining a first image inserted between at least one group of adjacent two highlight frames in N highlight frames based on event camera stream information includes:
[0016] Based on the motion speed of the target object between two adjacent wonderful frames in the N wonderful frames, the number of first images inserted between at least one group of two adjacent wonderful frames in the N wonderful frames is determined.
[0017] In the present application, the speed at which the target object moves at each moment may be different. When the target object moves faster, within a certain period of time, the target object has more actions, the action richness is higher, and these actions are difficult to be captured more clearly by the first camera. Therefore, when generating a video, the electronic device can compensate for the motion details of the target object that are not captured by the first camera based on the event stream information captured by the event camera. Specifically, the electronic device can insert more images when the target object moves faster according to the speed of the target object. In this way, the continuity and smoothness of video generation can be improved, thereby improving the user experience.
[0018] When the target object moves slowly, within a certain period of time, the target object has fewer movements, the richness of the movements is low, and these movements are easily captured more clearly by the first camera. Therefore, when generating a video, the electronic device can basically obtain the movement details of the target object without compensating for the movement details of the target object. Specifically, the electronic device can insert fewer images when the target object moves slowly according to the movement speed of the target object. In this way, a video with higher continuity and smoothness can be generated, and the power consumption of the electronic device inserting images can be reduced, thereby improving the user experience.
[0019] In a possible implementation manner of the first aspect, the preview stream image is at a first resolution, and the N highlight frames of the generated video are at a second resolution; wherein the second resolution is greater than the first resolution.
[0020] That is, the resolution of the image displayed on the preview interface of the electronic device is lower than the resolution of the image of the video to be generated. When the electronic device determines the wonderful frame, it generates an image with a higher resolution, which can improve the viewing experience of the user.
[0021] In a possible implementation of the first aspect, determining a first image inserted between at least one group of adjacent two highlight frames in N highlight frames based on event camera stream information includes: determining motion information based on the event camera stream information, and determining the first image inserted between at least one group of adjacent two highlight frames in N highlight frames based on the motion information; wherein the motion information is used to represent the motion of a moving object captured by the event camera, and the moving object includes a target object.
[0022] The motion of the moving object includes the motion of the target object and the motion of objects in the background of the target object. Based on the motion of the moving object captured by the event camera, the electronic device can also accurately obtain a clearer image containing the target object when the target object is moving. Since the generated inserted image is clearer, the electronic device can obtain a video with higher picture quality, present a higher quality video to the user, and provide a higher user experience.
[0023] In a possible implementation manner of the first aspect, the motion information includes optical flow information, where the optical flow information is used to represent changes in pixel brightness in an image captured by the event camera.
[0024] Optical flow refers to the change in pixel brightness caused by light obstruction or other reasons in the surrounding environment of the object during the shooting time. Based on the change in pixel brightness, the electronic device can more accurately obtain a clearer image of the target object's movement. Since the generated inserted image is clearer, the electronic device can obtain a video with higher picture quality, present a higher quality video to the user, and provide a better user experience.
[0025] In a possible implementation manner of the first aspect, determining motion information based on event camera stream information includes: denoising the event camera stream information by using a preset denoising algorithm to obtain first event camera stream information; and determining motion information based on the first event camera stream information.
[0026] The electronic device denoises the event camera stream, and can more accurately obtain the movement of the target object, thereby generating a clearer image to be inserted. Since the generated image to be inserted is clearer, the electronic device can obtain a video with higher picture quality, present a higher quality video to the user, and provide a higher user experience.
[0027] In a possible implementation of the first aspect, a first image inserted between two adjacent wonderful frames in N wonderful frames is determined based on motion information, including: two adjacent wonderful frames include an i-th wonderful frame and an i+1-th wonderful frame, and M frames of images are inserted between the i-th wonderful frame and the i+1-th wonderful frame; wherein i is an integer from 1 to N-1; for the n-th frame of the M frames, wherein n is an integer from 1 to H in sequence: based on the i-th frame motion information corresponding to the shooting moment of the i-th wonderful frame, and the n-th frame motion information of the event camera corresponding to the n-th frame of the image at the shooting moment, first motion difference information is determined; wherein the first motion difference information represents the n-th frame of the M frames of images. The change between the i-th frame motion information and the n-th frame motion information; based on the n-th frame motion information of the event camera corresponding to the n-th frame image at the shooting moment, and the i+1-th frame motion information of the event camera corresponding to the i+1-th frame of the wonderful frame at the shooting moment, determine the second motion difference information; wherein the second motion difference represents the change between the first i-frame motion information and the n-th frame motion information; determine the first n-th frame image to be fused based on the i-th frame of the wonderful frame and the first motion difference information; determine the second n-th frame image to be fused based on the i+1-th frame of the wonderful frame and the second motion difference information; generate the n-th frame image based on the first n-th frame image to be fused and the second n-th frame image to be fused.
[0028] In a possible implementation of the first aspect, a target wonderful video is generated based on a start frame, an end frame and a first image, including: generating a target wonderful video based on a start frame, an end frame, a first image and a second image; wherein the second image is an image captured within a preset time period before the start frame is captured.
[0029] The electronic device generates a video together with the images captured within a preset time period before the start frame is captured, the start frame, the end frame, and the first image. While presenting the user with the exciting movement moments of the target object, it can also present the user with the behavior of the target object before the movement, so that a video with richer picture content can be obtained, thereby improving the user experience.
[0030] In a possible implementation manner of the first aspect, if before detecting that the target object satisfies a preset motion posture based on the preview stream information corresponding to the preview stream screen, the method also includes: detecting whether the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information; detecting that the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream screen includes: detecting that the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information.
[0031] The event camera can more accurately track the changes in the posture of the target object during the movement. The event camera stream information captured by the event camera can assist in previewing the stream information and more accurately determine whether there is a target object that meets the preset motion posture.
[0032] In a second aspect, an electronic device is provided, which includes a processor and a memory; the memory is used to store code instructions; the processor is used to run the code instructions to execute a method for adjusting an audio signal in any possible design manner as in the first aspect.
[0033] In a third aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes a method for adjusting an audio signal in any possible design manner as in the first aspect.
[0034] In a fourth aspect, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the method in any possible design manner in the first aspect.
[0035] Among them, the technical effects brought about by any design method in the second aspect, the third aspect and the fourth aspect can refer to the technical effects brought about by different design methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A schematic diagram showing two frames of images taken by an electronic device;
[0037] Figure 2 A schematic diagram showing the display of the interface during video playback;
[0038] Figure 3 A schematic diagram of three frames of images obtained based on shooting by an event camera and a first camera is shown;
[0039] Figure 4 A schematic structural diagram of a mobile phone 100 is shown;
[0040] Figure 5 A schematic diagram of a process of video generation method is shown;
[0041] Figure 6 A schematic diagram showing the changes in the operation interface of the mobile phone 100 when the camera is turned on and the automatic snapshot mode in the camera is shown;
[0042] Figure 7 A schematic diagram showing the changes in the operation interface of the mobile phone 100 when the camera is turned on and the automatic snapshot mode in the camera is shown;
[0043] Figure 8A schematic diagram of a wonderful frame recognition algorithm is shown;
[0044] Fig. 9 Another schematic diagram of a wonderful frame recognition algorithm is shown;
[0045] Fig.10 A schematic diagram showing a comparison of the shooting frequencies of the first camera and the event camera is shown;
[0046] Fig.11 A schematic diagram showing a process of the mobile phone 100 determining N wonderful frames of a generated video from a preview stream image;
[0047] Fig.12 A schematic diagram of inserting a slow motion detail image is shown;
[0048] Fig.13 A schematic diagram of image interpolation is shown;
[0049] Fig.14 A schematic diagram showing another image interpolation method is shown;
[0050] Fig.15 A schematic diagram of a process for generating an insert image is shown;
[0051] Fig.16 A schematic diagram of a process for generating an interpolated image is shown;
[0052] Fig.17 A schematic diagram showing changes in the interface of a gallery is shown;
[0053] Fig.18 A schematic diagram of a process of a video generation method is shown. DETAILED DESCRIPTION
[0054] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0055] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.
[0056] In the embodiments of the present application, the words "exemplarily" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific way.
[0057] In order to solve the technical problems in the background technology, an embodiment of the present application proposes a video generation method.
[0058] It is understandable that with the development of photography technology, people have higher and higher requirements for photography quality, especially the clarity of photography. At present, when electronic devices such as mobile phones take photos of moving target objects, the target object and the electronic device have relative motion. Affected by factors such as camera exposure time and sampling frequency, the captured results are generally the wonderful moments of the target object without motion (discrete).
[0059] The image sensor corresponding to the event-based vision (EVS) of the electronic device can detect in real time the brightness changes of the pixels caused by the movement of the target object (the change is referred to as a motion event). Even if the relative movement speed between the camera (first camera) and the photographed target object is fast, the posture changes of the target object during the movement can be determined more accurately.
[0060] Therefore, the electronic device can use the event camera to compensate for the problem that the first camera cannot capture the slow-motion details of the target object when the target object moves faster. If the target object is moving, the brightness change of the pixel points of the target object detected by the event camera due to the movement can be used to obtain a multi-frame slow-motion detail image of the target object during the movement. Then, the electronic device inserts the multi-frame slow-motion detail image of the target object during the movement between the multi-frame images (such as RGB images) collected by the first camera to generate a video.
[0061] In some embodiments, the posture of the target object during the motion process meets certain conditions, and the photos of the target object taken by the electronic device at these moments are called wonderful frames of the target object.
[0062] In the embodiment of the present application, the video generated according to the above scheme has high image quality, less pseudo texture, better continuity and higher fluency. Moreover, the electronic device can play each frame of the image according to the arranged image, which is convenient for users to watch the wonderful moments of slow motion at a slow speed, and the user's viewing experience is better.
[0063] For example, Figure 3FIG. 4 shows a schematic diagram of three frames of images obtained by shooting based on an event camera and a first camera. Figure 3 As shown, the electronic device uses a camera (first camera) to obtain the A-frame image and the C-frame image. The event camera can capture the event flow information between the shooting time corresponding to the A-frame image and the shooting time corresponding to the C-frame image, and based on the event flow information and the A-frame image and the C-frame image, generate the B-frame image corresponding to the motion details (slow motion details) of the target object between the aforementioned two moments, that is, the B-frame image presents the slow motion details of the hurdle jump of user A during the hurdle jump process. Then, the electronic device can insert the B-frame image between the A-frame image and the C-frame image, and generate video 2 with the A-frame image, the B-frame image and the C-frame image.
[0064] In such Figure 2 In the example, the display of the interface during the playback of video 2 is shown. The electronic device can display all the images in sequence according to the order of the images included in the video 1. Exemplarily, when the electronic device starts to play video 2, the electronic device can display interface 113, and the display interface 113 includes an A frame image, and the A frame image is displayed from the 1st second to the 5th second. When the display reaches the 5th second, the electronic device can display interface 114, and the display interface 114 includes a B frame image, and the B frame image is displayed from the 5th second to the 15th second. When video 1 is about to finish playing, for example, at the low 15 seconds, the electronic device can display interface 115, and the display interface 115 includes a C frame image.
[0065] In this way, during the playback of video 2, frame A presents the scene when user A just starts to jump over the hurdles, frame B presents the scene when user A is about to complete the hurdles, and frame C presents the scene when the user has completed the hurdles and is running towards the destination. The user's actions presented in these three frames are a relatively continuous process, with high image quality, good continuity, less pseudo-texture, and better user experience.
[0066] Exemplarily, the electronic device in the embodiments of the present application may be a mobile phone, a tablet computer, a desktop, a laptop, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, and other devices with a wireless network connection function. The embodiments of the present application do not impose any special restrictions on the specific form of the electronic device.
[0067] Figure 4 FIG. 1 shows a schematic diagram of the structure of a mobile phone 100. Figure 4 As shown, the mobile phone 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0068] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0069] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0070] The controller may be the nerve center and command center of the mobile phone 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0071] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or circulated. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. Repeated access is avoided, the waiting time of the processor 110 is reduced, and the efficiency of the system is improved. The processor is used to execute the video generation method provided in the embodiment of the present application.
[0072] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0073] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present application is only a schematic illustration and does not constitute a structural limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0074] The charging management module 140 is used to receive charging input from a charger. The charger may be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 may receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 may receive wireless charging input through a wireless charging coil of the mobile phone 100. While the charging management module 140 is charging the battery 142, it may also power the mobile phone 100 through the power management module 141.
[0075] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), etc. In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0076] The wireless communication function of the mobile phone 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0077] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0078] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the mobile phone 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0079] The modem processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be sent into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After the low-frequency baseband signal is processed by the baseband processor, it is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker 170A, a receiver 170B, etc.), or displays an image or video through a display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0080] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the mobile phone 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, modulates the frequency of the electromagnetic wave signal and performs filtering, and sends the processed signal to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, modulate the frequency of it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0081] In some embodiments, the antenna 1 of the mobile phone 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the mobile phone 100 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0082] The mobile phone 100 implements the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.
[0083] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0084] The mobile phone 100 can realize the shooting function through ISP, camera 193, video codec, GPU, display screen 194 and application processor.
[0085] ISP is used to process the data fed back by camera 193. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 193.
[0086] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the mobile phone 100 may include 1 or N cameras 193, where N is a positive integer greater than 1. In an embodiment of the present application, the camera 193 may include a first camera and an event camera, and the first camera may be an RGB camera or a black and white camera.
[0087] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the mobile phone 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0088] Video codecs are used to compress or decompress digital videos. Mobile phone 100 may support one or more video codecs. Thus, mobile phone 100 may play or record videos in various coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0089] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function, such as storing music, video and other files in the external memory card.
[0090] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the mobile phone 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the mobile phone 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0091] The mobile phone 100 can implement audio functions such as music playing and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor.
[0092] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or some functional modules of the audio module 170 can be arranged in the processor 110.
[0093] The speaker 170A, also called a "speaker", is used to convert an audio electrical signal into a sound signal. The mobile phone 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0094] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the mobile phone 100 receives a call or voice message, the voice can be received by placing the receiver 170B close to the ear.
[0095] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The mobile phone 100 can be provided with at least one microphone 170C. In other embodiments, the mobile phone 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the mobile phone 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, realize directional recording function, etc.
[0096] The earphone interface 170D is used to connect a wired earphone and can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0097] The key 190 includes a power key, a volume key, etc. The key 190 can be a mechanical key or a touch key. The mobile phone 100 can receive key input and generate key signal input related to the user settings and function control of the mobile phone 100.
[0098] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0099] Indicator 192 may be an indicator light, which may be used to indicate charging status, power changes, messages, missed calls, notifications, etc.
[0100] The SIM card interface 195 is used to connect the SIM card. The SIM card can be connected to and separated from the mobile phone 100 by inserting it into the SIM card interface 195 or pulling it out from the SIM card interface 195. The mobile phone 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The mobile phone 100 interacts with the network through the SIM card to realize functions such as calls and data communications. In some embodiments, the mobile phone 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the mobile phone 100 and cannot be separated from the mobile phone 100.
[0101] Figure 5 FIG. 1 shows a flow chart of a video generation method. Figure 5 As shown, the process includes the following steps:
[0102] S501. In response to a user's operation of turning on a camera function, the mobile phone 100 turns on a first camera and displays a first interface; wherein the first interface includes a preview window, the preview window is used to display a preview stream image, and the preview stream image is an image captured by the first camera in real time.
[0103] It is understandable that when the mobile phone 100 turns on the camera function, the picture taken by the first camera in real time, ie, the preview stream picture, will be displayed for the user to watch. The first camera can be an RGB camera or a black and white camera.
[0104] If the first camera is an RGB camera, the preview stream image is a color image. Exemplarily, the format of the preview stream image can be an RGB image, a RAW image, or a YUV image.
[0105] Figure 6 1 shows a schematic diagram of the operation interface changes of the mobile phone 100 when the camera is turned on and the automatic snapshot mode in the camera. Figure 6 As shown in FIG. 4( a ), a camera icon 4011 is displayed in the interface 401 of the mobile phone 100. In response to the user clicking the camera icon 4011, the mobile phone 100 turns on the first camera and displays the interface 402. Figure 6 As shown in FIG. 4( b ), the interface 402 includes a preview window 4022 , and the preview window 4022 is used to display the picture taken by the first camera.
[0106] S502: When the automatic snapshot mode in the camera function is turned on, the mobile phone 100 turns on the event camera after turning on the first camera; wherein the event camera is used to capture and obtain event camera stream information.
[0107] It can be understood that the first camera (frame-based image sensor) outputs an overall image including the object and the background, such as an RGB stream, according to the frame rate and at a certain time interval. The event camera (event-based visual sensor) detects brightness changes in an asynchronous manner, and outputs pixel data with changes in combination with coordinate and time information, that is, it captures the trajectory information of the moving object with a very high temporal resolution and outputs the EVS stream. The event camera can capture the action details of the moving target object during the movement. In this way, when the mobile phone 100 turns on the first camera and the event camera, the mobile phone 100 is aimed at the current scene, continuously outputs the preview stream (RGB stream + EVS stream), and uses the EVS stream to guide the interpolation. In this way, the mobile phone 100 can capture the action details of the moving target object during the movement, which is convenient for the subsequent generation of a video with higher continuity of image content.
[0108] like Figure 6 As shown in FIG. 4( b ), the interface 402 also includes a setting control 4021. In response to the user clicking the setting control 4021, the mobile phone 100 displays the setting interface. Figure 6 As shown in FIG. 4( c ), the setting interface includes a smart photo option. In response to the user clicking on the expansion control 4031 corresponding to the smart photo option, the mobile phone 100 displays the smart photo interface. Figure 6 As shown in FIG. 4(d), the mobile phone 100 displays an intelligent photo taking interface 404. The intelligent photo taking interface 404 includes an automatic snapshot option and a switch 4041 for triggering the automatic snapshot mode to be turned on or off. The mobile phone 100 enters the automatic snapshot mode in response to the user turning on the switch 4041. The automatic snapshot mode can be used to automatically take photos when the automatic snapshot is turned on, and intelligently recognizes that there is a picture in which the target object meets the preset motion posture. For example, when intelligently recognizing the wonderful moments of people smiling, jumping, running, cats, dogs, etc., photos are automatically taken.
[0109] As a possible implementation, after the automatic snapshot mode is turned on in the settings, the mobile phone 100 turns off the camera and then turns it on again, the camera automatically enters the automatic snapshot mode. If the automatic snapshot mode is turned on in the above manner, the mobile phone 100 may not execute S502 after executing S501 and automatically enters the automatic snapshot mode.
[0110] In some other embodiments, the mobile phone 100 turns on the camera and sets an automatic snapshot control in the camera interface. The mobile phone 100 can respond to the user's operation on the automatic snapshot control to turn on the automatic snapshot mode. For example, the mobile phone 100 responds to the user clicking the automatic snapshot control in the camera interface to turn on the automatic snapshot mode. Figure 7 1 shows a schematic diagram of the operation interface changes of the mobile phone 100 when the camera is turned on and the automatic snapshot mode in the camera. Figure 7 As shown in FIG. 1 (a), the interface 402 also includes an automatic snapshot control 4023. When no operation is performed, the automatic snapshot control 4023 is displayed in the first display mode, for example, only the outline is displayed. After the mobile phone 100 responds to the user clicking the automatic snapshot control 4023, the automatic snapshot control 4023 changes from the first display mode to the second display mode. Figure 7 As shown in FIG. 4( b ), the outline area of the automatic snapshot control 4023 is filled with color. Of course, the first display mode may also be that the automatic snapshot control 4023 is displayed in a dark color, and the second display mode may be that the automatic snapshot control 4023 is displayed in a bright color.
[0111] As a possible implementation, after the automatic capture mode is turned on from the camera interface, the mobile phone 100 needs to operate the automatic capture control again after turning off the camera and turning on the camera again, so that the camera will be in the automatic capture mode again. If the automatic capture mode is turned on in the above manner, the mobile phone 100 can also execute S502 to enter the automatic capture mode after executing S501.
[0112] S503: The mobile phone 100 detects whether the target object satisfies a preset motion posture based on the preview stream information corresponding to the preview stream image.
[0113] If the mobile phone 100 detects that the target object does not meet the preset motion posture based on the preview stream information corresponding to the preview stream screen, it is not necessary to generate a picture in the shooting mode, and the process ends. If the mobile phone 100 detects that the target object meets the preset motion posture based on the preview stream information corresponding to the preview stream screen, it can generate a picture in the shooting mode as the material for generating a wonderful frame video.
[0114] If a target object exists in a preview frame of the preview stream, and the target object satisfies a preset motion posture, the preview frame is determined to be a wonderful frame. That is, the image refers to an image frame with meaningful image content. Meaningful image content includes clear image content in the image, such as people, animals, etc. It can be understood that in some embodiments, clear image content also includes people smiling with eyes open, or people smiling, laughing, smiling with eyes closed, blinking, and pouting. Furthermore, meaningful image content also includes image content making preset specific movements. For example, people jumping up, athletes jumping over hurdles, puppies running, etc.
[0115] The target object can be a person, an animal, a vehicle (such as a vehicle in a racing scene), etc.
[0116] The preset motion gestures may be motion gestures in various sports scenes, for example, the motions of people in various sports game scenes. For example, in a basketball game scene, the motions of athletes waving, jumping, shooting, and flying. For another example, in a racing scene, the scenes of vehicles driving fast and turning. For another example, in an animal sports scene, the scenes of horses running fast and dogs running fast.
[0117] It should be noted that the above examples are only examples and are not exhaustive. In other embodiments of the present application, the target object may also include other objects, and the preset motion gesture may also be other gestures.
[0118] In some embodiments, the mobile phone 100 can detect whether there is a picture in the preview stream in which the target object satisfies a preset motion posture through a wonderful frame recognition algorithm.
[0119] Figure 8 FIG. 1 shows a schematic diagram of a wonderful frame recognition algorithm. Figure 8 As shown, the wonderful frame recognition algorithm includes a target object motion posture recognition network 1, which is used to recognize whether there is a target object in the image and whether the target object satisfies a preset motion posture. If the target object motion posture recognition network recognizes that there is a target object in a preview frame of the preview stream image, and the target object satisfies the preset motion posture, the preview frame is determined to be a wonderful frame.
[0120] The mobile phone 100 may input the preview stream information corresponding to the preview stream screen into the target object motion posture recognition network 1, and the target object motion posture recognition network 1 may output a recognition result of whether it is a wonderful frame.
[0121] It can be understood that the event camera can more accurately track the changes in the posture of the target object during the movement. The event camera stream information captured by the event camera can assist the preview stream information to more accurately determine whether there is a target object that meets the preset motion posture. Therefore, in some other embodiments, the mobile phone 100 can also detect whether the target object meets the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information. If the mobile phone 100 detects that the target object does not meet the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information, there is no need to generate a picture in the shooting mode, and the process ends. If the mobile phone 100 detects that the target object meets the preset motion posture based on the preview stream information corresponding to the preview stream screen and the event camera stream information, a picture in the shooting mode can be generated as a material for generating a wonderful frame video.
[0122] As a possible implementation, the mobile phone 100 may also input the preview stream information and event camera stream information corresponding to the preview stream image into a target object motion posture recognition network, which is used to recognize whether there is a target object in the image and whether the target object satisfies a preset motion posture. If the target object motion posture recognition network recognizes that there is a target object in a preview frame of the preview stream image and the target object satisfies the preset motion posture, the preview frame is determined to be a wonderful frame.
[0123] For example, Fig. 9 FIG. 2 shows another schematic diagram of a wonderful frame recognition algorithm. Fig. 9 As shown, the wonderful frame recognition algorithm includes a target object motion posture recognition network 2, which is used to recognize whether there is a target object in the image and whether the target object meets a preset motion posture.
[0124] The mobile phone 100 may input the preview stream information and the event camera stream information corresponding to the preview stream screen into the target object motion posture recognition network 2. The target object motion posture recognition network 2 may output a recognition result of whether it is a wonderful frame.
[0125] Generally, the shooting frequency of the event camera is greater than the shooting frequency of the first camera. For example, the shooting frequency of the event camera is 1000 frames per second, and the shooting frequency of the first camera is 30 frames per second. 1000 frames per second is greater than 30 frames per second.
[0126] Fig.10 FIG. 4 shows a schematic diagram comparing the shooting frequencies of the first camera and the event camera. Fig.10 As shown, the shooting frequency of the first camera is 30 frames per second, approximately once every 33ms, and the shooting frequency of the event camera is 1000 frames per second, approximately once every 1ms.
[0127] Combination Fig. 9 and Fig.10 It can be seen that the mobile phone 100 can input the preview information corresponding to the first shooting of the first camera and the one frame of event camera information corresponding to the event camera, for example, the one frame of event camera information shot by the event camera for the first time, to the target object motion posture recognition network 2. The target object motion posture recognition network 2 can output the recognition result of whether the preview information corresponding to the first shooting of the first camera is a wonderful frame.
[0128] The mobile phone 100 can input the preview information corresponding to the second shooting of the first camera and a frame of event camera information corresponding to the shooting of the event camera, for example, a frame of event camera information shot for the 33rd time by the event camera, to the target object motion posture recognition network 2. The target object motion posture recognition network 2 can output a recognition result of whether the preview information corresponding to the second shooting of the first camera is a wonderful frame.
[0129] The mobile phone 100 can input the preview information corresponding to the third shot of the first camera and a frame of event camera information corresponding to the event camera, for example, a frame of event camera information shot for the 66th time by the event camera, to the target object motion posture recognition network 2. The target object motion posture recognition network 2 can output a recognition result of whether the preview information corresponding to the third shot of the first camera is a wonderful frame.
[0130] If the preview information corresponding to the 1st to 3rd shots of the first camera are all identified as wonderful frames, the mobile phone 100 can insert an image between the wonderful frame corresponding to the 1st shot of the first camera and the wonderful frame corresponding to the 2nd shot, and insert an image between the wonderful frame corresponding to the 2nd shot of the first camera and the wonderful frame corresponding to the 2nd shot, so as to generate a wonderful frame video.
[0131] If the preview information corresponding to the first and third shots of the first camera are both identified as highlight frames, the mobile phone 100 may insert an image between the highlight frame corresponding to the first shot of the first camera and the highlight frame corresponding to the third shot to generate a highlight frame video.
[0132] It can be understood that the multiple frames of preview information obtained by the first camera during multiple consecutive shootings can be referred to as preview stream information. The multiple frames of event camera information obtained by the event camera during multiple consecutive shootings can be referred to as event camera stream information. For example, Fig.10 As shown, the 2 frames of preview information corresponding to the first camera's 1st to 2nd shooting can be called preview stream information. The 3 frames of preview information corresponding to the first camera's 1st to 3rd shooting can also be called preview stream information. The 33 frames of event camera information corresponding to the event camera's 1st to 33rd shooting can be called event camera stream information. The 66 frames of event camera information corresponding to the event camera's 1st to 66th shooting can be called event camera stream information.
[0133] S504. The mobile phone 100 determines to generate N wonderful frames of the video based on the preview stream information; wherein the N wonderful frames include the start frame and the end frame of the video to be generated.
[0134] That is, the mobile phone 100 can determine N wonderful frames of the video to be generated based on the preview stream information of the target object satisfying the preset motion posture. The number of wonderful frames of the wonderful frame video to be generated can be multiple frames, and N is an integer greater than 1, for example, N is 2, 3, 4, 5, etc.
[0135] Fig.11 FIG. 1 shows a schematic diagram of a process in which the mobile phone 100 determines N wonderful frames of a generated video from preview stream information. Fig.11As shown, if the target object motion posture recognition network recognizes that a target object exists in a preview frame of the preview stream information, and the target object satisfies a preset motion posture, the preview frame is determined to be a wonderful frame. If the preview frame is a wonderful frame, the preview frame is taken out, and the preview frame is subjected to image quality processing to obtain an image with higher image quality. Higher image quality can be higher image clarity and / or higher image resolution.
[0136] Specifically, the mobile phone 100 can input the preview frame into a neural network, and the neural network can be used for image quality processing to obtain an image with higher image quality.
[0137] It can be understood that indicators of image quality include but are not limited to clarity, noise, color, etc.
[0138] In this way, when the mobile phone 100 determines a wonderful frame, it generates an image with higher resolution and / or clarity, which can enhance the user's viewing experience.
[0139] On the other hand, the mobile phone 100 can also take out the event camera stream information corresponding to the shooting moment of the preview frame, denoise the event camera stream information, and obtain the denoised event camera stream information. It should be noted that since the shooting frequency of the event camera is greater than the shooting frequency of the first camera, between the preview frames at the two adjacent shooting moments, there are also multiple frames of event camera stream information shot by the event camera. The mobile phone 100 can take out the event camera stream information corresponding to the shooting time period between two adjacent wonderful frame preview frames, denoise all event camera stream information within the time period, and obtain the denoised event camera stream information. The denoising method can refer to the relevant technology, which will not be repeated here. The denoising algorithm of event camera information or event camera stream information can refer to the relevant technology, which will not be repeated here.
[0140] For example, Fig.10 As shown, if the preview information corresponding to the 1st to 3rd shooting of the first camera is all identified as wonderful frames, the mobile phone 100 can denoise the event camera information obtained by the 1st shooting of the event camera, the event camera information obtained by the 66th shooting of the event camera, and the event camera stream information obtained from the 1st to 66th shooting of the event camera.
[0141] If the preview information corresponding to the first and second shots of the first camera are both identified as highlight frames, the mobile phone 100 may denoise the event camera information obtained by the first shot of the event camera and the event camera information obtained by the 33rd shot of the event camera.
[0142] The mobile phone 100 denoises the event camera stream and can accurately obtain a clearer image of the target object's motion. Since the generated inserted image is clearer, the mobile phone 100 can obtain a video with higher picture quality, present a higher quality video to the user, and provide a better user experience.
[0143] S505 . The mobile phone 100 obtains event camera stream information between the shooting time of the start frame and the shooting time of the end frame, and determines motion information based on the event camera stream information.
[0144] The motion information is used to represent the motion of the moving object photographed by the event camera. The moving object includes at least the target object and may also include objects in the background of the target object.
[0145] The motion of the moving object includes the motion of the target object and the motion of the objects in the background of the target object. Based on the motion of the moving object captured by the event camera, the mobile phone 100 can also accurately obtain a clearer image containing the target object when the target object is moving. Since the generated inserted image is clearer, the mobile phone 100 can obtain a video with higher picture quality, present a higher quality video to the user, and provide a higher user experience.
[0146] The motion information may include optical flow information, which is used to indicate changes in pixel brightness in images captured by the event camera. Optical flow refers to changes in pixel brightness caused by light obstruction or other reasons in the surrounding environment of the object during the capture time. Based on the changes in pixel brightness, the mobile phone 100 can more accurately obtain a clearer image of the target object's motion. Since the generated inserted image is clearer, the mobile phone 100 can obtain a video with higher picture quality, present a higher quality video to the user, and provide a higher user experience.
[0147] The mobile phone 100 may refer to related technologies to determine motion information based on event camera stream information, which will not be described in detail here.
[0148] It will be appreciated that the event camera stream information is used to determine the inserted image. For example, Fig.12 FIG. 4 shows a schematic diagram of inserting a slow motion detail image. Fig.12 As shown in the figure, if an image is inserted at the middle time T / 2 between the shooting time 0 and the shooting time T, then the image I to be played corresponding to the shooting time 0 can be inserted. 0 , the image to be played corresponding to time T I t , the event camera stream information from time 0 to the intermediate time T / 2 and the event camera stream information from the intermediate time T / 2 to time T obtain the image inserted at the intermediate time T / 2.
[0149] Specifically, the mobile phone 100 can obtain the event camera stream information between the shooting time of the start frame and the shooting time of the end frame, and determine the motion information in the aforementioned time period based on the aforementioned event camera stream. In this way, the mobile phone 100 can insert a slow-motion detail image of the target object moving in the aforementioned time period based on the motion information in the aforementioned time period.
[0150] S506. The mobile phone 100 determines a first image to be inserted between each two adjacent wonderful frames in the N wonderful frames based on the motion information; wherein the first image includes slow motion details of the target object between each two adjacent wonderful frames in the N wonderful frames.
[0151] The mobile phone 100 may insert a slow-motion detail image, ie, a first image, of the target object moving within the aforementioned time period based on the motion information within the aforementioned time period.
[0152] It can be understood that, in some embodiments, the number of first images to be inserted between every two adjacent wonderful frames in the N second wonderful frames is the same.
[0153] For example, taking N as 4, Fig.13 FIG. 2 shows a schematic diagram of an image interpolation process. Fig.13 As shown, wonderful frame 1, wonderful frame 2, wonderful frame 3 and wonderful frame 4 are determined 4 wonderful frames. In order to form a video with continuous and smooth picture content, M frames of images can be inserted between wonderful frame 1 and wonderful frame 2, M frames of images can be inserted between wonderful frame 2 and wonderful frame 3, and M frames of images can be inserted between wonderful frame 3 and wonderful frame 4.
[0154] Among them, the M frame images inserted between the wonderful frame 1 and the wonderful frame 2 are the slow-motion details of the target object's movement from the shooting moment 1 of the wonderful frame 1 to the shooting moment 2 of the wonderful frame 2, and the content of each frame of the M frame images may be different, and the specific content is the actual movement picture of the target object from the shooting moment 1 to the shooting moment 2.
[0155] The M frames of images inserted between wonderful frame 2 and wonderful frame 3 are slow-motion details of the target object's movement from shooting moment 2 of wonderful frame 2 to shooting moment 3 of wonderful frame 3, and the content of each frame of the M frames of images may be different, and the specific content is the actual movement picture of the target object from shooting moment 2 to shooting moment 3.
[0156] The M-frame images inserted between wonderful frame 3 and wonderful frame 4 are slow-motion details of the target object's movement from shooting moment 3 of wonderful frame 3 to shooting moment 3 of wonderful frame 4, and the content of each frame of the M-frame images may be different, and the specific content is the actual movement picture of the target object from shooting moment 3 to shooting moment 4.
[0157] In addition to the above-mentioned method in which the number of first images to be inserted between every two adjacent wonderful frames in N frames is the same, in some other embodiments, it is also possible that the number of first images to be inserted between not all groups of two adjacent wonderful frames in N frames is the same, but the number of first images to be inserted between two or more groups of two adjacent wonderful frames in N frames is the same.
[0158] In some embodiments, the number of first images to be inserted between each adjacent two wonderful frames in the second wonderful frame of the M frame is not the same. Specifically, based on the motion speed of the target object between the adjacent two wonderful frames in the second wonderful frame of the M frame, the number of first images to be inserted between each adjacent two wonderful frames in the second wonderful frame of the M frame is determined. If the motion speed of the target object between the adjacent two wonderful frames in the second wonderful frame of the M frame is faster, the target object moves more and has more motion details between the adjacent two wonderful frames. In order to improve the continuity and fluency of the image playback content in the video, it can be determined that the number of first images to be inserted between the adjacent two wonderful frames in the second wonderful frame of the M frame is larger. If the motion speed of the target object between the adjacent two wonderful frames in the second wonderful frame of the M frame is slower, the target object moves less and has less motion details between the adjacent two wonderful frames. It is determined that the number of first images to be inserted between the adjacent two wonderful frames in the second wonderful frame of the M frame is smaller, which can also ensure the continuity and fluency of the image playback content in the video.
[0159] In an embodiment of the present application, the speed at which the target object moves at each moment may be different. When the target object moves faster, within a certain period of time, the target object has more actions, the action richness is high, and these actions are difficult to be captured by the first camera. Therefore, when generating a video, the mobile phone 100 can compensate for the motion details of the target object that are not captured by the first camera based on the event stream information obtained by the event camera. Specifically, the mobile phone 100 can insert more images when the target object moves faster according to the speed of the target object. In this way, the continuity and fluency of video generation can be improved, thereby improving the user experience.
[0160] When the target object moves slowly, within a certain period of time, the target object has fewer movements, the richness of the movements is low, and these movements are easily captured by the first camera. Therefore, when generating a video, the mobile phone 100 can basically obtain the movement details of the target object without compensating for the movement details of the target object. Specifically, the mobile phone 100 can insert fewer images when the target object moves slowly according to the movement speed of the target object. In this way, a video with higher continuity and fluency can be generated, and the power consumption of the mobile phone 100 when inserting images can be reduced, thereby improving the user experience.
[0161] In addition to the above-mentioned manner that the number of first images to be inserted between each two adjacent wonderful frames in N wonderful frames is different, in some other embodiments, it is also possible that the number of first images to be inserted between not all groups of two adjacent wonderful frames in N wonderful frames is different, but the number of first images to be inserted between two or more groups of two adjacent wonderful frames in N wonderful frames is different. For example, the number of first images to be inserted between the first group of two adjacent wonderful frames in N wonderful frames is 2 frames, the number of first images to be inserted between the second group of two adjacent wonderful frames in N wonderful frames is 2 frames, and the number of first images to be inserted between the third group of two adjacent wonderful frames in N wonderful frames is 0 frames, that is, no image is inserted between the third group of two adjacent wonderful frames in N wonderful frames. For another example, the number of first images to be inserted between the first two adjacent wonderful frames in the N-frame wonderful frames is 3 frames, the number of first images to be inserted between the second two adjacent wonderful frames in the N-frame wonderful frames is 2 frames, and the number of first images to be inserted between the third two adjacent wonderful frames in the N-frame wonderful frames is 2 frames, that is, no image is inserted between the third two adjacent wonderful frames in the N-frame wonderful frames.
[0162] It can be understood that the playback time interval between every two frames of images in the generated video can be the same. In this way, if the number of images in the generated video is large, the video playback time can be longer; if the number of images in the generated video is small, the video playback time can be shorter. Therefore, the mobile phone 100 can control the playback time of the target object's action between every two adjacent wonderful frames by adjusting the number of images inserted between each two adjacent wonderful frames.
[0163] Specifically, the mobile phone 100 may insert the same number of images between every two adjacent wonderful frames by default. For example, taking N as 5, Fig.14 FIG. 4 shows another schematic diagram of image interpolation. Fig.14 As shown, wonderful frame 1, wonderful frame 2, wonderful frame 3, wonderful frame 4 and wonderful frame 5 are the determined 5 wonderful frames.
[0164] The playing time between wonderful frame 1 and wonderful frame 2 is T, the playing time between wonderful frame 2 and wonderful frame 3 is T, the playing time between wonderful frame 3 and wonderful frame 4 is T, and the playing time between wonderful frame 4 and wonderful frame 5 is also T.
[0165] Of course, in order to improve the user experience, the playing time between two adjacent wonderful frames can also be adjusted under the operation of the user, that is, by adjusting the insertion of different numbers of images between two adjacent wonderful frames, the purpose of achieving different playing time between two adjacent wonderful frames after adjustment can be achieved. Fig.14As shown, the playing time between wonderful frame 1 and wonderful frame 2 is T12, the playing time between wonderful frame 2 and wonderful frame 3 is T23, the playing time between wonderful frame 3 and wonderful frame 4 is T34, and the playing time between wonderful frame 4 and wonderful frame 5 is also T45. T12, T23, T34 and T45 can all be different.
[0166] The mobile phone 100 can determine the first image to be inserted between every two adjacent wonderful frames in N wonderful frames based on motion information by relevant technical means. An implementation scheme is briefly introduced below as an example.
[0167] Two adjacent wonderful frames include the ith wonderful frame and the i+1th wonderful frame, which can also be referred to as a group of two adjacent wonderful frames. M frames of images are inserted between the ith wonderful frame and the i+1th wonderful frame; wherein i takes N frame values from integers between 1 and N-1, and M takes an integer greater than or equal to 0.
[0168] Exemplarily, N is 5, and the 5 wonderful frames include the 1st wonderful frame, the 2nd wonderful frame, the 3rd wonderful frame, the 4th wonderful frame and the 5th wonderful frame. The 1st wonderful frame, the 2nd wonderful frame, the 3rd wonderful frame, the 4th wonderful frame and the 5th wonderful frame are arranged in sequence according to the sequence of the shooting time of the electronic device, or the 1st wonderful frame, the 2nd wonderful frame, the 3rd wonderful frame, the 4th wonderful frame and the 5th wonderful frame are arranged in sequence in the reverse sequence of the shooting time of the electronic device. Then the 1st wonderful frame and the 2nd wonderful frame are two adjacent wonderful frames, the 2nd wonderful frame and the 3rd wonderful frame are two adjacent wonderful frames, the 3rd wonderful frame and the 4th wonderful frame are two adjacent wonderful frames, and the 4th wonderful frame and the 5th wonderful frame are two adjacent wonderful frames.
[0169] For example, still taking N as 5, if M is 2, the mobile phone 100 can insert 2 frames of images between the 1st wonderful frame and the 2nd wonderful frame, insert 2 frames of images between the 2nd wonderful frame and the 3rd wonderful frame, insert 2 frames of images between the 3rd wonderful frame and the 4th wonderful frame, and insert 2 frames of images between the 4th wonderful frame and the 5th wonderful frame.
[0170] Of course, in addition to the above-mentioned method of inserting the same number of images between the i-th wonderful frame and the (i+1)-th wonderful frame, in some other embodiments, it is also possible that the number of images (i.e., first images) inserted between not all groups of adjacent two wonderful frames in N frames is the same, but the number of images inserted between two or more groups of adjacent two wonderful frames in N frames is the same.
[0171] For example, still taking N as 5, M can be 2, 2, 0 and 1. Then the mobile phone 100 can insert 2 frames of images between the 1st wonderful frame and the 2nd wonderful frame, insert 2 frames of images between the 2nd wonderful frame and the 3rd wonderful frame, insert 0 frames of images between the 3rd wonderful frame and the 4th wonderful frame, and insert 1 frame of image between the 4th wonderful frame and the 5th wonderful frame.
[0172] For example, still taking N as 5, M can be 2, 2, 2 and 1. Then the mobile phone 100 can insert 2 frames of images between the 1st wonderful frame and the 2nd wonderful frame, insert 2 frames of images between the 2nd wonderful frame and the 3rd wonderful frame, insert 2 frames of images between the 3rd wonderful frame and the 4th wonderful frame, and insert 1 frame of image between the 4th wonderful frame and the 5th wonderful frame.
[0173] In some other embodiments, the number of images inserted between every two adjacent highlight frames in the N highlight frames is different.
[0174] For example, still taking N as 5, M can be 2, 0, 1 and 3. Then the mobile phone 100 can insert 2 frames of images between the 1st wonderful frame and the 2nd wonderful frame, insert 0 frames of images between the 2nd wonderful frame and the 3rd wonderful frame, insert 1 frame of images between the 3rd wonderful frame and the 4th wonderful frame, and insert 3 frames of images between the 4th wonderful frame and the 5th wonderful frame.
[0175] In some other embodiments, the numbers of images inserted between two adjacent highlight frames in not all groups of N highlight frames are different, but the numbers of images inserted between two or more adjacent highlight frames in N highlight frames are different.
[0176] For example, still taking N as 5, M can be 2, 1, 0 and 0. Then the mobile phone 100 can insert 2 frames of images between the 1st wonderful frame and the 2nd wonderful frame, insert 1 frame of image between the 2nd wonderful frame and the 3rd wonderful frame, insert 0 frames of image between the 3rd wonderful frame and the 4th wonderful frame, and insert 0 frames of image between the 4th wonderful frame and the 5th wonderful frame.
[0177] For the nth frame image in M frames, where n is an integer from 1 to H, Fig.15 A schematic diagram of a process for generating an insertion image is shown.
[0178] S1501. The mobile phone 100 determines first motion difference information based on the i-th frame motion information corresponding to the shooting moment of the i-th wonderful frame and the n-th frame motion information corresponding to the n-th frame image at the shooting moment of the event camera.
[0179] The first motion difference information is used to indicate the change between the motion information of the i-th frame and the motion information of the n-th frame.
[0180] For example, Fig.10 As shown, taking i as 1 as an example, if the preview information corresponding to the first to second shots of the first camera are all identified as wonderful frames, the wonderful frame corresponding to the first shot of the first camera is the first wonderful frame. The wonderful frame corresponding to the second shot of the first camera is the second wonderful frame.
[0181] Taking i as 2 as an example, if the preview information corresponding to the first to third shots of the first camera are all identified as wonderful frames, the wonderful frame corresponding to the second shot of the first camera is the second wonderful frame, and the wonderful frame corresponding to the third shot of the first camera is the third wonderful frame.
[0182] Please continue with parameters Fig.10 , taking i as 1, M as 2, and n as 1 as an example, the mobile phone 100 can use the first frame motion information corresponding to the shooting moment of the first shooting of the event camera and the first frame motion information corresponding to the shooting moment of the 22nd shooting of the event camera to determine the motion difference information between the two frames. Similarly, n is 2 and so on.
[0183] The motion information is determined based on the event camera information. For example, the first frame of motion information is determined based on the first frame of event camera information corresponding to the first shooting moment of the event camera.
[0184] Fig.16 FIG. 2 shows a schematic diagram of a process for generating an inserted image. Fig.16 As shown, if at the shooting time t 0 Time to t 1 If you insert an image at any time t between the times, you can 0 The image to be played corresponding to the time t 1 The image to be played corresponding to the time t 0 Event camera stream information from time to time t and from time t to t 1 Event camera stream information at the moment Get the image I inserted at time t t .
[0185] Specifically, the mobile phone 100 can 0 Event camera stream information from time to time t and from time t to t 1 Event camera stream information at the moment Input to the optical flow estimation network to generate t 0 Movement information from time to time t and from time t to t1 Sports information at all times
[0186] Mobile phone 100 can be t 0 The difference between the motion information at time t and time t is calculated to obtain the motion difference information
[0187] S1502. The mobile phone 100 determines second motion difference information based on the nth frame motion information corresponding to the insertion moment of the nth frame image and the i+1th frame motion information at the shooting moment of the event camera corresponding to the i+1th wonderful frame.
[0188] The second motion difference is used to indicate the change between the motion information of the first i-th frame and the motion information of the n-th frame.
[0189] For example, Fig.10 As shown, taking i as 1 as an example, if the preview information corresponding to the first to third shots of the first camera are all identified as wonderful frames, the wonderful frame corresponding to the first shot of the first camera is the first wonderful frame. The wonderful frame corresponding to the second shot of the first camera is the second wonderful frame.
[0190] Please continue with parameters Fig.10 , taking i as 1, M as 2, and n as 1 as an example, the mobile phone 100 can use the motion information of the second frame corresponding to the shooting moment of the second shooting of the event camera and the motion information of the first frame corresponding to the shooting moment of the 22nd shooting of the event camera to determine the motion difference information between the two frames. Similarly, n is 2 and so on.
[0191] The motion information is determined based on the event camera information. For example, the first frame of motion information is determined based on the first frame of event camera information corresponding to the first shooting moment of the event camera.
[0192] For example, Fig.16 As shown, the mobile phone 100 can 1 The difference between the motion information at time t and time t is calculated to obtain the motion difference information
[0193] S1503: The mobile phone 100 determines the nth frame image to be first merged based on the i-th wonderful frame and the first motion difference information.
[0194] For example, Fig.16 As shown, the mobile phone 100 can be based on the image and motion difference information Get the image at time t
[0195] S1504: The mobile phone 100 determines the second n-th frame image to be merged based on the (i+1)-th wonderful frame and the second motion difference information.
[0196] For example, Fig.16 As shown, the mobile phone 100 can be based on the image I t and motion difference information Get the image at time t
[0197] S1505. The mobile phone 100 generates an nth frame image based on the first nth frame image to be fused and the second nth frame image to be fused.
[0198] For example, Fig.16 As shown, the mobile phone 100 can be based on the image at time t and images The fusion is performed to obtain an image with the wonderful frame inserted.
[0199] S507 . The mobile phone 100 generates a target wonderful video based on the start frame, the end frame and the first image.
[0200] It can be understood that the target wonderful video is used to record fleeting wonderful moments, and the target wonderful video can facilitate users to watch slow-motion wonderful moments at a slow speed.
[0201] In some embodiments, in order to improve the continuity of the target object's environment when playing the target wonderful frame video, the mobile phone 100 can add several frames of preparation pictures before the target object moves before the start frame. Specifically, the mobile phone 100 can generate the target wonderful video based on the start frame, the end frame, the first image and the second image; wherein the second image is an image captured within a preset time period before the start frame is captured.
[0202] The mobile phone 100 generates a video together with the images captured within a preset time period before the start frame is captured, the start frame, the end frame, and the first image. While presenting the wonderful movement moments of the target object to the user, it can also present the behavior of the target object before the movement to the user, so that a video with richer picture content can be obtained, thereby improving the user experience.
[0203] In some embodiments, the playing time of each frame of the target wonderful video is the same.
[0204] The target wonderful video refers to a video frame with meaningful image content. Meaningful image content includes clear image content in the image, such as people, animals, etc. Preferably, clear image content also includes a person smiling with eyes open, or a person smiling, laughing, smiling with eyes closed, blinking, or pouting. Furthermore, meaningful image content also includes image content making preset specific actions. For example, a person jumping up, an athlete hurdling, a puppy running, etc. For images that do not include people and animals that can perform specific actions, meaningful image content also includes plants, buildings that are complete and upright, etc.
[0205] In some embodiments, if the photos taken by the mobile phone 100 are to be automatically generated into videos, the "geolocation" switch in the setting interface of the camera application of the mobile phone 100 can be turned on, and the "network connection" switch in the gallery application can be turned on. In this way, the mobile phone 100 can place the video data generated by the mobile phone 100 in the gallery application.
[0206] In some embodiments, the generated target wonderful video can be placed in a gallery. For example, Fig.17 A schematic diagram showing the interface changes in a gallery is shown. Fig.17 As shown in Figure (a), this figure is an interface 405 corresponding to the moment option in the gallery, and the interface 405 includes a thumbnail area 4051 of the target wonderful video. It can be understood that there can be multiple target wonderful videos and corresponding thumbnail areas of multiple target wonderful videos. In order to briefly introduce the solution, only a thumbnail area of one target wonderful video is shown in the interface 405.
[0207] The mobile phone 100 may respond to the user's operation on the thumbnail area 4051 of the target wonderful video. For example, the mobile phone 100 may respond to the user's click operation on the thumbnail area 4051 of the target wonderful video by displaying the following information: Fig.17 The interface 406 shown in FIG. (b) in FIG. 406 includes a thumbnail window 4061 for dynamically playing the target wonderful video. The thumbnail window 4061 for dynamically playing the target wonderful video includes a video play button 4062. In response to the user clicking the video play button 4062, the mobile phone 100 plays the target wonderful video in a larger window. Exemplarily, the mobile phone 100 can play the target wonderful video in a full-screen manner.
[0208] Fig.18 FIG. 2 shows a schematic diagram of a video generation method. Fig.18 As shown, the mobile phone 100 is ready to take a photo, aims at the target scene, and displays the preview stream screen. Then, the mobile phone 100 performs wonderful frame recognition on the preview stream information, and automatically recognizes a series of continuous wonderful frames, which are wonderful frame segments. Then, the mobile phone 100 performs an automatic interpolation operation: use two adjacent wonderful frames to synthesize dense N images. Compress all interpolated images into a short video with continuous action. Then, the mobile phone can synthesize a video based on the aforementioned images: compress all interpolated images into a short video with continuous action.
[0209] It is understood that the mobile phone 100 can generally default the arrangement order of the wonderful frames. For example, the wonderful frames are arranged in the order of shooting by default. And the mobile phone 100 can generally default the interval between two adjacent wonderful frames. Of course, in order to facilitate the user's flexible operation and generate a more satisfactory wonderful video for the user, the mobile phone 100 can rearrange the order of the wonderful frames and readjust the interval between two adjacent wonderful frames under the operation of the user.
[0210] In the embodiment of the present application, when the target object and the camera have relative motion, the first camera cannot capture the slow motion details of the target object due to factors such as exposure time, while the data obtained by the event camera can obtain the slow motion details of the target object. The video generated by the video generation method provided by the embodiment of the present application has high image quality, less pseudo-texture, better continuity, and higher fluency. And the mobile phone 100 can play each frame of the image according to the arranged images, which is convenient for users to watch the slow motion wonderful moments slowly, and the user's viewing experience is better.
[0211] The embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned mobile terminal, the mobile terminal executes each function or step executed by the mobile phone 100 in the above-mentioned method embodiment.
[0212] The present application also provides a computer program product, which, when executed on a computer, enables the computer to execute the functions or steps executed by the mobile phone 100 in the above method embodiment. The computer may be the above mobile terminal (such as the mobile phone 100).
[0213] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.
[0214] Program code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0215] Program code can be implemented with high-level programming language or object-oriented programming language to communicate with the processing system. When necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.
[0216] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable storage media. Therefore, a machine-readable storage medium may include any mechanism for storing or disseminating information in a machine (e.g., computer) readable form, including but not limited to, floppy disks, optical disks, optical disks, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROM), random access memories (RAM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), magnetic cards or optical cards, flash memory, or a tangible machine-readable memory for disseminating information (e.g., carrier waves, infrared signal digital signals, etc.) based on the Internet in electrical, optical, acoustic, or other forms of propagation signals. Accordingly, machine-readable storage media include any type of machine-readable storage media suitable for storing or propagating electronic instructions or information in a form readable by a machine (eg, a computer).
[0217] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not mean that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0218] It can be understood that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.
[0219] It is understood that in the examples and descriptions of this patent, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises", or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. In the absence of further restrictions, an element defined by the statement "comprises a" does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.
[0220] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it will be apparent to those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present application.
Claims
1. A video generation method, It is characterized in that The method is applied to an electronic device, the electronic device comprising an event camera and a first camera, and the method comprises: In response to the user's operation of turning on the camera function, the first camera is turned on and a first interface is displayed; wherein the first interface includes a preview window, the preview window is used to display a preview stream image, and the preview stream image is an image captured by the first camera in real time; When the automatic snapshot mode in the camera function is turned on, the event camera is turned on; wherein the event camera is used to capture and obtain event camera stream information; If it is detected based on the preview stream information corresponding to the preview stream screen that the target object satisfies the preset motion posture, N frames of wonderful frames of the generated video are determined based on the preview stream information; wherein the N frames of wonderful frames include the start frame and the end frame of the video to be generated, and N is an integer greater than 1; Acquire event camera stream information between the shooting time of the start frame and the shooting time of the end frame, and determine a first image inserted between at least one group of adjacent two wonderful frames in the N wonderful frames based on the event camera stream information; wherein the first image includes slow motion details of the target object between at least one group of adjacent two wonderful frames in the N wonderful frames; A target wonderful video is generated based on the start frame, the end frame and the first image.
2. The method according to claim 1, It is characterized in that The number of first images inserted between at least two adjacent groups of wonderful frames in the N wonderful frames is the same.
3. The method according to claim 1, It is characterized in that The numbers of first images inserted between at least two adjacent groups of wonderful frames in the N wonderful frames are different.
4. The method according to claim 3, It is characterized in that Before determining the first image to be inserted between at least one group of adjacent two wonderful frames in the N wonderful frames based on the event camera stream information, the method includes: Based on the motion speed of the target object between two adjacent wonderful frames in the N wonderful frames, the number of first images inserted between at least one group of two adjacent wonderful frames in the N wonderful frames is determined.
5. The method according to any one of claims 1 to 4, It is characterized in that The preview stream image is of a first resolution, and the N highlight frames of the generated video are of a second resolution; wherein the second resolution is greater than the first resolution.
6. The method according to any one of claims 1 to 5, It is characterized in that The determining, based on the event camera stream information, a first image inserted between at least one group of adjacent two wonderful frames in the N wonderful frames comprises: Motion information is determined based on the event camera stream information, and a first image inserted between at least one group of adjacent two wonderful frames in the N wonderful frames is determined based on the motion information; wherein the motion information is used to represent the motion of a moving object photographed by the event camera, and the moving object includes a target object.
7. The method according to claim 6, It is characterized in that The determining motion information based on the event camera stream information comprises: De-noising the event camera stream information using a preset de-noising algorithm to obtain first event camera stream information; Motion information is determined based on the first event camera stream information.
8. The method according to claim 6, It is characterized in that The step of determining the first image to be inserted between two adjacent wonderful frames in the N wonderful frames based on the motion information comprises: The two adjacent wonderful frames include the i-th wonderful frame and the i+1-th wonderful frame, and M frames of images are inserted between the i-th wonderful frame and the i+1-th wonderful frame; wherein i is an integer between 1 and N-1; For the nth frame image in the M frames, where n is an integer from 1 to H: Determine first motion difference information based on the i-th frame motion information corresponding to the shooting moment of the i-th wonderful frame and the n-th frame motion information corresponding to the n-th frame image at the shooting moment of the event camera; wherein the first motion difference information represents the change between the first i-frame motion information and the n-th frame motion information; Determine second motion difference information based on the nth frame motion information of the event camera corresponding to the nth frame image at the shooting moment and the i+1th frame motion information of the event camera corresponding to the i+1th frame image at the shooting moment; wherein the second motion difference represents the change between the first i-frame motion information and the n-frame motion information; Determine the first n-th frame image to be fused based on the i-th wonderful frame and the first motion difference information; Determine the nth frame image to be merged based on the i+1th wonderful frame and the second motion difference information; An n-th frame image is generated based on the first n-th frame image to be fused and the second n-th frame image to be fused.
9. The method according to claim 1, It is characterized in that The step of generating a target wonderful video based on the start frame, the end frame and the first image comprises: A target wonderful video is generated based on the start frame, the end frame, the first image and the second image; wherein the second image is an image captured within a preset time period before the capture moment of the start frame.
10. The method according to claim 1, It is characterized in that If the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream picture, the method further comprises: detecting whether the target object satisfies the preset motion posture based on the preview stream information corresponding to the preview stream picture and the event camera stream information; Detecting that the target object satisfies a preset motion posture based on the preview stream information corresponding to the preview stream picture includes: Based on the preview stream information corresponding to the preview stream picture and the event camera stream information, it is detected that the target object satisfies a preset motion posture.
11. An electronic device, It is characterized in that The electronic device comprises a processor and a memory; the memory is used to store code instructions; the processor is used to run the code instructions, so that the electronic device executes the method as described in any one of claims 1-10.
12. A computer-readable storage medium, It is characterized in that The method comprises computer instructions, and when the computer instructions are executed on an electronic device, the electronic device executes the method as claimed in any one of claims 1 to 10.
Citation Information
Cited By
Video data processing method and system and computing equipment
CN121239897A