Photographing processing method and related device
By capturing image streams in motion scenes and using a scoring model to select reference image frames, high-quality and high-precision action photos and videos are generated, solving the problems of low photo accuracy and poor video quality in existing technologies, and improving shooting efficiency and storage utilization.
Patent Information
- Application Number
- CN202410465699.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-17
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-04-17
AI Technical Summary
In shooting action scenes, photos of exciting actions taken by taking pictures have low accuracy, while videos of exciting moments taken by shooting videos have poor image quality and take up storage space on electronic devices.
By capturing image streams through a camera, generating image sequences, and selecting reference image frames based on a scoring model, high-quality, high-precision photos and videos of exciting action scenes are produced. The videos are short in length, reducing storage requirements.
It enables the generation of high-precision photos and videos of exciting action, reduces storage space usage, and improves the efficiency of image and video generation.
Smart Images

Figure CN119255094B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart terminal technology, and in particular to a shooting and processing method and related equipment. Background Technology
[0002] With the development of terminal technology, smartphones, tablets, and other smart terminal devices all have shooting functions, allowing users to record various scenes by taking photos or videos. When shooting sports scenes, users often need to capture exciting movements during the action. However, exciting movements are fleeting, and there is a delay between when a user sees the exciting movement and when they take a picture. Therefore, photos of exciting movements obtained through photography cannot meet users' needs, while videos of exciting movements obtained through video recording have lower resolution and poorer image quality. Summary of the Invention
[0003] In view of the above, it is necessary to provide a shooting and processing method and related equipment to solve the problem that the photos of wonderful actions obtained by taking pictures cannot meet the needs of users, and the video quality of wonderful moments obtained by shooting videos is poor.
[0004] In a first aspect, this application provides a shooting processing method applied to an electronic device, the method comprising: in response to a user's operation of opening a camera application, acquiring an image stream through a camera; in response to the user's shooting operation, generating a first image sequence based on the image stream; acquiring multiple image frames from the first image sequence to generate a second image sequence; selecting a reference image frame from the second image sequence based on a score of each image frame in the second image sequence; acquiring multiple image frames from the first image sequence based on the reference image frame to generate a third image sequence; generating a target video based on the multiple image frames in the third image sequence; acquiring multiple image frames from the first image sequence based on the reference image frame to generate a fourth image sequence; and generating a target image based on the multiple image frames in the fourth image sequence.
[0005] Based on the above technical solution, during the process of shooting motion scenes in response to the user's shooting operation, high-quality and high-precision target photos, i.e., exciting action photos, can be automatically generated according to the image stream captured by the camera, as well as target videos, i.e., exciting action videos. The videos can record dynamic and exciting actions, and the video length is also relatively short, reducing the occupation of electronic device storage space.
[0006] In one possible implementation, the step of responding to a user's shooting operation and generating a first image sequence based on the image stream includes: obtaining image frames captured by the camera within a first preset time period before the user performs the shooting operation and image frames captured by the camera within a second preset time period after the user performs the shooting operation from the image stream, and generating the first image sequence.
[0007] Based on the above technical solution, by acquiring multiple image frames as candidate images for generating highlight videos and action images, the accuracy of the generated highlight videos and action images can be improved.
[0008] In one possible implementation, obtaining multiple image frames from the first image sequence to generate a second image sequence includes: obtaining one image frame from every first preset number of consecutive image frames in the first image sequence to generate the second image sequence.
[0009] Based on the above technical solution, a small number of image frames are selected from multiple image frames by frame extraction for subsequent image optimization, thereby improving the efficiency of image scoring and optimization.
[0010] In one possible implementation, selecting a reference image frame from the second image sequence based on the score of each image frame in the second image sequence includes: inputting each image frame in the second image sequence into at least one scoring model to obtain at least one score for each image frame; and selecting the reference image frame from the second image sequence based on the at least one score for each image frame.
[0011] Based on the above technical solution, a reference image frame with a higher score can be accurately selected from multiple image frames using at least one scoring model. Based on the reference image frame with a higher score, high-precision highlight videos and action images can be generated, which can also improve the generation efficiency of highlight videos and action images.
[0012] In one possible implementation, the at least one scoring model includes an action scoring model, an image sharpness scoring model, and an expression scoring model.
[0013] Based on the above technical solution, multiple image frames can be scored from multiple dimensions such as action, image quality, and expression, thereby accurately selecting reference image frames.
[0014] In one possible implementation, the method further includes: identifying a target object in a plurality of image frames of the first image sequence; and calculating the motion amplitude of the target object based on the plurality of image frames in the first image sequence.
[0015] Based on the above technical solution, motion detection can be performed on multiple image frames in an image sequence, which facilitates the subsequent determination of the number of images and frame rate during video interpolation, so that the generated highlight video can reflect the motion amplitude of the target object.
[0016] In one possible implementation, identifying the target object in multiple image frames of the first image sequence includes: inputting multiple image frames of the first image sequence into a target detection model, and outputting the target object in multiple image frames of the first image sequence through the target detection model.
[0017] Based on the above technical solution, the target object in the image frame can be accurately identified through the target detection model, thereby performing motion detection on the target object and determining the motion amplitude of the target object.
[0018] In one possible implementation, calculating the motion amplitude of the target object based on multiple image frames in the first image sequence includes: calculating the motion amplitude of the target object using optical flow based on every two adjacent image frames in the first image sequence.
[0019] Based on the above technical solution, the optical flow method can accurately perform motion detection on multiple image frames in an image sequence.
[0020] In one possible implementation, the step of obtaining multiple image frames from the first image sequence based on the reference image frame to generate a third image sequence includes: if the motion amplitude of the target object is greater than or equal to a preset motion amplitude threshold, obtaining multiple image frames within a third preset time period from the first image sequence based on the reference image frame to generate the third image sequence; or if the motion amplitude of the target object is less than the preset motion amplitude threshold, obtaining multiple image frames within a fourth preset time period from the first image sequence based on the reference image frame to generate the third image sequence, wherein the third preset time period is less than the fourth preset time period.
[0021] Based on the above technical solution, different numbers of image frames can be selected according to the motion amplitude of the target object to generate a video of the highlight moment, thereby improving the generation efficiency of the highlight moment video and enabling the highlight moment video to reflect the motion amplitude of the target object.
[0022] In one possible implementation, the step of obtaining multiple image frames within a third preset time period from the first image sequence based on the reference image frame and generating the third image sequence includes: obtaining multiple image frames within a fifth preset time period before the reference image frame from the first image sequence, and obtaining multiple image frames within a sixth preset time period after the reference image frame, to generate the third image sequence, wherein the sum of the fifth preset time period and the sixth preset time period is the third preset time period.
[0023] Based on the above technical solution, multiple image frames can be selected from before and after the reference image frame to generate a highlight video, making the generated highlight video more accurate.
[0024] In one possible implementation, generating a target video based on multiple image frames in the third image sequence includes: if the motion amplitude of the target object is greater than or equal to the preset motion amplitude threshold, interpolating multiple image frames in the third image sequence according to a first frame rate to generate the target video of a ninth preset duration; or if the motion amplitude of the target object is less than the preset motion amplitude threshold, interpolating multiple image frames in the third image sequence according to a second frame rate to generate the target video of a ninth preset duration, wherein the first frame rate is greater than the second frame rate.
[0025] Based on the above technical solution, video interpolation can be performed at different frame rates according to the motion amplitude of the target object, so that the video of the exciting moment can reflect the motion amplitude of the target object.
[0026] In one possible implementation, the step of obtaining multiple image frames within a fourth preset time period from the first image sequence based on the reference image frame and generating the third image sequence includes: obtaining multiple image frames within a seventh preset time period before the reference image frame from the first image sequence, and obtaining multiple image frames within an eighth preset time period after the reference image frame, to generate the third image sequence, wherein the sum of the seventh preset time period and the eighth preset time period is the fourth preset time period.
[0027] Based on the above technical solution, multiple image frames can be selected from before and after the reference image frame to generate a highlight video, making the generated highlight video more accurate.
[0028] In one possible implementation, the method further includes: storing the target video to a gallery application of the electronic device; optimizing the reference image frame; and using the optimized reference image frame as the cover image of the target video.
[0029] Based on the above technical solution, by setting a cover image for the target video, users can easily view the highlights of the video.
[0030] In one possible implementation, the step of obtaining multiple image frames from the first image sequence based on the reference image frame and generating a fourth image sequence includes: obtaining the reference image frame and a second preset number of image frames preceding the reference image frame from the first image sequence, and generating the fourth image sequence.
[0031] Based on the above technical solution, multiple image frames can be selected as candidate images for generating exciting action images, thereby improving the accuracy of the generated exciting action images.
[0032] In one possible implementation, generating a target image based on multiple image frames in the fourth image sequence includes: fusing the multiple image frames in the fourth image sequence to obtain a fused image; and optimizing the fused image to obtain the target image.
[0033] Based on the above technical solution, multiple image frames can be fused and optimized to generate high-quality, exciting action images.
[0034] In one possible implementation, the image stream includes a preview image stream and a captured image stream, wherein the resolution of the image frames in the captured image stream is greater than the resolution of the image frames in the preview image stream.
[0035] Based on the above technical solution, motion detection and video generation can be performed on the preview image stream, improving the efficiency of generating videos of exciting moments and generating high-quality images of exciting action based on the captured image stream.
[0036] Secondly, this application provides an electronic device, which includes a memory and a processor: wherein the memory is used to store program instructions; and the processor is used to read and execute the program instructions stored in the memory, and when the program instructions are executed by the processor, the electronic device performs the above-described shooting processing method.
[0037] Thirdly, this application provides a chip coupled to a memory in an electronic device, the chip being used to control the processor of the electronic device to execute the above-described image processing method.
[0038] Fourthly, this application provides a computer storage medium storing program instructions that, when executed on an electronic device, cause the processor of the electronic device to perform the aforementioned image processing method.
[0039] Furthermore, the technical effects brought about by the second to fourth aspects can be found in the descriptions of the methods in the above-mentioned method section, and will not be repeated here. Attached Figure Description
[0040] Figure 1 This is a schematic diagram of the interface for recording and taking pictures on an electronic device.
[0041] Figure 2 This is a schematic diagram of the interface for an electronic device to record multiple data points at once.
[0042] Figure 3 This is a schematic diagram of the interface for electronic devices to perform eagle-eye capture.
[0043] Figure 4 This is a schematic diagram of the interface of an electronic device that can simultaneously capture photos and animated GIFs.
[0044] Figure 5 This is a software architecture diagram of an electronic device provided in an embodiment of this application.
[0045] Figure 6 This is a flowchart of a shooting and processing method provided in an embodiment of this application.
[0046] Figure 7 This is a software architecture diagram of a camera service for an electronic device provided in an embodiment of this application.
[0047] Figure 8 This is a schematic diagram of a preview stream and a photo stream provided in an embodiment of this application.
[0048] Figure 9 This is a schematic diagram of a camera application interface provided in an embodiment of this application.
[0049] Figure 10 This is a schematic diagram of generating a third image sequence according to an embodiment of this application.
[0050] Figure 11 This is a schematic diagram of the interface of a gallery application provided in an embodiment of this application.
[0051] Figure 12 This is a flowchart of selecting a reference image frame from a third image sequence according to an embodiment of this application.
[0052] Figure 13 This is a flowchart of generating a fourth image sequence provided in one embodiment of this application.
[0053] Figure 14 This is a schematic diagram of generating a fourth image sequence according to an embodiment of this application.
[0054] Figure 15 This is a flowchart of generating a target video provided in one embodiment of this application.
[0055] Figure 16 This is a schematic diagram of multiple image frames in a fourth image sequence provided in an embodiment of this application.
[0056] Figure 17 This is a schematic diagram of multiple image frames after interpolation processing provided in an embodiment of this application.
[0057] Figure 18 This is a schematic diagram of generating a fifth image sequence according to an embodiment of this application.
[0058] Figure 19 This is a flowchart of generating a target image provided in an embodiment of this application.
[0059] Figure 20 This is a hardware architecture diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0060] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used in this application's specification is for the purpose of describing particular embodiments only and is not intended to limit the application. It should be understood that, unless otherwise stated, " / " in this application means "or". For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. "At least one" refers to one or more. "More than one" refers to two or more. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, and a, b, and c. Where there is no conflict, the following embodiments and features described herein can be combined with each other.
[0062] With the development of terminal technology, smartphones, tablets, and other smart terminal devices all have shooting functions, allowing users to record various scenes by taking photos or videos. When shooting sports scenes, users often need to capture exciting movements during the action. However, exciting movements are fleeting, and there is a delay between when a user sees the exciting movement and when they take a picture. Therefore, photos of exciting movements obtained through photography cannot meet user needs, while videos of exciting movements obtained through video recording have low resolution and poor image quality. Therefore, how to accurately capture and record exciting movements during sports scenes using terminal devices is a problem that urgently needs to be solved in the industry.
[0063] See Figure 1 The image shows a schematic diagram of an electronic device's interface for taking photos while recording. Electronic devices can provide a photo-taking function during video recording. For example, the device can provide a photo-taking control on the recording interface. When recording a moving scene, the user can record the movement and, upon noticing an exciting action, manually trigger the photo-taking control to capture a photo of that action, thus simultaneously recording both the video and the photo. However, videos recorded in this way are often long, consuming significant storage space, and have low resolution and poor image quality. Furthermore, since photos of exciting actions are taken manually by the user, there is a delay between when the user sees the exciting action and when they take the photo, resulting in a delayed image that may not capture the exact moment of the action.
[0064] See Figure 2 The image shows a schematic diagram of an electronic device's interface for recording multiple actions simultaneously. Electronic devices can also provide this functionality; for example, when filming a moving scene, the user can record the movement using the device, and after recording, the device can automatically generate photos of the exciting actions based on the captured video frames. However, videos recorded in this way are often long, consuming significant storage space, and have low resolution and poor image quality. Furthermore, photos of exciting actions generated from video frames also have low resolution and poor image quality.
[0065] See Figure 3The image shown is a schematic diagram of the interface for eagle-eye capture on an electronic device. The electronic device can also provide eagle-eye capture functionality. For example, when capturing a moving scene, the device can identify specific exciting actions in the scene based on preview image frames and automatically capture photos of those actions when they are detected. However, this method can only record photos of exciting actions, not videos, and it often only recognizes specific exciting actions, resulting in a limited range of exciting action types that can be captured.
[0066] See Figure 4 The image shown is a schematic diagram of an electronic device simultaneously capturing photos and animated GIFs. The electronic device can also provide the function of recording animated GIFs while taking photos. For example, when filming a moving scene, the user can manually capture a photo of an exciting action using the electronic device, which simultaneously generates an animated GIF of the action based on the preview image frame. However, animated GIFs recorded in this way have low resolution and poor image quality. Furthermore, since the photo of the exciting action is manually captured by the user, there is a delay between the user seeing the exciting action and executing the shooting operation, resulting in a delayed photo that may not accurately represent the exciting action. Consequently, the animation in the generated GIF may have poor motion continuity.
[0067] To address the issues of low accuracy in photos of exciting actions captured by photography, which fails to meet user needs, and poor video quality in capturing exciting moments, this application provides a shooting processing method. During the process of capturing a moving scene in response to a user's shooting operation, this method can automatically generate high-quality and high-precision photos of exciting actions based on the image stream captured by the camera, as well as automatically generate videos of exciting actions. The videos can record dynamic and exciting actions, and their shorter duration reduces the storage space occupied by electronic devices.
[0068] See Figure 5 The diagram shown is a software architecture diagram of an electronic device provided in an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. For example, the Android system is divided into four layers, from top to bottom: application layer 101, framework layer 102, Android runtime and system library 103, hardware abstraction layer 104, kernel layer 105, and hardware layer 106.
[0069] Application layer 101 may include a series of application packages. For example, application packages may include applications such as camera, gallery, calendar, calling, map, navigation, WLAN, Bluetooth, music, video, SMS, device control services, etc.
[0070] The framework layer 102 provides an Application Programming Interface (API) and programming framework for applications in the application layer. The application framework layer includes predefined functions. For example, it may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0071] The window manager manages window programs. It can obtain screen size, determine the presence of a status bar, lock the screen, and capture screenshots. The content provider stores and retrieves data, making it accessible to applications. This data can include videos, images, audio, made and received calls, browsing history and bookmarks, phone books, etc. The view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon can include views for displaying text and views for displaying images. The phone manager provides communication functionality for electronic devices, such as managing call status (including connection and disconnection). The resource manager provides applications with various resources, such as localized strings, icons, images, layout files, and video files. The notification manager allows applications to display notifications in the status bar, conveying informational messages that disappear automatically after a short pause without user interaction. For example, the notification manager is used to notify of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the system's top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting alert sounds, causing electronic devices to vibrate, and flashing indicator lights.
[0072] The Android Runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for the scheduling and management of the Android system. The core libraries consist of two parts: one part contains the functionalities that the Java language needs to call, and the other part contains the core Android libraries.
[0073] Application layer 101 and framework layer 102 run in a virtual machine. The virtual machine executes the Java files of the application layer and framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0074] System library 103 may include multiple functional modules. For example, a surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0075] The Surface Manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The Media Library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. The 3D Graphics Processing Library implements 3D graphics drawing, image rendering, compositing, and layer processing. The 2D Graphics Engine is the drawing engine for 2D graphics.
[0076] Hardware Abstraction Layer 104 runs in user space, encapsulates kernel-level drivers, and provides calling interfaces to the upper layers.
[0077] Kernel layer 105 is the layer between hardware and software. Kernel layer 105 contains at least the display driver, camera driver, audio driver, and sensor driver.
[0078] Kernel layer 105 is the core of the operating system for electronic devices. It is the first layer of software extension based on the hardware, providing the most basic functions of the operating system. It is the foundation for the operation of the operating system, responsible for managing system processes, memory, device drivers, files, and network systems, and determining the system's performance and stability. For example, the kernel can determine the timing of an application's operation on a certain part of the hardware.
[0079] Kernel layer 105 includes hardware-dependent programs such as interrupt handlers and device drivers, as well as basic, common, and frequently running modules such as clock management and process scheduling modules, and critical data structures. The kernel layer can be located within the processor or embedded in internal memory.
[0080] Hardware layer 106 includes the hardware of electronic devices, such as displays, buttons, cameras, etc.
[0081] See Figure 6 The diagram shown is a flowchart of a shooting processing method provided in an embodiment of this application. The method is applied in an electronic device, and the shooting processing method includes:
[0082] S101 responds to the user's action of opening the camera application and captures an image stream through the camera.
[0083] In one embodiment of this application, a user can open the camera application by clicking the camera application icon on the main interface of the electronic device, or by clicking the camera activation control on the lock screen of the electronic device, or by touching or pressing a preset physical button. In response to the user's action of opening the camera application, the electronic device captures an image stream of the scene in front of it using the camera at a preset sampling frequency. For example, the sampling frequency can be 30Hz, 32Hz, or 60Hz.
[0084] See Figure 7 The diagram shown is a software architecture diagram of the camera service of an electronic device provided in this application embodiment. In response to a user's operation to open the camera application, the camera application creates a camera process, instantiates the camera service (CameraService), the cache queue (CameraBufferQueue), the cache reading service (BlastBufferQueue, BBQ), the compositing service (SurfaceFlinger), the camera provider (CameraProvider), the hardware compositor (HWC), the signal center (DispSync), etc., and creates layers of controls to be displayed within the camera process, such as... Figure 7 The layers include Control 1, Control 2, Control 3, Control 4, etc. Control 1, Control 2, Control 3, and Control 4 can be controls for selecting shooting modes, shooting, setting shooting parameters, launching the gallery application, or switching lenses.
[0085] In one embodiment of this application, HWC periodically generates VSync-HW signals according to the refresh rate of the display screen. DispSync allocates corresponding VSync-APP signals to the user interface (UI) thread of the camera application according to the periodically generated VSync-HW signals, and periodically sends the generated VSync-APP signals to the UI thread of the camera application, so that the UI thread triggers the drawing and rendering process after receiving the VSync-APP signals.
[0086] In one embodiment of this application, DispSync further allocates a corresponding VSync-SF signal to SurfaceFlinger based on the periodically generated VSync-HW signal, and periodically sends the generated VSync-SF signal to SurfaceFlinger. Upon receiving the VSync-SF signal, SurfaceFlinger retrieves image frames from CameraBufferQuene via BBQ, and then performs composite processing on the retrieved image frames. The image frames cached in CameraBufferQuene can be obtained by CameraProvider processing image frames from the raw image stream captured by the camera. The raw image stream can be continuously captured by the camera based on a preview shooting request or a shooting request generated by the camera application. The preview shooting request or shooting request can be generated by the camera application based on the current shooting mode after the camera application is launched.
[0087] In one embodiment of this application, after receiving a preview shooting request or shooting request transmitted by CameraService, CameraProvider transmits the preview shooting request or shooting request to the camera driver, which then sends the preview shooting request or shooting request to the camera and drives the camera to capture image frames according to a preset sampling frequency.
[0088] In one embodiment of this application, after the camera captures an image frame, it continuously reports the image frames in the original image stream to the camera driver. The camera driver then transmits the image frames to the CameraProvider, which processes them accordingly. The CameraProvider then transmits the processed image frames to the CameraService, which adds them to the CameraBufferQueue.
[0089] In one embodiment of this application, SurfaceFlinger performs composite processing on image frames cached in CameraBufferQuene based on a VSync-SF signal periodically generated by DispSync. Upon receiving the VSync-SF signal, SurfaceFlinger retrieves the cached image frames from CameraBufferQuene via BBQ. BBQ also synchronizes the timestamp of each retrieved image frame added to CameraBufferQuene to SurfaceFlinger.
[0090] If SurfaceFlinger retrieves a cached image frame from CameraBufferQuene via BBQ, it performs compositing on that image frame to obtain a composite image, which is then transmitted to HWC. Simultaneously, SurfaceFlinger sends a command to BBQ to destroy the image frame. BBQ then transmits this destruction command to CameraBufferQuene, which instructs CameraService to call the dequeue function to remove the cached image frame from CameraBufferQuene.
[0091] The `dequeue` function is used to remove the first image frame from the specified queue for each matching element and execute the removed function, such as deleting the image frame at the head of the queue. When SurfaceFlinger retrieves cached image frames from CameraBufferQuene via BBQ, it can retrieve them from the head of the queue. Therefore, after compositing the image, calling the `dequeue` function as described above will remove the used image frames from CameraBufferQuene.
[0092] HWC transmits the composite image sent by SurfaceFlinger to the display driver, which then drives the display screen to show the corresponding image.
[0093] In one embodiment of this application, the image stream includes a preview stream and a capture stream. The preview stream includes multiple preview image frames, and the capture stream includes multiple capture image frames. The resolution of the capture image frames is greater than that of the preview image frames. For example, the resolution of the capture image frames is 8000*6000, and the resolution of the preview image frames is 4000*3000. After the camera application is started, the camera service (CameraService) in the framework layer and the camera provider (CameraProvider) in the hardware abstraction layer are initialized. After initialization, the camera application sends a preview capture request to the camera service, which in turn sends the preview capture request to the camera provider. The camera provider then sends the preview capture request to the camera driver in the kernel layer. The camera driver can drive the camera to capture image frames and return the captured multiple image frames to the camera provider. The camera provider returns the multiple image frames to the camera service, which then sends them to the rendering service in the framework layer. The rendering service displays the multiple image frames, generating the preview stream of the camera application. The camera service also stores the multiple image frames in a buffer, generating the capture stream of the camera application.
[0094] In another embodiment of this application, the image stream may include a photo capture stream, in which case the image frames cached and displayed by the camera application are all photo capture image frames from the photo capture stream.
[0095] S102, in response to the user's shooting operation, generates a first image sequence based on the image stream.
[0096] In one embodiment of this application, after the camera application is launched, the user can perform a shooting operation by clicking the shooting control on the camera application interface. The camera application responds to the user's shooting operation by generating and sending a shooting request to the camera service. The camera service then sends the shooting request to the camera provider, which in turn sends the shooting request to the camera driver. The camera driver drives the camera to capture image frames and returns the captured image frames to the camera provider. The camera provider returns the image frames to the camera service, which then sends them to the gallery application for storage. Then, the camera application continues to send preview shooting requests, the camera continues to capture image frames, and generates a preview stream and a photo stream for the camera application based on the captured image frames.
[0097] See Figure 8 The diagram illustrates a preview stream and a photo stream according to an embodiment of this application. In one embodiment, both the preview stream and the photo stream of the camera application can be stored in a cache, for example, in a round-robin cache. The round-robin cache has a preset capacity, which can be the number of image frames, for example, 60, 70, 80, or other values. A round-robin cache is a data structure used to store and manage data in a first-in, first-out (FIFO) order. A round-robin cache is typically a fixed-size circular queue composed of a series of fixed-size buffers. When data is written to the round-robin cache, it is written sequentially to the currently active buffer, progressing forward as the data is written. When the buffer is full, the oldest data is overwritten or deleted to make room for new data. Thus, the round-robin cache caches image frames according to the FIFO principle; for example, if the number of images cached in the round-robin cache exceeds the preset capacity, the oldest cached image frame is deleted.
[0098] In one embodiment of this application, a first image sequence can be generated by acquiring image frames captured in response to a shooting operation, a first preset number of image frames preceding the image frames captured in response to the shooting operation, and a second preset number of image frames following the image frames captured in response to the shooting operation from an image stream. For example, Figure 8Image frame t0 in the image is an image frame obtained in response to a shooting operation. The sum of a first preset quantity and a second preset quantity is the preset capacity of the round-robin buffer. For example, the preset capacity of the round-robin buffer is 60, the first preset quantity is 50, and the second preset quantity is 10. The image streams stored in the round-robin buffers form an image sequence; for example, a preview stream stored in one round-robin buffer forms a first image sequence, and a shooting stream stored in another round-robin buffer forms a second image sequence.
[0099] In another embodiment of this application, if the image stream includes a preview stream and a capture stream, in response to a user's capture operation, the time T at which the user performs the capture operation is determined, and based on the time T at which the user performs the capture operation, preview image frames captured by the camera within a first preset duration before time T, preview image frames corresponding to time T, and preview image frames captured by the camera within a second preset duration after time T are obtained from the preview stream to generate a first image sequence. For example, Figure 8 Image frame t0 in the preview stream is the preview image frame corresponding to time T. The sum of a first preset duration and a second preset duration can be set to a preset duration threshold. For example, the preset duration threshold is 4 / 3 seconds, the first preset duration is 1 second, and the second preset duration is 1 / 3 second. In other embodiments, the preset duration threshold can also be other durations; for example, the preset duration threshold can be determined based on the first preset duration, the second preset duration, and a fixed duration.
[0100] In another embodiment of this application, based on the time T when the user performs the shooting operation, image frames captured by the camera within a first preset time period before time T, image frames corresponding to time T, and image frames captured by the camera within a second preset time period after time T are obtained from the image stream to generate a second image sequence. For example, Figure 8 The image frame t0 in the image stream shown is the image frame captured at time T.
[0101] In another embodiment of this application, if the image stream includes a photo capture stream, in response to the user's shooting operation, the time T at which the user performs the shooting operation is obtained, and based on the time T at which the user performs the shooting operation, the photo capture image frames captured by the camera within a first preset time period before time T, the photo capture image frame corresponding to time T, and the photo capture image frames captured by the camera within a second preset time period after time T are obtained from the photo capture stream to generate a first image sequence.
[0102] In one embodiment of this application, when a user uses an electronic device to capture a moving scene, they can perform a shooting operation at the moment they believe the target in the moving scene is performing an exciting action. This causes the electronic device to respond to the shooting operation and generate a first image sequence corresponding to a preview stream and a second image sequence corresponding to a captured image stream, based on the image stream. See also... Figure 9The diagram shown is a schematic of a camera application interface provided in an embodiment of this application. For example, when a user uses an electronic device to film an athlete jumping, they can trigger the shooting control A when they see the athlete reach the highest point of the jump to perform a shooting operation. The electronic device responds to the shooting operation and generates a first image sequence and a second image sequence based on the image stream.
[0103] S103, acquire multiple image frames from the first image sequence to generate a third image sequence.
[0104] See Figure 10 The diagram illustrates the generation of a third image sequence according to an embodiment of this application. In one embodiment, a second image sequence is generated by acquiring one image frame from every third preset number of consecutive image frames in a first image sequence (e.g., the first image sequence corresponding to a preview stream). For example, the third preset number is 2, 3, 5, or other values. Specifically, the third image sequence is generated by acquiring one image frame from every third preset number of consecutive image frames, based on the earliest captured image frame in the first image sequence. For example, taking a third preset number of 3 as an example, one image frame is acquired from every 3 image frames in the first image sequence. Assuming a total of 30 image frames are acquired, the third image sequence is generated based on these 30 acquired image frames.
[0105] S104, Select a reference image frame from the third image sequence based on the score of each image frame in the third image sequence.
[0106] In one embodiment of this application, each image frame in the third image sequence is input into at least one preset scoring model to obtain at least one score for each image frame. Based on the at least one score for each image frame, a reference image frame is selected from the third image sequence. The at least one scoring model includes, but is not limited to, an action scoring model, an image sharpness scoring model, and an expression scoring model.
[0107] In one embodiment of this application, if there is at least one scoring model, each image frame in the third image sequence is input into the scoring model to obtain a score for each image frame, and the image frame with the highest score is selected from the third image sequence as a reference image frame. If there are multiple scoring models, each image frame in the third image sequence is input into multiple scoring models respectively to obtain a score for each image frame corresponding to each scoring model, thereby obtaining multiple scores for each image frame. The comprehensive score of each image frame is determined by summing or weighted summing the multiple scores, and the image frame with the highest comprehensive score is selected from the third image sequence as a reference image frame.
[0108] In one embodiment of this application, by selecting the image frame with the highest comprehensive score from the third image sequence as the reference image frame, and based on the reference image frame, videos of exciting moments in the captured motion scene and images of exciting actions can be accurately generated.
[0109] S105, obtain multiple image frames from the first image sequence based on the reference image frames, and generate a fourth image sequence.
[0110] In one embodiment of this application, a target object in multiple image frames of a first image sequence is identified, motion detection is performed on the target object in the multiple image frames, the motion amplitude of the target object is calculated based on the multiple image frames in the first image sequence, and multiple image frames are obtained from the first image sequence based on the motion amplitude of the target object and a reference image frame to generate a fourth image sequence.
[0111] S106, Generate a target video based on multiple image frames in the fourth image sequence.
[0112] In one embodiment of this application, a target video is generated based on the motion amplitude of the target object and multiple image frames in a fourth image sequence. The target video is a video of exciting actions in the shooting scene, and this video can be a slow-motion video or a time-lapse video.
[0113] S107, multiple image frames are obtained from the second image sequence based on the reference image frame to generate the fifth image sequence.
[0114] In one embodiment of this application, a reference image frame and a fourth preset number of image frames preceding the reference image frame are obtained from the second image sequence corresponding to the image stream to generate a fifth image sequence.
[0115] S108, Generate a target image based on multiple image frames in the fifth image sequence.
[0116] In one embodiment of this application, multiple image frames in the fifth image sequence are fused to obtain a fused image. A preset image-taking algorithm is then used to optimize the fused image to obtain a target image. The target image is an image of exciting action within the shooting scene.
[0117] In another embodiment of this application, multiple image frames in the fifth image sequence may be optimized first, and then the optimized images may be fused to obtain the target image.
[0118] The embodiments described above in this application continuously capture image streams through the camera when the user opens the camera application to shoot a moving scene, and automatically generate videos and images of exciting actions in the moving scene in response to the user's shooting operation. The images of exciting actions have high image quality and precision, and the videos can record dynamic and exciting actions. The video duration is also short, reducing the storage space occupied by electronic devices.
[0119] In one embodiment of this application, the method further includes: storing the target video to a gallery application of an electronic device; optimizing a reference image frame and using the optimized reference image frame as the cover image of the target video. The target video with the reference image frame as the cover image is displayed in the thumbnail interface provided by the gallery application. In one embodiment of this application, the optimization processing of the reference image frame includes, but is not limited to, noise reduction processing and tone mapping processing, wherein tone mapping processing refers to converting a standard dynamic range image into a high dynamic range image.
[0120] In one embodiment of this application, the method further includes: storing the target image to a gallery application of an electronic device. See also... Figure 11 The image shown is a schematic diagram of the interface of a gallery application provided in an embodiment of this application. When a user performs a shooting operation in the camera application, the electronic device responds to the shooting operation by simultaneously generating a video and an image of a highlight action in the shooting scene, and storing them in the gallery application. The user can view the video of the highlight action by clicking on the video cover, or view the image of the highlight action by clicking on the thumbnail of the image.
[0121] See Figure 12 The diagram shown is a flowchart of selecting a reference image frame from a third image sequence according to an embodiment of this application.
[0122] S1041, input each image frame in the third image sequence into the action scoring model, and output the action score of each image frame through the action scoring model.
[0123] In one embodiment of this application, the action scoring model can be a convolutional neural network model. The action scoring model is trained using multiple highlight action images from multiple sports scenarios as training data. These multiple sports scenarios include, but are not limited to, playing basketball, soccer, badminton, table tennis, running, and jumping. Highlight actions in basketball include, but are not limited to, shooting; highlight actions in soccer include, but are not limited to, shooting; highlight actions in badminton include, but are not limited to, spiking; highlight actions in table tennis include, but are not limited to, serving; and highlight actions in running include, but are not limited to, sprinting. The above sports scenarios and highlight actions within them are merely illustrative examples, and actual applications may not be limited to these examples. This application embodiment does not impose any limitations on these aspects. For example, the action scoring model can extract key point features from image frames and calculate the similarity between different actions and individual image frames, obtaining action scoring results based on these similarities.
[0124] In one embodiment of this application, each image frame in the third image sequence is input into the action scoring model. The action scoring model extracts features from the image frames. Based on the extracted features, it is analyzed whether the image frame contains exciting actions. The result of determining whether the image frame contains exciting actions and the probability that the image frame contains exciting actions are output. The probability that the image frame contains exciting actions is used as the action score of the image frame.
[0125] S1042, input each image frame in the third image sequence into the image sharpness scoring model, and output the sharpness score of each image frame through the image sharpness scoring model.
[0126] In one embodiment of this application, the image sharpness scoring model can be the Brenner gradient function. The sharpness factor of an image frame is calculated using the image sharpness scoring model, and this sharpness factor is used as the sharpness score of the image frame. The formula for calculating the Brenner gradient function is as follows:
[0127] D(f)=∑ y ∑ x |f(x+2,y)-f(x,y)| 2 (1).
[0128] In the calculation formula (1), (x,y) is the pixel value of each pixel in the image frame, f(x,y) is the gray value of each pixel, and D(f) is the sharpness factor of the image frame.
[0129] In other embodiments of this application, the image sharpness scoring model may also be the Tenengrad gradient function, the Laplacian operator gradient function, the gray-level variance function, the gray-level variance product function, the energy gradient function, etc.
[0130] In another embodiment of this application, the target object in each image frame of the third image sequence can also be identified, the sharpness score of the target object can be obtained, and the sharpness score of the image frame can be obtained, thereby performing a local blur evaluation on the image frame.
[0131] S1043, input each image frame in the third image sequence into the expression scoring model, and output the expression score of each image frame through the expression scoring model.
[0132] In one embodiment of this application, the expression scoring model can be a convolutional neural network model, which is trained and generated using facial images corresponding to multiple preset expressions as training data. These multiple preset expressions include, but are not limited to, smiling, crying, excitement, and tension. These preset expressions are merely illustrative examples, and in actual applications, they may not be limited to these examples; this application embodiment does not impose any limitations on them.
[0133] In one embodiment of this application, each image frame in the third image sequence is input into the expression scoring model. The features of the image frame are extracted by the expression scoring model. Based on the extracted features, it is analyzed whether the image frame contains a preset expression. The determination result of whether the image frame contains a preset expression and the probability that the image frame contains a preset expression are output. The probability that the image frame contains a preset expression is used as the expression score of the image frame.
[0134] S1044. Determine the comprehensive score for each image frame based on the action score, image sharpness score, and expression score for each image frame.
[0135] In one embodiment of this application, the sum of the motion score, image clarity score, and expression score for each image frame is calculated to obtain a comprehensive score for each image frame.
[0136] In another embodiment of this application, the motion score, image sharpness score, and expression score of each image frame are weighted and summed to obtain a comprehensive score for each image frame. The sum of the weights of the motion score, image sharpness score, and expression score is 1.
[0137] S1045, the image frame with the highest comprehensive score in the third image sequence is determined as the reference image frame.
[0138] In one embodiment of this application, the comprehensive scores of all image frames in the third image sequence are compared, and the image frame with the highest comprehensive score is determined as the reference image frame.
[0139] The above embodiments of this application can use multiple scoring models to select the best image frames in an image stream, thereby filtering out image frames containing exciting actions as reference image frames, which facilitates the generation of accurate videos and images of exciting actions based on the reference image frames.
[0140] See Figure 13 The diagram shown is a flowchart of generating a fourth image sequence according to an embodiment of this application.
[0141] S1051, Identify the target object in multiple image frames of the first image sequence.
[0142] In one embodiment of this application, multiple image frames of a first image sequence are input into a target detection model, and the target object in the multiple image frames of the first image sequence is determined by the target detection model. The target detection model can be a model established based on the principle of an instance segmentation algorithm.
[0143] S1052, calculate the motion amplitude of the target object based on multiple image frames in the first image sequence.
[0144] In one embodiment of this application, an optical flow method is used to calculate the motion vector of the second image frame relative to the first image frame in every two adjacent image frames of the first image sequence, wherein the second image frame is captured later than the first image frame. The optical flow method can be the Lucas-Kanade optical flow algorithm, the KLT optical flow algorithm, the Farneback optical flow algorithm, a deep learning-based optical flow algorithm, etc.
[0145] In one embodiment of this application, after calculating the motion vector of the second image frame relative to the first image frame, the motion vectors of the target object and all pixels in its neighborhood are obtained according to a pixel window of a preset size. The mean of the motion vectors of the target object and all pixels in its neighborhood is calculated to obtain the motion amplitude of the second image frame relative to the first image frame. For example, the preset size can be 3*3, 5*5, or other sizes.
[0146] In one embodiment of this application, the average motion amplitude of the second image frame relative to the first image frame in all adjacent image frames of the first image sequence is calculated to obtain the motion amplitude of the target object. The motion amplitude of the target object is expressed in pixels and is used to describe the degree and direction of motion exhibited by the target object in the image.
[0147] S1053, determine whether the motion amplitude of the target object is greater than or equal to the preset motion amplitude threshold. If the motion amplitude of the target object is greater than or equal to the preset motion amplitude threshold, the process proceeds to S1054; if the motion amplitude of the target object is less than the preset motion amplitude threshold, the process proceeds to S1055.
[0148] S1054, Based on the reference image frame, obtain multiple image frames within a third preset time period from the first image sequence to generate a fourth image sequence.
[0149] See Figure 14The diagram illustrates the generation of a fourth image sequence according to an embodiment of this application. In one embodiment, if the motion amplitude LV1 of the target object is greater than or equal to a preset motion amplitude threshold, a reference image frame, multiple image frames within a fifth preset time period before the reference image frame, and multiple image frames within a sixth preset time period after the reference image frame are obtained from the first image sequence to generate a fourth image sequence. For example, the preset motion amplitude threshold is 10 pixel units. The sum of the fifth and sixth preset time periods is the third preset time period. For example, the third preset time period is 0.75 seconds, the fifth preset time period is 0.5 seconds, and the sixth preset time period is 0.25 seconds.
[0150] S1055, Based on the reference image frame, obtain multiple image frames within a fourth preset time period from the first image sequence to generate a fourth image sequence.
[0151] In one embodiment of this application, if the motion amplitude LV2 of the target object is less than a preset motion amplitude threshold, a fourth image sequence is generated by acquiring a reference image frame, multiple image frames within a seventh preset time period before the reference image frame, and multiple image frames within an eighth preset time period after the reference image frame from the first image sequence. The fourth preset time period is greater than the third preset time period, and the sum of the seventh and eighth preset time periods is the fourth preset time period. For example, the fourth preset time period is 1.5 seconds, the seventh preset time period is 1 second, and the sixth preset time period is 0.5 seconds.
[0152] The target video generated in this application embodiment is a time-lapse video or slow-motion video of exciting actions. Therefore, if the motion amplitude of the target object in the image frame is large, fewer image frames are obtained for interpolation processing; if the motion amplitude of the target object in the image frame is small, more image frames are obtained for interpolation processing.
[0153] In another embodiment of this application, in S1053 and S1054, multiple image frames can be obtained from the second image sequence corresponding to the image stream using the same method to generate a fourth image sequence, thereby generating a video with a higher resolution.
[0154] The above embodiments of this application can estimate the motion speed of a target moving object by using multiple image frames in an image stream, and obtain different numbers of image frames according to the motion speed of the target moving object, so that the generated video of exciting action is adapted to the motion speed of the target moving object.
[0155] See Figure 15 The diagram shown is a flowchart of generating a target video according to an embodiment of this application.
[0156] S1061, determine whether the motion amplitude of the target object is greater than or equal to the preset motion amplitude threshold. If the motion amplitude of the target object is greater than or equal to the preset motion amplitude threshold, the process proceeds to S1062; if the motion amplitude of the target object is less than the preset motion amplitude threshold, the process proceeds to S1063.
[0157] S1062, interpolate multiple image frames in the fourth image sequence according to the first frame rate to generate a target video of the ninth preset duration.
[0158] In one embodiment of this application, the first frame rate is 240 fps (frames per second), and the ninth preset duration is 6 seconds. Based on the first frame rate, a preset interpolation algorithm is used to interpolate multiple image frames in the fourth image sequence to generate a target video of the ninth preset duration.
[0159] In one embodiment of this application, the preset interpolation algorithm can be a linear interpolation algorithm, a bilinear interpolation algorithm, or an optical flow-based interpolation algorithm. A linear interpolation algorithm, based on a linear relationship, generates an intermediate image frame by calculating the linear interpolation of pixels between two adjacent image frames. For example, it calculates the position ratio of each pixel in two adjacent image frames, and then performs a weighted average of the pixel values of the two pixels based on the position ratio to generate the pixel value of the corresponding pixel in the intermediate frame. A bilinear interpolation algorithm generates an intermediate image frame by performing a weighted average of the pixel values of the four nearest pixels. An optical flow-based interpolation algorithm uses an optical flow estimation method to predict the pixel displacement of the intermediate image frame, and then calculates the pixel value of the corresponding pixel in the intermediate image frame through interpolation.
[0160] S1063, interpolate multiple image frames in the fourth image sequence according to the second frame rate to generate a target video of a ninth preset duration. The first frame rate is greater than the second frame rate; for example, the second frame rate is 120fps.
[0161] The above embodiments of this application can interpolate multiple image frames at different frame rates according to the movement speed of the target moving object to generate a video of exciting action, so that the generated video of exciting action is adapted to the movement speed of the target moving object.
[0162] For example, see Figure 16 As shown, this is a sequence of multiple image frames in the fourth image sequence, including image frame a, image frame b, and image frame c. (See also...) Figure 17As shown, multiple image frames are generated after interpolation. Interpolation generates multiple image frames between image frame a and image frame b, and between image frame b and image frame c. The target video is then generated based on these generated image frames, image frame a, image frame b, and image frame c. Compared to the actions in the multiple image frames in the fourth image sequence, the actions in the multiple image frames of the target video are richer and more coherent, making it suitable for showcasing exciting actions.
[0163] See Figure 18 The diagram illustrates the generation of a fifth image sequence according to an embodiment of this application. In one embodiment, a reference image frame and a fourth preset number of image frames preceding the reference image frame are obtained from the second image sequence to generate the fifth image sequence. For example, the fourth preset number is 3. A preset image-capturing algorithm is executed on multiple image frames in the fifth image sequence to obtain a capture result. For example, the preset image-capturing algorithm includes, but is not limited to, a registration algorithm, a noise reduction algorithm, and a fusion algorithm. In another embodiment, a reference image frame, a fifth preset number of image frames preceding the reference image frame, and a sixth preset number of image frames following the reference image frame can also be obtained from the second image sequence to generate the fifth image sequence. For example, the fifth preset number is 3, and the sixth preset number is 2.
[0164] See Figure 19 The diagram shown is a flowchart of generating a target image according to an embodiment of this application.
[0165] S1081, register multiple image frames in the fifth image sequence.
[0166] In one embodiment of this application, a Scale-Invariant Feature Transform (SIFT) algorithm is used to register multiple image frames in a fifth image sequence. First, scale-space extremum detection is performed. A Gaussian filter is used to construct a scale-space pyramid for each image frame in the fifth image sequence. Keypoints are detected by Gaussian differences at different scales to extract stable scale-invariant features. Next, keypoint localization is performed. In each scale space, keypoints are selected by comparing the gradient and curvature of pixels with their surrounding pixels. Local extrema of the image are used for precise localization and suppression of edge responses. Then, orientation assignment is performed. A dominant orientation is assigned to each keypoint to ensure rotation invariance of the descriptor. A gradient orientation histogram is calculated, and the dominant orientation or multiple orientations are selected as the orientation of the keypoint. Next, descriptor generation is performed. Using the scale and orientation information of the keypoint, a descriptor is constructed in the neighborhood of the keypoint. During descriptor extraction, a multi-dimensional (e.g., 128-dimensional) feature vector is generated by combining the local gradient and gradient orientation of the image. Finally, feature matching is performed, matching the feature descriptors of every two images. For example, the similarity between feature vectors can be measured based on Euclidean distance or cosine similarity to select the best match. Finally, registration is performed. Based on the matched feature point pairs, an appropriate registration algorithm (such as the Random Sample Consensus Algorithm RANSAC) is used to calculate the transformation matrix between the images to achieve image alignment and registration.
[0167] S1082, noise reduction processing is performed on multiple image frames in the fifth image sequence.
[0168] In one embodiment of this application, a Laplacian pyramid algorithm is used to denoise multiple image frames in a fifth image sequence. First, images of different scales are generated by downsampling and Gaussian filtering the original image multiple times. From each level of the Gaussian pyramid, a corresponding Laplacian pyramid is constructed through upsampling and interpolation operations. The Laplacian pyramid contains detailed information at various image scales. Mean filtering or other filters are used to smooth the pixel values at each pyramid level, thereby performing denoising at each level of the Laplacian pyramid. The denoised Laplacian pyramid is combined with the Gaussian pyramid, and the denoised image frame is reconstructed through layer-by-layer upsampling and weighted summation.
[0169] S1083, multiple image frames in the fifth image sequence are fused to obtain a fused image.
[0170] In one embodiment of this application, an image fusion algorithm is used to fuse multiple image frames in the fifth image sequence after noise reduction to obtain a fused image. For example, the image fusion algorithm may be a Laplacian pyramid fusion algorithm, a wavelet transform fusion algorithm, a region segmentation and weighted average fusion algorithm, a neural network-based image fusion algorithm, etc.
[0171] By performing registration and noise reduction preprocessing before image fusion through the above embodiments of this application, the quality of the fused image can be improved.
[0172] In another embodiment of this application, key features of multiple image frames in the fifth image sequence can be extracted, and image fusion can be performed based on the key features, thereby highlighting the key features in the fused target image and making the target image present a blurring effect.
[0173] This application also provides an electronic device 100, see reference. Figure 20 As shown, the electronic device 100 may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, in-vehicle device, smart home device and / or smart city device. The specific type of electronic device 100 is not specifically limited in the embodiments of this application.
[0174] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, Universal Serial Bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and Subscriber Identification Module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0175] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0176] Processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0177] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0178] The processor 110 may also include a memory for storing instructions and data. In one embodiment of this application, the memory in the processor 110 is a cache memory. The memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instructions or data again, it can directly retrieve them from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0179] In one embodiment of this application, the processor 110 may include one or more interfaces. These interfaces may include an Inter-integrated Circuit (I2C) interface, an Inter-integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI) interface, a General-Purpose Input / Output (GPIO) interface, a Subscriber Identity Module (SIM) interface, and / or a Universal Serial Bus (USB) interface, etc.
[0180] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In one embodiment of this application, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.
[0181] The I2S interface can be used for audio communication. In one embodiment of this application, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to realize communication between the processor 110 and the audio module 170. In one embodiment of this application, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to realize the function of answering phone calls through a Bluetooth headset.
[0182] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In one embodiment of this application, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In another embodiment of this application, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0183] The UART interface is a universal serial data bus used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In one embodiment of this application, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In one embodiment of this application, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback via Bluetooth headphones.
[0184] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a Camera Serial Interface (CSI) and a Display Serial Interface (DSI). In one embodiment of this application, the processor 110 and the camera 193 communicate via the CSI interface to realize the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate via the DSI interface to realize the display function of the electronic device 100.
[0185] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In one embodiment of this application, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0186] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. Furthermore, the interface can be used to connect other electronic devices 100, such as AR devices.
[0187] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0188] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via a USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device 100 via the power management module 141.
[0189] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0190] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0191] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0192] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In one embodiment of this application, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In another embodiment of this application, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0193] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In one embodiment of this application, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and housed within the same device as the mobile communication module 150 or other functional modules.
[0194] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including Wireless Local Area Networks (WLANs) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0195] In one embodiment of this application, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the Beidou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0196] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0197] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a Liquid Crystal Display (LCD), an Organic Light-Emitting Diode (OLED), an Active-Matrix Organic Light-Emitting Diode (AMOLED), a Flexible Light-Emitting Diode (FLED), a Minied, Microled, Micro-OLED, or a Quantum Dot Light-Emitting Diode (QLED), etc. In one embodiment of this application, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0198] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0199] The ISP is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In one embodiment of this application, the ISP can be set in the camera 193.
[0200] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In one embodiment of this application, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0201] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0202] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0203] NPU stands for Neural Network (NN) computing processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0204] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0205] Random access memory can include static random-access memory (SRAM), dynamic random-access memory (DRAM), synchronous dynamic random-access memory (SDRAM), and double data rate synchronous dynamic random-access memory (DDR SDRAM, such as fifth-generation DDR SDRAM, which is generally called DDR5 SDRAM).
[0206] Non-volatile memory can include disk storage devices and flash memory.
[0207] Flash memory can be classified according to its operating principle, including NOR FLASH, NAND FLASH, 3D NAND FLASH, etc.; according to the level of the storage cell, including single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc.; and according to the storage specification, including universal flash storage (UFS) and embedded multi-media card (eMMC), etc.
[0208] The random access memory can be directly read and written by the processor 110. It can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data.
[0209] Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into random access memory for direct reading and writing by the processor 110.
[0210] The external memory interface 120 can be used to connect to external non-volatile memory, thereby expanding the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to perform data storage functions. For example, music, video, and other files can be stored in the external non-volatile memory.
[0211] Internal memory 121 or external memory interface 120 is used to store one or more computer programs. The one or more computer programs are configured to be executed by processor 110. The one or more computer programs include multiple instructions, which, when executed by processor 110, can implement the screen display detection method executed on electronic device 100 in the above embodiments, so as to realize the screen display detection function of electronic device 100.
[0212] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0213] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In one embodiment of this application, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0214] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0215] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0216] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0217] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0218] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0219] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0220] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0221] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In one embodiment of this application, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100. This application also provides a computer storage medium storing computer instructions. When the computer instructions are executed on the electronic device 100, the electronic device 100 performs the above-mentioned related method steps to implement the shooting processing method in the above embodiments.
[0222] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the shooting processing method described in the above embodiments.
[0223] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory; wherein, the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the imaging processing methods in the above-described method embodiments.
[0224] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.
[0225] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0226] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0227] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0228] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0229] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0230] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.
Claims
1. A method for processing images, applied to electronic devices, characterized in that, The method includes: In response to a user activating the camera application, the electronic device captures an image stream through its camera. In response to the user's shooting operation, a first image sequence is generated based on the image stream; Multiple image frames are obtained from the first image sequence to generate a second image sequence; A reference image frame is selected from the second image sequence based on the score of each image frame in the second image sequence; Identify the target object in multiple image frames of the first image sequence; The motion amplitude of the target object is calculated based on multiple image frames in the first image sequence; If the motion amplitude of the target object is greater than or equal to a preset motion amplitude threshold, multiple image frames within a third preset time period are obtained from the first image sequence based on the reference image frame to generate a third image sequence; Generate a target video based on multiple image frames in the third image sequence; A fourth image sequence is generated by acquiring multiple image frames from the first image sequence based on the reference image frame; The target image is generated based on multiple image frames in the fourth image sequence.
2. The shooting and processing method as described in claim 1, characterized in that, The step of responding to the user's shooting operation and generating a first image sequence based on the image stream includes: Based on the time when the user performs the shooting operation, the image frames captured by the camera within a first preset time period before the time and the image frames captured by the camera within a second preset time period after the time are obtained from the image stream to generate the first image sequence.
3. The shooting and processing method as described in claim 1, characterized in that, The step of obtaining multiple image frames from the first image sequence to generate a second image sequence includes: The second image sequence is generated by acquiring one image frame from every first preset number of consecutive image frames in the first image sequence.
4. The shooting and processing method as described in claim 1, characterized in that, The step of selecting a reference image frame from the second image sequence based on the score of each image frame in the second image sequence includes: Each image frame in the second image sequence is input into at least one scoring model to obtain at least one score for each image frame; The reference image frame is selected from the second image sequence based on at least one score for each image frame.
5. The shooting and processing method as described in claim 4, characterized in that, The at least one scoring model includes an action scoring model, an image sharpness scoring model, and an expression scoring model.
6. The shooting and processing method as described in claim 1, characterized in that, The identification of target objects in multiple image frames of the first image sequence includes: Multiple image frames from the first image sequence are input into the target detection model, and the target object in the multiple image frames of the first image sequence is output by the target detection model.
7. The shooting and processing method as described in claim 6, characterized in that, The step of calculating the motion amplitude of the target object based on multiple image frames in the first image sequence includes: The motion amplitude of the target object is calculated using optical flow based on each pair of adjacent image frames in the first image sequence.
8. The shooting and processing method as described in claim 1, characterized in that, The method further includes: If the motion amplitude of the target object is less than the preset motion amplitude threshold, multiple image frames within a fourth preset duration are obtained from the first image sequence based on the reference image frame to generate the third image sequence, wherein the third preset duration is less than the fourth preset duration.
9. The shooting and processing method as described in claim 8, characterized in that, The step of obtaining multiple image frames within a third preset time period from the first image sequence based on the reference image frame and generating a third image sequence includes: The third image sequence is generated by obtaining multiple image frames within a fifth preset time period before the reference image frame from the first image sequence, and multiple image frames within a sixth preset time period after the reference image frame, wherein the sum of the fifth preset time period and the sixth preset time period is the third preset time period.
10. The image processing method as described in claim 9, characterized in that, The step of generating the target video based on multiple image frames in the third image sequence includes: If the motion amplitude of the target object is greater than or equal to the preset motion amplitude threshold, interpolation processing is performed on multiple image frames in the third image sequence according to the first frame rate to generate the target video of the ninth preset duration; or If the motion amplitude of the target object is less than the preset motion amplitude threshold, interpolation processing is performed on multiple image frames in the third image sequence according to the second frame rate to generate the target video of the ninth preset duration, wherein the first frame rate is greater than the second frame rate.
11. The image processing method as described in claim 8, characterized in that, The step of obtaining multiple image frames within a fourth preset time period from the first image sequence based on the reference image frame and generating the third image sequence includes: The third image sequence is generated by obtaining multiple image frames within a seventh preset time period before the reference image frame from the first image sequence, and multiple image frames within an eighth preset time period after the reference image frame, wherein the sum of the seventh preset time period and the eighth preset time period is the fourth preset time period.
12. The image processing method as described in claim 9, characterized in that, The method further includes: Store the target video to the gallery application of the electronic device; The reference image frame is optimized, and the optimized reference image frame is used as the cover image of the target video.
13. The shooting and processing method as described in claim 1, characterized in that, The step of obtaining multiple image frames from the first image sequence based on the reference image frame to generate a fourth image sequence includes: The fourth image sequence is generated by obtaining the reference image frame and a second preset number of image frames preceding the reference image frame from the first image sequence.
14. The image processing method as described in claim 13, characterized in that, The step of generating a target image based on multiple image frames in the fourth image sequence includes: Multiple image frames in the fourth image sequence are fused to obtain a fused image; The fused image is optimized to obtain the target image.
15. The imaging processing method according to any one of claims 1 to 14, characterized in that, The image stream includes a preview image stream and a captured image stream, wherein the resolution of the image frames in the captured image stream is greater than the resolution of the image frames in the preview image stream.
16. An electronic device, characterized in that, The electronic device includes a memory and a processor: The memory is used to store program instructions; The processor is configured to read and execute the program instructions stored in the memory, and when the program instructions are executed by the processor, the electronic device performs the shooting processing method as described in any one of claims 1 to 15.
17. A chip coupled to a memory in an electronic device, characterized in that, The chip is used to control the electronic device to perform the shooting processing method as described in any one of claims 1 to 15.
18. A computer storage medium, characterized in that, The computer storage medium stores program instructions that, when executed on the electronic device, cause the processor of the electronic device to perform the image processing method as described in any one of claims 1 to 15.
Citation Information
Patent Citations
Image acquisition method and device, terminal and storage medium
CN108198177A
Image processing method and device
CN117479024A
Intelligent video editing method and system
US20220059133A1
Video processing method, and electronic device and readable medium
WO2023160142A1