Image processing method and electronic equipment

By detecting motion in electronic devices and adjusting gyroscope data using a target parameter tuning model for image stabilization, the problem of image jitter caused by user hand tremors or device shaking is solved, improving the stability of video recording and user experience, while saving resources.

CN121908133APending Publication Date: 2026-04-21HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-10-11
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

During video recording, image shakiness caused by user hand shaking or device movement can affect the recording quality.

Method used

By acquiring initial gyroscope data from electronic devices, detecting the device's motion state, and adjusting the gyroscope data using a target parameter tuning model during motion to perform image stabilization, the stabilization operation is only performed when necessary to reduce resource waste.

Benefits of technology

It improves the image stabilization effect of video images, enhances the user's shooting experience, and reduces unnecessary power consumption and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908133A_ABST
    Figure CN121908133A_ABST
Patent Text Reader

Abstract

The embodiment of the invention is applied to the technical field of electronics, and provides an image processing method and electronic equipment. In response to a first operation of a user, the electronic device acquires initial gyroscope data when acquiring a first image frame. Wherein the first operation is used for triggering the electronic equipment to display a preview interface corresponding to the camera application. Afterwards, the electronic equipment can detect whether the electronic equipment is in a motion state or not according to the initial gyroscope data. Afterwards, under the condition that the electronic equipment is in a motion state, the electronic equipment adjusts the initial gyroscope data by utilizing the target parameter adjustment model to obtain target gyroscope data. Afterwards, the electronic device can perform anti-shake processing on the first image frame according to the target gyroscope data to obtain a target image frame. And then, the electronic equipment can display the target image frame in a preview interface. According to the invention, the image stabilization effect of the video image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technology, and in particular to an image processing method and an electronic device. Background Technology

[0002] With the development of shooting technology, users have increasingly higher requirements for video recording quality. During video recording, factors such as user hand tremors or shaking of electronic devices can easily cause the electronic device to move, resulting in shaky images. Therefore, how to reduce image shakiness caused by electronic device shaking to improve video recording quality is an urgent problem to be solved. Summary of the Invention

[0003] This application provides an image processing method and an electronic device to solve image jitter caused by user manual operation or electronic device shaking, improve the image stabilization effect of video images, and thus enhance the user's shooting experience.

[0004] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0005] Firstly, an image processing method is provided for an electronic device. In this method, in response to a user's first operation, the electronic device acquires initial gyroscope data when capturing a first image frame. This first operation triggers the electronic device to display a preview interface corresponding to a camera application. Subsequently, the electronic device can detect whether it is in motion based on the initial gyroscope data. Then, if the electronic device is in motion, it adjusts the initial gyroscope data using a target parameter tuning model to obtain target gyroscope data. Next, the electronic device performs image stabilization processing on the first image frame based on the target gyroscope data to obtain the target image frame. Finally, the electronic device can display the target image frame in the preview interface.

[0006] In this application, since the target image frame is obtained by image processing based on target gyroscope data, and this target gyroscope data is obtained by adjusting the initial gyroscope data, the target image frame can be a stabilized image frame. This reduces image jitter caused by user actions or electronic device shaking, improving the image stabilization effect of the video image and thus enhancing the user's shooting experience. Furthermore, the electronic device only performs image stabilization when it is in motion. This reduces the waste of resources caused by frame-by-frame image stabilization, improves the utilization of computing resources, and reduces unnecessary power consumption.

[0007] In one possible implementation of the first aspect, the user's first operation may include any of the following: the user opening the camera application, the user controlling the electronic device to enter recording mode when the camera application is open, or the user clicking the recording control when the camera application is open.

[0008] In one possible implementation of the first aspect, the process for detecting whether the electronic device is in motion may specifically include: the electronic device calculating the average value of the initial gyroscope data based on multiple gyroscope data points. Then, the electronic device calculates a target value of the initial gyroscope data based on the average value. This target value includes a standard deviation and / or a variance. Subsequently, if the target value is less than a preset value, the electronic device determines that the electronic device is stationary; or, if the target value is greater than or equal to the preset value, the electronic device determines that the electronic device is in motion.

[0009] In this application, if the target value is less than the preset value, it indicates that the electronic device is relatively stable, meaning that the electronic device is not shaking or jittering. Therefore, the electronic device can determine that it is stationary and can directly refrain from performing image stabilization, thereby reducing unnecessary power consumption and extending the device's lifespan. Conversely, if the target value is greater than or equal to the preset value, it indicates that the electronic device may be shaking or jittering. Therefore, the electronic device can determine that it is in motion and can continue to detect motion to match the corresponding target parameter tuning model, providing a basis for subsequent adjustments to the camera pose.

[0010] In one possible implementation of the first aspect, the method further includes: when the electronic device is stationary, the electronic device does not perform anti-shake operation.

[0011] In this application, if the electronic device is stationary, it means there is no possibility of it shaking, implying it is likely housed in a fixed device. Therefore, to reduce unnecessary power consumption, the electronic device can forgo image stabilization, allowing it to directly output a video stream composed of multiple first image frames. This extends the device's lifespan and improves the user experience. Furthermore, since image stabilization is not performed, subsequent detection and pose adjustment are unnecessary, thus improving resource utilization and reducing waste.

[0012] In one possible implementation of the first aspect, the process of stabilizing the first image frame by the electronic device may specifically include: the electronic device adjusting the initial gyroscope data using a target parameter tuning model to obtain target gyroscope data. The target parameter tuning model is one of multiple candidate parameter tuning models, and the model parameters corresponding to different candidate parameter tuning models are different. Then, the electronic device obtains a rotation matrix based on the target gyroscope data. Next, the electronic device adjusts the image matrix of the first image frame based on the rotation matrix to obtain a target image matrix. The target image matrix is ​​used to represent the target image frame.

[0013] In this application, since the target parameter tuning model is one of multiple candidate parameter tuning models, and the model parameters corresponding to different candidate parameter tuning models are different, the electronic device can select any parameter tuning model from multiple candidate parameter tuning models according to the shooting situation to adjust the initial gyroscope data. In this way, the optimal parameter tuning model can be selected for targeted parameter adjustment, thereby improving the parameter adjustment rate and providing convenient conditions for subsequent rapid image stabilization.

[0014] In one possible implementation of the first aspect, the process of determining the target hyperparameter tuning model may specifically include: an electronic device acquiring an optical flow image. This optical flow image is used to characterize the motion information of the same object in adjacent image frames, which include a first image frame and a second image frame, the second image frame being acquired before the first image frame. Subsequently, the electronic device can determine the target hyperparameter tuning model from multiple candidate hyperparameter tuning models based on the optical flow image.

[0015] In this application, the electronic device can select a candidate parameter tuning model that matches the motion speed of the object pixels carried in the optical flow image from multiple candidate parameter tuning models as the target parameter tuning model. This allows the electronic device to select the appropriate parameter tuning model based on the motion of objects in adjacent image frames, meaning it can selectively choose the optimal parameter tuning model for parameter adjustment, thereby improving the parameter adjustment rate and facilitating subsequent rapid image stabilization.

[0016] In one possible implementation of the first aspect, the process of the electronic device determining the target hyperparameter tuning model may specifically include: the electronic device detecting whether the object pixels in the first image frame satisfy a preset motion rule based on the optical flow image. Then, if the object pixels in the first image frame do not satisfy the preset motion rule, the electronic device uses the first hyperparameter tuning model among multiple candidate hyperparameter tuning models as the target hyperparameter tuning model.

[0017] In this application, if the object pixels in the first image frame do not meet the preset motion rules, it indicates that the motion of the object pixels in the first image frame is not regular, meaning that the image jitter generated when the electronic device captures the first image frame is not regular. Therefore, the mobile phone can use the common parameter tuning model (i.e., the first parameter tuning model) among multiple candidate parameter tuning models as the target parameter tuning model. In this way, the utilization rate of mobile phone resources can be improved, and resource waste can be reduced.

[0018] In one possible implementation of the first aspect, the detection process for whether object pixels in the first image frame satisfy a preset motion rule may specifically include: the electronic device acquiring a historical optical flow image. This historical optical flow image is an optical flow image acquired before the historical optical flow image during video recording, and it carries the historical motion speed of the object pixels. Then, if the speed difference between the current motion speed and the historical motion speed is less than a preset difference, the electronic device determines that the object pixels in the first image frame satisfy the preset motion rule; or, if the speed difference between the current motion speed and the historical motion speed is greater than or equal to the preset difference, the electronic device determines that the object pixels in the first image frame do not satisfy the preset motion rule.

[0019] In this application, the motion velocity carried by historical optical flow images is compared with the motion velocity carried by other optical flow images to determine whether the object pixels in the first image frame meet a preset motion pattern. This enables accurate detection of the motion pattern of objects in the image, thus providing a foundation for subsequent accurate determination of the target parameter tuning model.

[0020] In one possible implementation of the first aspect, the method further includes: if the object pixels in the first image frame satisfy a preset motion law, the electronic device detects whether the object motion in the first image frame is intense based on the optical flow image. Then, if the object motion in the first image frame is intense, the electronic device uses a second parameter tuning model from multiple candidate parameter tuning models as the target parameter tuning model; wherein the second parameter tuning model can improve the convergence speed compared to the first parameter tuning model.

[0021] In this application, after determining that the object pixels in the first image frame meet the preset motion rules, the electronic device can continue to detect whether the object motion in the first image frame is violent. If the object motion in the first image frame is violent, it indicates that the electronic device may have experienced significant jitter when capturing the first image frame, meaning the electronic device may have experienced severe jitter. Therefore, the electronic device can use the second parameter tuning model from multiple candidate parameter tuning models as the target parameter tuning model. In this way, a target parameter tuning model that is suitable for violently moving scenes can be selected from multiple candidate parameter tuning models, thereby improving the adaptability of the parameter tuning model. In addition, since the second parameter tuning model can improve the convergence speed compared to the first parameter tuning model, the electronic device can reduce the number of model iterations and improve the adjustment rate of the gyroscope data by inputting the initial gyroscope data into the second parameter tuning model.

[0022] In one possible implementation of the first aspect, the method further includes: when the object in the first image frame moves slowly, the electronic device uses a third parameter tuning model among multiple candidate parameter tuning models as the target parameter tuning model; wherein the third parameter tuning model can improve convergence accuracy and convergence stability compared to the first parameter tuning model.

[0023] In this application, if the object in the first image frame moves slowly, it indicates that the shaking degree when the phone captured the first image frame may be small, meaning the phone shaking may be relatively smooth. Therefore, the phone can call a third parameter tuning model from multiple candidate parameter tuning models and use this third parameter tuning model as the target parameter tuning model. In this way, a target parameter tuning model that is suitable for the smooth motion scene can be selected from multiple candidate parameter tuning models, thereby improving the adaptability of the parameter tuning model. In addition, since the third parameter tuning model can improve the convergence accuracy and convergence stability compared to the first parameter tuning model, the electronic device can improve the accuracy of data adjustment by inputting gyroscope data into the third parameter tuning model, thereby improving the image stabilization accuracy of the video image.

[0024] In one possible implementation of the first aspect, the detection process of whether the object in the first image frame is moving violently may specifically include: if the movement speed of the object pixel is greater than a preset speed, the electronic device determines that the object in the first image frame is moving violently; or, if the movement speed of the object pixel is less than or equal to the preset speed, the electronic device determines that the object in the first image frame is moving slowly.

[0025] In this application, the motion of an object in the first image frame is determined by judging whether the motion speed of the object's pixels is greater than a preset speed. This enables accurate detection of the intensity of object motion in the image, improves the accuracy of the third detection result, and provides a foundation for subsequent precise target determination and parameter tuning models.

[0026] In one possible implementation of the first aspect, before adjusting the initial gyroscope data using the target hyperparameter tuning model, the method further includes: the electronic device triggering a warm-start function to activate the target hyperparameter tuning model. This warm-start function is used to increase the data output speed of the target hyperparameter tuning model.

[0027] In this application, if the hot-start function of the target parameter tuning model is enabled, the electronic device can obtain the target gyroscope data more quickly, reduce the number of iterations, and accelerate the convergence speed of the model.

[0028] In one possible implementation of the first aspect, the method further includes: in response to a user's stop recording operation, the electronic device generates a video stream based on multiple target image frames acquired during video recording. The electronic device then saves the video stream.

[0029] In this application, in response to the user's stop recording operation, the electronic device can assemble multiple target image frames obtained during video recording into a video stream and store the video stream in the electronic device for later viewing by the user, thus improving the user experience. Furthermore, since the video stream is a stabilized video stream, precise image stabilization can be achieved, improving the image stabilization effect and further enhancing the user's shooting experience.

[0030] Secondly, this application provides an electronic device, the electronic device including a camera, a display screen, a gyroscope sensor, one or more processors, and one or more memories; the one or more processors are coupled to the camera, the display screen, the gyroscope sensor, and the one or more memories; the camera is used to acquire a first image frame, the display screen is used to display a target image frame, the gyroscope sensor is used to acquire initial gyroscope data, and the one or more memories are used to store computer program code, the computer program code including computer instructions, which, when the one or more processors execute the computer instructions, cause the electronic device to perform the method described above.

[0031] Thirdly, this application provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the method described above.

[0032] Fourthly, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the method described above.

[0033] Fifthly, a chip is provided, comprising: an input interface, an output interface, a processor, and a memory, wherein the input interface, the output interface, the processor, and the memory are connected via an internal connection path, and the processor is used to execute code in the memory, wherein when the code is executed, the processor is used to execute the method described above.

[0034] The beneficial effects that the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect can achieve can be referred to the beneficial effects of the first aspect and any of its possible design embodiments, and will not be repeated here. Attached Figure Description

[0035] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0036] Figure 2 A schematic diagram illustrating the front-facing camera and rear-facing camera of a mobile phone, provided as an embodiment of this application;

[0037] Figure 3 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application;

[0038] Figure 4 A flowchart illustrating an image processing method provided in an embodiment of this application;

[0039] Figure 5 A schematic diagram of a video recording process provided in an embodiment of this application;

[0040] Figure 6 A flowchart for detecting whether a mobile phone is in motion is provided in an embodiment of this application;

[0041] Figure 7 A flowchart illustrating another image processing method provided in this application embodiment;

[0042] Figure 8 A schematic diagram illustrating the output of target gyroscope data provided in an embodiment of this application;

[0043] Figure 9 This is a flowchart illustrating the determination of target gyroscope data, provided as an embodiment of this application. Detailed Implementation

[0044] The technical solutions of the embodiments of this application are described below with reference to the accompanying drawings. In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to limit the application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0045] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.

[0046] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0047] When recording video using electronic devices, shaky hands or movement of the device itself can cause motion in both the device and the camera module (hereinafter referred to as the camera), resulting in jittery and blurry images. For example, when the moving subject is close to the camera, significant relative movement will cause more noticeable jitter in the image, leading to a poor user experience. For instance, in a selfie recording scenario, the face is close to the camera, and facial movement will result in more pronounced jitter in the facial image.

[0048] In some embodiments, to reduce image jitter and blur, electronic devices can employ electronic image stabilization (EIS) technology to stabilize video. EIS technology calculates changes in camera rotation and dynamically crops the video frame, making the video appear more stable. In other words, EIS technology stabilizes the video image by cropping the video frame.

[0049] Specifically, electronic devices can perform image stabilization based on the acquired gyroscope data to eliminate image jitter caused by movement of the electronic device between image frames. However, since electronic devices using the aforementioned EIS technology perform image stabilization on each individual image frame—meaning they will still perform stabilization even when the electronic device is stationary—this not only causes unnecessary power consumption but also increases video processing time, thus affecting the user's shooting experience.

[0050] Therefore, in order to improve the utilization of computing resources while achieving image stabilization, this application provides an image processing method. In this method, in response to a user's shooting operation, an electronic device acquires initial gyroscope data, which is the gyroscope data acquired when the electronic device captures the first image frame. Then, the electronic device can detect whether it is in motion based on the initial gyroscope data. Subsequently, if the electronic device is in motion, it uses a target parameter tuning model to adjust the initial gyroscope data to obtain target gyroscope data. Then, the electronic device can perform image stabilization processing on the first image frame based on the target gyroscope data to obtain the target image frame.

[0051] In this embodiment, since the target image frame is obtained by image processing based on target gyroscope data, and this target gyroscope data is obtained by adjusting the initial gyroscope data, the target image frame can be a stabilized image frame. This reduces image jitter caused by user actions or electronic device shaking, improving the image stabilization effect of the video image and thus enhancing the user's shooting experience. Furthermore, the electronic device only performs image stabilization when it is in motion. This reduces the waste of resources caused by frame-by-frame image stabilization, improving the utilization of computing resources and reducing unnecessary power consumption.

[0052] It should be noted that the electronic device in this application embodiment may be a mobile phone, tablet computer, smartwatch, desktop, laptop, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) / virtual reality (VR) device, etc., which includes a camera. This application embodiment does not impose any special restrictions on the specific form of the electronic device.

[0053] For example, Figure 1 A schematic diagram of the structure of electronic device 200 is shown. For example... Figure 1 As shown, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 211, a power management module 212, a battery 213, an antenna 1, an antenna 2, a mobile communication module 240, a wireless communication module 250, an audio module 270, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0054] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0055] Processor 210 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0056] In some embodiments, the electronic device 200 may perform the image processing method provided in this application via the processor 210.

[0057] The wireless communication function of electronic device 200 can be implemented through antenna 1, antenna 2, mobile communication module 240, wireless communication module 250, modem processor, and baseband processor.

[0058] Electronic device 200 implements display functions through a GPU, a display screen 294, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 210 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0059] The display screen (or screen) 294 is used to display images, videos, etc. In some embodiments, the electronic device 200 may include one or N display screens 294, where N is a positive integer greater than 1. In the embodiments of this application, the display screen 294 can be used to display a preview interface and a shooting interface, etc., in video recording mode.

[0060] Electronic device 200 can perform shooting functions through ISP, camera 293, video codec, GPU, display screen 294 and application processor.

[0061] The ISP (Image Signal Processor) is used to process data fed back from the camera 293. For example, when an electronic device takes a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element (or image sensor). The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 293. In some embodiments, the camera 293 includes a shutter. The shutter is a device in the camera used to control the duration of light exposure to the photosensitive element.

[0062] Camera 293 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 200 may include one or N cameras 293, where N is a positive integer greater than 1.

[0063] In some embodiments, camera 293 may include a lens, which is an optical component for generating images.

[0064] For example, the aforementioned N cameras 293 may include one or more front-facing cameras and one or more rear-facing cameras. See, for example, [link to relevant documentation]. Figure 2 Taking the aforementioned electronic device 200 as an example, which is a mobile phone. Figure 2 The interface (a) shows a front-facing camera, such as front-facing camera 20. Figure 2 The interface in (b) shows three rear cameras, such as rear camera 21, rear camera 22, and rear camera 23. Of course, the number of cameras in the above-mentioned mobile phone includes, but is not limited to, the number described in the above embodiments.

[0065] Among them, the above N cameras 293 may include one or more of the following cameras: main camera, telephoto camera, wide-angle camera, ultra-wide-angle camera, macro camera, fisheye camera, infrared camera, depth camera and monochrome camera.

[0066] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 200 selects a frequency point, the DSP is used to perform Fourier transforms on the frequency energy.

[0067] Electronic device 200 can implement audio functions through audio module 270 and application processor, such as music playback and recording. The audio module 270 may include a speaker, receiver, microphone, and headphone jack.

[0068] Buttons 290 include a power button, volume buttons, etc. Indicator 292 may be an indicator light.

[0069] The sensor module 280 may include a gyroscope sensor, an accelerometer sensor, a touch sensor, etc.

[0070] A gyroscope sensor can be used to acquire initial gyroscope data to determine the motion attitude of the electronic device 200. This initial gyroscope data includes angular velocities around three axes. In some embodiments, the angular velocities of the electronic device 200 around the three axes (i.e., x, y, and z axes) can be determined using the gyroscope sensor, thus obtaining the rotation information of the electronic device 200, which is equivalent to obtaining the camera's rotation information. The gyroscope sensor can be used for image stabilization.

[0071] An accelerometer can detect the magnitude of acceleration of an electronic device 200 in various directions (typically three axes). When the electronic device 200 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of the electronic device, and is applied to applications such as screen orientation switching and pedometers.

[0072] A touch sensor, also known as a "touch panel," can be located on the display screen 294. The touch sensor and display screen 294 together form a touchscreen, also known as a "touch screen." The touch sensor detects touch operations applied to or near it. The touch sensor can then transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 294. In some embodiments, the touch sensor may also be located on the surface of the electronic device 200, in a different position than the display screen 294.

[0073] For example, the software system of the aforementioned electronic device 200 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to illustrate the software structure of the electronic device 200.

[0074] Figure 3 This is a software structure block diagram of an electronic device 200 according to an embodiment of this application.

[0075] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer (application layer), the application framework layer (framework layer), the Android runtime and system libraries, and the kernel layer (or driver layer).

[0076] The application layer can include a series of application packages. This application layer can include multiple application packages. For example... Figure 3 As shown, this application package can include applications such as gallery, maps, phone, video, calendar, SMS, and camera. It's understandable that a camera application can be used to trigger an electronic device to record video using its camera.

[0077] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example... Figure 3 As shown, the application framework layer may include a window manager, content provider, phone manager, notification manager, view system, etc.

[0078] The window manager manages the window programs. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture the screen. The content provider stores and retrieves data, making this data accessible to applications. This data may include video, images, audio, made and received phone calls, browsing history and bookmarks, and phonebook entries. The phone manager provides communication functions for the electronic device 200, such as managing call status (including connection and disconnection).

[0079] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0080] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0081] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.

[0082] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0083] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0084] The system library can include multiple functional modules. For example, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), and 2D graphics engines (e.g., SGL). The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. The 3D graphics processing library is used for 3D graphics drawing, image rendering, compositing, and layer processing. The 2D graphics engine is the drawing engine for 2D graphics.

[0085] The Hardware Abstraction Layer (HAL) is a wrapper around Linux kernel drivers, providing interfaces to higher-level systems. It hides the hardware interface details of specific platforms, thus providing a virtual hardware platform for the operating system. In this embodiment, the HAL includes a state detection module, a motion pattern detection module, a model selection module, and a intensity detection module.

[0086] The state detection module is used to detect the motion of the electronic device. In some embodiments, after acquiring initial gyroscope data, the state detection module can perform motion detection on the electronic device based on the initial gyroscope data to determine whether the electronic device is in motion.

[0087] The motion pattern detection module is used to detect the motion patterns in an image. In some embodiments, when an optical flow image is acquired, the motion pattern detection model can detect whether the object pixels in the first image frame have a preset motion pattern based on the optical flow image. The optical flow image is obtained by feature extraction from adjacent image frames and is used to characterize the motion information of the same object in adjacent image frames.

[0088] The intensity detection module is used to detect the intensity of motion in an image. In some embodiments, when an optical flow image is acquired, the intensity detection module can detect whether the motion of an object in the first image frame is intense based on the optical flow image.

[0089] The model selection module is used to select the target hyperparameter tuning model from multiple candidate hyperparameter tuning models. In some embodiments, the model selection module can select the target hyperparameter tuning model from multiple candidate hyperparameter tuning models based on the motion patterns and intensity of the motion in the image, in order to improve the data adjustment effect.

[0090] For example, the software architecture of the aforementioned electronic device 200 may further include a kernel layer. This kernel layer is the layer between the hardware and the software. In this embodiment, the kernel layer may include a display driver, a camera driver, an audio driver, a sensor driver, etc.

[0091] Understandable Figure 3 The layers in the illustrated structure and the components contained in each layer do not constitute a specific limitation on the electronic device 200, i.e., the mobile phone. In other embodiments of this application, the structure may include more or fewer layers than illustrated, and each layer may include more or fewer components; this application does not impose any limitations.

[0092] The image processing method of this application can be used in shooting scenarios of electronic devices. For example, in scenarios where the electronic device can record video using its front or rear camera. Another example is in preview scenarios where the electronic device displays a preview interface. The following embodiments use a mobile phone as an example to describe the method of this application.

[0093] This application provides an image processing method. This method can be applied to video recording scenarios. In response to a user's shooting operation, the mobile phone can acquire initial gyroscope data. Then, the mobile phone can perform motion detection based on this initial gyroscope data to determine whether the phone is in motion. If the phone is in motion, it can use a target parameter tuning model to adjust the initial gyroscope data to obtain target gyroscope data, thereby eliminating image jitter and improving the image stabilization effect of the video image. For example, as shown... Figure 4 As shown, the image processing method may include S401 to S407.

[0094] S401, in response to the user's shooting operation, the mobile phone acquires initial gyroscope data. This initial gyroscope data is the gyroscope data acquired when the mobile phone captures the first image frame.

[0095] It's understandable that a phone can activate its camera function upon detecting a user's shooting action. For example, the phone can launch its camera app or a third-party app with photo or video recording capabilities (such as a photography app), thereby activating the phone's camera function.

[0096] In some embodiments, the user's shooting operation described above can be an operation where the user clicks the record control in video recording mode. For example, please refer to [link to relevant documentation]. Figure 5 The phone is displaying Figure 5 The main interface is shown in (a). This main interface includes the camera application icon. Specifically, in response to the user's click on the camera application icon, the phone can activate the camera application's shooting function and display, as shown in (a). Figure 5 The camera interface shown in (b) is the preview interface in camera mode. Afterwards, if the user clicks the video recording control in the camera interface, the phone can display... Figure 5 The recording interface shown in (c) is a preview interface in recording mode. Afterwards, if the user clicks the "Record" control 510 in the recording interface, the phone can generate the corresponding video.

[0097] For example, when a mobile phone displays something like... Figure 5 In the case of the main interface shown in (a), if a user's voice command to record is detected, the phone can directly start the shooting function and enter the recording mode, that is, directly display as shown in (a). Figure 5 The video recording interface shown in (c) is shown in the middle.

[0098] It should be noted that the mobile phone can also enter the recording mode in response to other user touch operations, voice commands or quick gestures, and the embodiments of this application do not limit the operation that triggers the mobile phone to enter the recording mode.

[0099] In one implementation, after the phone activates its camera function, it can capture a first image frame in preview mode, acquiring initial gyroscope data during the capture. The first image frame is the image frame currently captured by the phone's camera. This gyroscope data is used to detect the camera's pose (or initial pose) when capturing the first image frame. Specifically, the phone can determine the initial pose of the first image frame based on gyroscope data (x, y, z) measured along three axes (X, Y, and Z). Then, the phone can perform image stabilization on the first image frame based on this initial gyroscope data. In other words, the phone can perform image stabilization in the preview interface, ensuring that the video stream displayed in the preview is stabilized, thereby improving the user's visual experience.

[0100] For example, such as Figure 5 As shown, in the display Figure 5 The camera interface shown in (b) or the display Figure 5 In the recording interface shown in (c), the mobile phone can directly perform image stabilization, so that the video stream displayed on the interface is the video stream after image stabilization, thereby improving the user's visual experience.

[0101] Specifically, after receiving a user's shooting command, the mobile phone can capture a video stream through its camera. This camera can be a front-facing camera or a rear-facing camera; there is no specific limitation. For example, the number of cameras can be set according to actual needs. For instance, there can be one, two, etc., without further limitation.

[0102] It is understandable that since the above video stream is composed of multiple consecutive image frames, the mobile phone can perform image processing on each image frame during the process of capturing the video stream, so that the processed image frames are image frames after image stabilization. In other words, the video stream composed of the processed image frames is a shake-free video stream, thereby improving the user's shooting experience.

[0103] It's important to note that before acquiring the initial gyroscope data, the phone can parse the video file. This video file is in Extensible Markup Language (XML) format. This means the phone can parse the video file to obtain relevant camera parameters, such as lens distortion parameters, shutter speed, and ISO value. These camera parameters provide a foundation for subsequent video stabilization. In other words, parsing the video file lays the groundwork for subsequent image stabilization operations.

[0104] S402: The phone performs motion detection based on the initial gyroscope data and obtains the first detection result.

[0105] The first detection result mentioned above is used to indicate whether the mobile phone is in motion.

[0106] In some embodiments, after acquiring the initial gyroscope data, the mobile phone can detect whether it is in motion based on this data. If the phone is in motion, it indicates a possibility of shaking, and the phone can execute S404 to further select an appropriate anti-shake strategy, enabling targeted anti-shake operations to improve the accuracy of video image stabilization. If the phone is not in motion, i.e., is stationary, it indicates no possibility of shaking, and the phone can execute S405, i.e., not perform anti-shake operations. This reduces unnecessary power consumption and improves the utilization of computing resources.

[0107] In one implementation, the mobile phone can obtain a first detection result regarding whether it is in motion based on the initial gyroscope data and statistical methods. These statistical methods can include at least one of the mean, standard deviation, and variance. This allows for accurate detection of motion, providing a basis for subsequent determination of whether to perform image stabilization.

[0108] Specifically, such as Figure 6 As shown, the mobile phone can obtain the average value of the initial gyroscope data based on the initial gyroscope data. This is understandable because the frame rate at which the phone captures the first image frame is different from the sampling frequency when collecting gyroscope data. For example, the phone might only capture 30 first image frames per second, while it might capture 100 gyroscope data points per second. In other words, the initial gyroscope data can include multiple gyroscope data points. Therefore, after obtaining multiple gyroscope data points, the phone can perform calculations on these multiple gyroscope data points to obtain the average value of the initial gyroscope data.

[0109] In one implementation, the average value of the initial gyroscope data can be calculated using the following expression:

[0110]

[0111] in, X is the average value of the initial gyroscope data; n is the total number of gyroscope data; i Let (x, y, z) be the data from the i-th gyroscope.

[0112] Specifically, such as Figure 6As shown, after obtaining the average value of the initial gyroscope data, the mobile phone can obtain a target value for the initial gyroscope data based on this average value. This target value may include the standard deviation and / or variance. That is, the mobile phone can calculate the standard deviation of the average value of the initial gyroscope data to obtain the standard deviation value; and / or, the mobile phone can calculate the variance of the average value of the initial gyroscope data to obtain the variance value.

[0113] In one implementation, the above standard deviation can be calculated using the following expression:

[0114]

[0115] Where σ is the standard deviation of the initial gyroscope data.

[0116] In another implementation, the above variance value can be calculated using the following expression three:

[0117]

[0118] Where, σ 2 This represents the variance of the initial gyroscope data.

[0119] Specifically, such as Figure 6 As shown, after obtaining the target value of the initial gyroscope data, the mobile phone can determine whether the target value is less than a preset value. This preset value can be pre-set according to actual conditions. It can be understood that if the target value is the standard deviation, the preset value can be a preset standard deviation. For example, the preset standard deviation can be 0.1, 0.15, etc., without specific limitations. If the target value is the variance, the preset value can be a preset variance. For example, the preset variance can be 0.01, 0.02, etc., without specific limitations.

[0120] In some embodiments, if the target value is less than a preset value, the phone can obtain a first detection result that the phone is in a stationary state. It can be understood that if the target value is less than the preset value, it indicates that the phone is relatively stable, meaning that the phone is not shaking or vibrating. Therefore, the phone can obtain a first detection result that the phone is in a stationary state, so that the phone does not directly perform image stabilization, thereby reducing unnecessary power consumption and increasing the phone's usage time.

[0121] In other embodiments, if the target value is not less than a preset value, that is, if the target value is greater than or equal to the preset value, the mobile phone can obtain a first detection result that the mobile phone is in motion. It can be understood that if the target value is greater than or equal to the preset value, it indicates that the mobile phone may be shaking or vibrating. Therefore, the mobile phone can obtain a first detection result that the mobile phone is in motion, so that the mobile phone can continue to detect the motion and match the corresponding target parameter tuning model, providing a basis for subsequent adjustment of the camera pose.

[0122] In one example, when the target value includes a standard deviation, if the standard deviation is less than a preset standard deviation, the phone can detect that the phone is stationary. If the standard deviation is greater than or equal to the preset standard deviation, the phone can detect that the phone is in motion. In another example, when the target value includes a variance, if the variance is less than a preset variance, the phone can detect that the phone is stationary. If the variance is greater than or equal to the preset variance, the phone can detect that the phone is in motion. In yet another example, when the target value includes both a standard deviation and a variance, if both the standard deviation and variance are less than a preset standard deviation, the phone can detect that the phone is stationary. If the standard deviation is greater than or equal to the preset standard deviation, or the variance is greater than or equal to the preset variance, the phone can detect that the phone is in motion.

[0123] S403: If the first detection result indicates that the phone is stationary, the phone will not perform image stabilization.

[0124] Specifically, after obtaining the first detection result, if the result indicates that the phone is stationary, it means there is no possibility of the phone shaking, implying it is likely placed in a fixed device. Therefore, to reduce unnecessary power consumption, the phone can forgo image stabilization and directly output a video stream composed of multiple first image frames. This extends the phone's battery life and improves the user experience. Furthermore, since image stabilization is not performed, subsequent detection and pose adjustment operations are unnecessary, thus improving resource utilization and reducing resource waste.

[0125] S404, when the first detection result indicates that the mobile phone is in motion, the mobile phone acquires an optical flow image.

[0126] Specifically, after obtaining the first detection result, if the first detection result indicates that the phone is in motion, it means that the phone is likely shaking, that is, the phone may be in the user's hand. Therefore, in order to achieve accurate image stabilization of the video stream, the phone can acquire an optical flow image to further detect the phone's motion, thereby enabling targeted image stabilization operations and improving the accuracy of video image stabilization. This optical flow image is obtained based on adjacent image frames. These adjacent image frames include the first image frame and the second image frame, where the second image frame is the previous image frame captured before the first image frame. In other words, the second image frame and the first image frame are captured adjacent to each other.

[0127] It should be noted that the aforementioned optical flow image is used to characterize the motion information of the same object in adjacent image frames. This motion information can include the object's velocity and direction of motion. In other words, the optical flow image is used to characterize the displacement of corresponding pixels between adjacent image frames. Specifically, the optical flow image carries the motion velocity of the corresponding pixels (or object pixels) of the same object in adjacent image frames. This motion velocity can include the object pixel's velocity along the x-axis and its velocity along the y-axis.

[0128] In one implementation, the aforementioned optical flow image can be obtained by detecting adjacent image frames using an optical flow method. This optical flow method infers the object's speed and direction of movement by detecting changes in the intensity of image pixels over time. In other words, this optical flow method utilizes the temporal changes in pixels of each image frame in the video stream and the correlation between adjacent image frames to determine the pixel correspondence between the previous image frame (or second image frame) and the current image frame (or first image frame), thereby calculating the motion information of the object between adjacent image frames.

[0129] S405: The mobile phone determines the target parameter tuning model from multiple candidate parameter tuning models based on the optical flow image.

[0130] Specifically, after acquiring the aforementioned optical flow image, the mobile phone can select the candidate parameter tuning model that matches the motion speed of the object pixels carried in the optical flow image from multiple candidate parameter tuning models as the target parameter tuning model. In this way, the mobile phone can select the corresponding parameter tuning model according to the motion of objects in adjacent image frames, that is, it can selectively choose the optimal parameter tuning model for parameter adjustment, thereby improving the parameter adjustment rate and providing convenient conditions for subsequent rapid image stabilization.

[0131] In some embodiments, such as Figure 7 As shown, the process of determining the target parameter tuning model for the mobile phone can specifically include S4051 to S4056:

[0132] S4051, the mobile phone performs regularity detection on the first image frame based on the optical flow image to obtain the second detection result.

[0133] The second detection result is used to indicate whether the object pixels in the first image frame satisfy the motion law.

[0134] Specifically, after acquiring the optical flow image, the mobile phone can detect whether the object pixels in the first image frame satisfy a preset motion rule based on the optical flow image. This preset motion rule is a linear rule, meaning that the object pixels in the first image frame and their corresponding object pixels in the second image frame follow a linear function. For example, the preset motion rule could be that the object pixels in the first image frame tend to move horizontally compared to the object pixels in the second image frame, or that the object pixels in the first image frame tend to move vertically compared to the object pixels in the second image frame, or that the object pixels in the first image frame tend to move diagonally compared to the object pixels in the second image frame.

[0135] It is understandable that if the object pixels in the first image frame satisfy the preset motion pattern, it indicates that the shaking generated during mobile phone shooting has a certain regularity. Therefore, the mobile phone can execute S4053 to further determine whether the object pixels in the first image frame are moving violently, providing a basis for the subsequent accurate determination of the target parameter tuning model. If the object pixels in the first image frame do not satisfy the preset motion pattern, it indicates that the shaking generated during mobile phone shooting is not regular. Therefore, the mobile phone can directly execute S4052 to improve the calling efficiency of the target parameter tuning model, thereby improving the image stabilization efficiency of the video image.

[0136] In one implementation, the mobile phone can acquire a historical optical flow image. This historical optical flow image is an optical flow image acquired before the previously captured optical flow image during video recording. This historical optical flow image carries the historical motion velocity of the object pixels. The mobile phone can then determine whether the velocity difference between the historical motion velocity carried in the historical optical flow image and the motion velocity carried in the optical flow image is less than a preset difference. If the velocity difference is less than the preset difference, it indicates that the motion of the object pixels in the first image frame has a certain regularity. Therefore, the mobile phone can obtain a second detection result indicating that the object pixels in the first image frame satisfy the preset motion regularity. If the velocity difference is greater than or equal to the preset difference, it indicates that the motion of the object pixels in the first image frame is irregular. Therefore, the mobile phone can obtain a second detection result indicating that the object pixels in the first image frame do not satisfy the preset motion regularity.

[0137] In this embodiment, the motion velocity carried by the historical optical flow image is compared with the motion velocity carried by the optical flow image to determine whether the object pixels in the first image frame meet the preset motion law. This enables accurate detection of the motion law of objects in the image, improves the accuracy of the second detection result, and provides a foundation for subsequent accurate target determination and parameter tuning models.

[0138] S4052, if the second detection result indicates that the object pixels in the first image frame do not meet the preset motion law, the mobile phone will use the first parameter tuning model among multiple candidate parameter tuning models as the target parameter tuning model.

[0139] Specifically, after obtaining the second detection result, if the second detection result indicates that the object pixels in the first image frame do not meet the preset motion rules, it means that the motion of the object pixels in the first image frame is not regular, that is, the image jitter generated when the mobile phone captures the first image frame is not regular. Therefore, the mobile phone can use the common parameter tuning model (i.e., the first parameter tuning model) among multiple candidate parameter tuning models as the target parameter tuning model. The model parameters in this first parameter tuning model are default parameters. It should be noted that the convergence speed, convergence accuracy, and convergence stability of this first parameter tuning model are all in a stable state.

[0140] Among these, several candidate parameter tuning models are operator splitting quadratic program (OSQP) models. The OSQP model is a solver for convex quadratic programming problems, which decomposes the quadratic programming problem into a series of smaller subproblems and solves them iteratively to obtain the optimal solution. In this embodiment, the OSQP model is used to adjust the initial gyroscope data so that the adjusted gyroscope data (or target gyroscope data) reduces jitter in the first image frame. In other words, the target image frame obtained through this adjusted gyroscope data is the first image frame after image stabilization. The OSQP model includes the alternating direction method of multipliers (ADMM) algorithm. The ADMM algorithm is used to solve distributed convex optimization problems.

[0141] It is understood that the aforementioned candidate hyperparameter tuning models are OSQP models with different model parameters, meaning that the numerical values ​​of the model parameters for each candidate model are different. The model parameters of this OSQP model can include the step size ρ and the relaxation amount α. The step size is the step size in the iterative solution process of the OSQP model, used to adjust the magnitude of variable (i.e., the gyroscope data input to the model) updates during each model iteration in the ADMM algorithm. The relaxation amount is the relaxation variable in the iterative solution process of the OSQP model, used to adjust the updates of the original variable m and the dual variable n in the ADMM algorithm. For example, the step size ρ can be set to 25, and the relaxation amount α can be set to 1.6.

[0142] It should be noted that because the OSQP model parameters include the aforementioned relaxation parameters, the OSQP model appropriately relaxes the variable updates (i.e., the gyroscope data input to the model) in each iteration, thereby balancing convergence speed and stability. The step size mentioned above can affect the speed of Lagrange multiplier updates, thus affecting the overall convergence speed and stability of the algorithm. Specifically, the smaller the step size, the slower the convergence speed of the OSQP model; the larger the step size, the faster the convergence speed. However, if the step size is extremely large, the model will converge very quickly, leading to poor stability, i.e., the model is more oscillating. The impact of different step sizes on the iteration time of the OSQP model will be explained below with reference to Table 1.

[0143] Table 1

[0144] Step length 25 30 20 15 10 5 Iteration time (ms) 22.70 22.12 23.87 24.03 20.82 23.57

[0145] As shown in Table 1, the iteration time of the OSQP model is 22.70 ms when the step size is 25. The iteration time is 22.12 ms when the step size is 30. The iteration time is 23.87 ms when the step size is 20. The iteration time is 24.03 ms when the step size is 15. The iteration time is 20.82 ms when the step size is 10. The iteration time is 23.57 ms when the step size is 5. Here, step size 25 is the initial step size of the OSQP model. However, as can be seen from Table 1, the iteration time of the OSQP model is shortest when the step size is 10. In other words, the iteration time of the OSQP model with a step size of 25 is slower than that with a step size of 10. Therefore, setting the step size of the OSQP model to 10 can improve the convergence speed of the OSQP model, thereby improving the parameter tuning efficiency.

[0146] Optionally, during the iterative solution process, the OSQP model can calculate the error at preset intervals to adjust model parameters in a timely manner and improve the model's tuning efficiency. The preset number of iterations can be set according to actual needs. For example, the preset number of iterations can be 25. It's understandable that too many preset numbers may lead to excessive iterations, thus affecting the OSQP model's tuning efficiency. Too few preset numbers may cause the model to frequently calculate errors, increasing performance overhead and resulting in unnecessary resource waste.

[0147] S4053, when the second detection result indicates that the object pixels in the first image frame meet the preset motion law, the mobile phone enables the hot start function of the target parameter tuning model.

[0148] Specifically, after obtaining the second detection result, if the second detection result indicates that the object pixels in the first image frame satisfy a motion law, it means that the motion of the object pixels in the first image frame has a certain regularity. In other words, the image jitter generated when the mobile phone captures the first image frame may have a certain regularity. Therefore, the mobile phone can trigger the warm start function to be enabled. This warm start function is a feature of the aforementioned OSQP model, used to enable the OSQP model to quickly output target gyroscope data, that is, to improve the data output speed, thereby improving the model's convergence speed.

[0149] It should be noted that the settings parameters for the aforementioned warm-start function can include the initial primitive variable *m* and the initial dual variable *n*. The initial primitive variable *m* represents the initial solution to the optimization problem; that is, it is the standard gyroscope data predicted based on the optical flow image. This standard gyroscope data serves as a benchmark for adjusting the initial gyroscope data, determining where the model can quickly obtain the target gyroscope data. In other words, by using this standard gyroscope data to adjust the initial gyroscope data, the phone can obtain the target gyroscope data more quickly, reducing the number of iterations and accelerating the model's convergence speed.

[0150] Here, the initial dual variable *n* is a variable related to the constraints. That is, in the ADMM algorithm, the initial dual variable *n* is used to update the initial original variable *m* so that the updated original variable satisfies the constraints. In other words, the dual variable *n* helps the model find the optimal solution that satisfies the constraints more quickly, thereby improving the model's convergence speed.

[0151] In one implementation, when the hot-start function of the aforementioned OSQP model is enabled, the phone can save the settings parameters for this function. This allows the phone to directly access these settings to adjust the gyroscope data when the hot-start function is enabled later. This saves time in determining the target gyroscope data and improves the efficiency of gyroscope data adjustment.

[0152] In some embodiments, the mobile phone may not enable the hot start function of the parameter tuning model. That is, after determining that the object pixels in the first image frame meet the motion law, the mobile phone can directly execute S4054 to further determine whether the object in the first image frame moves violently, which provides a basis for the subsequent accurate determination of the target parameter tuning model.

[0153] S4054, the mobile phone performs displacement detection on the first image frame based on the optical flow image to obtain the third detection result.

[0154] The third detection result mentioned above is used to indicate whether the object in the first image frame is moving violently.

[0155] Specifically, after the warm-up function of the parameter tuning model is enabled, the mobile phone can detect whether the object in the first image frame is moving violently based on the aforementioned optical flow image. Vigorous movement refers to the object pixels moving relatively quickly, that is, the object pixels moving to a relatively far position.

[0156] It can be understood that when the displacement of the object pixel exceeds a preset displacement, meaning the object pixel's movement speed exceeds a preset speed, the phone can obtain a third detection result indicating that the object in the first image frame is moving rapidly. This displacement is the distance between the object pixel's position in the second image frame and its position in the first image frame. This displacement can be calculated based on the object pixel's movement speed. The preset displacement and preset speed can be set according to actual conditions.

[0157] In some embodiments, if the object in the first image frame is moving violently, it indicates that the phone may be shaking significantly when capturing the first image frame, meaning the phone may be shaking violently. Therefore, the phone can execute S4055 to select a target parameter tuning model from multiple candidate parameter tuning models that is suitable for the violently moving scene, thereby improving the fit of the parameter tuning model. If the object in the first image frame is moving slowly, it indicates that the phone may be shaking less violently when capturing the first image frame, meaning the phone may be shaking relatively smoothly. Therefore, the phone can execute S4056 to select a target parameter tuning model from multiple candidate parameter tuning models that is suitable for the smoothly moving scene, thereby improving the fit of the parameter tuning model.

[0158] In this embodiment, the mobile phone can select a corresponding candidate parameter tuning model from multiple candidate models as the target parameter tuning model based on whether the object in the first image frame is moving violently, in order to adjust the initial gyroscope data. This ensures that the target parameter tuning model is adapted to the shooting scene, enabling targeted data adjustments. This not only improves the accuracy of data adjustment but also increases the speed of data adjustment, thereby improving the stabilization efficiency of the video stream and enhancing the image stabilization effect of the video image.

[0159] In one implementation, the mobile phone can determine whether the movement speed of the object pixel is greater than a preset speed. If the movement speed of the object pixel is greater than the preset speed, it indicates that the movement of the object pixel in the first image frame is relatively rapid. Therefore, the mobile phone can obtain a third detection result indicating that the object in the first image frame is moving rapidly. If the movement speed of the object pixel is less than or equal to the preset speed, it indicates that the movement of the object pixel in the first image frame is relatively stable. Therefore, the mobile phone can obtain a third detection result indicating that the object in the first image frame is moving slowly.

[0160] In this embodiment, the motion of an object in the first image frame is determined by judging whether the motion speed of the object's pixels is greater than a preset speed. This allows for accurate detection of the intensity of object motion in the image, improving the accuracy of the third detection result and providing a foundation for subsequent precise target parameter tuning.

[0161] It should be noted that the activation process of the aforementioned warm start function and the displacement detection process of the first image frame can be executed simultaneously or sequentially, without any specific limitation. For example, the phone can execute steps S4053 and S4054 simultaneously. Alternatively, the phone can execute step S4054 first, and then execute step S4053.

[0162] S4055, when the third detection result indicates that the object in the first image frame is moving violently, the mobile phone selects the second parameter tuning model from multiple candidate parameter tuning models as the target parameter tuning model.

[0163] Specifically, after obtaining the third detection result, if the third detection result indicates that the object in the first image frame is moving violently, it means that the phone may have been shaking significantly when capturing the first image frame, i.e., the phone may have been shaking violently. Therefore, the phone can call the second parameter tuning model from multiple candidate parameter tuning models and use the second parameter tuning model as the target parameter tuning model. The model parameters in the second parameter tuning model can improve the convergence speed; that is, inputting the initial gyroscope data into the second parameter tuning model can reduce the number of model iterations and increase the adjustment rate of the gyroscope data.

[0164] It is understandable that the second parameter tuning model converges faster than the first parameter tuning model. However, the second parameter tuning model has lower convergence accuracy and poorer convergence stability compared to the first parameter tuning model.

[0165] S4056, if the third detection result indicates that the object in the first image frame is moving slowly, the mobile phone selects the third parameter tuning model from multiple candidate parameter tuning models as the target parameter tuning model.

[0166] Specifically, after obtaining the third detection result, if the third detection result indicates that the object in the first image frame moves slowly, it means that the degree of shaking when the phone captured the first image frame may be small, that is, the phone shaking may be relatively gentle. Therefore, the phone can call the third parameter tuning model from multiple candidate parameter tuning models and use the third parameter tuning model as the target parameter tuning model. The model parameters in this third parameter tuning model can improve convergence accuracy and convergence stability. In other words, inputting gyroscope data into this third parameter tuning model can improve the accuracy of data adjustment, thereby improving the image stabilization accuracy of the video image.

[0167] It is understandable that the third parameter tuning model has higher convergence accuracy and better convergence stability compared to the first parameter tuning model. However, the third parameter tuning model has a slower convergence speed compared to the first parameter tuning model.

[0168] S406: The mobile phone uses a target parameter tuning model to adjust the initial gyroscope data to obtain the target gyroscope data.

[0169] Specifically, after obtaining the target parameter tuning model, the mobile phone can directly input the initial gyroscope data into the target parameter tuning model to adjust the initial gyroscope data. This adjusted gyroscope data (i.e., the target gyroscope data) is the gyroscope data that reduces jitter in the first image frame. In other words, the target image frame obtained through this adjusted gyroscope data is the first image frame after image stabilization. This reduces image jitter caused by user actions or electronic device movement, improves the image stabilization effect of video images, and thus enhances the user's shooting experience.

[0170] It is understandable that, since the above target parameter tuning model is selected based on the motion velocity of object pixels carried by the optical flow image, and this motion velocity of object pixels is used to characterize the motion of the object in the first image frame, that is, the parameter tuning model corresponding to different motion situations is different. For example, please refer to... Figure 8The phone can input initial gyroscope data into a first parameter tuning model to obtain target gyroscope data. This first parameter tuning model is selected when the object's motion is erratic, its warm-up is not enabled, and its model parameters are at default values. Alternatively, the phone can input initial gyroscope data into a second parameter tuning model to obtain target gyroscope data. This second parameter tuning model is selected when the object's motion is regular and vigorous, its warm-up is enabled, and its model parameters are values ​​that improve the model's convergence speed. Alternatively, the phone can input initial gyroscope data into a third parameter tuning model to obtain target gyroscope data. This third parameter tuning model is selected when the object's motion is regular and slow, its warm-up is enabled, and its model parameters are values ​​that improve the model's convergence accuracy.

[0171] In some embodiments, the mobile phone can input multiple gyroscope data points into the target parameter tuning model to adjust each gyroscope data point individually, thereby obtaining multiple target gyroscope data points. This provides a foundation for subsequent precise image stabilization. Alternatively, the mobile phone can input the average value corresponding to the multiple gyroscope data points into the target parameter tuning model to adjust the average value, thereby obtaining the target average value of the first image frame, and using this target average value as the target gyroscope data. This improves image stabilization efficiency, thereby enabling faster video stream output.

[0172] In other embodiments, to improve the stabilization efficiency of the video stream, that is, to enable the phone to quickly output the stabilized video stream, the phone can store the initial gyroscope data until a preset number of first image frames are captured. The phone can then input multiple gyroscope data points acquired during the capture of this preset number of first image frames into the target parameter tuning model to obtain multiple target gyroscope data points. In other words, the phone can obtain target gyroscope data corresponding to multiple first image frames at once. This reduces the video stream processing time and improves the stabilization efficiency.

[0173] It should be noted that the target parameter tuning model mentioned above can be any one of multiple candidate parameter tuning models. That is, the target parameter tuning model can be the first parameter tuning model, the second parameter tuning model, or the third parameter tuning model. In contrast, in related technologies, the mobile phone only calls a general parameter tuning model to adjust the initial gyroscope data. In other words, compared with related technologies, this implementation selects the appropriate parameter tuning model based on the motion of the object in the first image frame. This means that the optimal parameter tuning model can be selected for targeted parameter adjustment, thereby improving the parameter adjustment rate and providing convenient conditions for subsequent rapid image stabilization.

[0174] Specifically, such as Figure 9 As shown, in response to the user's shooting operation, the mobile phone parses the video file. The video file is in XML format. Then, the mobile phone can obtain initial gyroscope data and camera parameters. These camera parameters may include at least one of lens distortion correction (LDC), field of view (FOV), and spatial alignment transform (SAT). Next, the mobile phone can perform image processing on the first image frame to obtain a standard image frame. This image processing is called image warping, which refers to correcting distorted image frames so that the processed image frame (i.e., the standard image frame) can effectively reduce parallax effects and improve image quality. Finally, the mobile phone can input the initial gyroscope data, camera parameters, and the standard image frame into a parameter tuning model to obtain target gyroscope data.

[0175] S407, the mobile phone performs image processing on the first image frame based on the target gyroscope data to obtain the target image frame.

[0176] Specifically, after obtaining the target gyroscope data, the mobile phone can perform image processing on the first image frame based on the target gyroscope data to obtain the target image frame. This target image frame is an image frame that has undergone image stabilization processing. This reduces image shake caused by user actions or electronic device movement, improves the stabilization effect of video images, and thus enhances the user's shooting experience.

[0177] In one implementation, the mobile phone can obtain the rotation matrix of the first image frame based on the target gyroscope data. Then, the mobile phone can adjust the image matrix of the first image frame based on the rotation matrix to obtain the target image matrix. This target image matrix represents the target image frame. Finally, the mobile phone can generate the target image frame based on this target image matrix.

[0178] It can be understood that image frames can be represented by matrices. That is, the image matrix of the first image frame can characterize the first image frame. Therefore, the mobile phone can perform image stabilization on the first image frame by adjusting its image matrix to obtain the stabilized first image frame, which is the target image frame. In this way, precise image stabilization can be achieved, improving the image stabilization effect of video images and thus enhancing the user's shooting experience.

[0179] Specifically, after obtaining the target image frames, the mobile phone can assemble multiple target image frames into a video stream. In other words, this video stream is a video stream that has undergone image stabilization. This enables precise image stabilization of the video stream, improving its image stabilization effect and thus enhancing the user's shooting experience.

[0180] In some embodiments, in response to a user's action to stop recording, the mobile phone can assemble multiple target image frames obtained during video recording into a video stream and store the video stream on the phone for later viewing by the user, thus improving the user experience. The user stopping recording can be done by clicking the stop control in recording mode.

[0181] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform the various functions or steps described in the method embodiments.

[0182] This application also provides a computer program product, including a computer program that, when run on an electronic device, causes the electronic device to perform the various functions or steps described in the above method embodiments.

[0183] This application provides a chip for executing instructions. When the chip is running, it executes the technical solutions described in the above embodiments. Its implementation principle and technical effects are similar and will not be repeated here.

[0184] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0185] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0186] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0187] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0188] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0189] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image processing method, characterized in that, Applied to electronic devices, the method includes: In response to the user's first operation, the electronic device acquires initial gyroscope data when capturing the first image frame; wherein, the first operation is used to trigger the electronic device to display a preview interface corresponding to the camera application; The electronic device detects whether it is in motion based on the initial gyroscope data; When the electronic device is in motion, the electronic device performs image stabilization processing on the first image frame based on the initial gyroscope data to obtain the target image frame; The electronic device displays the target image frame in the preview interface.

2. The method according to claim 1, characterized in that, The electronic device performs image stabilization processing on the first image frame based on the initial gyroscope data to obtain the target image frame, including: The electronic device uses a target parameter tuning model to adjust the initial gyroscope data to obtain target gyroscope data; wherein, the target parameter tuning model is one of multiple candidate parameter tuning models, and the model parameters corresponding to different candidate parameter tuning models are different; The electronic device obtains a rotation matrix based on the target gyroscope data; The electronic device adjusts the image matrix of the first image frame according to the rotation matrix to obtain the target image matrix; wherein the target image matrix is ​​used to characterize the target image frame.

3. The method according to claim 2, characterized in that, The method further includes: The electronic device acquires an optical flow image; wherein the optical flow image is used to characterize the motion information of the same object in adjacent image frames, the adjacent image frames include the first image frame and the second image frame, and the second image frame is acquired before the first image frame; The electronic device determines the target parameter tuning model from the plurality of candidate parameter tuning models based on the optical flow image.

4. The method according to claim 3, characterized in that, The electronic device determines the target hyperparameter tuning model from the plurality of candidate hyperparameter tuning models based on the optical flow image, including: The electronic device detects whether the object pixels in the first image frame satisfy a preset motion law based on the optical flow image; If the object pixels in the first image frame do not meet the preset motion rules, the electronic device will use the first parameter tuning model among the multiple candidate parameter tuning models as the target parameter tuning model.

5. The method according to claim 4, characterized in that, The optical flow image carries the motion velocity of the object pixels. The electronic device detects whether the object pixels in the first image frame satisfy a preset motion rule based on the optical flow image, including: The electronic device acquires historical optical flow images; wherein, the historical optical flow images are optical flow images acquired before the optical flow images during video recording, and the historical optical flow images carry the historical motion speed of the object pixels; If the speed difference between the current motion speed and the historical motion speed is less than a preset difference, the electronic device determines that the object pixels in the first image frame satisfy a preset motion rule. If the speed difference between the current motion speed and the historical motion speed is greater than or equal to the preset difference, the electronic device determines that the object pixels in the first image frame do not satisfy the preset motion pattern.

6. The method according to claim 4 or 5, characterized in that, The method further includes: If the object pixels in the first image frame meet the preset motion rules, the electronic device detects whether the object in the first image frame is moving violently based on the optical flow image. When the object in the first image frame is moving violently, the electronic device uses the second parameter tuning model among the multiple candidate parameter tuning models as the target parameter tuning model; wherein, the second parameter tuning model can improve the convergence speed compared to the first parameter tuning model; When the object in the first image frame moves slowly, the electronic device uses the third parameter tuning model among the multiple candidate parameter tuning models as the target parameter tuning model; wherein, the third parameter tuning model can improve convergence accuracy and convergence stability compared with the first parameter tuning model.

7. The method according to claim 6, characterized in that, The electronic device detects whether the object in the first image frame is moving violently based on the optical flow image, including: If the movement speed of the object pixel is greater than a preset speed, the electronic device determines that the object in the first image frame is moving violently; or, If the motion speed of the object pixel is less than or equal to the preset speed, the electronic device determines that the object in the first image frame is moving slowly.

8. The method according to any one of claims 2-7, characterized in that, Before the electronic device adjusts the initial gyroscope data using a target parameter tuning model to obtain the target gyroscope data, the method further includes: The electronic device triggers the activation of the hot start function of the target parameter tuning model; wherein, the hot start function is used to enable the target parameter tuning model to increase the data output speed.

9. The method according to any one of claims 1-8, characterized in that, The initial gyroscope data includes multiple gyroscope data points. The electronic device uses this initial gyroscope data to detect whether it is in motion, including: The electronic device calculates the average value of the initial gyroscope data based on the multiple gyroscope data. The electronic device calculates a target value for the initial gyroscope data based on the average value; wherein the target value includes the standard deviation and / or variance. If the target value is less than a preset value, the electronic device determines that it is in a stationary state. If the target value is greater than or equal to a preset value, the electronic device determines that it is in motion.

10. The method according to any one of claims 1-9, characterized in that, The user's first operation includes any one of the following: the user opening the camera application, the user controlling the electronic device to enter recording mode when the camera application is open, or the user clicking the recording control when the camera application is open.

11. The method according to claim 10, characterized in that, The method further includes: In response to the user's stop recording operation, the electronic device generates a video stream based on multiple target image frames acquired during the video recording process; The electronic device stores the video stream.

12. An electronic device, characterized in that, include: A camera, a display screen, a gyroscope sensor, one or more processors, and one or more memories; The one or more processors are coupled to the camera, the display screen, the gyroscope sensor, and the one or more memories; The camera is used to capture a first image frame, the display screen is used to display a target image frame, the gyroscope sensor is used to acquire initial gyroscope data, and the one or more memories are used to store computer program code, the computer program code including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1-11.

13. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 11.

14. A computer program product, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 11.