Motion vector detection method, camera module and electronic device

By introducing dynamic visual sensors and motion vector detection methods, the drift problem of gyroscope sensors was solved, improving the imaging quality and optical image stabilization performance of camera modules and electronic devices.

WO2026157273A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-09-17
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Gyroscope sensors suffer from long-term drift issues in optical image stabilization, resulting in blurred images and poor image quality.

Method used

By introducing a dynamic vision sensor combined with a motion vector detection method, motion vector detection accuracy is improved by detecting the edge vectors and gradients of events, clustering and filtering motion vectors.

Benefits of technology

It improves optical image stabilization performance and image quality, enhancing the imaging effect of camera modules and electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025121840_30072026_PF_FP_ABST
    Figure CN2025121840_30072026_PF_FP_ABST
Patent Text Reader

Abstract

A motion vector detection method, a camera module and an electronic device. The method comprises: obtaining a first image at a first time; determining a first edge vector of each pixel in the first image, the first edge vector of each pixel being a vector from the pixel to an edge in the first image; obtaining an event generated within a first target time period, the first target time period being after the first time; and, on the basis of the first edge vectors of pixels corresponding to the event, determining a motion vector of the first target time period, the motion vector of the first target time period being used for optical image stabilization of the camera module. The embodiments of the present application can improve the optical image stabilization performance, thereby improving the imaging quality of electronic devices.
Need to check novelty before this filing date? Find Prior Art

Description

Motion vector detection methods, camera modules and electronic devices

[0001] This invention claims priority to Chinese Patent Application No. 202510127430.8, filed on January 27, 2025, entitled "Method for Motion Vector Detection, Camera Module and Electronic Device", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of optical image stabilization technology, and in particular to a motion vector detection method, a camera module, and an electronic device. Background Technology

[0003] Optical image stabilization (OIS) uses physical methods to counteract image shake caused by hand tremors or other external factors during shooting. In OIS technology, a gyroscope sensor typically detects the amount of shaking in the electronic device, and a servo system quickly adjusts the position of the lens (or lens group) or sensor based on this shaking to compensate for the motion, thereby keeping the image of the subject in a constant position on the image plane. However, gyroscope sensor detection results suffer from long-term drift and other problems, making it impossible to provide accurate compensation to the servo system during image capture. This results in poor optical image stabilization performance, blurred images from electronic devices, and poor image quality. Summary of the Invention

[0004] This application provides a motion vector detection method, a camera module, and an electronic device, which can improve optical image stabilization performance and thus improve the imaging quality of the electronic device.

[0005] Firstly, this application provides a motion vector detection method, comprising a camera module, the method including: acquiring a first image at a first time; determining a first edge vector for each pixel in the first image; the first edge vector for each pixel being a vector from each pixel to an edge in the first image; acquiring events generated within a first target time period; the first target time period being after the first time; determining the motion vector of the first target time period based on the first edge vector of the pixel corresponding to the event, the motion vector of the first target time period being used for optical image stabilization of the camera module. This method achieves event-based motion vector detection, improves the detection accuracy of motion vectors, thereby improving the optical image stabilization performance of the camera module and enhancing the imaging quality of the electronic device.

[0006] In one possible implementation, obtaining the first image at a first instant includes: obtaining a raw image captured by a CMOS image sensor located in the camera module at the first instant; or, obtaining a target format image generated at the first instant, converting the target format image into a grayscale image, and using it as the first image; the target format image is generated based on the raw image captured by the image sensor; or, obtaining an event accumulation frame as the first image; the event accumulation frame is generated based on events generated within a preset first time period, the first time period being before the first instant or ending at the first instant. This provides multiple possible implementations for obtaining the first image at a first instant, making the acquisition of the first image more flexible and expanding the possible applicable scenarios of the method.

[0007] In one possible implementation, determining the motion vector of the first target time period based on the first edge vector of the pixel corresponding to the event includes: for each event within the first target time period, determining the motion vector of the event based on the first edge vector of the pixel corresponding to the event; and determining the motion vector of the first target time period based on the motion vector of the event within the first target time period. This method first determines the motion vector of the event, and then determines the motion vector of the first target time period based on the motion vector of the event within the first target time period, thereby improving the detection accuracy of the motion vector of the first target time period, and thus improving the optical image stabilization performance of the camera module and the imaging quality of the electronic device.

[0008] In one possible implementation, determining the motion vector of an event based on the first edge vector of the pixel corresponding to the event includes: determining the motion vector of the event based on the first edge vector of the pixel corresponding to the event and the first edge gradient of the pixel corresponding to the event, wherein the first edge gradient of the pixel corresponding to the event is the gradient of the nearest neighbor match of the pixel corresponding to the event in the edge of the first image. This method further combines the edge gradient of the pixel with the first edge vector of the pixel to determine the motion vector of the event, thereby improving the accuracy of determining the motion vector of the event, and thus improving the detection accuracy of the motion vector, improving the optical image stabilization performance of the camera module, and improving the imaging quality of the electronic device.

[0009] In one possible implementation, determining the motion vector for a target time period based on the motion vectors of events within a first target time period includes: clustering the motion vectors of events within the first target time period to obtain multiple categories of motion vectors; and determining the motion vector for the first target time period based on the motion vector of the target category among the multiple categories. This method clusters the motion vectors of events within the first target time period into multiple categories, and determines the motion vector for the first target time period based on the motion vector of the target category. This allows for the selection of the target category in practical applications based on different application scenarios and different optical image stabilization effects, thereby improving the detection accuracy of motion vectors, enhancing the optical image stabilization performance of the camera module, and improving the imaging quality of the electronic device.

[0010] In one possible implementation, determining the motion vector of the first target time period based on the motion vector of events within the first target time period includes: calculating the average value of the motion vectors of events within the first target time period as the motion vector of the first target time period.

[0011] In one possible implementation, before determining the motion vector for the first target time period based on the motion vectors of events within that time period, the method further includes: filtering the motion vectors of events within the first target time period to obtain valid motion vectors within that time period; and determining the target motion vector for the first target time period based on the motion vectors of events within that time period, which includes: determining the motion vector for the first target time period based on the valid motion vectors within that time period. This method, by filtering the motion vectors within the first target time period, can identify valid motion vectors, thus making the motion vector for the first target time period obtained based on these valid motion vectors more accurate. This improves the detection accuracy of the motion vectors, thereby enhancing optical image stabilization performance and the imaging quality of electronic devices.

[0012] In one possible implementation, the motion vectors of events within a target time period are filtered for validity to obtain valid motion vectors within that time period. This includes: for each event's motion vector within the first target time period, if the event's polarity information and the polarity information of the previous event for the corresponding pixel are different, the event's motion vector is determined to be a valid motion vector. This method, by filtering the validity of time-based motion vectors based on the event's device model information, improves the detection accuracy of motion vectors, thereby enhancing optical image stabilization performance and the imaging quality of electronic devices.

[0013] In one possible implementation, before acquiring the first image, the method further includes: receiving a startup request from a camera application; or receiving a photo-taking request; or receiving a video recording request. Therefore, this method can be applied to scenarios such as taking photos, recording video in response to a user's video recording request, or capturing preview video in response to a startup request from a camera application, improving the detection accuracy of motion vectors in these scenarios, thereby improving optical image stabilization performance and the imaging quality of electronic devices.

[0014] In one possible implementation, when using a camera module to capture video images, the method further includes: obtaining a second image when a change in the subject in the video is detected at a second time relative to the subject at a first time in the video; determining a second edge vector for each pixel in the second image; the second edge vector for each pixel is a vector from each pixel to an edge in the second image; obtaining events generated within a second target time period; the second target time period is located after the second time; determining a motion vector for the second target time period based on the second edge vector of the pixel corresponding to the event, and the motion vector for the second target time period is used for optical image stabilization of the camera module. In this method, when using a camera module to capture video images, the first edge vector of a pixel can be updated to a second edge vector, and then the motion vector for the second target time period can be determined based on the second edge vector of the pixel corresponding to the event. This allows the edge vector of the pixel to change with the change of the subject, improving the detection accuracy of motion vectors in the captured video image scene, thereby improving optical image stabilization performance and the imaging quality of the electronic device.

[0015] In a second aspect, embodiments of this application provide a camera module, including: a processor, which is configured to execute computer program instructions stored in a memory, wherein when the computer program instructions are executed by the processor, the camera module is triggered to execute the method of any one of the first aspects.

[0016] Thirdly, embodiments of this application provide an electronic device including the camera module of the second aspect.

[0017] Fourthly, embodiments of this application provide an electronic device, including: a processor and a memory; wherein one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the processor, cause the electronic device to perform the method of any one of the first aspects.

[0018] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method of any one of the first aspects.

[0019] Fifthly, embodiments of this application provide a computer program product, which includes a computer program that, when run on a computer, causes the computer to perform the method of any one of the first aspects. Attached Figure Description

[0020] Figure 1 is a schematic diagram of an electronic device provided in an embodiment of this application;

[0021] Figure 2 is a schematic diagram of the software structure of an electronic device provided in an embodiment of this application;

[0022] Figure 3 is a schematic diagram of an optical image stabilization system provided in an embodiment of this application;

[0023] Figure 4 is a schematic diagram of an imaging sensor provided in an embodiment of this application;

[0024] Figure 5 is a schematic diagram of a second structure of the optical image stabilization system provided in an embodiment of this application;

[0025] Figure 6A is a schematic diagram of a third structure of the optical image stabilization system provided in the embodiments of this application;

[0026] Figure 6B is a schematic diagram of the fourth structure of the optical image stabilization system provided in the embodiments of this application;

[0027] Figure 7A is a schematic diagram of the fifth structure of the optical image stabilization system provided in the embodiments of this application;

[0028] Figure 7B is a schematic diagram of the sixth structure of the optical image stabilization system provided in the embodiments of this application;

[0029] Figure 8 is a schematic flowchart of a motion vector detection method provided in an embodiment of this application;

[0030] Figure 9 is a schematic diagram of the working principle of the dynamic vision sensor provided in the embodiment of this application;

[0031] Figure 10 is another schematic flowchart of the motion vector detection method provided in the embodiments of this application;

[0032] Figure 11 is a schematic diagram of the third process of the motion vector detection method provided in the embodiment of this application;

[0033] Figure 12 is a schematic diagram comparing the edges and edge feature points provided in the embodiments of this application;

[0034] Figure 13 is a schematic diagram of the fourth process of the motion vector detection method provided in the embodiments of this application;

[0035] Figure 14 is a schematic diagram of the fifth type of motion vector detection method provided in the embodiments of this application;

[0036] Figure 15 is a schematic diagram of an application scenario of the motion vector detection method provided in the embodiments of this application;

[0037] Figures 16A and 16B are schematic diagrams of another interface of the motion vector detection method provided in the embodiments of this application;

[0038] Figure 17 is a schematic diagram of a processing module provided in an embodiment of this application;

[0039] Figure 18 is a schematic diagram of the sixth type of motion vector detection method provided in the embodiments of this application;

[0040] Figure 19 is a schematic diagram of the seventh type of motion vector detection method provided in the embodiments of this application. Detailed Implementation

[0041] The terminology used in the implementation section of this application is for the purpose of explaining specific embodiments of this application only, and is not intended to limit this application.

[0042] Optical image stabilization (OIS) is a physical method to counteract image shake caused by hand tremors or other external factors during image or video capture. Motion detection technology based on gyroscope sensors is currently a key technology in OIS. Electronic devices can be equipped with gyroscope sensors and camera modules. The gyroscope sensor detects the amount of shake in the electronic device and provides this information to a servo system in the camera module. The servo system then quickly adjusts the position of the lens (or lens group) and / or sensors in the camera module based on the amount of shake to compensate for the motion, thereby keeping the image of the subject constant on the image plane, achieving optical image stabilization.

[0043] However, gyroscope sensors suffer from long-term drift and other issues, making it impossible to detect accurate jitter during shooting, especially during long-exposure photography. Inaccurate gyroscope sensor results will cause the image of the photographed object to move on the image plane, resulting in blurred images and poor image quality of electronic devices.

[0044] To this end, this application also provides a motion vector detection method, a camera module, and an electronic device. By introducing a dynamic vision sensor (DVS), and combining the high dynamic range and low latency characteristics of the DVS, more accurate motion vectors can be detected, thereby improving the performance of optical image stabilization and improving image quality.

[0045] In this embodiment, the motion vector to be detected is used to describe the motion state of the camera module (or electronic device) on the image plane. The image plane is the plane on which an object can be clearly imaged through the lens; it can also be considered as the plane where the CMOS image sensor and the dynamic vision sensor are located.

[0046] Dynamic vision sensors, also known as event cameras or event-based dynamic vision sensors, can asynchronously output brightness change information of the image plane with high temporal resolution. They are especially suitable for low-light and fast-moving scenes, such as motion vector detection in optical image stabilization in low-light long-exposure scenes.

[0047] The motion vector detection method provided in this application can be applied to camera modules with OIS function, as well as to electronic devices such as mobile phones, cameras, and gimbals with OIS function.

[0048] Figure 1 shows a schematic diagram of the structure of electronic device 100. Electronic device 100 may include processor 110, internal memory 121, camera 193, etc. As shown in Figure 1, electronic device 100 may also include external memory interface 120, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0049] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0050] Processor 110 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0051] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0052] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0053] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0054] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0055] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0056] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1. Camera 193 can be implemented using one or more camera modules.

[0057] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0058] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0059] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0060] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.

[0061] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.

[0062] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the HarmonyOS system with a layered architecture as an example to illustrate the software structure of electronic device 100.

[0063] Figure 2 shows a software structure block diagram of an electronic device provided in an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. This embodiment uses the HarmonyOS system as an example to illustrate the software structure of an electronic device based on the HarmonyOS system. In some embodiments, the HarmonyOS system is divided into four layers, from top to bottom: the application layer, the application framework layer, the basic services layer, and the kernel layer.

[0064] The application layer can include several applications (hereinafter referred to as applications), such as image capture applications, gallery, calendar, WLAN, etc. The aforementioned image capture application can be, for example, the camera application shown in Figure 2, or other applications that can perform image capture.

[0065] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer, including various components and services to support developers' HarmonyOS development.

[0066] Basic services are the core capabilities of the HarmonyOS system, providing services to applications through the framework layer.

[0067] The kernel layer is the layer between hardware and software. The kernel layer can include: display drivers, camera drivers, audio drivers, sensor drivers, etc.

[0068] The motion vector detection method of this application will be described in detail in the following embodiments in conjunction with the above-described structure of the electronic device.

[0069] Figure 3 is a schematic diagram of an optical image stabilization system applicable to the motion vector detection method of this application. As shown in Figure 3, it includes: an optical lens group 301, a CMOS image sensor (CIS) 302, a dynamic vision sensor 303, a processing module 304, and a servo system 305.

[0070] The optical lens group 301 is used to collect the imaging beam and transmit the beam to the CMOS image sensor 302 and the dynamic vision sensor 303 respectively.

[0071] The CMOS image sensor 302 is used to generate images based on a light beam and send the images to the processing module 304. Specifically, the CMOS image sensor 302 is used to acquire images in photography or video recording scenarios and send the acquired images to an image capture application to generate photos or video frames. For example, in a photography scenario, the CMOS image sensor 302 can acquire images with different exposure durations and send these images to the image capture application to generate a photograph; in a video preview or video recording scenario, the CMOS image sensor 302 can acquire images at a certain frequency and send these images to the image capture application to generate video frames (recorded or preview videos). Each frame sent by the CMOS image sensor 302 to the processing module 304 can have a timestamp to indicate the image generation time. Multiple consecutive frames of images transmitted by the CMOS image sensor 302 to the processing module 304 can be referred to as an image stream.

[0072] A CMOS image sensor, also known as a photosensitive element, is a semiconductor chip containing hundreds of thousands to millions of photodiodes on its surface. When illuminated, these photodiodes generate electrical charges. A CMOS image sensor can be a complementary metal-oxide semiconductor (CMOS) device. CMOS devices primarily utilize silicon and germanium to create semiconductors where N-type (negative) and P-type (positive) semiconductors coexist. The current generated by these complementary effects can be recorded and interpreted by the processing chip as an image.

[0073] The dynamic vision sensor 303 generates events based on light beams and sends these events to the processing module 304. A series of events output by the dynamic vision sensor 303 can constitute an event stream.

[0074] Unlike CMOS image sensors used for image acquisition, dynamic vision sensors can be simply understood as sensors that "only sense moving objects." Each pixel in a dynamic vision sensor has an independent photoelectric sensing module. When the brightness change at that pixel exceeds a set threshold, it generates and outputs event information (sometimes called pulse data). Furthermore, because all pixels operate independently, the data output of a dynamic vision sensor is asynchronous and spatially sparse. This is the biggest difference between dynamic vision sensors and conventional image sensors. Dynamic vision sensors are event-driven photoelectric sensors. Each pixel in a dynamic vision sensor independently senses changes in light intensity. Pixels whose light intensity changes exceed a threshold are considered active pixels, and then the row and column position, polarity position, timestamp, and other information of the active pixels are packaged, encoded, and output in real time.

[0075] Based on the working principle of dynamic vision sensors, when the logarithmic brightness change of a pixel exceeds a preset threshold, the dynamic vision sensor can generate an event. Each event generated by the dynamic vision sensor can include the following parameters: pixel position (x, y), event generation time t, and event polarity p. The pixel position (x, y) records the row and column position of the pixel. The event generation time t records the generation time of the event, also known as the event timestamp, and in some embodiments, it can be at the microsecond level of precision. The event polarity p records the polarity of the event; in other words, it records the direction of the logarithmic brightness change of the pixel. In some embodiments, when the logarithmic brightness of the pixel increases, the event polarity p is positive, and when the logarithmic brightness decreases, the event polarity p is negative. In some embodiments, the data format of each event can be (x, y, p, t), representing the event generated by the pixel at coordinate (x, y) at time t, with event polarity p.

[0076] The processing module 304 is used to calculate the motion vector of each detection cycle according to the image stream and event stream according to the preset detection cycle, and generate control commands for the servo system 305 based on the motion vector of each detection cycle, so as to achieve continuous optical image stabilization according to the detection cycle within a certain time.

[0077] The servo system 305 is used to move the optical lens group 301, or the CMOS image sensor 302 and the dynamic vision sensor 303, or the optical lens group 301, the CMOS image sensor 302 and the dynamic vision sensor 303, based on control commands, to compensate for the motion vector in each detection cycle and achieve optical image stabilization.

[0078] It should be noted that during the implementation of optical image stabilization, the physical relative positions between the CMOS image sensor 302 and the dynamic vision sensor 303 are fixed. In other words, the servo system 305 needs to ensure that both move by the same amount of displacement each time to maintain their physical relative positions. Therefore, for ease of explanation, the CMOS image sensor 302 and the dynamic vision sensor 303 will be collectively referred to as the imaging sensor in the following embodiments.

[0079] Specifically, the processing module 304 can project the motion vector onto the image plane in the x and y directions based on the motion vector of each detection cycle, thereby calculating the displacement of the optical lens group 301 and / or the imaging sensor on the image plane to compensate for the projection of the motion vector in the x and y directions. The displacement of the optical lens group 301 and / or the imaging sensor on the image plane is then sent to the servo system 305, which moves the optical lens group 301 and / or the imaging sensor according to the displacement.

[0080] In other embodiments, the x and y directions of the projection of the motion vector in each detection cycle can be distinguished according to the movement mode of the electronic device. For example, when the movement mode of the electronic device is translation, the x and y directions can be from left to right and from top to bottom, respectively; when the movement mode of the electronic device is rotation, the x and y directions can be the yaw direction and the pitch direction, respectively.

[0081] Based on the above optical image stabilization system, after the imaging beam is collected by the optical lens group 301, it is transmitted to the CMOS image sensor 302 and the dynamic vision sensor 303 respectively. The CMOS image sensor 302 generates an image stream and outputs it to the processing module 304. The dynamic vision sensor 303 generates an event stream and outputs it to the processing module 304. The processing module 304 calculates the motion vector for each detection cycle based on the event stream, or the event stream and the image stream. Based on the motion vector for each detection cycle, it controls the servo system 305, so that the servo system 305 moves the optical lens group 301 and / or the imaging sensor to achieve optical image stabilization.

[0082] It should be noted that in the above optical image stabilization system, the CMOS image sensor 302 is an optional device. When the processing module 304 calculates the motion vector of each detection cycle based only on the event stream, the CMOS image sensor 302 can be omitted.

[0083] In other embodiments provided in this application, the processing module 304 may also directly send the motion vector of each detection cycle to the servo system 305, which then moves the optical lens group 301 and / or the imaging sensor based on the motion vector of each detection cycle to achieve optical image stabilization.

[0084] In some embodiments, the CMOS image sensor 302 and the dynamic vision sensor 303 can be separately configured. In other words, the CMOS image sensor and the dynamic vision sensor operate independently. For example, the dynamic vision sensor can be configured around the CMOS image sensor. The CMOS image sensor and the dynamic vision sensor are controlled by two different IC chips. The photosensitive unit on the CMOS image sensor can only acquire RGB pixels and output image information, while the photosensitive unit on the dynamic vision sensor can only acquire DVS pixels and output event information or event data. For example, as shown in Figure 4(a), four dynamic vision sensors are configured around the CMOS image sensor. It is understood that Figure 4(a) is only a schematic diagram of a possible configuration of the CMOS image sensor and the dynamic vision sensor. In some possible implementations, one dynamic vision sensor can be configured around the CMOS image sensor to output event information, or multiple dynamic vision sensors can be configured. The dynamic vision sensors can be configured at any position around the CMOS image sensor. The embodiments of this application do not limit the position or number of dynamic vision sensors around the CMOS image sensor.

[0085] When the CMOS image sensor 302 and the dynamic vision sensor 303 are separately configured, in some embodiments, the servo system 305 may include motors for the optical lens group 301, the CMOS image sensor 302, and the dynamic vision sensor 303 respectively. Thus, the servo system 305 can control the movement of the optical lens group 301, or control the movement of the CMOS image sensor 302 and the dynamic vision sensor 303, or control the movement of the optical lens group 301, the CMOS image sensor 302, and the dynamic vision sensor 303 respectively, by driving different motors or combinations of motors. In other embodiments, as shown in FIG5, in order to achieve the same displacement movement of the CMOS image sensor 302 and the dynamic vision sensor 303, they can be configured on the same gimbal A. The servo system 305 may include gimbal A and motors for the optical lens group 301 and gimbal A respectively. Thus, the servo system 305 can also control the movement of the optical lens group 301 and / or control the movement of gimbal A by driving different motors or combinations of motors.

[0086] As shown in Figures 6A and 6B, when the CMOS image sensor 302 and the dynamic vision sensor 303 are separately configured, the optical lens group 301 can be implemented by a beam-splitting optical lens group in order to transmit the light beam to the image sensor 302 and the dynamic vision sensor 303 respectively. The beam-splitting optical lens group can split the received imaging light beam into two beams, which are then transmitted to the CMOS image sensor 302 and the dynamic vision sensor 303 respectively.

[0087] In other embodiments, as shown in Figures 7A and 7B, the CMOS image sensor 302 and the dynamic vision sensor 303 can be implemented using a fused sensor that integrates the functions of both the CMOS image sensor and the dynamic vision sensor. In this case, the optical lens group 301 can transmit the imaging beam to the fused sensor without beam splitting.

[0088] A fusion sensor refers to a sensor whose photosensitive unit can collect both red, green, and blue (RGB) pixels and DVS pixels; or, part of the sensor's photosensitive unit is used to collect RGB pixels and another part is used to collect DVS pixels.

[0089] Here, RGB pixels refer to the pixels required for a CMOS image sensor to output image information, while DVS pixels refer to the pixels required for a dynamic vision sensor to output event information based on brightness changes.

[0090] Figure 4(b) shows a schematic diagram of a fusion sensor. A fusion sensor can share pixels on the same integrated circuit (IC) chip to simultaneously acquire RGB pixels and DVS pixels, and output image information and event information. Figure 4(c) shows a schematic diagram of another type of fusion sensor. In a fusion sensor, the IC chip can also not share pixels. On each photosensitive unit, some pixels are used to acquire DVS pixels, and other pixels are used to acquire RGB pixels.

[0091] In some embodiments, the servo system 305 may provide motors for the optical lens group 301 and the fusion sensor respectively, so that the servo system 305 can control the movement of the optical lens group 301, or control the movement of the fusion sensor, or control the simultaneous movement of the optical lens group 301 and the fusion sensor by driving the motor or a combination of motors.

[0092] The processing module of the optical image stabilization system in this application embodiment can be located in the processor of the camera module or in the processor of the electronic device. Examples are as follows:

[0093] As shown in Figure 6A, the beam splitting optical lens group, CMOS image sensor, dynamic vision sensor and servo system in the above optical image stabilization system can be located in the camera module, and the processing module can be located in the processor of the camera module.

[0094] As shown in Figure 6B, the beam splitting optical lens group, CMOS image sensor, dynamic vision sensor and servo system in the above optical image stabilization system can be located in the camera module, and the processing module can be located in the processor of the electronic device.

[0095] As shown in Figure 7A, the optical lens group, fusion sensor and servo system in the above optical image stabilization system can be located in the camera module, and the processing module can be located in the processor of the camera module.

[0096] As shown in Figure 7B, the optical lens group, fusion sensor and servo system in the above optical image stabilization system can be located in the camera module, and the processing module can be located in the processor of the electronic device.

[0097] In some embodiments, the electronic device may include an image signal processor (ISP), and the processing module 303 may be located in the ISP processor of the electronic device.

[0098] Based on the software structure of the electronic device shown in Figure 2, the processing module in the optical image stabilization system of this application embodiment can be located in the kernel layer, framework layer, or application layer, etc., and this application embodiment is not limited thereto. For example, the processing module can be set in the camera driver, or it can be set in an image capturing application such as a camera application.

[0099] It should be noted that the implementation structure of the servo system 305 in the above embodiments is not intended to limit the specific implementation structure of the servo system 305 in this application embodiment, as long as it can move the optical lens group 301 and / or the imaging sensor according to the control instructions of the processing module. For example, the servo system 305 may also include a gimbal B and a motor corresponding to the gimbal B. The optical lens group 301, the CMOS image sensor 302, and the dynamic vision sensor 303 can be disposed on the gimbal B, so that the servo system 305 can control the movement of the gimbal B by driving the motor corresponding to the gimbal B, thereby realizing the simultaneous movement of the optical lens group 301, the CMOS image sensor 302, and the dynamic vision sensor 303 by the same amount of displacement, without having to drive the motors corresponding to the three separately, and so on.

[0100] The motor in this application embodiment is a device that converts electrical energy into mechanical energy. A motor can also be called a motor. In some related technologies, motors are mostly used in industries such as industry, transportation, and aviation that require higher power, while electric motors are mostly used in home appliances and electronic components that require lower power.

[0101] It should be noted that the above embodiments use an optical image stabilization system comprising one camera module, where the processing module controls the camera module based on motion vectors detected by the system to achieve optical image stabilization. In other embodiments, the optical image stabilization system may include multiple camera modules. The processing module can send control commands determined based on motion vectors to the servo system of each of the multiple camera modules, thereby achieving synchronous movement of the multiple camera modules. The multiple camera modules mentioned above can be the same or different camera modules. The same camera module refers to a camera module with the same lens group configuration and similar or identical optical performance. For example, camera modules can be divided into main camera modules, telephoto modules, wide-angle modules, etc. The multiple camera modules included in the optical image stabilization system of this application embodiment can be the same or different camera modules among the above three types.

[0102] Figure 8 is a schematic flowchart of a motion vector detection method according to an embodiment of this application. This method can be executed by the processing module in the aforementioned optical image stabilization system. As shown in Figure 8, the method may include:

[0103] Step 801: Obtain the first image at the first moment.

[0104] The first image can be a raw image generated by a CMOS image sensor, or it can be a grayscale image converted from an image of a certain preset format, or it can be an event accumulation frame generated based on events generated by a dynamic visual sensor over a period of time.

[0105] The following explanation of the event accumulation frame, using Figure 9 as an example, illustrates the working principle of the dynamic vision sensor. As shown in Figure 9, x and y represent the x-axis and y-axis of the captured image, and the t-axis represents time. As time changes, when the disk rotates clockwise, the position of the black dots on the disk changes, resulting in a change in pixel brightness. During this process, the dynamic vision sensor receives the brightness change of the pixels in the captured image and generates a series of events. These events are sequentially output to the processing module as an event stream. An event accumulation frame is an image obtained by projecting the three-dimensional point cloud of the event stream over a period of time onto the xy-plane along the t-axis dimension. The three-dimensional point cloud refers to a dataset of three-dimensional coordinate points arranged according to a regular grid. In other words, the event accumulation frame records whether an event has occurred at each pixel in the captured image over a period of time. In some embodiments, the pixel value of each pixel in the event accumulation frame can be 0 or 1. 0 and 1 are used to record whether an event has occurred at that pixel over a period of time; for example, 0 indicates that no event has occurred at that pixel over a period of time, and 1 indicates that an event has occurred at that pixel over a period of time.

[0106] In some embodiments, this step may specifically include: obtaining a raw image captured by a CMOS image sensor in a first moment as a first image.

[0107] In other embodiments, this step may specifically include: obtaining a target format image generated in a first time, and converting the target format image into a grayscale image as a first image.

[0108] The target format image can be generated based on one or more raw images captured by an image sensor. For example, the target format image can be a multi-channel image such as an RGB image, HSV image, or YUV image.

[0109] The target format image mentioned above could be, for example, a video frame from an image stream generated by an image capture application based on a CMOS image sensor.

[0110] In other embodiments, this step may specifically include: obtaining an event accumulation frame for a first time period T1 as a first image. The event accumulation frame for the first time period T1 refers to an event accumulation frame generated based on events occurring within the first time period. The first time period T1 may be located before the first time or may end at the first time.

[0111] Step 802: Determine the first edge vector of each pixel in the first image; the first edge vector of each pixel is the vector from each pixel to the edge in the first image.

[0112] In an image, an edge refers to a set of pixels where features such as brightness, color, or texture change abruptly. These changes typically represent the boundaries of different objects within the image. Image edges can be detected using edge detection methods. Edge detection methods generally calculate the difference between each pixel in the image and its neighboring pixels. This difference can be quantized using edge detection operators, and pixels whose differences from their neighbors exceed a threshold are identified as edges. Edge detection operators can include, but are not limited to, the Canny operator, Roberts operator, Prewitt operator, Sobel operator, and Laplacian operator.

[0113] Specifically, the first edge vector of a pixel can be the vector from the pixel to its nearest neighbor along the edges of the first image. The nearest neighbor along the edges of the first image refers to the pixel with the smallest distance from the pixel along the edges of the first image. In this case, the magnitude of the first edge vector can be the distance between the pixel and its nearest neighbor, and the direction can be the direction in which the pixel points towards its nearest neighbor.

[0114] The specific implementation of this step can be referred to the description of the embodiments shown in Figures 10 and 11, which will not be repeated here.

[0115] Step 803: Obtain the events generated within the first target time period; the first target time period is located after the first time.

[0116] In this step, events whose generation time t is within the first target time period can be obtained based on the event's generation time t, thus obtaining the events generated within the first target time period.

[0117] The first target time period in this step is the time period following the first time. In some embodiments, the first target time period may be a detection cycle of the motion vector.

[0118] The implementation of the motion vector detection period is not limited in this application embodiment; different detection periods can be the same or different in duration. The following is an exemplary description.

[0119] In some embodiments, the processing module can divide the detection cycle according to a preset duration. The specific value of the preset duration is not limited in this application embodiment. For example, it can be a value between 1 and 10 ms.

[0120] In other embodiments, the processing module may use the receipt of a preset number of events as a detection cycle. The specific value of the preset number is not limited in this embodiment; for example, the preset number could be 500, meaning that the processing module considers every 500 events received as a detection cycle. For instance, the processing module considers the period from the end of the previous detection cycle to the receipt of the 500th event as one detection cycle. It should be noted that the time it takes for the dynamic vision sensor to generate events is not fixed; therefore, the time it takes for the processing module to receive the same number of events is not fixed.

[0121] Step 804: Determine the motion vector of the first target time period based on the first edge vector of the pixel corresponding to the event. The motion vector of the first target time period is used for optical image stabilization of the camera module.

[0122] The implementation of this step can be referred to the description of the embodiment shown in Figure 13, which will not be repeated here.

[0123] It is understood that steps 801-802 can be considered as the steps for determining the first edge vector of the pixel, and steps 803-804 are the steps for determining the motion vector. In some embodiments, after executing steps 801-802, the processing module can execute steps 803-804 multiple times to determine multiple motion vectors for the first target time period, and then implement optical image stabilization of the camera module based on the motion vector of each first target time period. For example, when the first target time period is a detection cycle of the motion vector, steps 803-804 can be executed according to the detection cycle, thereby achieving optical image stabilization of the camera module according to the detection cycle.

[0124] In the method shown in Figure 8, the first edge vector of each pixel is first calculated based on the first image. Then, based on the events generated within the first target time period, the motion vector of the first target time period is determined according to the first edge vector of the pixel corresponding to the event. This achieves event-based motion vector detection, improves the detection accuracy of motion vectors, and thus improves the optical image stabilization performance of the camera module and the imaging quality of the electronic device.

[0125] The specific implementation of step 802 will be illustrated below with reference to Figures 10 and 11.

[0126] As shown in Figure 10, step 802 may include:

[0127] Step 1001: Extract the edges of the first image to obtain the edges of the first image.

[0128] When the first image is a raw image or a grayscale image obtained by converting a target format image, edge extraction of the first image can be achieved through relevant edge detection algorithms.

[0129] The basic principle of edge detection algorithms is to identify pixels whose difference from neighboring pixels exceeds a difference threshold as pixels included in the image's edge. Pixels detected based on this principle may form complete edge contours or discontinuous edge segments. Therefore, some edge detection algorithms can further perform edge connection processing on discontinuous edge segments detected based on the above principle, so that the image edges include complete edge contours. The specific value of the aforementioned difference threshold is not limited in the embodiments of this application.

[0130] Therefore, depending on the edge detection algorithm used, the edges extracted in this step of the first image may include: complete edge contours, and / or, discontinuous edge fragments.

[0131] When the first image is an event accumulation frame, edge extraction of the first image can be achieved through morphological processing. For details, please refer to relevant technologies. This application will not elaborate further in its embodiments.

[0132] Step 1002: Determine the first edge vector of each pixel in the first image based on the edges of the first image.

[0133] In this step, for each pixel in the first image, the nearest neighbor match of the pixel in the edge of the first image can be determined first, that is, the pixel with the smallest distance from the pixel in the edge of the first image can be determined. Then, the vector from the pixel to the nearest neighbor match is determined as the first edge vector of the pixel.

[0134] Step 1003: Store the first edge vector of each pixel in the first image.

[0135] The storage location of the first edge vector of each pixel in the first image in this step is not limited in this embodiment of the application, as long as the subsequent processing module can read the first edge vector of each pixel in the first image.

[0136] As shown in Figure 11, in order to reduce the data processing volume of calculating the first edge vector of each pixel in this embodiment and improve the calculation speed of the first edge vector of each pixel in the first image, the above step 1002 can also be replaced by the following steps 1101 to 1102.

[0137] Step 1101: Extract edge feature points from the edges of the first image.

[0138] In some embodiments, pixels with gradient values ​​exceeding a preset gradient threshold can be extracted from the edges of the first image as edge feature points.

[0139] The gradient of a pixel refers to the change in brightness of that pixel, which can be obtained by calculating the brightness difference between surrounding pixels. Common gradient operators include the Sobel operator and the Scharr operator. The calculation of the gradient value of each pixel in the edge during this step can be implemented using relevant methods, and this embodiment does not impose any limitations. The specific value of the preset gradient threshold is also not limited in this embodiment.

[0140] It should be noted that in some edge detection algorithms, the gradient of each pixel can be calculated first based on the pixel brightness, and then the edges in the image can be identified based on the pixel gradients. If the gradient of each pixel included in the edge has already been calculated when performing edge extraction on the first image in the aforementioned step 1001, then the calculation result in step 1001 can be directly obtained in this step without recalculation.

[0141] Step 1102: Determine the first edge vector of each pixel in the first image based on the edge feature points.

[0142] In this step, for each pixel in the first image, the nearest neighbor match of the pixel in the edge feature points of the first image can be determined first, that is, the edge feature point with the smallest distance to the pixel in the edge feature points of the first image can be determined. Then, the vector from the pixel to the above nearest neighbor match is determined as the first edge vector of the pixel.

[0143] Figure 12 shows a comparison diagram between the edges of the first image and the edge feature points of the first image. It can be seen that by extracting the edge feature points in step 1101, the edges of the first image can be further reduced from complete edge contours and / or discontinuous edge fragments to discrete pixels. This reduces the amount of data processing required to determine the nearest neighbor matching of each pixel in step 1102, and thus reduces the amount of data processing required for the processing module to determine the first edge vector of each pixel.

[0144] The specific implementation of step 804 described above will be illustrated below with reference to Figure 13. As shown in Figure 13, step 604 may specifically include:

[0145] Step 1301: For each event within the first target time period, determine the motion vector of the event based on the first edge vector of the pixel corresponding to the event.

[0146] In some embodiments, the motion vector of an event can be determined based on the first edge vector and the first edge gradient of the pixel corresponding to the event. Specifically, the motion vector of the event can be obtained by calculating the dot product of the first edge vector and the first edge gradient of the pixel corresponding to the event.

[0147] Specifically, the first edge gradient of a pixel can be the gradient of its nearest neighbor match within the edges of the first image; in other words, it is the gradient of the pixel closest to the pixel within the edges of the first image. The calculation method for the pixel gradient can be implemented using relevant technologies, and this application embodiment does not impose any limitations. It should be noted that if the gradient of each pixel included in the edge is already calculated during edge extraction of the first image in step 1001, the calculation result from step 1001 can be directly obtained in this step without recalculation.

[0148] Step 1302: Determine the motion vector of the first target time period based on the motion vector of the events within the first target time period.

[0149] In some embodiments, the average or median value of the motion vectors of events within a first target time period can be calculated as the motion vector for the first target time period. The implementation of calculating the average or median value of multiple motion vectors can be found in related technologies, and will not be elaborated here.

[0150] In other embodiments, the motion vectors of events within a first target time period can be clustered to obtain multiple categories of motion vectors; the motion vector of the first target time period can be determined based on the motion vector of the target category among the multiple categories.

[0151] When performing clustering, the number of clusters and the criteria for dividing the clusters can be set independently according to the use case.

[0152] It is understandable that when a user takes a picture using an electronic device, the captured image can typically be divided into a foreground image and a background image. Based on this, in some embodiments, the images can be clustered into two categories: foreground and background. In this case, the motion vectors of events whose corresponding pixels are located in the foreground image can be clustered into the foreground category, and the motion vectors of events whose corresponding pixels are located in the background image can be clustered into the background category, thereby dividing the motion vectors of the events into two categories.

[0153] The target categories mentioned above can be determined based on different application scenarios and different optical image stabilization effects. Taking clustering into foreground and background as an example, in portrait photography or panning shots, the foreground category can be used as the target category to solve the problem of blurred subjects in these scenarios. In photo or video shooting scenarios, the background category can be used as the target category to solve the problem of camera shake during shooting.

[0154] In some embodiments, when determining the motion vector of the first target time period based on the motion vector of the target category, the motion vector of the first target time period can be determined by methods such as calculating the average value or median value of the motion vector of the target category.

[0155] In some embodiments, the validity of the motion vectors of events within the first target time period can be verified in this step to filter out the valid motion vectors within the first target time period. Accordingly, in step 1302, the motion vectors of the first target time period can be determined based on the valid motion vectors within the first target time period. In this case, the implementation of step 1302 can refer to the foregoing description, the main difference being that the motion vectors of events within the first target time period are replaced with the valid motion vectors within the first target time period.

[0156] In some embodiments, the validity of the motion vector of each event can be verified based on the polarity information of the event. Specifically, for the motion vector of each event within a first target time period, if the polarity information of the event is different from the polarity information of the previous event of the corresponding pixel, the motion vector of the event is determined to be a valid motion vector; if the polarity information of the event is the same as the polarity information of the previous event of the corresponding pixel, the motion vector of the event is determined to be an invalid motion vector. For example, if an event is (x1, y1, p1, t1), when determining the validity of the motion vector of event (x1, y1, p1, t1), the event of the pixel at position (x1, y1) before time t1 can be obtained, assuming it is (x1, y1, p0, t0). If the polarities of p1 and p0 are the same, the motion vector of event (x1, y1, p1, t1) is invalid; if the polarities of p1 and p0 are different, the motion vector of event (x1, y1, p1, y1) is valid.

[0157] By filtering the validity of motion vectors within the first target time period, valid motion vectors can be selected. As a result, the motion vectors obtained for the first target time period based on the valid motion vectors are more accurate, thereby improving the detection accuracy of motion vectors and thus improving the optical image stabilization performance and the imaging quality of electronic devices.

[0158] In other embodiments, to improve the motion vector detection accuracy of the motion vector detection method in this application during the process of acquiring video images using a camera module, and thereby improve the optical image stabilization performance of the camera module, the following steps 1401 to 1405 can be added to the above embodiments in other embodiments provided in this application. Figure 14 shows an example of adding steps 1401 to 1405 to the embodiment shown in Figure 8.

[0159] In this embodiment, the use of a camera module to capture video images can be, for example, in a scene where an image capture application displays a preview video to the user or in a scene where a video is recorded in response to a user's recording request. In the following embodiments, the video recorded in response to a user's recording request is also referred to as a video recording.

[0160] Step 1401: Acquire the first video frame captured at the first moment and detect the subject in the first video frame; acquire video frames captured after the first moment and detect the subject in each video frame.

[0161] The video frame in this step can be a video frame (preview video or recorded video) generated by the image capture application based on the CMOS image sensor in the camera module; in other words, it can be a video frame in the video that the image capture application displays to the user.

[0162] The first video frame captured at the first moment can be, for example, a video frame generated by the image capture application at the first moment, or a video frame displayed to the user by the image capture application at the first moment. Correspondingly, video frames captured after the first moment can be, for example, video frames generated by the image capture application after the first moment, or video frames displayed to the user by the image capture application after the first moment.

[0163] The subject in a video frame refers to the object whose image is located within the video frame.

[0164] In some embodiments, an image of the subject can be obtained from a video frame, and semantic recognition can be performed on the image to obtain semantic information about the subject.

[0165] It is understandable that the number of objects in a video frame can be one or more.

[0166] It is understandable that the acquisition of the first video frame captured at the first moment can be performed during or after step 801, while the acquisition of video frames captured after the first moment can be performed separately after each video frame is generated.

[0167] Step 1402: Compare the subject in each video frame with the subject in the first video frame to obtain the comparison result.

[0168] In some embodiments, in this step, the semantic information of the subject in each video frame after the first time can be compared with the semantic information of the subject in the first video frame. If the comparison result of the semantic information of the subject in a certain video frame and the semantic information of the subject in the first video frame is different, it means that the subject in the video frame has changed relative to the subject in the first video frame. If the comparison result is the same, it means that the subject in the video frame has not changed relative to the subject in the first video frame.

[0169] In some embodiments, when the semantic information of the subject in each video frame after the first time can be matched with the semantic information of the subject in the first video frame, if for each video frame there is a subject whose semantic information cannot be matched with the semantic information of the subject in the first video frame, the comparison result is different; otherwise, it is the same.

[0170] The processing in steps 1401 to 1402 can detect whether the subject in the video has changed relative to the subject in the first moment of the video.

[0171] Step 1403: When the subject in the second video frame changes relative to the subject in the first video frame, a second image is obtained.

[0172] The second video frame is one of the video frames captured after the first time in steps 1401 and 1402.

[0173] Similar to obtaining the first image, in some embodiments, obtaining the second image in this step may specifically include: obtaining a raw image captured by the CMOS image sensor at a second time as the second image. The aforementioned second time may, for example, be the time when the second video frame is captured.

[0174] In other embodiments, obtaining the second image in this step may specifically include: obtaining a target format image generated at a second time, converting the target format image into a grayscale image, and using it as the second image.

[0175] The target format image can be generated based on one or more raw images acquired by a CMOS image sensor. The target format image can be, for example, a multi-channel image such as an RGB image, HSV image, or YUV image. In some embodiments, the second image can be, for example, a second video frame.

[0176] In other embodiments, obtaining the second image in this step may specifically include: obtaining an event accumulation frame for a second time period T2 as the second image. The event accumulation frame for the second time period T2 refers to an event accumulation frame generated based on events occurring within the second time period. The second time period T2 may be located before the second time or may end at the second time.

[0177] Step 1404: Determine the second edge vector for each pixel in the second image; the second edge vector for each pixel is the vector from each pixel to the edge in the second image.

[0178] Step 1405: Obtain the events generated within the second target time period, which is after the second time period.

[0179] Step 1406: Determine the motion vector of the second target time period based on the second edge vector of the pixel corresponding to the event. The motion vector of the second target time period is used for optical image stabilization of the camera module.

[0180] The specific implementation of steps 1404 to 1406 above can refer to the implementation of steps 802 to 804. The main difference is that the corresponding replacement of the processing object is performed, such as replacing the first image with the second image, replacing the first edge vector with the second edge vector, and replacing the first target time period with the second target time period, etc.

[0181] It is understood that in other embodiments, the processing module can further acquire video frames captured after the second time period, compare the subject in each video frame with the subject in the second video frame, and when the subject in the third video frame captured at the third time period changes relative to the subject in the second video frame, obtain a third image, determine the third edge vector of each pixel in the third image, and then determine the motion vector of the third target time period based on the third edge vector of the pixel corresponding to the event within the third target time period within the third time period. This process is repeated until the video image acquisition ends. Specifically, when the video is a preview video, the end of video image acquisition may be, for example, when the image capture application detects a user's photo or video recording request; when the video is a recorded video, the end of video image acquisition may be, for example, when the image capture application detects a user's request to end recording.

[0182] The implementation principle of the motion vector detection method of the present application embodiment shown above will be described by way of example.

[0183] Motion-generated event model based on motion vision sensors As shown by ΔL = p·C, events generated by characteristic motion mostly occur at locations with large gradients, i.e., edge locations. Here, Δt is a short time interval, and v is the velocity on the image plane. Let ΔL be the logarithmic brightness gradient and p be the logarithmic brightness change. When ΔL reaches a threshold C, an event is generated, and p is the polarity of the event. Motion vision sensors can acquire high temporal resolution edge position information by collecting events during motion. By aligning or matching the edge position information obtained from the motion vision sensor with that obtained from the CMOS image sensor, the motion vector on the image plane can be obtained.

[0184] Based on the above principles, in the motion vector detection method of this application embodiment, the first edge vector of each pixel is calculated based on the first image to obtain the edge position information obtained by the CMOS image sensor; then, events generated within the first target time period are obtained to obtain the edge position information obtained by the motion vision sensor within the first target time period; finally, the motion vector of the first target time period is determined according to the first edge vector of the pixel corresponding to the event, thereby realizing the alignment or matching processing of the edge position information obtained by the motion vision sensor and the edge position information obtained by the CMOS image sensor to obtain the motion vector on the image plane as the motion vector of the first target time period. Accordingly, the processing module can control the servo system to perform compensating motion based on the motion vector, realizing optical image stabilization. Thus, the motion vector detection method of this application embodiment utilizes the above implementation principle, following the forward processing approach, to obtain the motion vector on the image plane by obtaining the edge position information obtained by the motion vision sensor and the edge position information obtained by the CMOS image sensor respectively, and then performing alignment or matching processing, for optical image stabilization.

[0185] In some related technologies, the first-order time derivative based on the brightness field is used. With scene gradient pre-constructed field The relationship between motion vectors v Examining events generated by a dynamic visual sensor within a very short time interval Δt, we obtain information on brightness changes occurring within Δt, and preconstruct a field based on known scene gradients. Inversely solve for the direction of the motion vector v occurring within Δt, and estimate the magnitude of the motion vector v. Wherein, It is a scene gradient pre-constructed field The matrix after taking the logarithm.

[0186] Compared to the inverse motion vector detection methods in the aforementioned related technologies, the motion vector detection method of this application requires less data processing and has a faster processing speed. Moreover, when calculating the motion vector within a certain target time period (e.g., a first target time period or a second target time period), the validity of the motion vector of each event within the target time period can be verified based on the polarity information of the event, and the motion vector can be calculated based on the valid motion vector, thereby making the motion vector more accurate and reliable.

[0187] The motion vector detection method of this application embodiment can be applied to improve the optical image stabilization performance of a camera module when acquiring images. The motion vector detection method of this application embodiment will be further illustrated below with reference to the application scenarios shown in Figures 15-16B and the structure of the processing module shown in Figure 17.

[0188] In some embodiments, as shown in FIG17, the processing module of this application embodiment may specifically include: an edge vector generation module 1701, a storage module 1702, and a motion vector calculation module 1703, wherein...

[0189] The edge vector generation module 1701 is used to generate the edge vector of a pixel.

[0190] Storage module 1702 is used to store the edge vectors of pixels.

[0191] The motion vector calculation module 1703 is used to calculate the motion vector for each detection cycle based on the pixel edge vector.

[0192] Figure 18 is a flowchart illustrating a motion vector detection method according to an embodiment of this application based on the processing module structure shown in Figure 17. As shown in Figure 18, the method may include:

[0193] Step 1801: The camera application detects the startup request and sends the first instruction to the edge vector generation module and the motion vector calculation module.

[0194] Referring to Figure 15, when the user clicks the camera application icon 1511 on the device's main interface 1510, the electronic device detects the user's request to launch the camera application and launches it. The camera application detects the launch request, displays the photo preview interface 1520, triggers the camera module to capture preview video images, and sends a first instruction to the edge vector generation module and the motion vector calculation module.

[0195] The aforementioned first instruction is used to instruct the edge vector generation module and the motion vector calculation module to perform motion vector detection during the video image acquisition process of the camera module.

[0196] Step 1802: The edge vector generation module obtains image 1 and determines the edge vector 1 of each pixel in image 1.

[0197] The edge vector generation module can obtain a raw image frame captured by the CMOS image sensor in the camera module as Image 1 from the camera driver, or it can obtain a frame image of the preview video from the camera application as Image 1.

[0198] Step 1803: The edge vector generation module sends the edge vector 1 of each pixel to the storage module.

[0199] Step 1804: The storage module stores the edge vector 1 for each pixel.

[0200] Step 1805: The motion vector calculation module obtains the events generated within the detection period according to the detection period.

[0201] Step 1806: The motion vector calculation module obtains the edge vector 1 of the pixel corresponding to each event from the storage module, and determines the motion vector of the detection period based on the edge vector 1 of the corresponding pixel.

[0202] Step 1807: The edge vector generation module detects whether the subject in each video frame of the preview video has changed relative to the subject in the first video frame. If so, proceed to step 1808; otherwise, return to step 1807 and continue to detect whether the subject in the next video frame of the preview video has changed relative to the subject in the first video frame.

[0203] The first video frame could be, for example, the first frame of a preview video.

[0204] Referring to Figures 16A and 16B, when the subject in the preview video displayed to the user by the camera application changes, such as the change of the person in Figure 16A or the change of the object in Figure 16B, the judgment result in step 1807 may be yes, thus executing step 1808.

[0205] Step 1808: The edge vector generation module obtains image 2 and determines the edge vector 2 of each pixel in image 2.

[0206] Step 1809: The edge vector generation module sends the edge vector 2 of each pixel to the storage module.

[0207] Step 1810: The storage module updates the edge vector 1 of each pixel to the edge vector 2.

[0208] Step 1811: The motion vector calculation module obtains the events generated within the detection period according to the detection period.

[0209] Step 1812: The motion vector calculation module obtains the edge vector 2 of the pixel corresponding to each event from the storage module, and determines the motion vector of the detection period based on the edge vector 2 of the corresponding pixel.

[0210] Step 1813: The camera application detects the photo-taking request and sends a second instruction to the edge vector generation module and the motion vector calculation module.

[0211] Referring to Figure 15, when the user clicks the camera control 1521 in the camera preview interface 1520, the camera application can detect the camera request. The second instruction is used to instruct the edge vector generation module and the motion vector calculation module to perform motion vector detection during the camera capture.

[0212] Step 1814: The edge vector generation module obtains image 3 and determines the edge vector 3 of each pixel in image 3.

[0213] In some embodiments, after the camera application detects a photo-taking request, it may not immediately trigger the camera module to acquire the raw image required for the photo according to the exposure duration. Instead, it may first send a second instruction to the edge vector generation module and the motion vector calculation module, thereby triggering the execution of steps 1814 to 1818. At this time, the edge vector generation module can first control the camera module to acquire a frame of raw image as image 3, and then the edge vector generation module or the camera application can control the camera module to acquire the image required for the photo according to different exposure durations. Through this processing, it can be ensured that image 3 is closer to the captured image, improving the optical image stabilization performance during the photo-taking process.

[0214] Step 1815: The edge vector generation module sends the edge vector 3 of each pixel to the storage module.

[0215] Step 1816: The storage module stores the edge vector 3 for each pixel.

[0216] Step 1817: The motion vector calculation module obtains the events generated within the detection period according to the detection period.

[0217] Step 1818: The motion vector calculation module obtains the edge vector 3 of the pixel corresponding to each event from the storage module, and determines the motion vector of the detection period based on the edge vector 3 of the corresponding pixel.

[0218] Referring to Figures 15 and 18, when the user clicks the camera application icon 1511 on the device's main interface 1510, the camera application starts and displays the photo preview interface 1520. At this time, the camera application can trigger the camera module to capture preview video images, and trigger the processing module to execute the methods shown in steps 1802 to 1806, so that the camera module performs optical image stabilization based on the motion vectors detected by the motion vector detection method during the process of capturing preview video images.

[0219] Furthermore, when the subject changes in the preview video displayed to the user by the camera application (e.g., the subject changes in Figure 16A or the subject changes in Figure 16B), the processing module can be triggered to execute the steps shown in steps 1808 to 1812 to update the edge vector of each pixel stored in the storage module, thereby improving the accuracy of motion vector detection by the processing module during the acquisition of preview video images by the camera module.

[0220] When a user clicks the camera control 1521 in the photo preview interface 1520, the camera application can trigger steps 1814 to 1818 to achieve optical image stabilization of the camera module during the photo taking process.

[0221] In some other embodiments, as shown in FIG19, steps 1813 to 1818 in FIG18 are replaced by steps 1901 to 1903 below.

[0222] Step 1901: The camera application detects the photo-taking request and sends a second instruction to the motion vector calculation module.

[0223] Step 1902: The motion vector calculation module obtains the events generated within the detection period according to the detection period.

[0224] Step 1903: The motion vector calculation module obtains the edge vector 2 of the pixel corresponding to each event from the storage module, and determines the motion vector of the detection period based on the edge vector 2 of the corresponding pixel.

[0225] Referring to Figures 15 and 19, when the user clicks the photo-taking control 1521 in the photo preview interface 1520, the camera application detects the photo-taking request and triggers the camera module to acquire the required images for taking the photo according to different exposure durations. Furthermore, by sending a second instruction, the motion vector calculation module in the processing module can be triggered to execute steps 1902 to 1903. In steps 1902 to 1903, the motion vector for each detection cycle during the photo-taking process is calculated based on the most recent image acquisition (e.g., image 2 mentioned above) before the user clicks the photo-taking control 1521 and the calculated edge vector of each pixel in the image (e.g., edge vector 2 mentioned above). This achieves optical image stabilization for the camera module during the photo-taking process. In this method, the edge vector generation module does not need to calculate the edge vector during the photo-taking process, thereby shortening the calculation time of the motion vector in the first detection cycle and thus shortening the photo-taking time.

[0226] It is understood that the photo preview interface in the above embodiments can be further extended to preview interfaces of other shooting modes such as video preview interface, portrait photo preview interface, and night scene photo preview interface. Thus, when the camera application triggers the camera module to acquire the preview video image that needs to be displayed in the above interface, the processing module is triggered to execute the motion vector detection method shown in steps 1802 to 1812, so that the camera module performs optical image stabilization based on the motion vector detected by the motion vector detection method during the acquisition of the preview video image.

[0227] It is understood that the above embodiments can also be further extended to the point that when the camera application is triggered to record video, or when the camera application triggers the camera module to acquire video images, the trigger processing module executes the motion vector detection method shown in steps 1802 to 1812, so that the camera module performs optical image stabilization based on the motion vector detected by the motion vector detection method during the acquisition of video images.

[0228] It is understood that the photo-taking request in the above embodiments can be further extended to portrait photo-taking requests, night scene photo-taking requests, etc.

[0229] It is understood that the camera application in the above embodiments can also be extended to other image capture applications that support functions such as taking photos and / or recording videos.

[0230] This application provides a camera module, including a processor for executing the method provided in this application.

[0231] This application also provides an electronic device, including the aforementioned camera module.

[0232] This application also provides an electronic device, including a processor for executing the method provided in this application.

[0233] This application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to execute the method provided in this application.

[0234] This application also provides a computer program product, which includes a computer program that, when run on a computer, causes the computer to perform the method provided in this application.

[0235] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0236] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0237] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0238] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0239] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A motion vector detection method, characterized in that, The method includes the following: A camera module is provided. Obtain the first image in the first instance; Determine the first edge vector for each pixel in the first image; the first edge vector for each pixel is the vector from each pixel to an edge in the first image; Obtain events that occur within a first target time period; the first target time period is after the first time. The motion vector of the first target time period is determined based on the first edge vector of the pixel corresponding to the event, and the motion vector of the first target time period is used for optical image stabilization of the camera module.

2. The method according to claim 1, characterized in that, The process of obtaining the first image at the first moment includes: The first image is obtained by capturing a raw image from a CMOS image sensor located within the camera module; or... Obtain the target format image generated in the first instant, convert the target format image into a grayscale image, and use it as the first image; the target format image is generated based on the raw image acquired by the image sensor; or... An event accumulation frame is obtained as the first image; the event accumulation frame is generated based on events generated within a preset first time period, the first time period being before the first time or ending at the first time.

3. The method according to claim 1, characterized in that, Determining the motion vector of the first target time period based on the first edge vector of the pixel corresponding to the event includes: For each event within the first target time period, the motion vector of the event is determined based on the first edge vector of the pixel corresponding to the event. The motion vector of the first target time period is determined based on the motion vector of the events within the first target time period.

4. The method according to claim 3, characterized in that, Determining the motion vector of the event based on the first edge vector of the pixel corresponding to the event includes: The motion vector of the event is determined based on the first edge vector of the pixel corresponding to the event and the first edge gradient of the pixel corresponding to the event, wherein the first edge gradient of the pixel corresponding to the event is the gradient of the nearest neighbor match of the pixel corresponding to the event in the edge of the first image.

5. The method according to claim 3, characterized in that, Determining the motion vector of the target time period based on the motion vector of events within the first target time period includes: The motion vectors of events within the first target time period are clustered to obtain multiple categories of motion vectors; the motion vector of the first target time period is determined based on the motion vector of the target category among the multiple categories.

6. The method according to claim 3, characterized in that, Determining the motion vector of the first target time period based on the motion vector of events within the first target time period includes: Calculate the average value of the motion vectors of events within the first target time period, and use it as the motion vector of the first target time period.

7. The method according to any one of claims 3 to 6, characterized in that, Before determining the motion vector of the first target time period based on the motion vector of events within the first target time period, the method further includes: The motion vectors of events within the first target time period are filtered for validity to obtain the valid motion vectors within the first target time period. Determining the target motion vector for the first target time period based on the motion vector of events within the first target time period includes: The motion vector for the first target time period is determined based on the valid motion vector within the first target time period.

8. The method according to claim 4, characterized in that, The step of filtering the motion vectors of events within the target time period to obtain valid motion vectors within the target time period includes: For the motion vector of each event within the first target time period, if the polarity information of the event is different from the polarity information of the previous event of the corresponding pixel, the motion vector of the event is determined to be a valid motion vector.

9. The method according to any one of claims 1 to 8, characterized in that, Before obtaining the first image at the first moment, the process also includes: A request to launch the camera app has been received; or, A photo-taking request has been received; or, A recording request has been received.

10. The method according to any one of claims 1 to 9, characterized in that, When using the camera module to capture video images, the method further includes: When a change is detected in the subject in the video relative to the subject in the video at the first time, a second image is obtained; Determine the second edge vector for each pixel in the second image; the second edge vector for each pixel is the vector from each pixel to an edge in the second image; Obtain events that occur within a second target time period; the second target time period is located after the second time period. The motion vector of the second target time period is determined based on the second edge vector of the pixel corresponding to the event, and the motion vector of the second target time period is used for optical image stabilization of the camera module.

11. A camera module, characterized in that, include: A processor for executing computer program instructions stored in a memory, wherein when the computer program instructions are executed by the processor, the camera module is triggered to perform the method according to any one of claims 1 to 10.

12. An electronic device, characterized in that, Includes the camera module as described in claim 11.

13. An electronic device, characterized in that, include: Processor, memory; One or more computer programs are stored in the memory, the one or more computer programs including instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a computer, causes the computer to perform the method described in any one of claims 1 to 10.

15. A computer program product, characterized in that, The computer program product includes a computer program that, when run on a computer, causes the computer to perform the method described in any one of claims 1 to 10.